Video processing method, video processing device, equipment and storage medium
By scaling and quantizing the motion vector difference in the video block, the problem of low processing efficiency of motion vector difference under multiple reference frames is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202510279260.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2022-10-13
- Publication Date
- 2025-06-06
AI Technical Summary
In video encoding, when the motion vector difference (MVD) is signaled by a joint signal for multiple reference frames, the prior art is difficult to effectively scale and process, resulting in insufficiency of encoding.
A joint MVD scaling method for composite inter prediction of video blocks is proposed, by scaling and quantizing a single MVD, predicted motion vectors are derived, and quantized according to the predicted pixel resolution to generate quantized MVDs.
Through this method, it is possible to effectively scale and process motion vector differences, improve video encoding and decoding efficiency, reduce the amount of encoded data, and improve the compression performance of the video stream.
Smart Images

Figure CN120111252A_ABST
Abstract
Description
[0001] Incorporated by Reference
[0002] This application is based on and claims the benefit of priority of U.S. non-provisional patent application No. 17 / 962,028 filed on October 7, 2022, which is based on and claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 281,010 filed on November 18, 2021 and U.S. Provisional Application No. 63 / 289,017 filed on December 13, 2021, both of which are entitled "MVD Scaling for Joint MVD Coding". These prior patent applications are incorporated herein by reference in their entirety.
[0003] This application files a divisional application for the Chinese patent application with application number 202280009026.5, application date October 13, 2022, and invention name “MVD scaling for joint MVD coding”. Technical Field
[0004] The present disclosure relates generally to video coding, and in particular, to a method for processing video blocks of a video stream, a video processing apparatus, a device, and a non-transitory computer-readable medium. Background Art
[0005] This background description is provided herein for the purpose of generally presenting the context of the present disclosure. To the extent the work described in this background section, the work of the presently named inventors and aspects of the description that may not otherwise be qualified as prior art at the time of filing this application are neither explicitly nor implicitly admitted as prior art to the present disclosure.
[0006] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each of which has a spatial dimension of, for example, 1920×1080 luminance samples and associated full or sub-sampled chrominance samples. The series of pictures can have a fixed or variable picture rate (alternatively referred to as a frame rate) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bit rate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a chrominance subsampling of 4:2:0 at 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.
[0007] One purpose of video encoding and decoding can be to reduce the redundancy of uncompressed input video signals by compression. Compression can help reduce the bandwidth and / or storage space requirements mentioned above, in some cases reducing by two orders of magnitude or more. Both lossless compression and lossy compression and their combinations can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully retained during encoding and is not fully restored during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough to present a reconstructed signal useful for the intended application, although there is some information loss. In the case of video, lossy compression is widely used in many applications. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio that can be achieved by a specific encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows encoding algorithms that produce higher losses and higher compression ratios.
[0008] Video encoders and decoders may utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.
[0009] Video codec techniques may include techniques known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-mode, the picture may be referred to as an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, may be used to reset decoder states, and may therefore be used as the first picture in an encoded video bitstream and video session or as a still image. The samples of the block after intra-prediction may then be transformed to the frequency domain, and the transform coefficients thus generated may be quantized before entropy coding. Intra-prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficient, the fewer bits are required to represent the block after entropy coding at a given quantization step size.
[0010] Conventional intra-frame coding, such as that known from, for example, MPEG (Moving Picture Experts Group, MPEG)-2 generation coding techniques, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to encode / decode blocks based on metadata and / or surrounding sample data obtained during encoding and / or decoding of data blocks that are spatially adjacent and precede the data blocks being intra-coded or decoded in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and does not use reference data from other reference pictures.
[0011] There may be many different forms of intra-frame prediction. When more than one such technique may be used in a given video coding technique, the techniques used may be referred to as intra-frame prediction modes. One or more intra-frame prediction modes may be provided in a particular codec. In some cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and intra-frame coding parameters for a video block may be encoded separately or included together in a mode codeword. Which codeword is used for a given mode, sub-mode, and / or parameter combination may have an impact on the coding efficiency gain through intra-frame prediction, and therefore may have an impact on the entropy coding technique used to convert the codeword into a bitstream.
[0012] Certain modes of intra prediction were introduced by H.264, refined in H.265, and further refined in newer coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Typically, for intra prediction, the predictor block can be formed using the values of neighboring samples that have become available. For example, the available values of a specific set of neighboring samples along a specific direction and / or line can be copied to the predictor block. A reference to the direction used can be encoded in the bitstream or it can itself be predicted.
[0013] Reference Figure 1A , a subset of nine predictor directions specified in the 33 possible intra-frame predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes specified in H.265) is depicted at the bottom right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample at 101 is predicted using neighboring samples. For example, arrow (102) indicates that sample (101) is predicted based on one or more neighboring samples at an angle of 45 degrees to the horizontal direction to the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more neighboring samples at an angle of 22.5 degrees to the horizontal direction to the lower left of sample (101).
[0014] Still refer to Figure 1A , a square block (104) of 4×4 samples is depicted at the upper left (indicated by bold dashed lines). The square block (104) includes 16 samples, each of which is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). Since the size of the block is 4×4 samples, S44 is at the lower right. Further shown are example reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, prediction samples that are adjacent to the block being reconstructed are used.
[0015] The intra picture prediction of block 104 may start by copying reference sample values from neighboring samples according to a signaled prediction direction. For example, assume that the encoded video bitstream includes signaling indicating, for this block 104, the prediction direction of arrow (102) - i.e., the samples are predicted according to one or more prediction samples to the upper right at a 45 degree angle to the horizontal direction. In such a case, samples S41, S32, S23, and S14 are predicted according to the same reference sample R05. Then, sample S44 is predicted according to reference sample R08.
[0016] In some cases, the values of multiple reference samples may be combined, such as by interpolation, in order to calculate the reference sample; in particular, when the direction is not divisible by 45 degrees.
[0017] As video coding techniques continue to advance, the number of possible directions is also increasing. For example, in H.264 (2003), there are nine different directions that can be used for intra-frame prediction. In H.265 (2013), the number of different directions increased to 33, and at the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most appropriate intra-frame prediction directions, and certain techniques in entropy coding can be used to encode those most appropriate directions with a small number of bits, thereby accepting a specific bit penalty for the direction. In addition, the direction itself can sometimes be predicted based on the neighboring directions used in the intra-frame prediction of the decoded neighboring blocks.
[0018] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to JEM is shown to illustrate that the number of prediction directions increases in various encoding techniques that evolve over time.
[0019] The manner in which bits representing intra-prediction directions are mapped to prediction directions in the coded video bitstream may vary from one video coding technique to another; and may range, for example, from simple direct mappings of prediction directions to intra-prediction modes, to codewords, to complex adaptive schemes involving most probable modes, and the like. However, in all cases, there may be certain directions used for leading prediction that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, those less probable directions will likely be represented by a larger number of bits than more probable directions.
[0020] Inter-picture prediction or inter-prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) - after spatial shifting in the direction indicated by a motion vector (hereinafter MV) - may be used to predict a newly reconstructed picture or picture portion (e.g., block). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (similar to the time dimension).
[0021] In some video compression techniques, the current MV applicable to a specific area of sample data can be predicted based on other MVs, for example, based on other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and that precede the current MV in the decoding order. This can significantly reduce the total amount of data required to encode the MV by eliminating redundancy in the relevant MVs, thereby improving compression efficiency. For example, since there is a statistical possibility that a larger area than the area applicable to a single MV moves in a similar direction in the video sequence when encoding an input video signal derived from a camera (referred to as natural video), similar motion vectors derived from MVs of neighboring areas can be used for prediction in some cases, so MV prediction can work effectively. This causes the actual MV for a given area to be similar or identical to the MV predicted based on the surrounding MVs. Such an MV can be represented with fewer bits after entropy coding than the number of bits that would be used if the MV was directly encoded rather than predicted based on neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lossy due to, for example, round-off errors when computing the predictor from several surrounding MVs.
[0022] Various MV prediction mechanisms are described in H.265 / HEVC (High Efficiency Video Coding, HEVC) (ITU (International Telecommunication Union, ITU)-T H.265 Recommendation, "High Efficiency Video Coding", December 2016). Among the multiple MV prediction mechanisms specified by H.265, the following is a technique referred to as "spatial merging".
[0023] Specifically, refer to Figure 2, the current block (201) includes samples that have been found by the encoder during the motion search process to be predictable from a previous block of the same size that has been spatially shifted. Instead of encoding the MV directly, the same MV as denoted by A can be used. 0 , A 1 and B 0 , B 1 , B 2 The MV associated with any of the five surrounding samples (corresponding to 202 to 206, respectively) is derived from metadata associated with one or more reference pictures, for example, from the latest (in decoding order) reference picture. In H.265, MV prediction can use a predictor from the same reference picture used by a neighboring block.
[0024] When motion vector differences (MVDs) are jointly signaled for multiple reference frames and the distances between the reference frames and the current frame are unequal, there is a need to provide a technical solution for scaling the MVD for at least one reference frame. Summary of the invention
[0025] The present disclosure relates generally to video coding, and in particular, to methods and systems for deriving and scaling motion vector differences (MVDs) in joint MVD scaling for composite inter prediction of video blocks and their signaling. In one example implementation, a method and video apparatus for processing video blocks of a video stream are disclosed. For example, the method may include scaling the joint motion vector differences and quantizing the scaled motion vector differences to derive a predicted motion vector according to a pixel resolution of the predicted motion vector.
[0026] In an example implementation, a method for processing a video block of a video stream is disclosed. The method may include: extracting at least one flag from the video stream; determining based on the at least one flag: the video block is inter-predicted by at least a first reference block located by a first motion vector in a first reference frame and a second reference block located by a second motion vector in a second reference frame; the first motion vector is to be predicted by a first motion vector difference (MVD) relative to the first reference motion vector; and the second motion vector is to be predicted by a second MVD relative to the second reference motion vector; wherein the first MVD and the second MVD are derived from a single MVD in the video stream. The method may also include: receiving a single MVD from a video stream; scaling the single MVD to generate a scaled MVD; quantizing the scaled MVD using the MVD pixel resolution to generate a quantized MVD based on an allowed accuracy of the first motion vector or the second motion vector; generating one of the first motion vector or the second motion vector based on the quantized MVD and a corresponding reference motion vector in the first reference motion vector and the second reference motion vector; and reconstructing one of the first reference block or the second reference block for inter-frame prediction of the video block from one of the first reference frame or the second reference frame based on the first motion vector or the second motion vector.
[0027] In the above implementation, the method may further include: extracting a first frame index for identifying the first reference frame and a second frame index for identifying the second reference frame from the video stream.
[0028] In any of the above implementations, at least one flag is signaled in the video stream before the first frame index and the second frame index.
[0029] In any of the above implementations, the at least one flag includes an indication of a joint MVD mode in a set of composite inter prediction modes.
[0030] In any of the above-mentioned implementations, the set of composite inter-frame prediction modes may include: a joint MVD mode, in which the first motion vector and the second motion vector are jointly predicted by a signaled MVD; a NEAR-NEAR inter-frame prediction mode, in which the first motion vector and the second motion vector are both signaled without any MVD; a NEAR-NEW inter-frame prediction mode, in which the first motion vector is signaled without any MVD and the second motion vector is predicted by the signaled MVD; a NEW-NEAR inter-frame prediction mode, in which the second motion vector is signaled without any MVD and the first motion vector is predicted by the signaled MVD; and a NEW-NEW inter-frame prediction mode, in which the first motion vector and the second motion vector are individually predicted by individually signaled MVDs.
[0031] In any of the above-mentioned implementations, the joint MVD mode includes a sub-mode of a composite NEW-NEW inter-frame prediction mode, wherein both the first motion vector and the second motion vector are predicted rather than being directly signaled; and the at least one flag is included in the video stream, and in response to the video block being predicted in the composite NEW-NEW inter-frame prediction mode, the at least one flag is extracted from the video stream.
[0032] In any of the above implementations, the MVD pixel resolution of one of the first MVD or the second MVD includes one of 1 / 64 pixel, 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel or integer pixel resolution.
[0033] In any of the above implementations, the MVD pixel resolution is preconfigured or adaptively signaled in the video stream.
[0034] In any of the above implementations, the MVD pixel resolution is lower than the scaled pixel resolution of the scaled MVD; and the scaled MVD is quantized by reducing the scaled pixel resolution of the scaled MVD to the MVD pixel resolution to generate a quantized MVD.
[0035] In any of the above implementations, the single MVD is scaled based on a first distance between a current frame of the video block and a first reference frame or a second distance between the current frame of the video block and a second reference frame to generate a scaled MVD.
[0036] In any of the above implementations, one of the first motion vector or the second motion vector generated based on the quantized MVD corresponds to a first reference frame or a second reference frame belonging to a predefined reference frame list.
[0037] In any of the above implementations, when the first distance is equal to the second distance, the scaling factor used for scaling is 1.
[0038] In any of the above implementations, one of the first motion vector or the second motion vector generated based on the quantized MVD corresponds to a reference frame of the first reference frame and the second reference frame having a smaller distance from the current frame. Alternatively, in any of the above implementations, one of the first motion vector or the second motion vector generated based on the quantized MVD corresponds to a reference frame of the first reference frame and the second reference frame having a larger distance from the current frame.
[0039] In any of the above implementations, the scaling is performed based on an additional scaling factor in addition to the first distance and the second distance, and the additional scaling factor is between -1 and 1.
[0040] In any of the above implementations, the method may also include: extracting additional scaling factors from the video stream as part of a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, or a coding tree unit header.
[0041] In any of the above implementations, quantizing the scaled MVD includes removing N least significant bits from the scaled MVD, N being an integer derived based on a difference between a MVD pixel resolution and a scaled pixel resolution of the scaled MVD. Alternatively, in any of the above implementations, quantizing the scaled MVD includes adding a rounding factor and then removing the N least significant bits from the scaled MVD, N being an integer derived based on a difference between the MVD pixel resolution, the scaled pixel resolution of the scaled MVD, and the rounding factor.
[0042] Aspects of the present disclosure also provide a video encoding device or a video decoding device, which includes a memory for storing instructions and a processor. The processor is configured to execute the instructions to perform the above-mentioned video processing method.
[0043] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform a method for video decoding and / or encoding.
[0044] Aspects of the present disclosure also provide a video processing device. The device includes: an extraction module configured to extract at least one flag from a video stream; a determination module configured to determine based on the at least one flag: a video block is inter-frame predicted by at least a first reference block located by a first motion vector in a first reference frame and a second reference block located by a second motion vector in a second reference frame, the first motion vector is predicted by a first MVD relative to the first reference motion vector, and the second motion vector is predicted by a second MVD relative to the second reference motion vector, wherein the first MVD and the second MVD are jointly signaled as a single MVD in the video stream; a receiving module configured to receive the single MVD signaled in the video stream; and a scaling module. A module configured to scale a single MVD to generate a scaled MVD; a quantization module configured to quantize the scaled MVD using the MVD pixel resolution according to an allowed accuracy of the first motion vector or the second motion vector to generate a quantized MVD; a generation module configured to generate one of a first motion vector and a second motion vector based on the quantized MVD and a corresponding reference motion vector of the first reference motion vector and the second reference motion vector; and a reconstruction module configured to reconstruct one of a first reference block or a second reference block for inter-frame prediction of a video block from one of a first reference frame or a second reference frame based on one of the first motion vector and the second motion vector.
[0045] According to the video processing method and video processing device of the present disclosure, when the motion vector difference (MVD) is jointly signaled for multiple reference frames and the distances between the reference frames and the current frame are not equal, a technical solution is provided for scaling the MVD for at least one reference frame, thereby improving the video encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0047] Figure 1A A schematic diagram showing an exemplary subset of intra prediction direction modes;
[0048] Figure 1B A diagram showing exemplary intra prediction directions;
[0049] Figure 2 A schematic diagram showing a current block and its surrounding spatial merging candidates for motion vector prediction in one example is shown;
[0050] Figure 3 A schematic diagram showing a simplified block diagram of a communication system (300) according to an example implementation;
[0051] Figure 4 A schematic diagram showing a simplified block diagram of a communication system (400) according to an example implementation;
[0052] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment;
[0053] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example implementation;
[0054] Figure 7 shows a block diagram of a video encoder according to another example embodiment;
[0055] Figure 8 shows a block diagram of a video decoder according to another example embodiment;
[0056] Fig. 9 A scheme for dividing coding blocks according to an exemplary embodiment of the present disclosure is shown;
[0057] Fig.10 Another scheme of coding block partitioning according to an example implementation of the present disclosure is shown;
[0058] Fig.11 Another scheme of coding block partitioning according to an example implementation of the present disclosure is shown;
[0059] Fig.12 shows an example partitioning of a base block into coding blocks according to an example partitioning scheme;
[0060] Fig.13 An example ternary partitioning scheme is shown;
[0061] Fig.14 An example quadtree binary tree coding block partitioning scheme is shown;
[0062] Fig.15 A scheme for dividing a coding block into a plurality of transform blocks and encoding the order of the transform blocks according to an example implementation of the present disclosure is shown;
[0063] Fig.16 Another scheme for dividing a coding block into a plurality of transform blocks and encoding the order of the transform blocks according to an example implementation of the present disclosure is shown;
[0064] Fig.17 Another scheme for dividing a coding block into a plurality of transform blocks according to an example implementation of the present disclosure is shown;
[0065] Fig.18 A flowchart showing a method according to an example implementation of the present disclosure;
[0066] Fig.19 A schematic diagram of a computer system according to an example implementation of the present disclosure is shown. DETAILED DESCRIPTION
[0067] Throughout the specification and claims, terms may have subtly different meanings that are suggested or implied in the context beyond the explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, it is intended that the claimed subject matter includes combinations of all or part of the exemplary embodiments / implementations.
[0068] Typically, terms can be understood at least in part from usage in context. For example, terms such as "and", "or" or "and / or" as used herein can include various meanings, which can depend at least in part on the context in which such terms are used. Typically, if "or" is used for an association list, such as A, B or C, then "or" is intended to represent A, B and C, which are used here in an inclusive sense, and A, B or C, which are used here in an exclusive sense. In addition, at least in part depending on the context, the terms "one or more" or "at least one" as used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Similarly, terms such as "one", "an" or "the" can also be understood to convey singular usage or to convey plural usage, which depends at least in part on the context. In addition, the term "based on" or "determined by..." can be understood to not necessarily be intended to convey an exclusive set of factors, and can alternatively allow the presence of additional factors that are not necessarily explicitly described, which also depends at least in part on the context. Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3In an example, a first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via a network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video picture, and display the video picture based on the recovered video data. The unidirectional data transmission can be implemented in a media service application, etc.
[0069] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which bidirectional transmission can be implemented, for example, during a video conferencing application. For the bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture at an accessible display device based on the restored video data.
[0070] exist Figure 3 In the example of , terminal devices (310), (320), (330) and (340) can be implemented as servers, personal computers and smart phones, but the applicability of the basic principles of the present disclosure may not be limited to this. Implementations of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (350) represents any number or type of network that transmits encoded video data among terminal devices (310), (320), (330) and (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching channels, packet switching channels and / or other types of channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless explicitly stated herein, the architecture and topology of network (350) may be unimportant to the operation of the present disclosure.
[0071] As examples of applications of the disclosed subject matter, Figure 4The placement of the video encoder and video decoder in a video streaming environment is shown. The disclosed subject matter may be equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., etc.
[0072] The video streaming system may include a video capture subsystem (413), which may include a video source (401), such as a digital camera, for creating an uncompressed video picture or image stream (402). In an example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402) is depicted as a thick line to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), and the video picture stream (402) may be processed by an electronic device (420) coupled to the video source (401) including a video encoder (403). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize the lower amount of data when compared to the uncompressed video picture stream (402), and the encoded video data (404) (or encoded video bitstream (404)) can be stored on a streaming server (405) for future use or directly to a downstream video device (not shown). Figure 4 One or more streaming client subsystems (406) and (408) in the video transmission system can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in the electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an outgoing video picture stream (411) that is uncompressed and can be presented on a display (412) (e.g., a display screen) or another rendering device (not depicted). The video decoder 410 can be configured to perform some or all of the various functions described in the present disclosure. In some streaming systems, the encoded video data (404), (407) and (409) (e.g., video bitstreams) can be encoded according to certain video encoding standards / video compression standards. Examples of these standards include ITU-T Recommendation H.265. In an example, the developing video coding standard is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC and other video coding standards.
[0073] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0074] Figure 5 A block diagram of a video decoder (510) according to any of the following embodiments of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used to replace Figure 4 A video decoder (410) in an example of FIG.
[0075] The receiver (531) can receive one or more encoded video sequences to be decoded by the video decoder (510). In the same embodiment or another embodiment, one encoded video sequence can be decoded at a time, wherein the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence can be associated with multiple video frames or images. The encoded video sequence can be received from a channel (501), which can be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective processing circuit systems (not depicted). The receiver (531) can separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) can be set between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) can be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be external to the video decoder (510) and separate from the video decoder (510) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (510) for purposes such as preventing network jitter, and there may be another additional buffer memory (515) internal to the video decoder (510), for example to handle playback timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may not be required, or the buffer memory (515) may be small. For use on a best-effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and its size may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (510).
[0076] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510) and potential information for controlling a presentation device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530), but the display (512) may be coupled to the electronic device (530), such as Figure 5As shown. The control information for (one or more) rendering devices may be in the form of a Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the encoded video sequence received by the parser (520). The entropy encoding of the encoded video sequence may be performed according to a video coding technique or a video coding standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Picture (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) may also extract information such as transform coefficients (eg, Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0077] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).
[0078] Depending on the type of the coded video picture or part of the coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol (521) may involve multiple different processing units or functional units. The units involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For simplicity, this subgroup control information flow between the parser (520) and the following multiple processing units or functional units is not described.
[0079] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a plurality of functional units as described below. In a practical implementation operating under commercial constraints, many of these functional units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.
[0080] The first unit may include a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) may receive quantized transform coefficients as (one or more) symbols (521) and control information from the parser (520), including information indicating which type of inverse transform to use, block size, quantization factors / parameters, quantization scaling matrix, etc. The sealer / inverse transform unit (551) may output blocks including sample values that may be input into an aggregator (555).
[0081] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may use surrounding block information that has already been reconstructed and stored in a current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) caches a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add the prediction information that has been generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0082] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to a block that has been inter-coded and possibly motion compensated. In such a case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for inter-picture prediction. After motion compensation of the obtained samples according to the symbols (521) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) (the output of unit 551 may be referred to as residual samples or residual signal) to generate output sample information. The address within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by a motion vector, which is used by the motion compensated prediction unit (553) in the form of a symbol (521), which may have, for example, an X component, a Y component (shift) and a reference picture component (time). Motion compensation may also include interpolation of sample values retrieved from a reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with motion vector prediction mechanisms, etc.
[0083] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) that are available to the loop filter unit (556) as symbols (521) from the parser (520), but the video compression techniques may also be responsive to meta information obtained during decoding of a previous (in decoding order) portion of the coded picture or coded video sequence and to previously reconstructed and loop filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in more detail below.
[0084] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for future inter-picture prediction.
[0085] Once fully reconstructed, certain coded pictures can be used as reference pictures for future inter-picture prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.
[0086] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows both the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as tools that can only be used under the profile. In order to conform to the standard, the complexity of the encoded video sequence may be within a range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0087] In some example embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence (one or more). The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0088] Figure 6 A block diagram of a video encoder (603) according to an example implementation of the present disclosure is shown. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may also include a transmitter (640) (e.g., a transmission circuit system). The video encoder (603) may be used to replace Figure 4 A video encoder (403) in an example.
[0089] The video encoder (603) can be used to obtain the video source (601) (the video source (601) is not Figure 6 In an example of an electronic device (620) receiving video samples, a video source (601) can capture (one or more) video images to be encoded by a video encoder (603). In another example, the video source (601) can be implemented as a part of the electronic device (620).
[0090] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, XYZ, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures or images that are given motion when viewed sequentially. The picture itself may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by one of ordinary skill in the art. The following description focuses on samples.
[0091] According to some example embodiments, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Implementing a suitable encoding speed constitutes a function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units as described below. For simplicity, the coupling is not depicted. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions that belong to the video encoder (603) optimized for certain system designs.
[0092] In some example embodiments, the video encoder (603) may be configured to operate in an encoding loop. As an oversimplified description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder would create sample data, even if the embedded decoder 633 processes the encoded video stream produced by the source encoder 630 without entropy coding (because in the video compression techniques considered in the disclosed subject matter, any compression between the encoded video bitstream and the symbols in entropy coding can be lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors) is used to improve encoding quality.
[0093] The operation of the "local" decoder (633) can be combined with the above Figure 5 The operation of the "remote" decoder described in detail is identical to that of the video decoder (510). However, reference is also briefly made to Figure 5, when symbols are available and the entropy encoder (645) and parser (520) encoding / decoding the symbols into an encoded video sequence can be lossless, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633) in the encoder.
[0094] At this point, it can be observed that any decoder technology other than parsing / entropy decoding that may only exist in the decoder may also necessarily need to exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on the decoder operation related to the decoding portion of the encoder. Since encoder technology is mutually inverse to the decoder technology that has been fully described, the description of encoder technology can be simplified. Only in certain areas or aspects, a more detailed description of the encoder is provided below.
[0095] During operation, in some example implementations, the source encoder (630) may perform motion compensated predictive encoding that predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures". In this manner, the encoding engine (632) encodes the differences (or residuals) in color channels between pixel blocks of the input picture and pixel blocks of (one or more) reference pictures that may be selected as prediction references (one or more) for the input picture. The term "residual" and its adjective form "residual" may be used interchangeably.
[0096] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be lossy processing. When the encoded video data can be decoded at the video decoder ( Figure 6 When the video encoder (603) is decoded at a local video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture, and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the far-end (remote) video decoder.
[0097] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (635) may operate on a pixel block by pixel block basis to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0098] The controller (650) can manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0099] The outputs of all the above-mentioned functional units may be subjected to entropy coding in an entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing them according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0100] The transmitter (640) can buffer the encoded video sequence(s) created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0101] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain coded picture type to each coded picture, which may affect the coding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned one of the following picture types:
[0102] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those of ordinary skill in the art are aware of those variations of I pictures and their corresponding applications and features.
[0103] A predictive picture (P picture) may be a picture that may be encoded and decoded using intra prediction or inter prediction, which predicts sample values of each block using at most one motion vector and a reference index.
[0104] Bidirectional predictive pictures (B pictures) can be pictures that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.
[0105] The source picture may typically be spatially subdivided into a plurality of coding blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively), and encoded on a block-by-block basis. These blocks may be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding allocations applied to the corresponding pictures of the blocks. For example, blocks of an I picture may be non-predictively encoded, or blocks of an I picture may be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture may be predictively encoded via spatial prediction or via temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture may be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures. Source pictures or intermediately processed pictures may be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same approach, as described in further detail below.
[0106] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0107] In some example embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0108] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often shortened to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits temporal or other correlations between pictures. For example, a particular picture in encoding / decoding, which is referred to as the current picture, can be divided into blocks. The blocks in the current picture can be encoded by a vector called a motion vector when they are similar to reference blocks in a reference picture that has been previously encoded and still buffered in the video. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.
[0109] In some example embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both subsequent to the current picture in the video in decoding order (but may be in the past or future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be jointly predicted by a combination of the first reference block and the second reference block.
[0110] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.
[0111] According to some example embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a video picture sequence is divided into coding tree units (Coding Tree Unit, CTU) for compression, and the CTUs in the picture may have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU may include three parallel coding tree blocks (Coding Tree Block, CTB): a luminance CTB and two chrominance CTBs. Each CTU may be recursively partitioned into one or more coding units (CUs) with a quadtree. For example, a 64×64 pixel CTU may be partitioned into a 64×64 pixel CU or four 32×32 pixel CUs. Each of one or more of the 32×32 blocks may be further partitioned into four 16×16 pixel CUs. In some example embodiments, each CU may be analyzed during encoding to determine a prediction type for the CU among various prediction types, such as an inter-prediction type or an intra-prediction type. Depending on temporal and / or spatial predictability, a CU may be partitioned into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in decoding (encoding / decoding) are performed in units of prediction blocks. Partitioning a CU into PUs (or PBs of different color channels) may be performed in various spatial modes. For example, a luma or chroma PB may include a matrix of values (e.g., luma values) for samples such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, and the like.
[0112] Figure 7 A diagram of a video encoder (703) according to another example implementation of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and encode the processed block into an encoded picture that is part of an encoded video sequence. The example video encoder (703) may be used instead of Figure 4 The video encoder (403) in the example.
[0113] For example, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) then uses, for example, Rate-Distortion Optimization (RDO) to determine whether to best encode the processing block using intra mode, inter mode, or bidirectional prediction mode. In the case where it is determined that the processing block is encoded in intra mode, the video encoder (703) can encode the processing block into an encoded picture using intra prediction techniques; and in the case where it is determined that the processing block is encoded in inter mode or bidirectional prediction mode, the video encoder (703) can encode the processing block into an encoded picture using inter prediction or bidirectional prediction techniques, respectively. In some example embodiments, merge mode can be used as a sub-mode of inter-picture prediction, wherein a motion vector is derived based on the predictor without utilizing an encoded motion vector component external to one or more motion vector predictors. In some other example embodiments, there may be a motion vector component applicable to the subject block. Therefore, the video encoder (703) may include Figure 7 Components not explicitly shown, such as a mode decision module for determining a prediction mode for a processing block.
[0114] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721) and an entropy encoder (725) coupled together are shown in the example arrangement.
[0115] The inter-frame encoder (730) is configured to: receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order), generate inter-frame prediction information (e.g., description of redundant information according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is embedded in Figure 6 The decoding unit 633 (shown as Figure 7 The residual decoder 728, as described in further detail below, decodes the decoded reference picture based on the encoded video information.
[0116] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with an already encoded block in the same picture, generate quantization coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) can calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0117] The general controller (721) can be configured to: determine general control data, and control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines a prediction mode of a block, and provides a control signal to a switch (726) based on the prediction mode. For example, when the prediction mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in the bitstream; and when the prediction mode of the block is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in the bitstream.
[0118] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate a transform coefficient. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate the transform coefficient. Then, the transform coefficient is subjected to quantization to obtain a quantized transform coefficient. In various example embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are appropriately processed to generate decoded pictures, and these decoded pictures may be buffered in a memory circuit (not shown) and used as reference pictures.
[0119] The entropy encoder (725) may be configured to format the bitstream to include the encoded blocks and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. When the block is encoded in the merge sub-mode of the inter-frame mode or the bi-prediction mode, there may be no residual information.
[0120] Figure 8 A diagram of an example video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (810) may be used instead of Figure 4 A video decoder (410) in an example of FIG.
[0121] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874) and an intra-frame decoder (872) coupled together are shown in the example arrangement of.
[0122] The entropy decoder (871) can be configured to reconstruct certain symbols representing syntax elements constituting the encoded picture from the encoded picture. Such symbols may include, for example, a mode for encoding a block (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata used by the intra decoder (872) or the inter decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, etc. In an example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be subjected to inverse quantization and provided to the residual decoder (873).
[0123] The inter-frame decoder (880) may be configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.
[0124] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0125] The residual decoder (873) can be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) can also utilize certain control information (to include quantizer parameters (Quantizer Parameter, QP)), which can be provided by the entropy decoder (871) (since this may only be low data volume control information, the data path is not depicted).
[0126] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) in the spatial domain to form a reconstructed block, which forms part of the reconstructed picture as part of the reconstructed video. Note that other suitable operations such as deblocking operations may also be performed to improve visual quality.
[0127] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In some example embodiments, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more processors executing software instructions.
[0128] Turning to the block partitioning for encoding and decoding, the general partitioning can start from the base block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After following any example partitioning process or other processes described below or a combination thereof to split or divide the base block, a final set of partitions or coding blocks can be obtained. Each of these partitions can be in one of the various partitioning levels in the partitioning hierarchy and can have various shapes. Each of the partitions can be referred to as a coding block (CB). For various example partitioning implementations further described below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because they can form the following units, some basic encoding / decoding decisions can be made for the units, and the encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. The coding block can be a luminance coding block or a chrominance coding block. The CB tree structure for each color can be referred to as a coding block tree (CBT).
[0129] The coding blocks of all color channels may be collectively referred to as coding units (CUs). The hierarchical structures of all color channels may be collectively referred to as coding tree units (CTUs). The division modes or structures of various color channels in a CTU may be the same or different.
[0130] In some implementations, the partitioning tree schemes or structures for the luma channel and the chroma channel may not necessarily be the same. In other words, the luma channel and the chroma channel may have respective coding tree structures or patterns. In addition, whether the luma channel and the chroma channel use the same or different coding partitioning tree structures and the actual coding partitioning tree structures to be used may depend on whether the slice being encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chroma channel and the luma channel may have respective coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luma channel and the chroma channel may share the same coding partitioning tree scheme. When respective coding partitioning tree structures or patterns are applied, the luma channel may be divided into CBs by one coding partitioning tree structure, and the chroma channel may be divided into chroma CBs by another coding partitioning tree structure.
[0131] In some example implementations, a predetermined partitioning pattern may be applied to a base block. Fig. 9As shown, the example 4-way partitioning tree can start from a first predefined level (e.g., 64×64 block level or other size, as a base block size), and the base block can be hierarchically partitioned down to a predefined lowest level (e.g., 4×4 level). For example, the base block can be subjected to four predefined partitioning options or modes indicated by 902, 904, 906 and 908, where the partition designated as R is allowed for recursive partitioning because Fig. 9 The same partitioning options shown in can be repeated at lower scales until the lowest level (e.g., 4×4 level). Fig. 9 Additional restrictions apply to the partitioning scheme. Fig. 9 In the implementation of , rectangular partitions (e.g. 1:2 / 2:1 rectangular partitions) can be allowed, but they cannot be recursive, while square partitions can be recursive. If necessary, follow Fig. 9 The recursive partitioning generates the final set of coding blocks. The coding tree depth can be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root block, such as a 64×64 block, can be set to 0, and the root block follows Fig. 9 After being further split once, the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from the 64×64 base block to the 4×4 minimum partition will be 4 (starting from level 0). Such a partitioning scheme can be applied to one or more of the color channels. Each color channel can follow Fig. 9 The scheme of being divided independently (for example, the division mode or option in the predefined mode can be determined independently for each of the color channels at each hierarchical level). Alternatively, two or more of the color channels can share Fig. 9 The same hierarchical pattern tree (e.g., the same partitioning pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).
[0132] Fig.10 Another example predefined partitioning pattern that allows recursive partitioning to form a partitioning tree is shown. Fig.10 As shown, the structure or pattern may be divided in a predefined manner 10. The root block may start at a predefined level (eg, starting from a base block at a 128x128 level or a 64x64 level). Fig.10 Example partition structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Fig.10 The partition type with 3 sub-partitions indicated by 1002, 1004, 1006, and 1008 in the second row of can be referred to as a "T-type" partition. The "T-type" partitions 1002, 1004, 1006, and 1008 can be referred to as a left T-type, a top T-type, a right T-type, and a bottom T-type. In some example implementations, Fig.10 The rectangular partitions are not allowed to be further subdivided. The coding tree depth can also be defined to indicate the partition depth starting from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and the root block follows Fig.10 After being further split once, the coding tree depth increases by 1. In some implementations, only full square partitions in 1010 may be allowed to follow Fig.10 In other words, for the square partitions within the T-patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If necessary, follow Fig.10 The partitioning process performed in a recursive manner generates the final set of coding blocks. Such a scheme can be applied to one or more of the color channels. In some implementations, more flexibility can be added to the use of partitions below the 8×8 level. For example, in some cases, 2×2 chroma inter-frame prediction can be used.
[0133] In some other example implementations of coding block partitioning, a quadtree structure can be used to partition a base block or intermediate block into quadtree partitions. Such quadtree partitioning can be applied hierarchically and recursively to any square-shaped partition. Whether a base block or intermediate block or partition is further quadtree partitioned can be adapted to various local characteristics of the base block or intermediate block / partition. The quadtree partitioning at the picture boundary can be further adjusted. For example, implicit quadtree partitioning can be performed at the picture boundary so that the block will remain quadtree partitioned until the size fits the picture boundary.
[0134] In some other example implementations, a hierarchical binary partition from a base block may be used. For such a scheme, a base block or an intermediate level block may be divided into two partitions. The binary partition may be horizontal or vertical. For example, a horizontal binary partition may divide a base block or an intermediate block into equal right and left partitions. Similarly, a vertical binary partition may divide a base block or an intermediate block into equal upper and lower partitions. Such a binary partition may be hierarchical and recursive. A decision may be made at each of the base block or the intermediate block as to whether the binary partition scheme should continue, and if the scheme continues further, a decision may be made as to whether horizontal binary partition or vertical binary partition should be used. In some implementations, further partitioning may stop at a predefined minimum partition size (in one or two dimensions). Alternatively, further partitioning may stop once a predefined partition level or depth from the base block is reached. In some implementations, the aspect ratio of the partition may be limited. For example, the aspect ratio of the partition may not be less than 1:4 (or greater than 4:1). In this way, a vertical strip partition having a vertical to horizontal aspect ratio of 4:1 can only be further vertically binary divided into an upper partition and a lower partition each having a vertical to horizontal aspect ratio of 2:1.
[0135] In still other examples, a ternary partitioning scheme may be used to partition the base block or any intermediate block, such as Fig.13 The ternary pattern can be Fig.13 1302 is implemented vertically or as shown in Fig.13 The 1304 level is achieved. Fig.13 The example partition ratio (vertically or horizontally) in is shown as 1:2:1, but other ratios may also be predefined. In some implementations, two or more different ratios may be predefined. Such a ternary partitioning scheme may be used to complement a quadtree or binary partitioning structure, because such a ternary tree partitioning is able to capture objects located at the center of a block in one continuous partition, while quadtrees and binary trees are always partitioned along the center of the block, thus partitioning the objects into separate partitions. In some implementations, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transformations.
[0136] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree partitioning scheme and binary partitioning scheme can be combined to partition the base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition can be a quadtree partition or a binary partition, if specified, subject to a set of predefined conditions. Fig.14 A specific example is shown in FIG. Fig.14In the example of , the base block is first quadtree-partitioned into four partitions, as shown by 1402, 1404, 1406, and 1408. Thereafter, each of the resulting partitions is quadtree-partitioned into four further partitions (e.g., 1408), or binary-partitioned at the next level into two further partitions (horizontally or vertically, such as 1402 or 1406, for example, both are symmetrical), or not partitioned (e.g., 1404). Binary or quadtree partitioning can be allowed to be recursively used for square-shaped partitions, as shown by the overall example partitioning pattern of 1410 and the corresponding tree structure / representation in 1420, where the solid lines represent quadtree partitioning and the dashed lines represent binary partitioning. A flag can be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, consistent with the partitioning structure of 1410, the flag "0" can represent horizontal binary partitioning, and the flag "1" can represent vertical binary partitioning. For quadtree partitioning, there is no need to indicate the partition type, because quadtree partitioning always partitions a block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some implementations, a flag "1" may represent a horizontal binary partition, and a flag "0" may represent a vertical binary partition.
[0137] In some example implementations of QTBT, the quadtree and binary segmentation rule set may be represented by the following predefined parameters and corresponding functions associated therewith:
[0138] -CTU size: the root node size of the quadtree (the size of the base block)
[0139] -MinQTSize: Minimum allowed quadtree leaf node size
[0140] -MaxBTSize: Maximum allowed binary tree root node size
[0141] -MaxBTDepth: Maximum allowed binary tree depth
[0142] -MinBTSize: The minimum allowed binary tree leaf node size
[0143] In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128×128 luma samples with two corresponding 64×64 chroma sample blocks (when example chroma subsampling is considered and used), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes from their minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the node is 128×128, the node will not be split by the binary tree first because the size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes that do not exceed MaxBTSize can be split by the binary tree. Fig.14 In the example of , the base block is 128×128. According to a predefined set of rules, the base block can only be quadtree split. The base block has a partition depth of 0. Each of the four resulting partitions is 64×64 - not exceeding MaxBTSize and can be further quadtree or binary tree split at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of the binary tree node is equal to MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of the binary tree node is equal to MinBTSize, further vertical splitting is disregarded.
[0144] In some example implementations, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or separate QTBT structures for luma and chroma. For example, for P slices and B slices, the luma CTB and chroma CTB in one CTU can share the same QTBT structure. However, for I slices, the luma CTB can be divided into CBs by the QTBT structure, and the chroma CTB can be divided into chroma CBs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice, for example, an I slice can consist of coding blocks of a luma component or coding blocks of two chroma components, and a CU in a P slice or a B slice can consist of coding blocks of all three color components.
[0145] In some other implementations, the QTBT scheme can be supplemented with the above ternary scheme. Such an implementation can be called a multi-type-tree (MTT) structure. For example, in addition to the binary split of the node, you can also choose Fig.13In some implementations, only square nodes can be subjected to ternary partitioning. An additional flag can be used to indicate whether the ternary partition is horizontal or vertical.
[0146] The design of two or more levels of trees, such as the QTBT implementation and the QTBT implementation supplemented by ternary partitioning, is motivated primarily by complexity reduction. In theory, the complexity of traversing the tree is TD, where T represents the number of partition types and D is the depth of the tree. A trade-off can be made by using multiple types (T) while reducing the depth (D).
[0147] In some implementations, CB can be further divided. For example, for the purpose of intra-frame or inter-frame prediction during the encoding process and the decoding process, CB can be further divided into multiple prediction blocks (PB). In other words, CB can be further split into different sub-partitions, in which separate prediction decisions / configurations can be made. In parallel, for the purpose of depicting the level of transformation or inverse transformation of video data, CB can be further divided into multiple transform blocks (TB). The division schemes of CB to PB and TB can be the same or different. For example, each division scheme can be performed using its own process based on various characteristics of video data, for example. In some example implementations, PB division schemes and TB division schemes can be independent. In some other example implementations, PB division schemes and TB division schemes and boundaries can be related. In some implementations, for example, TB can be divided after PB division, and in particular, each PB can be further divided into one or more TBs after the division following the coding block is determined. For example, in some implementations, PB can be divided into one, two, four or other number of TBs.
[0148] In some implementations, in order to divide the base block into coding blocks and further into prediction blocks and / or transform blocks, the luma channel and the chroma channel may be processed differently. For example, in some implementations, the division of the coding block into prediction blocks and / or transform blocks may be allowed for the luma channel, while such division of the coding block into prediction blocks and / or transform blocks may not be allowed for the chroma channel. In such an implementation, the transformation and / or prediction of the luma block can therefore be performed only at the coding block level. For another example, the minimum transform block size of the luma channel and the chroma channel may be different, for example, the coding block of the luma channel may be allowed to be divided into smaller transform blocks and / or prediction blocks than the chroma channel. For yet another example, the maximum depth of dividing the coding block into transform blocks and / or prediction blocks may be different between the luma channel and the chroma channel, for example, the coding block of the luma channel may be allowed to be divided into deeper transform blocks and / or prediction blocks than (one or more) chroma channels. For a specific example, a luma coding block may be partitioned into transform blocks of various sizes, which may be represented by recursive partitioning down to up to 2 levels, and may allow transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes from 4×4 to 64×64. However, for chroma blocks, only the largest possible transform block specified for luma blocks may be allowed.
[0149] In some example implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partition may depend on whether the PB is intra-coded or inter-coded.
[0150] The partitioning of a coding block (or prediction block) into transform blocks may be implemented in various example schemes, including but not limited to recursively or non-recursively performing quadtree partitioning and predefined pattern partitioning, and additionally taking into account transform blocks at the boundaries of coding blocks or prediction blocks. In general, the resulting transform blocks may be at different partitioning levels, may not have the same size, and may not need to be square in shape (e.g., the resulting transform blocks may be rectangular with certain allowed sizes and aspect ratios). The following is about Fig.15 , Fig.16 and Fig.17 Other examples are described in further detail.
[0151] However, in some other implementations, the CB obtained via any of the above partitioning schemes can be used as a basic or minimum coding block for prediction and / or transformation. In other words, no further segmentation is performed for the purpose of performing inter-frame prediction / intra-frame prediction and / or for transformation purposes. For example, the CB obtained from the above QTBT scheme can be directly used as a unit for performing prediction. Specifically, such a QTBT structure removes the concept of multiple partition types, that is, it removes the separation of CU, PU and TU, and supports greater flexibility in the CU / CB partition shape as described above. In such a QTBT block structure, the CU / CB can have a square or rectangular shape. The leaf nodes of such a QTBT are used as units for prediction and transformation processing without any further division. This means that CU, PU and TU have the same block size in such an example QTBT coding block structure.
[0152] The above various CB partitioning schemes and further partitioning of CB to PB and / or TB (including no PB / TB partitioning) may be combined in any manner.The following specific implementations are provided as non-limiting examples.
[0153] The following describes a specific example implementation of coding block and transform block partitioning. In such an example implementation, the above-mentioned recursive quadtree partitioning or a predefined partitioning pattern (e.g., Fig. 9 and Fig.10 The base block is divided into coding blocks (those in ). At each level, whether further quadtree segmentation should be continued for a specific partition can be determined by local video data characteristics. The resulting CB can be at different quadtree segmentation levels and can have different sizes. A decision can be made at the CB level (or CU level, for all three color channels) about whether to use inter-frame picture (time) prediction or intra-frame picture (spatial) prediction to encode the picture area. According to the predefined PB segmentation type, each CB can be further divided into one, two, four or other number of PBs. Within a PB, the same prediction process can be applied, and relevant information can be transmitted to the decoder based on the PB. After obtaining the residual block by applying the prediction process based on the PB segmentation type, the CB can be divided into TBs according to another quadtree structure similar to the coding tree for the CB. In this specific implementation, the CB or TB can be square in shape, but it is not necessarily limited to a square shape. In addition, in this specific example, for inter-frame prediction, the PB can be square or rectangular in shape, but for intra-frame prediction, the PB can be only square. The coding block can be divided into, for example, four square-shaped TBs. Each TB can be further recursively partitioned (using quadtree partitioning) into smaller TBs, which are called Residual Quadtrees (RQTs).
[0154] Another example implementation for partitioning a base block into CBs, PBs, and / or TBs is further described below. Fig. 9 or Fig.10 In the example of a coding tree structure, a CB may have a square or rectangular shape. Specifically, a coding tree block (CTB) may be first partitioned by a quadtree structure. Then, the quadtree leaf nodes may be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary splitting may be used. Fig.11 shown in . Specifically, Fig.11 The example multi-type tree structure includes four types of splits, which are called vertical binary splits (SPLIT_BT_VER) (1102), horizontal binary splits (SPLIT_BT_HOR) (1104), vertical ternary splits (SPLIT_TT_VER) (1106) and horizontal ternary splits (SPLIT_TT_HOR) (1108). The CBs then correspond to the leaves of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, the split is used for both prediction and transform processing without any further partitioning. This means that: in most cases, CBs, PBs and TBs have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color components of the CB. In some implementations, in addition to binary or ternary splits, Fig.11 The nested mode can also include quadtree partitioning.
[0155] Fig.12 A specific example of a quadtree with a nested multi-type tree coding block structure (including a quadtree partitioning option, a binary partitioning option, and a ternary partitioning option) for block partitioning of a base block is shown. In more detail, Fig.12 The base block 1200 is shown to be divided into four square partitions 1202, 1204, 1206 and 1208 by a quadtree partition. For each of the quadtree partitions, a further Fig.11The multi-type tree structure and quadtree are used for further segmentation decisions. Fig.12 In the example of FIG. 1204, partition 1204 is not further partitioned. Partitions 1202 and 1208 each use another quadtree partition. For partition 1202, the second level quadtree partitions the upper left, upper right, lower left, and lower right partitions using the third level of the quadtree partition, Fig.11 Horizontal binary segmentation 1104, no segmentation and Fig.11 The horizontal ternary segmentation 1108. The partition 1208 uses another quadtree segmentation, and the second level quadtree segmentation upper left, upper right, lower left and lower right partitions are respectively used Fig.11 The third level segmentation of the vertical ternary segmentation 1106, no segmentation, no segmentation and Fig.11 The horizontal binary partition 1104. The two sub-partitions of the third level upper left partition 1208 are respectively based on Fig.11 The horizontal binary segmentation 1104 and the horizontal ternary segmentation 1108 are further segmented. The partition 1206 is adopted according to Fig.11 The second level segmentation mode of the vertical binary segmentation 1102 is divided into two partitions, and the two partitions are divided into two partitions according to Fig.11 The horizontal ternary segmentation 1108 and the vertical binary segmentation 1102 are further segmented at the third level. Fig.11 1104, a fourth level segmentation is further applied to one of them.
[0156] For the specific example above, the maximum luma transform size may be 64×64, and the maximum supported chroma transform size may be different from luma, for example, 32×32. Fig.12 The example CB in the example is usually not further split into smaller PBs and / or TBs, but when the width or height of the luminance coding block or the chrominance coding block is larger than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically split in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0157] In the specific example above where the base block is divided into CBs, and as described above, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P slices and B slices, the luma CTB and chroma CTB in one CTU can share the same coding tree structure. For example, for I slices, luma and chroma can have separate coding block tree structures. When a separate block tree structure is applied, the luma CTB can be divided into luma CBs by one coding tree structure, and the chroma CTB is divided into chroma CBs by another coding tree structure. This means that a CU in an I slice can consist of coding blocks of a luma component or coding blocks of two chroma components, and a CU in a P slice or B slice always consists of coding blocks of all three color components unless the video is monochrome.
[0158] When the coding block is further divided into multiple transform blocks, the transform blocks therein can be ordered in the bitstream in various orders or scanning modes. The following further describes in detail the example implementations of dividing the coding block or prediction block into transform blocks and the encoding order of the transform blocks. In some example implementations, as described above, the transform partitioning can support transform blocks of various shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the transform block size ranges from, for example, 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, then the transform block partitioning can be applied only to the luminance component, so that for the chrominance block, the transform block size is the same as the coding block size. In addition, if the coding block width or height is greater than 64, both the luminance coding block and the chrominance coding block can be implicitly split into multiple min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.
[0159] In some example implementations of transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coding block can be further partitioned into multiple transform blocks, where the partition depth reaches a predefined number of levels (e.g., 2 levels). The transform block partition depth and size may be related. For some example implementations, the mapping from the transform size of the current depth to the transform size of the next depth is shown below in Table 1.
[0160] Table 1: Change partition size settings
[0161]
[0162] Based on the example mapping of Table 1, for a 1:1 square block, the next level transform partitioning can create four 1:1 square sub-transform blocks. The transform partitioning can stop at, for example, 4×4. Therefore, the transform size of 4×4 at the current depth corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level transform partitioning can create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next level transform partitioning can create two 1:2 / 2:1 sub-transform blocks.
[0163] In some example implementations, for the luma component of an intra-coded block, additional restrictions may be applied with respect to transform block partitioning. For example, for each level of transform partitioning, all sub-transform blocks may be restricted to have the same size. For example, for a 32×16 coding block, a level 1 transform partition creates two 16×16 sub-transform blocks and a level 2 transform partition creates eight 8×8 sub-transform blocks. In other words, the second level partitioning must be applied to all first level sub-blocks to keep the transform unit sizes equal. Fig.15 1 shows an example of transform block partitioning of an intra-coded square block following Table 1 and the coding order indicated by the arrows. Specifically, 1502 shows a square coding block. The first level partitioning into 4 equal-sized transform blocks according to Table 1 and the coding order indicated by the arrows are shown in 1504. The second level partitioning into 16 equal-sized transform blocks according to Table 1 and the coding order indicated by the arrows are shown in 1506.
[0164] In some example implementations, the above restrictions for intra-coding may not apply to the luma component of inter-coded blocks. For example, after the first level of transform partitioning, any of the sub-transform blocks may be further partitioned independently at more than one level. Thus, the resulting transform blocks may or may not have the same size. Fig.16 An example partitioning of inter-coded blocks into transform blocks and their coding order is shown in FIG. Fig.16 In the example of FIG. 1 , an inter-coded block 1602 is partitioned into transform blocks at two levels according to Table 1. At the first level, the inter-coded block is partitioned into four transform blocks of equal size. Then, only one of the four transform blocks (not all four transform blocks) is further partitioned into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes as shown by 1604. The example coding order of these 7 transform blocks is shown by Fig.16 As shown by the arrow in 1604.
[0165] In some example implementations, for chroma components, some additional restrictions may be imposed on transform blocks. For example, for chroma components, the transform block size may be as large as the coding block size, but not smaller than a predefined size, such as 8×8.
[0166] In some other example implementations, for coding blocks with a width (W) or height (H) greater than 64, both luma coding blocks and chroma coding blocks may be implicitly split into multiple min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively. Here, in the present disclosure, "min(a,b)" may return the smaller value between a and b.
[0167] Fig.17 Another alternative example scheme for dividing a coding block or prediction block into transform blocks is further shown. Fig.17 As shown, instead of using recursive transform partitioning, a predefined set of partition types can be applied to the coding block according to the transform type of the coding block. Fig.17 In the specific example shown in , one of 6 example partition types can be applied to split the coding block into various numbers of transform blocks. Such a scheme for generating transform block partitions can be applied to coding blocks or prediction blocks.
[0168] In more detail, Fig.17 The partitioning scheme of provides up to 6 example partition types for any given transform type (transform type refers to, for example, the type of main transform, such as ADST and others). In this scheme, a transform partition type can be assigned to each coding block or prediction block based on, for example, rate-distortion cost. In an example, the transform partition type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. A specific transform partition type can correspond to a transform block partition size and mode, as determined by Fig.17 As shown in the 6 transform partition types shown in . The correspondence between various transform types and various transform partition types can be predefined. An example is shown below, where the uppercase label indicates the transform partition type that can be assigned to the coding block or prediction block based on the rate-distortion cost:
[0169] PARTITION_NONE: Allocate a transform size equal to the block size.
[0170] PARTITION_SPLIT: Allocate a transform size of 1 / 2 the block size in width and 1 / 2 the block size in height.
[0171] PARTITION_HORZ: Allocates a transform size with the same width as the block size and 1 / 2 the height of the block size.
[0172] PARTITION_VERT: Allocates a transform size with a width of 1 / 2 the block size and the same height as the block size.
[0173] PARTITION_HORZ4: Allocates a transform size with the same width as the block size and 1 / 4 the height of the block size.
[0174] PARTITION_VERT4: Allocates a transform size with a width of 1 / 4 the block size and the same height as the block size.
[0175] In the above example, Fig.17 The transform partition types shown all contain uniform transform sizes for the partitioned transform blocks. This is only an example and not a limitation. In some other implementations, mixed transform block sizes may be used for the partitioned transform blocks of a particular partition type (or mode).
[0176] Then, the PB (or CB, also referred to as PB when not further divided into prediction blocks) obtained from any of the above partitioning schemes can become a separate block for encoding via intra-frame prediction or inter-frame prediction. For inter-frame prediction of the current PB, the residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.
[0177] Inter prediction can be implemented, for example, in a single reference mode or a composite reference mode. In some implementations, a skip flag may first be included in the bitstream of the current block (or at a higher level) to indicate whether the current block is inter-coded and not skipped. If the current block is inter-coded, another flag may be further included in the bitstream as a signal indicating whether a single reference mode or a composite reference mode is used for the prediction of the current block. For a single reference mode, a reference block may be used to generate a prediction block for the current block. For a composite reference mode, two or more reference blocks may be used to generate a prediction block by, for example, weighted averaging. A composite reference mode may be referred to as more than one reference mode, two reference modes, or multiple reference modes. A reference block or multiple reference blocks may be identified using a reference frame index or multiple reference frame indexes and additionally using a corresponding motion vector or multiple motion vectors indicating the displacement between (one or more) reference blocks and the current block relative to the frame (e.g., horizontal pixels and vertical pixels). For example, an inter-frame prediction block of a current block may be generated from a single reference block identified by a motion vector in a reference frame as a prediction block in a single reference mode, while for a composite reference mode, a prediction block may be generated by a weighted average of two reference blocks indicated by two reference frame indices and two corresponding motion vectors in two reference frames. The motion vector may be encoded and included in the bitstream in various ways.
[0178] In some implementations, the encoding system or decoding system may maintain a decoded picture buffer (decoded picture buffer, DPB). Some images / pictures may be kept in the DPB waiting to be displayed (in the decoding system), and some images / pictures in the DPB may be used as reference frames for implementing inter-frame prediction (in the decoding system or encoding system). In some implementations, the reference frames in the DPB may be marked as short-term references or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame for inter-frame prediction of a block in a subsequent video frame that is a predefined number (e.g., 2) closest to the current frame in decoding order. A long-term reference frame may include a frame in the DPB, which may be used to predict an image block in a frame that is more than a predefined number of frames away from the current frame in decoding order. Such marking information about short-term reference frames and long-term reference frames may be referred to as a reference picture set (Reference Picture Set, RPS), and may be added to the header of each frame in the encoded bitstream. Each frame in the encoded video stream may be identified by a Picture Order Counter (POC), which is numbered according to the playback sequence, either in an absolute manner or relative to a group of pictures starting from, for example, an I-frame.
[0179] In some example implementations, one or more reference picture lists containing identifications of short-term reference frames and long-term reference frames for inter-frame prediction may be formed based on information in the RPS. For example, a single picture reference list may be formed for unidirectional inter-frame prediction, the single picture reference list being denoted as L0 reference (or reference list 0), while two picture reference lists may be formed for bidirectional inter-frame prediction, the two picture reference lists being denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 list and the L1 list may be ordered in various predetermined ways. The length of the L0 list and the length of the L1 list may be signaled in the video bitstream. When multiple references used to generate a prediction block by weighted averaging in a composite prediction mode are on the same side of the frame where the block to be predicted is located, the unidirectional inter-frame prediction may be in a single reference mode or in a composite reference mode. Bidirectional inter-frame prediction may be only a composite mode because bidirectional inter-frame prediction involves at least two reference blocks.
[0180] In some implementations, a merge mode (MM) for inter-frame prediction may be implemented. Typically, for merge mode, one or more of the motion vectors in the single reference prediction of the current PB or the motion vectors in the composite reference prediction may be derived from (one or more) other motion vectors, rather than being calculated and signaled independently. For example, in the encoding system, the (one or more) current motion vectors of the current PB may be represented by the difference between the (one or more) current motion vectors and one or more other already encoded motion vectors (referred to as reference motion vectors). Such a difference of (one or more) motion vectors, rather than the entirety of (one or more) current motion vectors, may be encoded and included in the bitstream, and may be linked to (one or more) reference motion vectors. Accordingly, in the decoding system, the (one or more) motion vectors corresponding to the current PB may be derived based on (one or more) decoded motion vector differences and (one or more) decoded reference motion vectors linked thereto. As a specific form of general merge mode (MM) inter-frame prediction, such inter-frame prediction based on (one or more) motion vector differences may be referred to as merge mode with motion vector difference (MMVD). Therefore, MM in general or MMVD in particular can be implemented to exploit the correlation between motion vectors associated with different PBs to improve coding efficiency. For example, adjacent PBs may have similar motion vectors, so the MVD may be small and can be efficiently encoded. For another example, for blocks that are similarly positioned / placed in space, the motion vectors may be correlated in time (between frames).
[0181] In some example implementations, an MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Additionally or alternatively, an MMVD flag may be included during the encoding process and signaled in the bitstream to indicate whether the current PB is in MMVD mode. MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, frame level, picture level, sequence level, etc. For a specific example, both an MM flag and an MMVD flag may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether the MMVD mode is for the current CU.
[0182] In some example implementations of MMVD, a list of reference motion vectors (RMV) or MV predictor candidates for motion vector prediction may be formed for a block being predicted. The list of RMV candidates may contain a predetermined number (e.g., 2) of MV predictor candidate blocks whose motion vectors may be used to predict the current motion vector. The RMV candidate blocks may include blocks selected from neighboring blocks and / or temporal blocks in the same frame (e.g., blocks at the same position in a previous frame or subsequent frame of the current frame). These options represent blocks at a spatial position or temporal position relative to the current block that may have motion vectors similar to or identical to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two or more candidates. In order to be on the list of RMV candidates, for example, the candidate block may be required to have the same reference frame (or frame) as the current block, must exist (e.g., when the current block is close to the edge of the frame, a boundary check needs to be performed), and must have been encoded during the encoding process and / or decoded during the decoding process. In some implementations, if available and the above conditions are met, the list of merge candidates can be first filled with spatially adjacent blocks (scanned in a specific predefined order), and then the list of merge candidates can be filled with temporal blocks if space is still available in the list. For example, adjacent RMV candidate blocks can be selected from the left block and the top block of the current block. The list of RMV predictor candidates can be dynamically formed as a dynamic reference list (DRL) at various levels (sequence, picture, frame, slice, super block, etc.). The DRL can be signaled in the bitstream.
[0183] In some implementations, the actual MV predictor candidate used as a reference motion vector for predicting the motion vector of the current block can be signaled. In the case where the RMV candidate list contains two candidates, a one-bit flag called a merge candidate flag can be used to indicate the selection of the reference merge candidate. For the current block being predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list. The encoder can determine which of the RMV candidates more closely predicts the MV of the current coding block and signal the selection to the DRL as an index.
[0184] In some example implementations of MMVD, after an RMV candidate is selected and used as a base motion vector predictor for a motion vector to be predicted, a motion vector difference (MVD or ΔMV, representing the difference between the motion vector to be predicted and a reference candidate motion vector) may be calculated in the encoding system. Such an MVD may include information representing the magnitude of the MV difference and the direction of the MV difference, both of which may be signaled in the bitstream. The motion difference magnitude and the motion difference direction may be signaled in various ways.
[0185] In some example implementations of the MMVD, a distance index may be used to specify the magnitude information of the motion vector difference, and a distance index may be used to indicate one of a set of predefined offsets representing a predefined motion vector difference from a starting point (reference motion vector). The MV offset according to the signaled index may then be added to the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset may be determined by the direction information of the MVD. An example predefined relationship between the distance index and the predefined offset is specified in Table 2.
[0186] Table 2 - Example relationship between distance index and predefined MV offsets
[0187]
[0188] In some example implementations of MMVD, a direction index may be further signaled, and the direction index may be used to represent the direction of the MVD relative to a reference motion vector. In some implementations, the direction may be limited to either a horizontal direction or a vertical direction. An example 2-bit direction index is shown in Table 3. In the example of Table 3, the description of the MVD may vary depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a unidirectional prediction block or to a bidirectional prediction block, where both reference frame lists point to the same side of the current picture (i.e., both reference pictures have POCs greater than the current picture's POC or are less than the current picture's POC), the symbol in Table 3 may specify the symbol (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bidirectional prediction block having two reference pictures at different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the symbols in Table 3 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 may have an opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the symbols in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset of the reference MV associated with picture reference list 0 has an opposite value.
[0189] Table 3 - Example implementation of the sign of the MV offset specified by the direction index
[0190] Direction Index 00 01 10 11 x-axis (horizontal) + -– N / A N / A y-axis (vertical) N / A N / A + -–
[0191] In some example implementations, the MVD may be scaled according to the difference in the POC in each direction. If the difference in the POC in both lists is the same, no scaling is required. Otherwise, if the difference in the POC in reference list 0 is greater than the difference in the POC of reference list 1, the MVD of reference list 1 is scaled. If the POC difference of reference list 1 is greater than list 0, the MVD of list 0 may be scaled in the same manner. If the starting MV is unidirectionally predicted, then the MVD is added to the available or reference MVs.
[0192] In some example implementations of MVD encoding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately encoding and signaling two MVDs, symmetric MVD encoding may be implemented such that only one MVD needs to be signaled and the other MVD may be derived from the signaled MVD. In such an implementation, motion information including reference picture indices for both list 0 and list 1 is signaled. However, only the MVD associated with, for example, reference list 0 is signaled, and the MVD associated with reference list 1 is not signaled but derived. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" may be included in the bitstream to indicate whether reference list 1 is not signaled in the bitstream. If the flag is 1, indicating that reference list 1 is equal to zero (and therefore not signaled), then a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, meaning that there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, BiDirPredFlag may be set to 1 if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag of 1 may indicate that a symmetric mode flag is additionally signaled in the bitstream. When BiDirPredFlag is 1, the decoder may extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag may be signaled at the CU level (if necessary), and the symmetric mode flag may indicate whether a symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD coding mode is used, and the reference picture indexes of only list 0 and list 1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list 0 (referred to as "MVD0") are signaled, and another motion vector difference "MVD1" is derived instead of signaling. For example, MVD1 can be derived as -MVD0. Thus, in the example symmetric MVD mode, only one MVD is signaled. In some other example implementations of MV prediction, for both single reference mode and composite reference mode MV prediction, a coordination scheme can be used to implement general merge mode MMVD and some other types of MV prediction. Various syntax elements can be used to signal the way to predict the MV of the current block.
[0193] For example, for a single reference mode, the following MV prediction modes may be signaled:
[0194] NEARMV - without any MVD, directly use one of the motion vector predictors (MVP) in the list indicated by the DRL (Dynamic Reference List) index.
[0195] NEWMV - uses one of the motion vector predictors (MVPs) in the list, signaled by the DRL index, as a reference, and applies a delta to the MVP (eg, using MVD).
[0196] GLOBALMV - Use motion vectors based on frame-level global motion parameters.
[0197] Likewise, for the composite reference inter prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes may be signaled:
[0198] NEAR_NEARMV - For each of the two MVs to be predicted, without MVD, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used.
[0199] NEAR_NEWMV - For predicting the first of the two motion vectors, in the absence of MVD, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as the reference MV; for predicting the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as the reference MV in combination with a ΔMV (MVD) additionally signaled.
[0200] NEW_NEARMV - For predicting the second of the two motion vectors, in the absence of MVD, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as the reference MV; for predicting the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as the reference MV in combination with a ΔMV (MVD) that is additionally signaled.
[0201] NEW_NEWMV - uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV, and uses this reference MV in conjunction with the additionally signaled ΔMV to predict each of the two MVs.
[0202] GLOBAL_GLOBALMV - Uses the MV from each reference based on its frame level global motion parameters.
[0203] Thus, the term "NEAR" above refers to MV prediction using a reference MV as a general merge mode without any MVD, whereas the term "NEW" refers to MV prediction involving using a reference MV as in MMVD mode and offsetting the reference MV using a signaled or derived MVD. For composite inter-frame prediction, both the reference base motion vector and the motion vector increment mentioned above may generally be different between two references or two MVDs or may generally be independent, even though, for example, the two MVDs may be correlated and such correlation may be exploited to reduce the amount of information required to signal the two motion vector increments. In order to exploit such correlation, joint signaling of the two MVDs may be implemented and indicated in the bitstream, as described in further detail below.
[0204] In some implementations, an optical flow-based approach may be implemented to refine MVs at the sub-block level for composite inter prediction. In particular, the optical flow equation may be applied to formulate a least squares problem, from which fine motions may be derived from the gradients of composite inter prediction samples. Using these fine motions, the MVs of each sub-block may be refined within the prediction block, which may enhance inter prediction quality. This coding feature reflects an extension of the concept of the bidirectional optical flow approach, as it supports MV refinement when two reference blocks have arbitrary temporal distances from the current block. In some example implementations, four additional inter composite modes may therefore be added, which are listed below:
[0205] NEAR_NEARMV_OPTFLOW;
[0206] NEAR_NEWMV_OPTFLOW;
[0207] NEW_NEARMV_OPTFLOW;
[0208] NEW_NEWMV_OPTFLOW.
[0209] These modes may be referred to as optical flow modes, and the reference MV type is defined as in the regular composite mode (e.g., NEAR_NEWMV_OPTFLOW has the same reference MV type as in NEAR_NEWMV), but composite prediction is performed based on the MV refined subblock-wise instead of the original MV.
[0210] In some example implementations of the MVD, a predefined pixel resolution of the MVD may be allowed. For example, a motion vector accuracy (or precision) of 1 / 8 pixel may be allowed. The MVD described above in various MV prediction modes may be constructed and signaled in various ways. In some implementations, various syntax elements may be used to signal the above-mentioned (one or more) motion vector differences in reference frame list 0 or list 1.
[0211] For example, a syntax element called "mv_joint" may specify which components of the motion vector difference associated with it are non-zero. For MVD, this is signaled jointly for all non-zero components. For example, the value of mv_joint is:
[0212] 0 may indicate that there is no non-zero MVD along the horizontal or vertical direction;
[0213] 1 may indicate that there is non-zero MVD only along the horizontal direction;
[0214] 2 may indicate that there is non-zero MVD only along the vertical direction;
[0215] 3 may indicate that there is non-zero MVD in both the horizontal and vertical directions.
[0216] When the "mv_joint" syntax element of the MVD signals that there are no non-zero MVD components, then no further MVD information may be signaled. However, if the "mv_joint" syntax signals that there are one or two non-zero components, then additional syntax elements may be further signaled for each of the non-zero MVD components as described below.
[0217] For example, a syntax element called "mv_sign" may be used to additionally specify whether the corresponding motion vector difference amount is positive or negative.
[0218] For another example, a syntax element called "mv_class" can be used to specify the class of motion vector differences in a predefined set of classes for the corresponding non-zero MVD component. For example, the predefined classes of motion vector differences can be used to divide the continuous amplitude space of motion vector differences into non-overlapping ranges, where each range corresponds to an MVD class. Therefore, the MVD class notified by the signal indicates the amplitude range of the corresponding MVD component. In the example implementation shown in Table 4 below, a higher class corresponds to a motion vector difference with a larger amplitude range. In Table 4, the symbol (n, m] is used to represent a range of motion vector differences greater than n pixels and less than or equal to m pixels.
[0219] Table 4: Motion vector difference magnitude categories
[0220]
[0221]
[0222] In some other examples, a syntax element called "mv_bit" may also be used to specify the integer portion of the offset between a non-zero motion vector difference component and the starting amplitude of the correspondingly signaled MV class amplitude range. In this way, mv_bit can indicate the amplitude or magnitude of the MVD. The number of bits required in "my_bit" to signal the entire range of each MVD category may vary depending on the MV category. For example, MV_CLASS_0 and MV_CLASS_1 in the implementation of Table 4 may require only a single bit to indicate an integer pixel offset of 1 or 2 from a starting MVD of 0; each higher MV_CLASS in the example implementation of Table 4 may progressively require one more bit for "mv_bit" compared to the previous MV_CLASS.
[0223] In some other examples, a syntax element called "mv_fr" may also be used to specify the first 2 fractional bits of the motion vector difference of the corresponding non-zero MVD component, while a syntax element called "mv_hp" may be used to specify the third fractional bit (high resolution bit) of the motion vector difference of the corresponding non-zero MVD component. The two bits of "mv_fr" essentially provide a 1 / 4 pixel MVD resolution, while the "mv_hp" bit may further provide a 1 / 8 pixel resolution. In some other implementations, more than one "mv_hp" bit may be used to provide an MVD pixel resolution finer than 1 / 8 pixel. In some example implementations, an additional flag may be signaled at one or more of the various levels to indicate whether an MVD resolution of 1 / 8 pixel or higher is supported. If the MVD resolution is not applied to a particular coding unit, the syntax element above for the corresponding unsupported MVD resolution may not be signaled.
[0224] In some of the example implementations above, fractional resolution may be independent of different categories of MVD. In other words, a predefined number of "mv_fr" and "mv_hp" bits for signaling fractional MVD for non-zero MVD components may be used to provide similar options for motion vector resolution regardless of the magnitude of the motion vector difference.
[0225] However, in some other example implementations, the resolution of motion vector differences in various MVD amplitude categories can be distinguished. Specifically, high-resolution MVD for large MVD amplitudes of higher MVD categories may not provide statistically significant improvements in compression efficiency or coding gain. In this way, for a larger MVD amplitude range corresponding to a higher MVD amplitude category, the MVD can be encoded with a reduced resolution (integer pixel resolution or fractional pixel resolution). Similarly, for generally larger MVD values, the MVD can be encoded with a reduced resolution (integer pixel resolution or fractional pixel resolution). Such MVD category-related or MVD amplitude-related MVD resolutions can generally be referred to as adaptive MVD resolutions, amplitude-related adaptive MVD resolutions, or amplitude-related MVD resolutions. The term "resolution" can also be referred to as "pixel resolution". In order to achieve better compression efficiency overall, adaptive MVD resolutions can be implemented in various ways as described by the following example implementations. In particular, the number of signaling bits reduced by for a less accurate MVD may be greater than the additional bits required to encode the inter-prediction residual due to such less accurate MVD, because it is statistically observed that processing the MVD resolution of a large amplitude or high-class MVD at a level similar to that of a low amplitude or low-class MVD in a non-adaptive manner may not significantly increase the inter-prediction residual coding efficiency of blocks with a large amplitude or high-class MVD. In other words, using a higher MVD resolution for a large amplitude or high-class MVD may not yield much coding gain compared to using a lower MVD resolution.
[0226] In some general example implementations, the pixel resolution or precision of the MVD may or may not increase as the MVD class increases. A decrease in the pixel resolution of the MVD corresponds to a coarser MVD (or a larger step size from one MVD level to the next). In some implementations, the correspondence between MVD pixel resolution and MVD class may be specified, predefined, or preconfigured, and thus may not need to be signaled in the coded bitstream.
[0227] In some example implementations, the MV categories of Table 3 may each be associated with a different MVD pixel resolution.
[0228] In some example implementations, each MVD category may be associated with a single allowed resolution. In some other implementations, one or more MVD categories may each be associated with two or more optional MVD pixel resolutions. Thus, signaling in the bitstream of a current MVD component having such an MVD category may be followed by additional signaling indicating the optional pixel resolutions selected for the current MVD component.
[0229] In some example implementations, the adaptively allowed MVD pixel resolutions may include, but are not limited to, 1 / 64 pixel (pixel), 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels ... (arranged in descending order of resolution). In this way, each of the ascending MVD categories may be associated with one of these MVD pixel resolutions in a non-ascending manner. In some implementations, the MVD category may be associated with two or more of the above-mentioned resolutions, and the higher resolution may be lower than or equal to the lower resolution of the previous MVD category. For example, if the MV_CLASS_3 of Table 4 is associated with optional 1 pixel and 2 pixel resolutions, the highest resolution that the MV_CLASS_4 of Table 4 may be associated with would be 2 pixels. In some other implementations, the highest allowable resolution of the MV category may be higher than the lowest allowable resolution of the previous (lower) MV category. However, the average value of the allowable resolutions of the ascending MV categories may only be non-ascending.
[0230] In some implementations, when fractional pixel resolutions higher than 1 / 8 pixel are allowed, the "mv_fr" and "mv_hp" signaling may be extended accordingly to a total of more than 3 fractional bits.
[0231] In some example implementations, fractional pixel resolution may be allowed only for MVD categories that are lower than or equal to a threshold MVD category. For example, fractional pixel resolution may be allowed only for MVD_CLASS_0 and not for all other MV categories in Table 4. Likewise, fractional pixel resolution may be allowed only for MVD categories that are lower than or equal to any of the other MV categories in Table 4. For other MVD categories that are higher than the threshold MVD category, only integer pixel resolutions of the MVD are allowed. In this way, for MVDs whose signaled MVD category is higher than or equal to the threshold MVD category, it may not be necessary to signal fractional resolution signaling such as one or more of the "mv-fr" and / or "mv-hp" bits. For MVD categories with resolutions lower than 1 pixel, the number of bits in the "mv-bit" signaling may be further reduced. For example, for MV_CLASS_5 in Table 4, the range of MVD pixel offsets is (32, 64], so 5 bits are required to signal the entire range with 1 pixel resolution. However, if MV_CLASS_5 is associated with 2-pixel MVD resolution (a lower resolution than 1 pixel resolution), 4 bits instead of 5 bits may be required for "mv-bit", and neither "mv-fr" nor "mv-hp" needs to be signaled after "mv_class" is signaled as MV_CLASS_5.
[0232] In some example implementations, fractional pixel resolution may be allowed only for MVDs with integer values below a threshold integer pixel value. For example, fractional pixel resolution may be allowed only for MVDs less than 5 pixels. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 of Table 4, but not for all other MV classes. For another example, fractional pixel resolution may be allowed only for MVDs less than 7 pixels. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 of Table 4 (with a range below 5 pixels), but not for MV_CLASS_3 and higher (with a range above 5 pixels). For an MVD belonging to MV_CLASS_2, whose pixel range contains 5 pixels, fractional pixel resolution of the MVD may or may be allowed depending on the "mv-bit" value. If the "mv-bit" value is signaled as 1 or 2 (so that the integer part of the signaled MVD is 5 or 6, calculated as the start of the pixel range of MV_CLASS_2, offset by 1 or 2, as indicated by "mv-bit"), then fractional pixel resolution may be allowed. Otherwise, if the "mv-bit" value is signaled as 3 or 4 (so that the integer part of the signaled MVD is 7 or 8), then fractional pixel resolution may not be allowed.
[0233] In some other implementations, only a single MVD value may be allowed for MV classes that are equal to or higher than a threshold MV class. For example, such a threshold MV class may be MV_CLASS_2. Therefore, MV_CLASS_2 and above may only be allowed to have a single MVD value and no fractional pixel resolution. Single allowed MVD values for these MV classes may be predefined. In some examples, the allowed single value may be a higher end value of the corresponding ranges for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be higher than or equal to the threshold class of MV_CLASS_2, and the single allowed MVD values for these classes may be predefined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048, respectively. In some other examples, the allowed single value may be an intermediate value of the corresponding ranges for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be above the class threshold, and the single allowed MVD values for these classes may be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536, respectively. Any other value within the range may also be defined as the single allowed resolution for the corresponding MVD class.
[0234] In the above implementation, when the signaled "mv_class" is equal to or above the predefined MVD class threshold, only the "mv_class" signaling is sufficient to determine the MVD value. Then "mv_class" and "mv_sign" will be used to determine the magnitude and direction of the MVD.
[0235] In this way, when the MVD is signaled for only one reference frame (from reference frame list 0 or list 1, but not both) or for two reference frames jointly, the accuracy (or resolution) of the MVD can depend on the associated category of the motion vector difference in Table 3 and / or the magnitude of the MVD.
[0236] In some other implementations, the pixel resolution or accuracy of the MVD may or may not increase as the MVD amplitude increases. For example, the pixel resolution may depend on the integer portion of the MVD amplitude. In some implementations, fractional pixel resolution may only be allowed for MVD amplitudes less than or equal to an amplitude threshold. For a decoder, the integer portion of the MVD amplitude may be first extracted from the bitstream. The pixel resolution may then be determined, and a decision may then be made as to whether any fractional MVD exists in the bitstream and needs to be parsed (e.g., if the fractional pixel resolution does not allow for a specific extracted MVD integer amplitude, then the fractional MVD bit may not be included in the bitstream that needs to be extracted). The example implementations related to the adaptive MVD pixel resolution associated with the MVD category above are applicable to the adaptive MVD pixel resolution associated with the MVD amplitude. For a specific example, an MVD category that is higher than or includes the amplitude threshold may be allowed to have only one predefined value.
[0237] Turning to the various composite inter prediction modes, in which each MV can be predicted by a reference motion vector and encoded by an MVD, the two MVDs can be signaled separately or jointly in the bitstream, as described above. Thus, in some example implementations, in addition to the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMW, NEW_NEWMV, and GLOBAL_GLOBALMV modes described above, another inter prediction mode called JOINT_NEWMV can be introduced for a mode in which the MVDs of reference list 0 and reference list 1 are jointly signaled. Specifically, if the inter prediction mode is indicated as NEW_NEWMV, the MVDs of reference list 0 and reference list 1 are signaled separately, while when the inter prediction mode is indicated as JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are jointly signaled. In particular, for joint MVD, only one MVD called joint_delta_mv may need to be signaled and transmitted in the bitstream, and the MVDs of reference list 0 and reference list 1 may be derived from joint_delta_mv. The derived MVD may then be combined with the reference motion vectors in reference list 0 or reference list 1 to generate two motion vectors for locating reference blocks for composite inter prediction.
[0238] In some implementations of composite inter prediction, the JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMW, NEW_NEWMV, and GLOBAL_GLOBALMV modes. In such implementations, syntax may be included in the bitstream at any of the various signaling levels (e.g., sequence level, picture level, frame level, slice level, tile level, super block level, etc.) for indicating any of these alternative composite inter prediction modes. Alternatively, the JOINT_NEWMV mode may be implemented as a sub-mode of the NEW_NEWMV mode. In other words, in the NEW_NEWMV mode, the two MVDs for the two reference blocks are signaled jointly (thus the JOINT_NEWMV sub-mode) or not signaled jointly (another sub-mode of the NEW_NEWMV mode). In such an implementation, a first syntax element may be included in the bitstream for indicating any one of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMW, NEW_NEWMV, and GLOBAL_GLOBALMV modes, and when the first syntax element indicates that the NEW_NEWMV mode is selected for the coding block, then a second syntax element may be further included in the bitstream and extracted by the decoder for indicating whether the MVDs for the coding block are signaled individually or jointly.
[0239] For a joint MVD implementation in composite inter prediction, the MVD associated with a reference MV may be derived from a signaled joint MVD (e.g., the joint_delta_mv described above) from the bitstream. For example, such a derivation may involve scaling the signaled joint MVD to obtain one or both of the two MVDs. In other words, the signaled joint MVD may be scaled before being added to the motion vector predictor(s) (MVP(s)) or the reference MV. As a result of the scaling, the precision or pixel resolution of the scaled MVD may differ from the allowed precision of the motion vector difference. In some example implementations, such MVD(s) scaled from the joint signaled MVD may first be quantized to the allowed precision of the MVD of the current picture or slice or tile or super block or coded block before being added to the reference MVD(s) used to generate the motion vector(s).
[0240] In some example implementations, a frame index of a reference frame in a composite inter prediction mode may be signaled in a bitstream. The frame index may correspond to a picture order counter (POC) associated with the reference frame. The distance between the reference frame and the current frame may be defined and represented by the difference between the corresponding POCs. The direction of the reference frame (before or after the current frame) may be represented by a sign. In this way, a signed distance may be used to represent the position of the reference frame relative to the current frame. The reference frames used for the composite inter prediction mode may be referred to as a first reference frame and a second reference frame.
[0241] In some example implementations, a reference frame index (or multiple reference frame indices) for indicating (one or more) reference frames may be signaled prior to syntax indicating whether (one or more) MVDs are signaled jointly. As further shown below, the distance (or distances) between the (one or more) reference frames and the current frame may be used to derive the (one or more) MVDs from the jointly signaled MVDs. By signaling a frame index or multiple frame indices for (one or more) reference frames prior to signaling whether to use a joint MVD or prior to signaling a joint MVD, the distance (or distances) between (one or more) reference frames and the current frame will be available in the buffer when needed.
[0242] In some example implementations, the allowed precision of the MVD(s) may include, but is not limited to, 1 / 64 pixel, 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, integer pixel (also referred to as full pixel), 2 pixels, 3 pixels, 4 pixels, or other integer precision. As described above, the allowed precision of the MVD(s), for example, may be predetermined or adaptively determined. For example, the adaptive MVD resolution may be determined by the corresponding MVD amplitude or MV category as described above.
[0243] In one specific example implementation, for the current coded block (or super block or picture), only full-pixel MVD may be allowed, and due to scaling, the precision of the scaled MVD from the jointly signaled MVD may become higher than full-pixels. For example, the scaled MVD may have a precision of 1 / 2 pixel. In that case, the precision of the scaled MVD may then need to be quantized to full pixels.
[0244] Likewise, in another specific example implementation, for the current coded block (or super block or picture), a 1 / 4 pixel MVD may be allowed, and due to scaling, the precision of the scaled MVD from the jointly signaled MVD may become higher than or become 1 / 4 pixel. For example, the scaled MVD may have a precision of 1 / 8 pixel. In that case, then the precision of the scaled MVD may need to be quantized to 1 / 4 pixel.
[0245] In some example implementations, when the distances between two reference frames and the current frame are the same, no scaling may be performed. In other words, the jointly signaled MVD may be directly used as the actual MVD for the two motion vectors with respect to the two reference frames.
[0246] In some example implementations, a joint MVD may be signaled such that the jointly signaled MVD may be used directly for one of the reference motion vectors and only needs to be scaled to derive another MVD for another reference motion vector. For example, in some implementations, an MVD associated with a reference frame from a predetermined reference list (e.g., one of reference list 0 or reference list 1) may be directly signaled, and another MVD associated with another reference list may be derived or scaled from the signaled MVD. Such derivation or scaling may be based on the distance between the two reference frames and the current frame. For a specific example, the MVD for a first reference frame of reference list 1 may be directly signaled, while the MVD for a second reference frame in reference list 0 may be derived / scaled from the signaled MVD. Scaling the signaled MVD to generate the MVD associated with the second reference frame in reference list 0 may be based on a first frame distance (between the first reference frame and the current frame) and a second frame distance (between the current frame and the second reference frame). For example, such scaling of the jointly signaled MVD to generate the MVD for the second reference frame may be based on a ratio between the first distance and the second distance.
[0247] In some other example implementations, a joint MVD may be signaled such that the jointly signaled MVD may be used directly for one of the reference motion vectors and only needs to be scaled to derive another MVD for another reference motion vector. For example, in some implementations, the MVD associated with the reference frame of the two reference frames having a shorter distance to the current frame may be directly signaled as a joint MVD without scaling, while the reference frame of the two reference frames having a longer distance to the current frame is derived / scaled from the jointly signaled MVD. Such derivation or scaling may be based on the distance between the two reference frames and the current frame. For example, scaling the signaled joint MVD to generate an MVD for the reference frame having the longer distance may be based on a ratio between the longer frame distance and the shorter frame distance.
[0248] Likewise, in some example implementations, a joint MVD may be signaled such that the jointly signaled MVD may be used directly for one of the reference motion vectors and need only be scaled to derive another MVD for another reference motion vector. For example, in some implementations, an MVD associated with a reference frame of two reference frames having a longer distance from the current frame may be directly signaled as a joint MVD without scaling, while a reference frame of two reference frames having a shorter distance from the current frame is derived / scaled from the jointly signaled MVD. Such derivation or scaling may be based on the distance between the two reference frames and the current frame. For example, scaling the signaled joint MVD to generate an MVD for a reference frame having a shorter distance may be based on a ratio between the shorter frame distance and the longer frame distance.
[0249] In some example implementations, the above-mentioned scaling may be additionally based on the orientation of the reference frame. For example, the direction of scaling (zooming in or out) may depend on the orientation of the reference frame (whether the reference frame has a higher POC or a lower POC).
[0250] In some other implementations, additional scaling factors or weighting factors may be applied directly to the jointly signaled MVD or to the jointly signaled MVD scaled as above to derive one or both of the MVDs. Such factors may be predefined or may be signaled at various levels (sequence, picture, frame, slice, tile, superblock, etc., as part of, for example, SPS, VPS, PPS, picture header, tile header, slice header, frame header, CTU (or superblock) header, block level). In some cases, the implementation of this additional or alternative scaling factor may help improve coding gain because the actual jointly signaled MVD may be represented by a smaller number of bits when the additional scaling factors or weighting factors are more sparsely decomposed and signaled at a higher syntax level (e.g., SPS, VPS, PPS, picture header, tile header, slice header, frame header, CTU (or superblock) header). In some example implementations, such scaling or weighting factor may be between -1 and 1. For example, such scaling or weighting factor may be 1 / 2.
[0251] Quantizing the scaled jointly signaled MVD to the allowed precision of the MVD or motion vector of the current picture, slice, tile, super block, or coding block can be implemented in various ways. For example, the scaled MVD value can be down-quantized to the next lower allowed value based on the allowed precision of the MVD or MV. For another example, the scaled MVD value can be up-quantized to the next higher allowed value based on the allowed precision of the MVD or MV.
[0252] In one specific example implementation, quantization of the scaled jointly signaled MVD may be performed by discarding N least significant bits of the scaled MVD to maintain precision consistent with allowed precision, where N is an integer and is determined by the difference between the precision of the scaled MVD and the allowed precision of the MVD.
[0253] In another alternative example implementation, quantization of the scaled jointly signaled MVD may be performed by adding a rounding factor and then discarding the N least significant bits to keep the precision consistent with the allowed precision. For example, both the rounding factor and the integer N may be determined based at least on the scaled MVD and the difference between the precision and the allowed precision of the MVD.
[0254] Fig.18 A flowchart 1800 of an example method for scaling a jointly signaled MVD and quantizing the scaled MVD following the principles underlying the above implementations is shown. The example method flow starts at 1801. In S1810, at least one flag is extracted from a video stream. In S1820, based on the extracted flag, it is determined that: a video block is inter-predicted by at least a first reference block located by a first motion vector in a first reference frame and a second reference block located by a second motion vector in a second reference frame; the first motion vector is to be predicted by a first motion vector difference (MVD) relative to the first reference motion vector; and the second motion vector is to be predicted by a second MVD relative to the second reference motion vector; wherein the first MVD and the second MVD are jointly signaled as a single MVD in the video stream. In S1830, a single MVD is received from the video stream. In S1840, the single MVD is scaled to generate a scaled MVD. In S1850, the scaled MVD is quantized to generate a quantized MVD having a MVD pixel resolution according to an allowed accuracy of the first motion vector or the second motion vector. In S1860, a first motion vector and a second motion vector are generated based on the quantized MVD and a corresponding reference motion vector of the first reference motion vector and the second reference motion vector. In S1870, a first reference block or a second reference block for inter-frame prediction of the video block is reconstructed from one of the first reference frame or the second reference frame based on one of the first motion vector or the second motion vector. The example method stops at S1899.
[0255] In the embodiments and implementations of the present disclosure, any step and / or operation can be combined or arranged in any amount or order as needed. Two or more of the steps and / or operations can be performed in parallel. The embodiments and implementations in the present disclosure can be used alone or in combination in any order. In addition, each of the method (or embodiment), encoder and decoder can be implemented by a processing circuit system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transient computer-readable medium. The embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term block can be interpreted as a prediction block, a coding block or a coding unit, i.e., a CU. The term block here can also be used to refer to a transform block. In the following items, when talking about block size, it can refer to block width or height, or the maximum value of width and height, or the minimum value of width and height, or area size (width*height), or the aspect ratio of the block (width:height or height:width).
[0256] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Fig.19 A computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0257] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code comprising instructions that may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.
[0258] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, etc.
[0259] Fig.19 The components for the computer system (1900) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (1900).
[0260] The computer system (1900) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0261] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard (1901), mouse (1902), touchpad (1903), touch screen (1910), data gloves (not shown), joystick (1905), microphone (1906), scanner (1907), camera (1908).
[0262] The computer system (1900) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include: tactile output devices (e.g., tactile feedback through a touch screen (1910), a data glove (not shown), or a joystick (1905), but there may also be tactile feedback devices that do not function as input devices); audio output devices (e.g., speakers (1909), headphones (not depicted)); visual output devices (e.g., screens (1910), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through, for example, stereo output; virtual reality glasses (not depicted); holographic displays and smoke generators (not depicted)); and printers (not depicted).
[0263] The computer system (1900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1920) with CD / DVD etc. media (1921), thumb drives (1922), removable hard drives or solid-state drives (1923), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), etc.
[0264] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0265] The computer system (1900) may also include an interface (1954) to one or more communication networks (1955). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include: local area networks (e.g., Ethernet, wireless LAN); cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicle and industrial networks including CAN buses, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1949) (e.g., a USB port of the computer system (1900)); other networks are typically integrated into the core of the computer system (1900) by attaching to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smart phone computer system). Using any of these networks, the computer system (1900) can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., a CAN bus to certain CAN bus devices), or bidirectional (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.
[0266] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core ( 1940 ) of the computer system ( 1900 ).
[0267] The core (1940) may include one or more central processing units (CPUs) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1943), hardware accelerators for certain tasks (1944), graphics adapters (1950), etc. These devices, along with read-only memory (ROM) (1945), random access memory (1946), internal mass storage devices (e.g., internal non-user accessible hard drives, SSDs, etc.) (1947), may be connected via a system bus (1948). In some computer systems, the system bus (1948) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (1948) directly or via a peripheral bus (1949). In an example, a screen (1910) may be connected to a graphics adapter (1950). The architecture of the peripheral bus includes PCI, USB, etc.
[0268] The CPU (1941), GPU (1942), FPGA (1943) and accelerator (1944) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1945) or RAM (1946). Transient data can also be stored in RAM (1946), while permanent data can be stored in, for example, an internal mass storage device (1947). Fast storage and retrieval of any of the memory devices can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROMs (1945), RAMs (1946), etc.
[0269] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0270] The embodiments of the present application also provide a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method provided in the above embodiments.
[0271] As a non-limiting example, a computer system (1900) having an architecture and in particular a core (1940) can provide functionality as a result of (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as described above and certain storage devices of the core (1940) having non-transitory properties, such as a core internal mass storage device (1947) or ROM (1945). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (1940). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core (1940) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM (1946) and modifying such data structures according to processing defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise implemented in circuits (e.g., accelerators (1944)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and references to logic may include software. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0272] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. Therefore, it should be recognized that those skilled in the art will be able to conceive of many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.
[0273] Appendix A: Acronyms
[0274] JEM: Joint Exploration Model
[0275] VVC: Versatile Video Coding
[0276] BMS: Benchmark Set
[0277] MV: Motion Vector
[0278] HEVC: High Efficiency Video Coding
[0279] SEI: Supplemental Enhancement Information
[0280] VUI: Video Availability Information
[0281] GOP: Group of Pictures
[0282] TU: Transform Unit
[0283] PU: prediction unit
[0284] CTU: Coding Tree Unit
[0285] CTB: Coding Tree Block
[0286] PB: prediction block
[0287] HRD: Hypothetical Reference Decoder
[0288] SNR: Signal to Noise Ratio
[0289] CPU: Central Processing Unit
[0290] GPU: Graphics Processing Unit
[0291] CRT: cathode ray tube
[0292] LCD: Liquid Crystal Display
[0293] OLED: Organic Light Emitting Diode CD: Compact Disc
[0294] DVD: Digital Video Disc
[0295] ROM: Read Only Memory
[0296] RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit
[0297] PLD: Programmable Logic Device
[0298] LAN: Local Area Network
[0299] GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus
[0300] PCI: Peripheral Component Interconnect
[0301] FPGA: Field Programmable Gate Array SSD: Solid State Drive
[0302] IC: Integrated Circuit
[0303] HDR: High Dynamic Range
[0304] SDR: Standard Dynamic Range
[0305] JVET: Joint Video Exploration Team
[0306] MPM: Most Probable Mode
[0307] WAIP: Wide Angle Intra Prediction
[0308] CU: Coding Unit
[0309] PU: prediction unit
[0310] TU: Transform Unit
[0311] CTU: Coding Tree Unit
[0312] PDPC: Position Dependent Prediction Combination
[0313] ISP: Intra-frame sub-partitioning
[0314] SPS: Sequence Parameter Set
[0315] PPS: Picture Parameter Set
[0316] APS: Adaptive Parameter Set
[0317] VPS: Video Parameter Set
[0318] DPS: Decoding Parameter Set
[0319] ALF: Adaptive Loop Filter
[0320] SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter
[0321] CCSO: Cross Component Sample Offset
[0322] LSO: Local Sample Offset
[0323] LR: Loop Restoration Filter
[0324] AV1: AOMedia Video 1 AV2: AOMedia Video 2 MVD: Motion Vector Difference
[0325] CfL: Predicting Chroma from Luma
[0326] SDT: Semi-Decoupled Tree
[0327] SDP: Semi-Decoupled Partitioning
[0328] SST: Semi-Separating Tree
[0329] SB: Super Block
[0330] IBC (or IntraBC): Intra Block Copy
[0331] CDF: Cumulative density function
[0332] SCC: Screen Content Coding
[0333] GBI: Generalized Bidirectional Prediction
[0334] BCW: Bidirectional Prediction with CU-Level Weights CIIP: Combined Intra-Inter Prediction
[0335] POC: Image Sequential Counting
[0336] RPS: Reference Picture Set
[0337] DPB: Decoded Picture Buffer
[0338] MMVD: Merge mode with motion vector difference
Claims
1. A method for processing video blocks of a video stream, It is characterized in that The method comprises: extracting at least one marker from the video stream; Based on the at least one flag, determining: The video block is inter-predicted by at least a first reference block located by a first motion vector in a first reference frame and a second reference block located by a second motion vector in a second reference frame; The first motion vector is to be predicted by a first motion vector difference (MVD) relative to a first reference motion vector; and The second motion vector is predicted by a second MVD relative to a second reference motion vector; wherein the first MVD and the second MVD are jointly signaled as a single MVD in the video stream; receiving a single MVD signaled in the video stream; scaling the single MVD to generate a scaled MVD; quantizing the scaled MVD using an MVD pixel resolution to generate a quantized MVD according to an allowed accuracy of the first motion vector or the second motion vector; generating one of the first motion vector and the second motion vector based on the quantized MVD and a corresponding reference motion vector of the first reference motion vector and the second reference motion vector; and One of the first reference block or the second reference block used for inter-frame prediction of the video block is reconstructed from one of the first reference frame or the second reference frame based on one of the first motion vector and the second motion vector.
2. The method according to claim 1, It is characterized in that The method further includes extracting, from the video stream, a first frame index for identifying the first reference frame and a second frame index for identifying the second reference frame.
3. The method according to claim 2, It is characterized in that The at least one flag is signaled in the video stream prior to the first frame index and the second frame index.
4. The method according to claim 1, It is characterized in that The at least one flag comprises an indication of a joint MVD mode among a set of composite inter prediction modes.
5. The method according to claim 4, It is characterized in that The set of composite inter prediction modes includes: the joint MVD mode, wherein both the first motion vector and the second motion vector are jointly predicted by a signaled MVD; NEAR-NEAR inter prediction mode, wherein both the first motion vector and the second motion vector are signaled without any MVD; NEAR-NEW inter prediction mode, wherein the first motion vector is signaled without any MVD and the second motion vector is predicted by the signaled MVD; NEW-NEAR inter prediction mode, wherein the second motion vector is signaled without any MVD and the first motion vector is predicted by the signaled MVD; and NEW-NEW inter prediction mode, wherein the first motion vector and the second motion vector are independently predicted by independently signaled MVDs.
6. The method according to claim 4, Features: The joint MVD mode comprises a sub-mode of a compound NEW-NEW inter prediction mode, wherein both the first motion vector and the second motion vector are predicted rather than directly signaled; and The at least one flag is included in the video stream and is extracted from the video stream in response to the video block being predicted in the composite NEW-NEW inter prediction mode.
7. The method according to any one of claims 1 to 6, It is characterized in that The MVD pixel resolution of one of the first MVD or the second MVD includes one of 1 / 64 pixel, 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel or integer pixel resolution.
8. The method according to claim 7, It is characterized in that The MVD pixel resolution is preconfigured or adaptively signaled in the video stream.
9. The method according to claim 7, Features: The MVD pixel resolution is lower than a scaled pixel resolution of the scaled MVD; as well as The scaled MVD is quantized by reducing the scaled pixel resolution of the scaled MVD to the MVD pixel resolution to generate a quantized MVD.
10. The method according to any one of claims 1 to 6, It is characterized in that The single MVD is scaled based on a first distance between a current frame of the video block and the first reference frame or a second distance between the current frame of the video block and the second reference frame to generate the scaled MVD.
11. The method according to claim 10, It is characterized in that One of the first motion vector and the second motion vector generated based on the quantized MVD corresponds to the first reference frame or the second reference frame belonging to a predefined reference frame list.
12. The method according to claim 10, It is characterized in that When the first distance is equal to the second distance, a scaling factor used for performing the scaling is 1.
13. The method according to claim 10, It is characterized in that One of the first motion vector or the second motion vector generated based on the quantized MVD corresponds to a reference frame of the first reference frame and the second reference frame that is smaller in distance from the current frame.
14. The method according to claim 10, It is characterized in that One of the first motion vector or the second motion vector generated based on the quantized MVD corresponds to a reference frame of the first reference frame and the second reference frame that has a larger distance from the current frame.
15. The method according to claim 10, It is characterized in that The scaling is performed based on an additional scaling factor in addition to the first distance and the second distance, the additional scaling factor being between -1 and 1.
16. The method according to claim 15, It is characterized in that The method further comprises extracting the additional scaling factor from the video stream as part of a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, or a coding tree unit header.
17. The method according to any one of claims 1 to 6, It is characterized in that Quantizing the scaled MVD includes removing N least significant bits from the scaled MVD, N being an integer derived based on a difference between the MVD pixel resolution and a scaled pixel resolution of the scaled MVD.
18. The method according to any one of claims 1 to 6, It is characterized in that Quantizing the scaled MVD includes adding a rounding factor and then removing N least significant bits from the scaled MVD, where N is an integer derived based on a difference between the MVD pixel resolution, a scaled pixel resolution of the scaled MVD, and the rounding factor.
19. A video processing device, It is characterized in that The apparatus comprises a memory for storing instructions and a processor, the processor being configured to execute the instructions to perform the method according to any one of claims 1 to 18 based on the first motion vector.
20. A non-transitory computer readable medium for storing computer instructions, which when executed by a processor of a video device are configured to cause the video device to perform the method according to any one of claims 1 to 18 based on the first motion vector.