Video decoding method and device, computer equipment and storage medium

By introducing the adaptive motion vector difference (ADAPTMV) mode in video encoding, the pixel resolution is adaptively adjusted according to the magnitude of MVD in the inter-frame prediction of video blocks, the problem of low encoding efficiency caused by fixed MVD resolution in the prior art is solved, and higher compression efficiency and encoding and decoding efficiency are achieved.

CN119996706APending Publication Date: 2025-05-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510159203.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-05-25
Filing Date
2022-05-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing video encoding technology, the pixel resolution of motion vector difference (MVD) is fixed, and it cannot be adaptively adjusted according to the magnitude of the MVD, resulting in low encoding efficiency in high MVD situations.

Method used

An adaptive motion vector difference (ADAPTMV) mode is proposed. By explicitly determining the pixel resolution of the MVD in the inter-frame prediction of the video block, the relevant syntax elements are extracted and decoded from the video code stream according to whether the ADAPTMV mode is notified to be used to optimize the accuracy settings of the MVD.

Benefits of technology

By adaptively adjusting the pixel resolution of MVD, the compression efficiency of video encoding is improved and the encoding and decoding efficiency is enhanced, especially in high MVD situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996706A_ABST
    Figure CN119996706A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video decoding method and device, computer equipment and a storage medium. The method comprises the following steps: receiving a video code stream; extracting an inter prediction syntax element from the video bitstream to determine whether to signal an ADAPTMV mode for at least one video block in the video bitstream, the ADAPTMV mode being a single reference inter prediction mode with adaptive motion vector difference (MVD) pixel resolution; determining a current MVD pixel resolution associated with the at least one video block based on whether the ADAPTMV mode is signaled in the inter prediction syntax element; and, based on whether the ADAPTMV mode is signaled in the inter prediction syntax element and the current MVD pixel resolution, extracting at least one MVD-related syntax element associated with the at least one video block from the video bitstream, and decoding the at least one MVD-related syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by Reference

[0002] This application claims priority to U.S. non-provisional application No. 17 / 824,248, filed on May 25, 2022, and entitled “Adaptive Precision for Single Reference Motion Vector Difference,” and priority to U.S. provisional application No. 63 / 282,549, filed on November 23, 2021, and entitled “Adaptive MVD for Single Reference,” the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The embodiments of the present application relate to video coding, and more particularly to a video decoding method, apparatus, computer equipment and storage medium. Background Art

[0004] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors described in the background section and various aspects of this specification is performed does not indicate that it is prior art at the time of filing this application, and it is never explicitly or implicitly admitted that it is prior art for this application.

[0005] Video encoding and decoding are made possible by inter-picture prediction techniques with motion compensation. An uncompressed digital video may include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures has a fixed or variable picture rate (also informally referred to as a frame rate), for example, 60 pictures per second or 60Hz. Uncompressed video has very large bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution, 60Hz frame rate) with 8 bits per sample requires close to 1.5Gbit / s bandwidth. One hour of such video requires more than 600GB of storage space.

[0006] One goal of video encoding and decoding is to reduce redundant information in the input video signal through compression. Video compression can help reduce the bandwidth or storage space requirements mentioned above, in some cases by two or more orders of magnitude. Both lossless and lossy compression, as well as combinations of the two, can be used. Lossless compression refers to techniques that reconstruct an exact replica of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of distortion allowed depends on the application. For example, users of some consumer streaming applications may tolerate higher distortion than users of television applications. The achievable compression ratio reflects that higher allowed / tolerable distortion results in higher compression ratios.

[0007] Video encoders and decoders may utilize several broad categories of techniques including, for example: motion compensation, transforms, quantization, and entropy coding.

[0008] Video codec techniques may include known intra-frame coding techniques. In intra-frame coding, sample values ​​are represented without reference to samples or other data of previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture can be an intra-frame picture. Intra-frame pictures and their derivatives (such as independent decoder refresh pictures) can be used to reset the decoder state, so they can be used as the first picture in the encoded video code stream and video session, or as a still image. Samples of intra-frame blocks can be used for transformation, and the transformation coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique to minimize sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficient, the fewer bits are needed to represent the block after entropy coding at a given quantization step size.

[0009] Conventional intra-frame coding, as known from coding techniques such as MPEG-2, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to derive data blocks from, for example, surrounding sample data and / or metadata, which are obtained during spatially adjacent encoding / decoding and before the decoding order. Such techniques have subsequently been referred to as "intra-frame prediction" techniques. It should be noted that, at least in some cases, intra-frame prediction uses only reference data of the current picture being reconstructed, and not reference data of reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technique, the techniques used can be encoded as intra prediction modes. In some cases, a mode can have sub-modes and / or parameters, and these modes can be encoded separately or contained in a mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination affects the codec efficiency gain through intra prediction, and therefore the entropy coding technique used to convert the codewords into the bitstream.

[0011] H.264 introduced an intra-frame prediction mode, which was improved in H.265 and further improved in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and BenchMark Set (BMS). The values ​​of neighboring samples belonging to existing samples can be used to form a prediction block. Depending on the direction, the sample values ​​of neighboring samples are copied to the prediction block. The reference of the direction used can be encoded in the bitstream or can be predicted by itself.

[0012] In the prior art, the pixel resolution or precision of MVD is the same for all MVD values ​​in an image. However, the researchers found that providing higher precision for larger MVD values ​​did not bring enough improvement in compression efficiency. Summary of the invention

[0013] Embodiments of the present application provide a method, apparatus, computer device, and storage medium for video decoding, which are capable of providing and signaling adaptive resolution of motion vector differences in inter-frame prediction of video blocks.

[0014] On the one hand, an embodiment of the present application provides a video decoding method, including:

[0015] Receive video stream;

[0016] Extracting an inter-frame prediction syntax element from the video bitstream to determine whether to signal an ADAPTMV mode for at least one video block in the video bitstream, wherein the ADAPTMV mode refers to a single reference inter-frame prediction mode with adaptive motion vector difference (MVD) pixel resolution;

[0017] determining a current MVD pixel resolution associated with the at least one video block based on whether the ADAPTMV mode is signaled in the inter prediction syntax element; and,

[0018] Based on whether the ADAPTMV mode is signaled in the inter-frame prediction syntax element and the current MVD pixel resolution, at least one MVD-related syntax element associated with the at least one video block is extracted from the video bitstream and the at least one MVD-related syntax element is decoded.

[0019] On the other hand, an embodiment of the present application further provides a video decoding device, including:

[0020] A receiving module, used for receiving a video code stream;

[0021] an extraction module, configured to extract inter-frame prediction syntax elements from the video bitstream to determine whether to signal an ADAPTMV mode for at least one video block in the video bitstream, wherein the ADAPTMV mode refers to a single reference inter-frame prediction mode with an adaptive motion vector difference (MVD) pixel resolution;

[0022] a determination module for determining a current MVD pixel resolution associated with the at least one video block based on whether the ADAPTMV mode is signaled in the inter prediction syntax element; and

[0023] A decoding module is used to extract at least one MVD-related syntax element associated with the at least one video block from the video bitstream based on whether the ADAPTMV mode is signaled in the inter-frame prediction syntax element and the current MVD pixel resolution, and decode the at least one MVD-related syntax element.

[0024] On the other hand, an embodiment of the present application further provides a computer device, including a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the video decoding method as described above.

[0025] On the other hand, an embodiment of the present application further provides a non-temporary computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the processor implements the video decoding method as described above.

[0026] On the other hand, the embodiment of the present application further provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned video decoding method.

[0027] It can be seen from the above technical solution that the method provided in the embodiment of the present application provides a new inter-frame coding mode, namely, the ADAPTMV mode, under the single reference mode; by explicitly determining the ADAPTMV mode, and then determining the current MVD pixel resolution and at least one MVD-related syntax element, it is possible to indicate whether the accuracy of the MVD depends on the associated MVD category and / or value, thereby optimizing the accuracy setting, achieving better compression efficiency overall, and improving encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Other features, properties and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, in which:

[0029] Figure 1A is a schematic diagram of an exemplary subset of intra prediction modes;

[0030] Figure 1B is a schematic diagram of an exemplary intra prediction direction;

[0031] Figure 2 A schematic diagram showing a current block and its surrounding spatial merging candidates for motion vector prediction according to an embodiment;

[0032] Figure 3 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment;

[0033] Figure 4 is a schematic diagram of a simplified block diagram of a communication system according to another embodiment;

[0034] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment;

[0035] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment;

[0036] Figure 7 shows a block diagram of an encoder according to another embodiment;

[0037] Figure 8 shows a block diagram of a decoder according to another embodiment;

[0038] Fig. 9 An example of coding block partitioning according to an embodiment of the present application is shown;

[0039] Fig.10 An example of coding block partitioning according to another embodiment of the present application is shown;

[0040] Fig.11 An example of coding block partitioning according to another embodiment of the present application is shown;

[0041] Fig.12 An example of partitioning a coding block into multiple transform blocks according to an embodiment of the present application is shown;

[0042] Fig.13 An example of a ternary partitioning scheme is shown;

[0043] Fig.14 An example of a quad binary tree coding block partitioning scheme is shown

[0044] Fig.15 An example of partitioning a coding block into multiple transform blocks and a coding order of the transform blocks according to an embodiment of the present application is shown;

[0045] Fig.16 An example of partitioning a coding block into multiple transform blocks and a coding order of the transform blocks according to another embodiment of the present application is shown;

[0046] Fig.17 An example of partitioning a coding block into multiple transform blocks according to another embodiment of the present application is shown;

[0047] Fig.18 A method flow chart according to an embodiment of the present application is shown;

[0048] Fig.19 is a schematic diagram of a computer device according to an embodiment. DETAILED DESCRIPTION

[0049] Throughout the specification and claims, terms may have subtle meanings that are suggested or implied by the context rather than the explicitly stated meaning. The phrase "in one embodiment" or "in some embodiments) as used herein does not necessarily refer to the same embodiment as the phrase used and the phrase "in another embodiment" or "implementation" or "in some implementations" as used herein does not necessarily refer to the same implementation and the phrase "in another implementation" or "in other implementations" as used herein does not necessarily refer to different implementations. For example, the claimed subject matter includes all or part of the combination of the exemplary embodiments / implementations.

[0050] In general, terms may be derived at least in part from the context. For example, as used herein, terms such as "and," "or," and "and / or" may include multiple meanings that may depend at least in part on the context in which the terms are used. In general, "or," if used in a list of associations, such as a, B, or C, means a, B, and C, where used in an inclusive sense, as well as a, B, or C, where used in an exclusive sense. Additionally, as used herein, the terms "at least one" or "at least one," depending at least in part on the context, may be used to describe any feature, structure, or characteristic in a singular sense, or may be used to describe a combination of features, structures, and characteristics in a plural sense. Furthermore, the terms "based on" or "based on" may be understood as not necessarily intended to convey an exclusive set of factors, but rather may allow for the presence of additional factors that are not necessarily explicitly described, again, depending at least in part on the context.

[0051] Reference Figure 1A, a subset of nine known prediction directions from the 33 possible prediction directions of H.265 (corresponding to 33 angular modes in 35 intra-frame prediction modes) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted based on at least one sample at the upper right that is at an angle of 45 degrees to the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted based on at least one sample at the lower left that is at an angle of 22.5 degrees to the horizontal.

[0052] Still reference Figure 1A , a square block (104) including 4×4 samples is shown in the upper left (indicated by the thick dashed line). The square block (104) consists of 16 samples, each of which is marked with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second (from the top) sample in the Y dimension and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the X and Y dimensions. Since the block is a 4×4 size sample, S44 is located in the lower right corner. Reference samples following a similar numbering scheme are also shown. Reference samples are marked with "R" and their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, the prediction samples and blocks are adjacent when reconstructed, so there is no need to use negative values.

[0053] The intra picture prediction of block 104 can be performed by copying reference sample values ​​from neighboring samples occupied by the signaled prediction direction. For example, assume that the encoded video bitstream includes signaling that indicates, for this block, a prediction direction consistent with arrow (102), i.e., predicting samples based on at least one prediction sample to the upper right at an angle of 45 degrees to the horizontal. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then, sample S44 is predicted based on reference sample R08.

[0054] In some cases, for example by interpolation, the values ​​of multiple reference samples may be combined in order to calculate the reference sample, particularly when the direction is not divisible by 45 degrees.

[0055] As video coding technology has developed, the number of directions has gradually increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013) and JEM / VVC / BMS, and at the time of this application, up to 65 directions can be supported. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding can be used to identify these most likely directions with a small number of bits, receiving the loss of some less likely directions. Further, these directions themselves can sometimes be predicted from adjacent directions used by adjacent, decoded blocks.

[0056] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to JEM is shown to illustrate the increasing number of prediction directions over time.

[0057] The mapping of intra prediction direction bits in the coded video bitstream to represent the direction can vary according to different video coding techniques; and can range, for example, from simple direct mapping of prediction directions to intra prediction modes, to codewords, to complex adaptive schemes involving most probable modes, and the like. However, in all cases, there are some directions that are statistically less likely to appear in the video content than others. Since the goal of video compression is to reduce redundancy, in a well-working video codec, these less likely directions are represented using more bits than more likely directions.

[0058] Motion compensation may be a lossy compression technique and may involve a technique where a block of sample data from a previously reconstructed picture or part of a reconstructed picture (reference picture) is used for prediction of a newly reconstructed picture or part of a picture after being spatially shifted in the direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, where the third dimension indicates the reference picture in use (the latter may indirectly be the temporal dimension).

[0059] In some video compression techniques, the MV applied to a region of sample data can be predicted based on other MVs, such as those MVs associated with another region of sample data that is spatially adjacent to the region being reconstructed and that precede the MV in decoding order. Doing so can greatly reduce the amount of data required to encode the MV, thereby eliminating redundant information and increasing the amount of compression. MV prediction can be done effectively, for example, when encoding an input video signal derived from a camera (called natural video), there is a statistical probability that regions larger than the region to which a single MV applies will move in a similar direction, so in some cases, similar motion vectors derived from MVs of neighboring regions can be used for prediction. This results in the MV found for a given region being similar or identical to the MV predicted from the surrounding MVs, and after entropy coding, it can be represented with fewer bits than the number of bits used when the MV is directly encoded. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself may be lossy, for example due to rounding errors when calculating predicted values ​​based on several surrounding MVs.

[0060] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T H.265 Recommendation, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms provided by H.265, this application describes a technique referred to below as "spatial merging".

[0061] Please refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on the previous block of the same size that has generated a spatial offset. In addition, the MV can be derived from metadata associated with one or at least two reference pictures instead of encoding the MV directly. For example, the MV is derived from the metadata of the nearest reference picture (in decoding order) using the MV associated with any of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively). In H.265, MV prediction can use the prediction value of the same reference picture that the neighboring blocks are also using.

[0062] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present application is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3In the example of , a first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of at least one encoded video code stream. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video picture, and display the video picture based on the restored video data. In media service applications and the like, unidirectional data transmission may be common.

[0063] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, such as may occur during a video conference. For bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture on an accessible display device based on the restored video data.

[0064] exist Figure 3 In the example of, terminal devices (310), (320), (330) and (340) can be shown as servers, personal computers and smart phones, but the principles of the present application may not be limited to this. Embodiments of the present application can be applied to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330) and (340), including, for example, wired (wireline / wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of this discussion, the architecture and topology of network (350) may be irrelevant to the operation of the present application, unless explained below.

[0065] As an example, Figure 4The video encoder and video decoder are shown in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0066] The streaming system may include an acquisition subsystem (413), which may include a video source (401), such as a digital camera, for creating an uncompressed video picture stream (402). In an embodiment, the video picture stream (402) includes samples captured by the digital camera of the video source (401). Compared to the encoded video data (404) (or the encoded video bitstream), the video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (402) can be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (402), the encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (404) (or the encoded video bitstream (404)), which can be stored on the streaming server (405) for future use. At least one streaming client subsystem, such as Figure 3 The client subsystem (406) and the client subsystem (408) in the streaming server (405) can access the streaming server (405) to retrieve the copy (407) and the copy (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in the electronic device (330). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), the video data (407), and the video data (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application can be used in the context of the VVC standard.

[0067] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0068] Figure 5 is a block diagram of a video decoder (510) according to an embodiment disclosed in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) of an embodiment.

[0069] The receiver (531) may receive at least one encoded video sequence to be decoded by the video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be provided external to the video decoder (510) (not shown). In other cases, a buffer memory (not shown) is provided outside the video decoder (510) to prevent network jitter, for example, and another buffer memory (515) may be configured inside the video decoder (510) to handle broadcast timing, for example. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. Of course, in order to use on a service packet network such as the Internet, a buffer memory (515) may also be required, and the buffer memory may be relatively large and have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not shown) outside the video decoder (510).

[0070] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The types of symbols include information used to manage the operation of the video decoder (510) and potential information used to control a display device such as a display device (512) (e.g., a display screen) that is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5 As shown in . The control information for the display device may be a parameter set fragment (not indicated) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (520) may extract a subgroup parameter set of at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0071] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).

[0072] Depending on the type of coded video picture or portion of coded video picture (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (520) and the multiple units below is not described.

[0073] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0074] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients as symbols (521) from the parser (520) and control information including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).

[0075] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from a current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0076] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (521), these samples may be added to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signals) by the aggregator (555) to generate output sample information. The acquisition of the predicted samples by the motion compensated prediction unit (553) from the address in the reference picture memory (557) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (553) in the form of the symbols (521), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0077] The output samples of the aggregator (555) may be employed by various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, as well as to previously reconstructed and loop filtered sample values.

[0078] The output of the loop filter unit (556) may be a sample stream that may be output to a display device (512) and stored in a reference picture memory (557) for subsequent inter-picture prediction.

[0079] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0080] The video decoder (510) may perform decoding operations according to a predetermined video compression technique, such as in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0081] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0082] Figure 6 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is provided in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace Figure 4 A video encoder (403) in an embodiment.

[0083] The video encoder (603) can be used to obtain the video source (601) (not Figure 6 In another embodiment, the video source (601) is a part of the electronic device (620) to receive video samples, and the video source can collect video images to be encoded by the video encoder (603). In another embodiment, the video source (601) is a part of the electronic device (620).

[0084] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0085] According to an embodiment, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, coupling is not indicated in the figure. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be used to have other suitable functions that are related to the video encoder (603) optimized for a certain system design.

[0086] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in embodiments, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression technology considered in this application, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, eg due to channel errors) is also used in some related techniques.

[0087] The operation of the "local" decoder (633) can be combined with the above Figure 5 The "remote" decoder described in detail for the video decoder (510) is identical. However, additional brief reference is made to Figure 5 , when symbols are available and the entropy encoder (645) and parser (520) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).

[0088] At this point, it can be observed that any decoder technology except the parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. For this reason, the application focuses on decoder operation. The description of encoder technology can be simplified because encoder technology is mutually reverse to the decoder technology described comprehensively. Only a more detailed description is needed in certain areas and is provided below.

[0089] During operation, in some embodiments, the source encoder (630) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to at least one previously encoded picture from a video sequence designated as a "reference picture." In this manner, the encoding engine (632) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a prediction reference for the input picture.

[0090] The local video decoder (633) may decode the encoded video data that may be designated as a reference picture based on the symbol created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 6 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0091] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (635) may perform operations on a sample block-by-pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).

[0092] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0093] The outputs of all the above functional units may be entropy encoded in an entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0094] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0095] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0096] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.

[0097] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0098] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0099] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block-wise. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0100] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0101] In an embodiment, the transmitter (640) may transmit additional data when transmitting the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0102] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0103] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.

[0104] In addition, merge mode technology can be used in inter-picture prediction to improve encoding and decoding efficiency.

[0105] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Furthermore, each CTU can be split into at least one coding unit (CU) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0106] Figure 7 is a diagram of a video encoder (703) according to another embodiment disclosed in the present application. The video encoder (703) is used to receive a processed block (e.g., a prediction block) of sample values ​​in a current video picture in a video picture sequence, and encode the processed block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used to replace Figure 4A video encoder (403) in an embodiment.

[0107] In an HEVC embodiment, a video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use intra mode, inter mode, or bidirectional prediction mode to encode the processing block. When encoding the processing block in intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when encoding the processing block in inter mode or bidirectional prediction mode, the video encoder (703) may use inter prediction or bidirectional prediction techniques to encode the processing block into an encoded picture, respectively. In some video encoding techniques, the merge mode may be an inter-picture prediction submode, in which a motion vector is derived from at least one motion vector predictor without the aid of an encoded motion vector component external to the predictor. In some other video encoding techniques, there may be a motion vector component applicable to the subject block. In an embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0108] exist Figure 7 In an embodiment of the present invention, the video encoder (703) includes: Figure 7 Shown are an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together.

[0109] The inter-frame encoder (730) is used to receive samples of a current block (e.g., a processing block), compare the block with at least one reference block in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., redundant information description according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0110] The intra-frame encoder (722) is used to receive samples of a current block (e.g., a processing block), compare the block with an encoded block in the same picture in some cases, generate quantization coefficients after transformation, and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to at least one intra-frame coding technique). In an embodiment, the intra-frame encoder (722) also calculates an intra-frame prediction result (e.g., a predicted block) based on the intra-frame prediction information and a reference block in the same picture.

[0111] The general controller (721) is used to determine the general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the mode is inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.

[0112] The residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra-frame encoder (722) or the inter-frame encoder (730). The residual encoder (724) is used to operate based on the residual data to encode the residual data to generate a transform coefficient. In an embodiment, the residual encoder (724) is used to convert the residual data from the time domain to the frequency domain and generate a transform coefficient. The transform coefficient is then processed by quantization to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra-frame encoder (722) and the inter-frame encoder (730). For example, the inter-frame encoder (730) can generate a decoded block based on the decoded residual data and the inter-frame prediction information, and the intra-frame encoder (722) can generate a decoded block based on the decoded residual data and the intra-frame prediction information. The decoded blocks are appropriately processed to generate a decoded picture, and in some embodiments, the decoded picture may be buffered in a memory circuit (not shown) and used as a reference picture.

[0113] The entropy encoder (725) is used to format the code stream to produce an encoded block. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (such as intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the code stream. It should be noted that according to the disclosed subject matter, when the block is encoded in the inter-frame mode or the merge sub-mode of the bidirectional prediction mode, there is no residual information.

[0114] Figure 8FIG. 8 is a diagram of a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is used to replace Figure 3 A video decoder (410) of an embodiment.

[0115] exist Figure 8 In one embodiment, the video decoder (810) includes: Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874) and an intra-frame decoder (872) coupled together are shown in FIG.

[0116] The entropy decoder (871) may be used to reconstruct certain symbols from the encoded picture, which represent syntax elements constituting the encoded picture. Such symbols may include, for example, a mode for encoding the block (e.g., intra mode, inter mode, bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that may identify certain samples or metadata for prediction by the intra decoder (872) or the inter decoder (880), respectively, residual information in the form of, for example, quantized transform coefficients, and the like. In an embodiment, when the prediction mode is inter or bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be inversely quantized and provided to the residual decoder (873).

[0117] The inter-frame decoder (880) is used to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0118] The intra-frame decoder (872) is used to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.

[0119] The residual decoder (873) is used to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require some control information (to obtain the quantizer parameter QP), and this information can be provided by the entropy decoder (871) (the data path is not indicated because this is only low-volume control information).

[0120] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which in turn can be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0121] It should be noted that the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using any suitable technology. In an embodiment, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using at least one integrated circuit. In another embodiment, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using at least one processor that executes software instructions.

[0122] Let's look at the block partitions for encoding and decoding. Generally, partitions can start from basic blocks and can follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitions can be hierarchical and recursive. After following any example partitioning procedure or other procedures described below or a combination thereof to divide or partition the basic block, a final set of partitions or coding blocks can be obtained. Each of these partitions can be in one of the partitioning levels in the partition hierarchy and can have various shapes. Each partition can be called a coding block (CB). For various example partition implementations further described below, each CB obtained can be of any allowed size and partition level. Such partitions are called coding blocks because they can form units for which some basic encoding / decoding decisions can be made, and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video code stream. The highest or deepest level in the final partition represents the depth of the coding block partition structure of the tree. The coding block can be a luminance coding block or a chrominance coding block. The CB tree structure for each color may be referred to as a coded block tree (CBT).

[0123] The coding blocks of all color channels may be collectively referred to as coding units (CUs). The hierarchical structure of all color channels may be collectively referred to as coding tree units (CTUs). The partitioning modes or structures of various color channels in a CTU may be the same or different.

[0124] In some implementations, the partition tree schemes or structures for the luma channel and the chroma channels may not necessarily be the same. In other words, the luma channel and the chroma channels may have separate coding tree structures or modes. Further, whether the luma channel and the chroma channels use the same or different coding partition tree structures and the actual coding partition tree structures to be used may depend on whether the stripe being encoded is a P, B, or I stripe. For example, for an I stripe, the chroma channel and the luma channel may have separate coding partition tree structures or coding partition tree structure modes, while for a P or B stripe, the luma channel and the chroma channels may share the same coding partition tree scheme. When a separate coding partition tree structure or mode is applied, the luma channel may be partitioned into CBs by one coding partition tree structure, and the chroma channel may be partitioned into chroma CBs by another coding partition tree structure.

[0125] In some example implementations, a predetermined partitioning scheme may be applied to basic blocks. Fig. 9 As shown, the example 4-way partition tree can start from a first predefined level (e.g., 64×64 block level or other size, as a basic block size), and the basic block can be partitioned hierarchically down to a predefined lowest level (e.g., 4×4 level). For example, the basic block can be subject to four predefined partitioning options or modes indicated by 902, 904, 906 and 908, where the partition designated as R is allowed for recursive partitioning because it can be repeated at a lower ratio such as Fig. 9 In some implementations, the same partitioning options as indicated in , up to the lowest level (e.g., 4×4 level). Fig. 9 Additional restrictions apply to partitioning schemes. Fig. 9 In the implementation of , rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they may not be recursive, whereas square partitions may be allowed to be recursive. If necessary, follow Fig. 9 The recursive partitioning of generates the final set of coding blocks. The coding tree depth can be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64×64 block) can be set to 0, and in accordance with Fig. 9 After further splitting the root block once, the coding tree depth increases by 1. For the above scheme, the maximum or deepest level of the minimum partition from the 64×64 basic block to 4×4 will be 4 (starting from level 0). This partitioning scheme can be applied to at least one color channel. It can be followed Fig. 9The scheme partitions each color channel independently (e.g., a partitioning pattern or option in a predefined pattern may be determined independently for each of the color channels at each hierarchical level). Optionally, two or more color channels may share Fig. 9 The same hierarchical pattern tree (for example, the same partitioning pattern or option in the predefined patterns can be selected for two or more color channels at each hierarchical level).

[0126] Fig.10 Another example predefined partitioning mode that allows recursive partitioning to form a partition tree is shown. Fig.10 As shown in , an example 10-way partition structure or pattern may be predefined. A root block may start at a predefined level (eg, from a basic block at a 128x128 level or a 64x64 level). Fig.10 Example partition structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Fig.10 The partition type with three sub-partitions indicated as 1002, 1004, 1006, and 1008 in the second row of the table may be referred to as a "T-type" partition. The "T-type" partitions 1002, 1004, 1006, and 1008 may be referred to as a left T-type, a top T-type, a right T-type, and a bottom T-type. In some example implementations, further subdivision is not allowed. Fig.10 The coding tree depth can be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and in accordance with Fig.10 After further splitting the root block once, the coding tree depth increases by 1. In some implementations, all square partitions in 1010 may only be allowed to follow Fig.10 In other words, for the square partitions within the T-patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If necessary, follow Fig.10 The recursive partitioning procedure of generates the final set of coding blocks. This scheme can be applied to at least one color channel. In some implementations, more flexibility can be added to the use of partitions below the 8×8 level. For example, 2×2 chroma inter-frame prediction can be used in some cases.

[0127] In some other example implementations for coding block partitioning, a quadtree structure can be used to partition a basic block or an intermediate block into quadtree partitions. This quadtree partitioning can be applied hierarchically and recursively to any square partition. Whether a basic block or an intermediate block / partition is further quadtree partitioned can be adapted to various local characteristics of the basic block or intermediate block / partition. The quadtree partitions at the picture boundaries can be further adjusted. For example, implicit quadtree partitioning can be performed at the picture boundaries so that the block will remain quadtree partitioned until the size fits the picture boundary.

[0128] In some other example implementations, a hierarchical binary partitioning from a basic block can be used. For this scheme, a basic block or an intermediate block can be partitioned into two partitions. The binary partitioning can be horizontal or vertical. For example, horizontal binary partitioning can split a basic block or an intermediate block into equal right and left partitions. Similarly, vertical binary partitioning can split a basic block or an intermediate block into equal upper and lower partitions. This binary partitioning can be hierarchical and recursive. For each of the basic block or the intermediate block, a decision can be made as to whether the binary partitioning scheme should continue, and if the scheme continues further, whether horizontal or vertical binary partitioning should be used. In some implementations, further partitioning can stop at a predefined minimum partition size (in one or two dimensions). Optionally, once a predefined partitioning level or depth starting from a basic block is reached, further partitioning can stop. In some implementations, the aspect ratio of the partition can be limited. For example, the aspect ratio of the partition can be no less than 1:4 (or greater than 4:1). As such, a vertical stripe partition having a vertical to horizontal aspect ratio of 4:1 may only be further vertically binary partitioned into upper and lower partitions each having a vertical to horizontal aspect ratio of 2:1.

[0129] In some other examples, such as Fig.13 As shown in , the ternary partitioning scheme can be used to partition the basic block or any intermediate block. The ternary pattern can be Fig.13 1302 is implemented vertically, or as shown in Fig.13 1304 is shown horizontally. Although Fig.13 The example partition ratio in , vertically or horizontally, is shown as 1:2:1, but other ratios can be predefined. In some implementations, two or more different ratios can be predefined. This ternary partitioning scheme can be used to supplement the quadtree or binary partitioning structure, because this ternary tree partitioning is able to capture objects located at the center of a block in one continuous partition, while the quadtree and binary tree always partition along the block center and thus partition the object into separate partitions. In some implementations, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transformations.

[0130] The above partitioning schemes can be combined in any manner at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition basic blocks into a quadtree-binary tree (QTBT) structure. In this scheme, a basic block or intermediate block / partition can be either quadtree partitioned or binary partitioned, if specified, subject to a predefined set of conditions. Fig.14 A specific example is illustrated in . Fig.14 In the example of , the basic block is first quadtree partitioned into four partitions, as shown in 1402, 1404, 1406 and 1408. Thereafter, each of the resulting partitions is either quadtree partitioned into four additional partitions (such as 1408), or binary partitioned into two additional partitions (horizontally or vertically, such as 1402 or 1406, for example, both are symmetrical) at the next level, or not partitioned (such as 1404). For square partitions, binary or quadtree partitioning can be recursively allowed, as shown in the overall example partitioning pattern of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag can be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partitioning is horizontal or vertical. For example, as shown in 1420, consistent with the partitioning structure of 1410, flag "0" can represent horizontal binary partitioning, and flag "1" can represent vertical binary partitioning. For quadtree partitioning, there is no need to indicate the partition type, because quadtree partitioning always partitions a block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some implementations, a flag "1" may indicate horizontal binary partitioning, and a flag "0" may indicate vertical binary partitioning.

[0131] In some example implementations of QTBT, the quadtree and binary segmentation rule set may be represented by the following predefined parameters and corresponding functions associated with them:

[0132] CTU size: the root node size of the quadtree (the size of the basic block)

[0133] MinQTSize: The minimum allowed quadtree leaf node size

[0134] MaxBTSize: Maximum allowed binary tree root node size

[0135] MaxBTDepth: Maximum allowed binary tree depth

[0136] MinBTSize: minimum allowed binary tree leaf node size

[0137] In some example implementations of the QTBT partition structure, the CTU size can be set to 128×128 luma samples with two corresponding 64×64 chroma sample blocks (when example chroma subsampling is considered and used), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes from their minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the node is 128×128, it will not be split by the binary tree first because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes not exceeding MaxBTSize can be partitioned by the binary tree. Fig.14 In the example, the basic block is 128×128. According to a predefined set of rules, the basic block can only be split by a quadtree. The partition depth of the basic block is 0. Each of the four resulting partitions is 64×64, not exceeding MaxBTSize, and can be further split by a quadtree or a binary tree at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splits can be disregarded. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal splits can be disregarded. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical splits are disregarded.

[0138] In some example implementations, the above QTBT scheme can be configured to provide flexibility for luma and chroma to have the same QTBT structure or separate QTBT structures. For example, for P and B slices, the luma and chroma CTBs in one CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CBs by the QTBT structure, and the chroma CTB can be partitioned into chroma CBs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice, for example, an I slice can consist of coding blocks of a luma component or coding blocks of two chroma components, and a CU in a P or B slice can consist of coding blocks of all three color components.

[0139] In some other implementations, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. Such an implementation can be referred to as a multi-type tree (MTT) structure. For example, in addition to the binary partitioning of the nodes, one can choose Fig.13One of the ternary partitioning modes. In some implementations, only square nodes can be subjected to ternary partitioning. An additional flag can be used to indicate whether the ternary partitioning is horizontal or vertical.

[0140] The design of two-level or multi-level trees such as the QTBT implementation and the QTBT implementation complemented by ternary partitioning can be motivated primarily by reducing complexity. In theory, the complexity of traversing the tree is T D , where T represents the number of split types and D is the depth of the tree. A trade-off can be made by using more types (T) while reducing the depth (D).

[0141] In some implementations, the CB can be further partitioned. For example, the CB can be further partitioned into multiple prediction blocks (PBs) for intra-frame or inter-frame prediction during encoding and decoding. In other words, the CB can be further divided into different sub-partitions, in which separate prediction decisions / configurations can be made. At the same time, for the purpose of depicting the level of transformation or inverse transformation of video data, the CB can be further partitioned into multiple transform blocks (TBs). The partitioning schemes of CB to PB and TB can be the same or different. For example, each partitioning scheme can be performed using its own program based on various characteristics of video data, for example. In some example implementations, the PB and TB partitioning schemes can be independent. In some other example implementations, the PB and TB partitioning schemes and boundaries can be related. In some implementations, for example, the TB can be partitioned after the PB partition, and specifically, each PB (after being determined after the partition of the coding block) can then be further partitioned into at least one TB. For example, in some implementations, the PB can be divided into one, two, four or other number of TBs.

[0142] In some implementations, in order to partition the basic block into coding blocks and further into prediction blocks and / or transform blocks, the luma channel and the chroma channel may be processed differently. For example, in some implementations, the partitioning of the coding block into prediction blocks and / or transform blocks may be allowed for the luma channel, while such partitioning of the coding block into prediction blocks and / or transform blocks may not be allowed for at least one chroma channel. In such an implementation, the transformation and / or prediction of the luma block can therefore be performed only at the coding block level. For another example, the minimum transform block size of the luma channel and at least one chroma channel may be different, for example, the partitioning of the coding block of the luma channel into transform blocks and / or prediction blocks that are smaller than the chroma channel may be allowed. For yet another example, the maximum depth of partitioning the coding block into transform blocks and / or prediction blocks may be different between the luma channel and the chroma channel, for example, the partitioning of the coding block of the luma channel into transform blocks and / or prediction blocks that are deeper than at least one chroma channel may be allowed. For a specific example, the luma coding block may be partitioned into transform blocks of multiple sizes, which may be represented by recursive partitioning down to up to 2 levels, and may allow transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes from 4×4 to 64×64. However, for chroma blocks, only the largest possible transform block specified for the luma block is allowed.

[0143] In some example implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partitions may depend on whether the PB is intra-coded or inter-coded.

[0144] Partitioning a coding block (or prediction block) into transform blocks may be implemented in various example schemes, including but not limited to recursively or non-recursively quadtree partitioning and predefined pattern partitioning, and additionally considering transform blocks at the boundaries of the coding block or prediction block. In general, the resulting transform blocks may be at different partitioning levels, may not be of the same size, and may not need to be square in shape (e.g., they may be rectangular with some allowed size and aspect ratio). The following is combined with Fig.15 , Fig.16 and Fig.17 Other examples are described in further detail.

[0145] However, in some other implementations, the CB obtained via any of the above partitioning schemes can be used as a basic or minimum coding block for prediction and / or transformation. In other words, no further segmentation is performed for the purpose of performing inter-frame prediction / intra-frame prediction and / or for the purpose of transformation. For example, the CB obtained from the above QTBT scheme can be used directly as a unit for performing prediction. Specifically, this QTBT structure removes the concept of multiple partition types, that is, it removes the separation of CU, PU and TU, and provides more flexibility for the above-mentioned CU / CB partition shape. In this QTBT block structure, the CU can have a square or rectangular shape. The leaf nodes of this QTBT are used as units of prediction and transformation processing without any further partitioning. This means that CU, PU and TU have the same block size in this example QTBT coding block structure.

[0146] The above various CB partitioning schemes and further partitioning of CB into PB and / or TB (excluding PB / TB partitioning) can be combined in any manner. The following specific implementations are provided as non-limiting examples.

[0147] A specific example implementation of coding block and transform block partitioning is described below. In this example implementation, recursive quadtree partitioning or the above-mentioned predefined partitioning patterns (such as Fig. 9 and Fig.10 The basic block is divided into coding blocks according to the mode in (in). At each level, whether the further quadtree segmentation of a specific partition should continue can be determined by the local video data characteristics. The resulting CB can be at various quadtree segmentation levels and have various sizes. A decision can be made at the CB level (or CU level, for all three color channels) about whether to use inter-frame picture (time) prediction or intra-frame picture (spatial) prediction to encode the picture area. Each CB can be further divided into one, two, four or other number of PBs according to the predefined PB segmentation type. Within a PB, the same prediction process can be applied, and the relevant information can be transmitted to the decoder on the basis of the PB. After obtaining the residual block by applying the prediction process based on the PB segmentation type, the CB can be partitioned into TBs according to another quadtree structure of the coding tree similar to the CB. In this specific implementation, the CB or TB can be, but not necessarily limited to, a square shape. Further, in this specific example, the PB can be a square or rectangular shape for inter-frame prediction and can be only a square for intra-frame prediction. The coding block can be divided into, for example, four square-shaped TBs. Each TB can be further recursively partitioned (using quadtree partitioning) into smaller TBs, referred to as residual quadtrees (RQTs).

[0148] Another example implementation for partitioning a basic block into CBs, PBs, and / or TBs is described further below. For example, a quadtree with nested multi-type trees based on a binary and ternary partitioning segmentation structure (e.g., QTBT or QTBT with ternary partitioning as described above) may be used instead of using a quadtree such as Fig. 9 or Fig.10 The multi-partition unit type shown. The separation of CB, PB and TB (i.e., partitioning CB into PB and / or TB, and partitioning PB into TB) can be abandoned, unless a CB with a size that is too large for the maximum transform length is required, in which case further segmentation of such CB may be required. The example partitioning scheme can be designed to provide more flexibility for the CB partition shape so that both prediction and transformation can be performed at the CB level without further partitioning. In this coding tree structure, the CB can have a square or rectangular shape. Specifically, the coding tree block (CTB) can be partitioned first by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by nesting a multi-type tree structure. Fig.11 An example of a nested multi-type tree structure using binary or ternary segmentation is shown in FIG. Fig.11 The example multi-type tree structure includes four types of segmentation, referred to as vertical binary segmentation (SPLIT_BT_VER) (1102), horizontal binary segmentation (SPLIT_BT_HOR) (1104), vertical ternary segmentation (SPLIT_TT_VER) (1106), and horizontal ternary segmentation (SPLIT_TT_HOR) (1108). The CB then corresponds to the leaves of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, the segment is used for prediction and transform processing without any further partitioning. This means that in most cases, CB, PB, and TB have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CB. In some implementations, in addition to binary or ternary segmentation, Fig.11 The nested pattern can further include quadtree partitioning.

[0149] Fig.12 A specific example of a quadtree with a nested multi-type tree coding block structure (including quadtree, binary and ternary partitioning options) for a basic block is shown. In more detail, Fig.12 The basic block 1200 is shown to be quad-tree partitioned into four square partitions 1202, 1204, 1206 and 1208. Further use is made for each quad-tree partition Fig.11 The multi-type tree structure and the decision of the quadtree for further segmentation. Fig.12In the example of , partition 1204 is not further partitioned. Partitions 1202 and 1208 each use another quadtree partition. For partition 1202, the upper left, upper right, lower left, and lower right partitions of the second level quadtree partitioning use the third level of quadtree partitioning, Fig.11 Horizontal binary segmentation 1104, non-segmentation and Fig.11 The horizontal ternary partition 1108 is divided into two parts. Partition 1208 adopts another quadtree partition, and the upper left, upper right, lower left and lower right partitions of the second level quadtree partition adopt Fig.11 The third level of vertical ternary segmentation 1106 is segmentation, non-segmentation, non-segmentation and Fig.11 Horizontal binary segmentation 1104. Fig.11 The horizontal binary segmentation 1104 and the horizontal ternary segmentation 1108 further segment the two sub-partitions of the third-level upper left partition 1208. Fig.11 The second level segmentation mode of the vertical binary segmentation 1102 divides the partition 1206 into two partitions according to Fig.11 The two partitions are further segmented at the third level by using horizontal ternary segmentation 1108 and vertical binary segmentation 1102. Fig.11 The horizontal binary segmentation 1104 further applies the fourth level segmentation to one of them.

[0150] For the specific example above, the maximum luma transform size may be 64×64, and the maximum supported chroma transform size may be different from luma at, for example, 32×32. Fig.12 The example CB in the example is further split into smaller PBs and / or TBs. When the width or height of the luminance coding block or the chrominance coding block is larger than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically split in the horizontal and / or vertical directions to meet the transform size limit in that direction.

[0151] In a specific example for partitioning a basic block into the above CBs, and as described above, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P and B slices, the luma and chroma CTBs in one CTU can share the same coding tree structure. For example, for I slices, luma and chroma can have separate coding block tree structures. When a separate block tree structure is applied, the luma CTB can be partitioned into luma CBs by one coding tree structure, and the chroma CTB can be partitioned into chroma CBs by another coding tree structure. This means that a CU in an I slice can consist of coding blocks of a luma component or coding blocks of two chroma components, and a CU in a P slice or a B slice always consists of coding blocks of all three color components unless the video is monochrome.

[0152] When the coding block is further partitioned into multiple transform blocks, the transform blocks therein can be ordered in the bitstream in various orders or scanning modes. The following is further described in detail an example implementation for partitioning a coding block or prediction block into transform blocks and the encoding order of the transform blocks. In some example implementations, as described above, transform partitioning can support transform blocks of multiple shapes (e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1), with transform block sizes ranging from, for example, 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, the transform block partitioning can be applied only to the luma component, so that for the chroma block, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, the luma coding block and the chroma coding block can be implicitly split into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.

[0153] In some example implementations of transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coded blocks may be further partitioned into multiple transform blocks with a partition depth up to a predetermined number of levels (e.g., 2 levels). The transform block partition depth and size may be related. For some example implementations, the mapping from the transform size of the current depth to the transform size of the next depth is shown in Table 1 below.

[0154] Table 1: Change partition size settings

[0155]

[0156] Based on the example mapping of Table 1, for a 1:1 square block, the next level transform partitioning can create four 1:1 square sub-transform blocks. The transform partitioning can stop at 4×4, for example. In this way, the transform size 4×4 at the current depth corresponds to the same size 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level transform partitioning can create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next level transform partitioning can create two 1:2 / 2:1 sub-transform blocks.

[0157] In some example implementations, for the luma component of an intra-coded block, additional restrictions may be applied with respect to transform block partitioning. For example, for each level of transform partitioning, all sub-transform blocks may be constrained to have equal size. For example, for a 32×16 coding block, a level 1 transform split creates two 16×16 sub-transform blocks and a level 2 transform split creates eight 8×8 sub-transform blocks. In other words, the second level split must be applied to all first level sub-blocks to keep the transform unit sizes equal. Fig.151506 shows an example of transform block partitioning for intra-coded square blocks of Table 1 below, and the coding order illustrated by arrows. Specifically, 1502 shows a square coding block. It has a coding order indicated by arrows. In 1506, it is shown that all first-level equal-sized blocks are divided into 16 equal-sized transform blocks in the second level according to Table 1, and the coding order is indicated by arrows.

[0158] In some example implementations, the above restrictions on intra-coding may not apply to the luma component of an inter-coded block. For example, after the first level of transform partitioning, any of the sub-transform blocks may be further partitioned one level independently. Thus, the resulting transform blocks may or may not have the same size. Fig.16 An example of segmenting an inter-frame coded block into transform locks with their coding order is shown in FIG. Fig.16 In the example of FIG. 1 , the inter-frame coded block 1602 is divided into two levels of transform blocks according to Table 1. At the first level, the inter-frame coded block is divided into four transform blocks of equal size. Then, as shown in 1604, only one of the four transform blocks (not all) is further divided into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes. The example coding order of these 7 transform blocks is as follows: Fig.16 As shown by the arrow in 1604.

[0159] In some example implementations, for at least one chroma component, some additional restrictions on transform blocks may be applied. For example, for at least one chroma component, the transform block size may be as large as the coding block size, but not smaller than a predefined size, such as 8×8.

[0160] In some other example implementations, for coding blocks with width (W) or height (H) greater than 64, the luma coding blocks and chroma coding blocks may be implicitly split into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively. Here, in the present disclosure, "min(a,b)" may return the smaller value between a and b.

[0161] Fig.17 Another optional example scheme for partitioning a coding block or prediction block into transform blocks is further shown. Fig.17 As shown in , depending on the transform type of the coding block, a predefined set of partition types can be applied to the coding block without using recursive transform partitioning. Fig.17 In the specific example shown in , one of the six example partition types can be applied to split a coding block into various numbers of transform blocks. This scheme for generating transform block partitions can be applied to either a coding block or a prediction block.

[0162] In more detail, Fig.17The partition scheme of provides up to 6 example partition types for any given transform type (transform type refers to, for example, a primary transform, such as ADST and other types). In this scheme, each coding block or prediction block can be assigned a transform partition type based on, for example, rate-distortion cost. In an example, the transform partition type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. Fig.17 As shown in the 6 transform partition types shown in the figure, a specific transform partition type can correspond to a transform block partition size and mode. The correspondence between various transform types and various transform partition types can be predefined. The following shows an example with uppercase marks, which indicate the transform partition type that can be assigned to a coding block or a prediction block based on the rate-distortion cost:

[0163] PARTITION_NONE: Allocate a transform size that is equal to the block size.

[0164] PARTITION_SPLIT: Allocate a transform size whose width is 1 / 2 of the block size and whose height is 1 / 2 of the block size.

[0165] PARTITION_HORZ: Allocate a transform size whose width is the same as the block size and whose height is 1 / 2 of the block size.

[0166] PARTITION_VERT: Allocate a transform size whose width is 1 / 2 of the block size and whose height is the same as the block size.

[0167] PARTITION_HORZ4: Allocates a transform size whose width is the same as the block size and whose height is 1 / 4 of the block size.

[0168] PARTITION_VERT4: Allocates a transform size whose width is 1 / 4 of the block size and whose height is the same as the block size.

[0169] In the above example, if Fig.17 The transform partition types shown all include a uniform transform size for the transform blocks of the partition. This is merely an example and not a limitation. In some other implementations, mixed transform block sizes may be used for the transform blocks of the partitions in a particular partition type (or mode).

[0170] The PB (or CB, also referred to as PB when not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can then become individual blocks for encoding via intra-frame prediction or inter-frame prediction. For inter-frame prediction for the current PB, the residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.

[0171] Inter prediction may be implemented, for example, in a single reference mode or a composite reference mode. In some implementations, a skip flag may first be included in the code stream for the current block (or at a higher level) to indicate whether the current block is inter-coded and not skipped. If the current block is inter-coded, another flag may also be included in the code stream as a signal to indicate whether a single reference mode or a composite reference mode is used for prediction of the current block. For a single reference mode, one reference block may be used to generate a prediction block for the current block. For a composite reference mode, two or more reference blocks may be used to generate a prediction block, for example, by weighted averaging. A composite reference mode may be referred to as more than one reference mode, two reference modes, or multiple reference modes. At least one reference block may be identified using a reference frame index or multiple reference frame indexes, and additionally using a corresponding motion vector or multiple motion vectors indicating at least one displacement in position (e.g., in horizontal and vertical pixels) between at least one reference block and the current block. For example, an inter-frame prediction block for the current block may be generated from a single reference block identified as a prediction block in a single reference mode by a motion vector in a reference frame, while for a composite reference mode, a prediction block may be generated by a weighted average of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. At least one motion vector may be encoded and included in a bitstream in various ways.

[0172] In some implementations, the encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be maintained in the DPB waiting to be displayed (in the decoding system), and some images / pictures in the DPB may be used as reference frames to implement inter-frame prediction (in the decoding system or the encoding system). In some implementations, the reference frames in the DPB may be marked as short-term references or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame for inter-frame prediction of a block in the current frame, or a frame for inter-frame prediction of a block in a predefined number (e.g., 2) of subsequent video frames closest to the current frame in decoding order. Long-term reference frames may include frames in the DPB that can be used to predict image blocks in frames that are more than a predefined number away from the current frame in decoding order. Information about such markings for short-term and long-term reference frames may be referred to as a reference picture set (RPS) and may be added to the header of each frame in the encoded code stream. Each frame in the encoded video stream can be identified by a picture order count (POC), which is numbered in an absolute manner according to the playback sequence or relative to a group of pictures, for example starting from an I frame.

[0173] In some example implementations, at least one reference picture list may be formed based on the information in the RPS, which includes identification of short-term and long-term reference frames for inter-frame prediction. For example, a single picture reference list may be formed for unidirectional inter-frame prediction, denoted as L0 reference (or reference list 0), while two picture reference lists may be formed for bidirectional inter-frame prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1), for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. When multiple references used to generate a prediction block by weighted averaging in a composite prediction mode are on the same side of the block to be predicted, the unidirectional inter-frame prediction may be in a single reference mode or in a composite reference mode. Bidirectional inter-frame prediction may be the only composite mode because bidirectional inter-frame prediction involves at least two reference blocks.

[0174] In some implementations, a merge mode (MM) for inter-frame prediction may be implemented. In general, for merge mode, a motion vector in a single reference prediction or at least one motion vector in a composite reference prediction for a current PB may be derived from at least one other motion vector, rather than being independently calculated and signaled. For example, in an encoding system, at least one current motion vector for a current PB may be represented by at least one difference between at least one current motion vector and at least one other encoded motion vector (referred to as a reference motion vector). Such at least one difference in at least one motion vector, rather than the entire at least one current motion vector, may be encoded and included in a bitstream and may be linked to at least one reference motion vector. Accordingly, in a decoding system, at least one motion vector corresponding to the current PB may be derived based on at least one decoded motion vector difference and at least one decoded reference motion vector linked to them. As a specific form of general merge mode (MM) inter-frame prediction, such inter-frame prediction based on at least one motion vector difference may be referred to as merge mode with motion vector difference (MMVD). Therefore, MM may be implemented in general, or MMVD may be implemented in particular, to improve coding efficiency by exploiting the correlation between motion vectors associated with different PBs. For example, neighboring PBs may have similar motion vectors, and thus the MVD may be small and may be efficiently encoded.For another example, for similarly positioned / located blocks in space, the motion vectors may be correlated in time (between frames).

[0175] In some example implementations, an MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Additionally, or optionally, an MMVD flag may be included in the bitstream and signaled during the encoding process to indicate whether the current PB is in MMVD mode. The MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. For a specific example, both the MM flag and the MMVD flag may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether the MMVD mode is for the current CU.

[0176] In some example implementations of MMVD, a merge candidate list for motion vector prediction may be formed for the block being predicted. The merge candidate list may contain a predetermined number (e.g., 2) of MV predictor candidate blocks whose motion vectors may be used to predict the current motion vector. The MVD candidate blocks may include blocks selected from adjacent blocks in the same frame and / or temporal blocks (e.g., blocks at the same position in the previous or next frame of the current frame). These options represent blocks that may have similar or identical motion vectors as the current block at a spatial or temporal position relative to the current block. The size of the MV predictor candidate list may be predetermined. For example, the list may contain two candidates. In order to be on the merge candidate list, for example, the candidate block may need to have the same reference frame (or frames) as the current block, must exist (e.g., when the current block is close to the edge of the frame, a boundary check needs to be performed), and must have been encoded during the encoding process, and / or decoded during the decoding process. In some implementations, the merge candidate list may first be filled with spatially adjacent blocks (scanned in a specific predefined order) (if available and the above conditions are met), and then filled with temporal blocks (if space is still available in the list). For example, the adjacent candidate blocks may be selected from the left block and the top block of the current block. The merged MV prediction value candidate list may be signaled in the bitstream.

[0177] In some implementations, the actual merge candidate may be signaled to be used as a reference motion vector to predict the motion vector of the current block. In the case where the merge candidate list contains two candidates, a one-bit flag called a merge candidate flag may be used to indicate the selection of the reference merge candidate. For the current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list.

[0178] In some example implementations of MMVD, after a merge candidate is selected and used as a base motion vector predictor for a motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) may be calculated in the encoding system. Such an MVD may include information representing the magnitude of the MV difference and the direction of the MV difference, both of which may be signaled in the bitstream. The motion difference magnitude and the motion difference direction may be signaled in various ways.

[0179] In some example implementations of the MMVD, a distance index may be used to specify magnitude information of a motion vector difference and indicate one of a predefined set of offsets representing a predefined motion vector difference from a starting point (reference motion vector). The MV offset according to the signaled index may then be added to the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset may be determined by the direction information of the MVD. An example predefined relationship between the distance index and the predefined offset is specified in Table 2.

[0180] Table 2 - Example relationship between distance index and predefined MV offsets

[0181]

[0182] In some example implementations of MMVD, a direction index may be further signaled and used to indicate the direction of the MVD relative to a reference motion vector. In some implementations, the direction may be limited to either the horizontal direction or the vertical direction. An example 2-bit direction index is shown in Table 3. In the example of Table 3, the interpretation of the MVD may vary depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a single-prediction block or corresponds to a dual-prediction block and the two reference frame lists point to the same side of the current picture (i.e., the POCs of the two reference pictures are both greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 3 may specify the symbol (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bi-predicted block and the two reference pictures are on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs of the offsets of the MVs corresponding to the reference pictures in picture reference list 0 and the reference pictures in picture reference list 1 may have opposite values ​​(opposite signs for the offsets). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the signs in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the signs for the offsets of the reference MVs associated with picture reference list 0 have opposite values.

[0183] Table 3 - Example implementation of symbols for MV offsets specified by direction index

[0184] Directions to IDX 00 01 10 11 X-axis (horizontal) + - not applicable not applicable Y axis (vertical) not applicable not applicable + -

[0185] In some example implementations, the MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, the MVD for reference list 1 is scaled. If the POC difference for reference list 1 is greater than the POC difference for reference list 0, the MVD for list 0 may be scaled in the same manner. If the starting MV is uni-predicted, the MVD is added to the available or reference MV.

[0186] In some example implementations of MVD encoding and signaling for bidirectional composite prediction, in addition to or alternatively to separately encoding and signaling two MVDs, a symmetric MVD encoding can be implemented so that only one MVD needs to be signaled, and the other MVD can be derived from the signaled MVD. In this implementation, motion information including reference picture indices of list-0 and list-1 is signaled. However, only the MVD associated with, for example, reference list-0 is signaled, and the MVD associated with reference list-1 is derived without signaling. Specifically, at the slice level, a flag "mvd_11_zero_flag" can be included in the codestream to indicate whether reference list-1 is not signaled in the codestream. If the flag is 1, indicating that reference list-1 is equal to zero (and therefore not signaled), the bidirectional prediction flag "BiDirPredFlag" can be set to 0, meaning that there is no bidirectional prediction. Otherwise, if mvd_11_zero_flag is zero, BiDirPredFlag can be set to 1 when the nearest reference picture in list-0 and the nearest reference picture in list-1 form the forward and backward directions, and both list-0 and list-1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag of 1 may indicate that the symmetric mode flag is additionally signaled in the bitstream. When BiDirPredFlag is 1, the decoder may extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag may be signaled at the CU level (if necessary), and it may indicate whether the symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, and only the reference picture indexes of both list-0 and list-1 (referred to as "mvp_10_flag" and "mvp_11_flag") are signaled with the MVD associated with list-0 (referred to as "MVD0"), and the other motion vector difference "MVD1" will be derived instead of being signaled. For example, MVD1 can be derived as -MVD0. In this way, only one MVD is signaled in the example symmetric MVD mode. In some other example implementations for MV prediction, a coordination scheme can be used to implement general merge mode, MMVD, and some other types of MV prediction for both single reference mode and composite reference mode MV prediction. Various syntax elements can be used to signal the manner in which the MV for the current block is predicted.

[0187] For example, for single reference mode, the following MV prediction modes may be signaled:

[0188] NEARMV - directly uses one of the motion vector predictors (MVPs) in the list indicated by the dynamic reference list (DRL) index without any MVD.

[0189] NEWMV - uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and applies the delta to the MVP (eg, using MVD).

[0190] GLOBALMV - Use motion vectors based on frame-level global motion parameters.

[0191] Likewise, for the composite reference inter prediction mode using two reference frames corresponding to two MVs to be predicted, the following MV prediction modes may be signaled:

[0192] NEAR_NEARMV - For each of the two MVs to be predicted, without MVD, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index.

[0193] NEAR_NEWMV - To predict the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as a reference MV without MVD; to predict the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as a reference MV in combination with an additionally signaled delta MV (MVD).

[0194] NEW_NEARMV - To predict the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as a reference MV without MVD; to predict the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as a reference MV in combination with an additionally signaled delta MV (MVD).

[0195] NEW_NEWMV - uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV, and uses it in conjunction with the additionally signaled delta MV to predict each of the two MVs.

[0196] GLOBAL_GLOBALMV - Uses MVs from each reference based on their frame-level global motion parameters.

[0197] Therefore, the term "NEAR" above refers to MV prediction using a reference MV without MVD as a general merge mode, while the term "NEW" refers to MV prediction involving using a reference MV and offsetting it with a signaled MVD as in MMVD mode. For composite inter-frame prediction, the above reference basic motion vector and motion vector increment can generally be different or independent between the two references, even though they can be correlated, and this correlation can be used to reduce the amount of information required to signal the two motion vector increments. In this case, joint signaling of the two MVDs can be implemented and indicated in the codestream.

[0198] The above dynamic reference list (DRL) can be used to store a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.

[0199] In some example implementations, a predefined resolution for the MVD may be allowed. For example, a motion vector precision (or accuracy) of 1 / 8 pixel may be allowed. The MVD in the various MV prediction modes described above may be constructed and signaled in various ways. In some implementations, various syntax elements may be used to signal the above at least one motion vector difference in reference frame list 0 or list 1.

[0200] For example, a syntax element called "mv_joint" may specify which components of the motion vector difference associated with it are non-zero. For MVD, this is signaled jointly for all non-zero components. For example, the mv_joint value is

[0201] 0 may indicate that there is no non-zero MVD in the horizontal or vertical direction;

[0202] 1 may indicate that there is non-zero MVD only along the horizontal direction;

[0203] 2 may indicate that there is non-zero MVD only along the vertical direction;

[0204] 3 may indicate that there is non-zero MVD in both the horizontal and vertical directions.

[0205] When the "mv_joint" syntax element for MVD signals that there are no non-zero MVD components, no further MVD information is signaled. However, if the "mv_joint" syntax signals that there are one or two non-zero components, additional syntax elements may be further signaled for each of the non-zero MVD components as described below.

[0206] For example, a syntax element called "mv_sign" may be used to additionally specify whether the corresponding motion vector difference amount is positive or negative.

[0207] For another example, a syntax element referred to as "mv_class" can be used to specify a class of motion vector differences for corresponding non-zero MVD components from a predefined set of categories. For example, predefined categories for motion vector differences can be used to divide the continuous magnitude space of motion vector differences into non-overlapping ranges, where each range corresponds to an MVD category. Therefore, the signaled MVD category indicates the magnitude range of the corresponding MVD component. In the example implementation shown in Table 4 below, a higher category corresponds to a motion vector difference with a larger magnitude range. In Table 4, the symbol (n, m] is used to represent a range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0208] Table 4: Magnitude categories for motion vector differences

[0209] MV Category MVD Magnitude MV_CLASS_0 (0,2] MV_CLASS_1 (2,4] MV_CLASS_2 (4,8] MV_CLASS_3 (8,16] MV_CLASS_4 (16,32] MV_CLASS_5 (32,64] MV_CLASS_6 (64,128] MV_CLASS_7 (128,256] MV_CLASS_8 (256,512] MV_CLASS_9 (512,1024] MV_CLASS_10 (1024,2048]

[0210] In some other examples, a syntax element referred to as "mv_bit" may be further used to specify the integer portion of the offset between a non-zero motion vector difference component and the starting magnitude of the correspondingly signaled MV class magnitude range. The number of bits required to signal the full range of each MVD class in "my_bit" may vary as a function of the MV class. For example, MV_CLASS 0 and MV_CLASS1 in the implementation of Table 4 may require only a single bit to indicate an integer pixel offset of 1 or 2 from a starting MVD of 0; each higher MV_CLASS in the example implementation of Table 4 may require progressively more than one bit for "mv_bit" than the previous MV_CLASS.

[0211] In some other examples, a syntax element called "mv_fr" may be further used to specify the first 2 fractional bits of the motion vector difference for the corresponding non-zero MVD component, while a syntax element called "mv_hp" may be used to specify the third fractional bit (high-resolution bit) of the motion vector difference for the corresponding non-zero MVD component. The 2-bit "mv_fr" essentially provides a 1 / 4 pixel MVD resolution, while the "mv_hp" bit may further provide a 1 / 8 pixel resolution. In some other implementations, more than one "mv_hp" bit may be used to provide an MVD pixel resolution finer than 1 / 8 pixel. In some example implementations, an additional flag may be signaled at at least one of the levels to indicate whether a 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, the syntax element above for the corresponding non-supported MVD resolution may not be signaled.

[0212] In some example implementations above, fractional resolution may be independent of different categories of MVD. In other words, similar options for motion vector resolution may be provided using a predefined number of "mv_fr" and "mv_hp" bits for signaling fractional MVDs of non-zero MVD components, regardless of the magnitude of the motion vector difference.

[0213] However, in some other example implementations, the resolutions for motion vector differences in various MVD magnitude categories may be distinguished. Specifically, high-resolution MVD for large MVD magnitudes of higher MVD categories may not statistically significantly improve compression efficiency. In this way, for larger MVD magnitude ranges corresponding to higher MVD magnitude categories, the MVD may be encoded with a reduced resolution (integer pixel resolution or fractional pixel resolution). Similarly, typically for larger MVD values, the MVD may be encoded with a reduced resolution (integer pixel resolution or fractional pixel resolution). This MVD category-related or MVD magnitude-related MVD resolution may generally be referred to as an adaptive MVD resolution. The adaptive MVD resolution may be implemented in various ways as described in the following example implementations to achieve overall better compression efficiency. In particular, due to the statistical observation that processing the MVD resolution of large magnitude or high category MVDs at a similar level as the resolution of low magnitude or low category MVDs in a non-adaptive manner may not significantly increase the coding efficiency of the inter-frame prediction residual for blocks with large magnitude or high category MVDs, the reduction in the number of signaling bits by targeting a less precise MVD may be greater than the additional bits required to encode the inter-frame prediction residual due to such less precise MVD. In other words, using a higher MVD resolution for large magnitude or high category MVDs may not produce more coding gain than using a lower MVD resolution.

[0214] In some general example implementations, the pixel resolution or precision used for the MVD may or may not increase as the MVD class increases. A reduced pixel resolution for the MVD corresponds to a coarser MVD (or a larger step size from one MVD class to the next). In some implementations, the correspondence between MVD pixel resolution and MVD class may be specified, predefined, or preconfigured, and thus may not need to be signaled in the encoded bitstream.

[0215] In some example implementations, the MV categories of Table 3 may each be associated with a different MVD pixel resolution.

[0216] In some example implementations, each MVD category may be associated with a single allowed resolution. In some other implementations, at least one MVD category may be associated with two or more selectable MVD pixel resolutions. Thus, a signal in the bitstream for the current MVD component having such an MVD category may be followed by additional signaling indicating which selectable pixel resolution is selected for the current MVD component.

[0217] In some example implementations, the MVD pixel resolutions adaptively allowed may include, but are not limited to, 1 / 64-pel (pixel), 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1-4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels ... (in descending order of resolution). In this way, each ascending MVD category may be associated with one of these resolutions in a non-ascending manner. In some implementations, an MVD category may be associated with two or more of the above resolutions, and the higher resolution may be lower than or equal to the lower resolution used for the previous MVD category. For example, if the MV_CLASS_3 of Table 4 may be associated with optional 1 pixel and 2 pixel resolutions, the highest resolution that the MV_CLASS_4 of Table 4 can be associated with will be 2 pixels. In some other implementations, the highest permissible resolution for an MV category may be higher than the lowest permissible resolution of the previous (lower) MV category. However, the average value of the allowed resolutions for ascending MV categories may only be non-ascending.

[0218] In some implementations, when fractional pixel resolutions higher than 1 / 8 pixel are allowed, the "mv_fr" and "mv_hp" signaling may be correspondingly extended to a total of more than 3 fractional bits.

[0219] In some example implementations, fractional pixel resolution may be allowed only for MVD classes that are lower than or equal to a threshold MVD class. For example, fractional pixel resolution may be allowed only for MVD-CLASS 0 and not for all other MV classes of Table 4. Similarly, fractional pixel resolution may be allowed only for MVD classes that are lower than or equal to any of the other MV classes of Table 4. For other MVD classes that are higher than the threshold MVD class, only integer pixel resolutions are allowed for the MVD. In this way, for MVDs signaled with MVD classes that are higher than or equal to the threshold MVD class, it may not be necessary to signal fractional resolution signaling such as at least one of the "mv-fr" and / or "mv-hp" bits. For MVD classes with resolutions lower than 1 pixel, the number of bits in the "mv-bit" signaling may be further reduced. For example, for MV_CLASS_5 in Table 4, the range of MVD pixel offsets is (32, 64], so 5 bits are required to signal the entire range at 1 pixel resolution. However, if MV_CLASS_5 is associated with 2-pixel MVD resolution (a lower resolution than 1 pixel resolution), then "mv-bit" may require 4 bits instead of 5 bits, and neither "mv-fr" nor "mv-hp" needs to be signaled after the signaling of "mv_class" as MV-CLASS_5.

[0220] In some example implementations, fractional pixel resolutions may only be allowed for MVDs whose integer values ​​are below a threshold integer number of pixel values. For example, fractional pixel resolutions may only be allowed for MVDs that are less than 5 pixels. Corresponding to this example, fractional resolutions may only be allowed for MV_CLASS_0 and MV_CLASS_1 of Table 4, and not for all other MV classes. For another example, fractional pixel resolutions may only be allowed for MVDs that are less than 7 pixels. Corresponding to this example, fractional resolutions may be allowed for MV_CLASS_0 and MV_CLASS_1 of Table 4 (ranges below 5 pixels), and not for MV_CLASS_3 and higher classes (ranges above 5 pixels). For an MVD belonging to MV_CLASS_2, whose pixel range contains 5 pixels, fractional pixel resolutions may or may be allowed for the MVD, depending on the "mv-bit" value. If the signaled "m-bit" value is 1 or 2 (so that the integer part of the signaled MVD is 5 or 6, calculated as the start of the pixel range with MV_CLASS_2MV_CLASS_2 indicated by "m-bit" for offset 1 or 2), then fractional pixel resolution may be allowed. Otherwise, if the signaled "mv-bit" value is 3 or 4 (so that the integer part of the signaled MVD is 7 or 8), then fractional pixel resolution may not be allowed.

[0221] In some other implementations, for MV classes equal to or above a threshold MV class, only a single MVD value may be allowed. For example, such a threshold MV class may be MV_CLASS2. Therefore, MV_CLASS_2 and above may only be allowed to have a single MVD value and no fractional pixel resolution. Single allowed MVD values ​​for these MV classes may be predefined. In some examples, the allowed single value may be a higher end value of the corresponding range for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be higher than or equal to the threshold class of MV_CLASS2, and the single allowed MVD values ​​for these classes may be predefined to be 8, 16, 32, 64, 128, 256, 512, 1024, and 2048, respectively. In some other examples, the allowed single value may be an intermediate value of the corresponding range for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be above the class threshold, and the single allowed MVD values ​​for these classes may be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536, respectively. Any other value within the range may also be defined as the single allowed resolution for the corresponding MVD class.

[0222] In the above implementation, when the signaled "mv_class" is equal to or above the predefined MVD class threshold, only the "mv_class" signaling is sufficient to determine the MVD value. Then "mv_class" and "mv_sign" will be used to determine the magnitude and direction of the MVD.

[0223] In this way, when the MVD is signaled for only one reference frame (from reference frame list 0 or list 1, but not both), or when the MVD is signaled for two reference frames together, the accuracy (or resolution) of the MVD can depend on the associated category of the motion vector difference and / or the magnitude of the MVD in Table 3.

[0224] In some other implementations, the pixel resolution or precision for MVD may decrease or not increase as the MVD magnitude increases. For example, the pixel resolution may depend on the integer portion of the MVD magnitude. In some implementations, fractional pixel resolution may only be allowed for MVD magnitudes that are less than or equal to a magnitude threshold. For a decoder, the integer portion of the MVD magnitude may first be extracted from the bitstream. The pixel resolution may then be determined, and a decision may then be made as to whether any fractional MVD exists in the bitstream and needs to be parsed (e.g., if the fractional pixel resolution does not allow for a particular extracted MVD integer magnitude, then the fractional MVD bit may not be included in the bitstream that needs to be extracted). The above example implementations associated with adaptive MVD pixel resolution associated with MVD categories are applicable to adaptive MVD pixel resolution associated with MVD magnitudes. For a specific example, an MVD category that is higher than or includes a magnitude threshold may be allowed to have only one predefined value.

[0225] The various example implementations above apply to single reference mode. These implementations also apply to the example NEW_NEARMV, NEAR_NEWMV and / or NEW_NEWMV modes in composite prediction under MMVD. These implementations are generally applicable to adaptive resolution for any MVD.

[0226] In some example implementations, MVD with adaptive resolution may be viewed as a separate inter-prediction single-reference mode, referred to herein as "ADAPTMV" or "ADAPTIVEMV" or "ADVANCEDMV" mode. Similar to conventional intra- or inter-prediction modes, such an adaptive single-reference inter-coding mode may be determined and specified at the frame level, picture level, coding block level, and other levels. The specification of such a mode indicates that (1) the corresponding coding block is intra-coded, (2) the coding block is predicted by a prediction block in a single reference frame, (3) the corresponding motion vector is also predicted via the reference motion vector and the MVD, and (4) adaptive resolution is applied to the encoding of the MVD. For example, as described above, the MVD pixel resolution may depend on the MVD level and / or the MVD magnitude.

[0227] In some examples, this single reference inter prediction mode with adaptive pixel resolution can be implemented as a sub-prediction mode of the regular inter prediction mode and can therefore be signaled after the regular inter prediction flag. In some specific implementations, the flag for signaling the ADAPTMV mode can follow the frame index of the reference frame used for inter prediction in the bitstream.

[0228] As described above, in the single reference inter prediction mode, MV can be predicted or constructed in various ways, and MV may or may not involve MVD. For example, in NEWMV mode, MV is predicted by a reference MV together with MVD, while in NEARMV mode, MV is predicted directly by a reference MV without any MVD. However, in GLOBALMV mode, MV is based on frame-level global motion parameters instead of any reference MV or MVD. In some example implementations, the added ADAPTMV mode can be implemented in parallel with these three modes. In this way, four sub-modes under the single reference inter prediction umbrella can be implemented:

[0229] ADAPTMV - uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and applies a delta to the MVP (e.g., using MVD) at an adaptive pixel resolution or precision;

[0230] NEWMV - uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and applies a delta to the MVP (e.g., using MVD) at a fixed or non-adaptive pixel resolution;

[0231] NEARMV - directly uses one of the motion vector predictors (MVPs) in the list indicated by the dynamic reference list (DRL) index without any MVD; and

[0232] GLOBAL MV - Uses motion vectors based on frame-level global motion parameters.

[0233] Therefore, these four single reference inter prediction modes can be signaled with one syntax in the codestream, in addition to other MVD and MV related syntax as described above.

[0234] In some example implementations, the context allocation for signaling the ADAPTMV mode can be the same as the context allocation for NEWMV, NEARMV, and GLOBALMV. In other words, regardless of whether these modes are signaled under the same syntax, the same probability model is used during entropy coding, such as using a context-based adaptive binary arithmetic coding (CABAC) algorithm to encode them.

[0235] In some example implementations, the context derivation for signaling various MVD-related syntaxes may depend on whether the MVD pixel resolution is adaptive. Examples of MVD-related syntaxes may include, but are not limited to, the above-mentioned mv_joint, mv_bit, mv_sign, mv_class, mv_fr, and mv_hp, as well as other MVD-related syntaxes. In this way, these MVD syntaxes may be signaled after the ADAPTMV mode is signaled, and a context or probability model may be derived by considering whether the ADAPTMV mode is signaled in the bitstream. Therefore, this implementation takes into account the differences in the symbol probability distributions for these MVD-related syntaxes in the ADAPTMV mode or the non-ADAPTMV mode.

[0236] For example, the probability distribution for mv joint or mv_class may be different when the ADAPTMV mode is used and when the ADAPTMV mode is not used. For a specific example, for mv_class distributions that are less concentrated on the lower classes of Table 4, it may be more likely to be encoded at an adaptive resolution and associated with the ADAPTMV mode. Thus, in some example implementations, if the current block is encoded in the ADAPTMV mode, a context may be derived for signaling mv_joint (or mv_class). Otherwise, at least one other different context may be used to signal mv_joint (or mv_class).

[0237] Fig.18 A flowchart 1800 of an example method following the above principles for implementing adaptive MVD resolution and signaling thereof is shown. The example decoding method flow starts at S1801. In S1810, a video stream is received. In S1820, an inter-frame prediction syntax element is extracted from the video stream to determine whether an ADAPTMV mode is signaled for at least one video block in the video stream, the ADAPTMV mode referring to a single reference inter-frame prediction mode with adaptive motion vector difference (MVD) pixel resolution. In S1830, based on whether the ADAPTMV mode is signaled in the inter-frame prediction syntax element, a current MVD pixel resolution associated with the at least one video block is determined. In S1840, based on whether the ADAPTMV mode is signaled in the inter-frame prediction syntax element and the current MVD pixel resolution, at least one MVD-related syntax element associated with the at least one video block is extracted from the video stream, and the at least one MVD-related syntax element is decoded.

[0238] In the embodiments and implementations of the present disclosure, any steps and / or operations may be combined or arranged in any number or order as required. Two or more steps and / or operations may be performed in parallel. The embodiments and implementations in the present disclosure may be used alone or in combination in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium. The embodiments in the present disclosure may be applied to luminance blocks or chrominance blocks. The term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU). The term "block" may also be used here to refer to a transform block. In the following terms, when referring to a block size, it may refer to a block width or height, or a maximum value of width and height, or a minimum value of width and height, or a zone size (width*height), or an aspect ratio of a block (width:height, or height:width).

[0239] The above techniques can be implemented as computer software through computer-readable instructions and physically stored in at least one computer-readable storage medium. Fig.19 A computer device (1900) is shown that is suitable for implementing certain embodiments of the disclosed subject matter.

[0240] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or the like to create a code comprising instructions, which may be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), or the like, or executed by decoding, microcode, or the like.

[0241] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, etc.

[0242] Fig.19 The components shown for the computer device (1900) are exemplary in nature and are not intended to limit the scope of use or functionality of computer software implementing embodiments of the present application. The configuration of the components should not be interpreted as having any dependency or requirement on any component or combination of components shown in the exemplary embodiment of the computer device (1900).

[0243] The computer device (1900) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from at least one human user through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0244] The human-computer interface input device may include at least one of the following (only one of which is drawn): keyboard (1901), mouse (1902), touchpad (1903), touch screen (1910), data gloves (not shown), joystick (1905), microphone (1906), scanner (1907), camera (1908).

[0245] The computer device (1900) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate at least one sense of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1910), a data glove (not shown), or a joystick (1905), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1909), headphones (not shown)), visual output devices (e.g., screens (1910) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have a touch screen input function, each of which has or does not have a tactile feedback function - some of which may output two-dimensional visual output or output of more than three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0246] The computer device (1900) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable compact disks (CD / DVD ROM / RW) (1920) with CD / DVD or similar media (1921), thumb drives (1922), removable hard disk drives or solid state drives (1923), traditional magnetic media such as tapes and floppy disks (not shown), special-purpose ROM / ASIC / PLD-based devices such as security software protectors (not shown), and the like.

[0247] Those skilled in the art should also understand that the term "computer-readable storage media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0248] The computer device (1900) may also include an interface (1954) to at least one communication network (1955). For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. The network also includes local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial networks (including CANBus), and the like. Some networks typically require an external network interface adapter for connecting to some universal data port or peripheral bus (1949) (for example, a USB port of the computer device (1900)); other systems are typically integrated into the core of the computer device (1900) by connecting to a system bus as described below (for example, an Ethernet interface integrated into a PC computer device or a cellular network interface integrated into a smart phone computer device). By using any of these networks, the computer device (1900) can communicate with other entities. The communication can be one-way, for receiving only (e.g., wireless television), one-way for sending only (e.g., CAN bus to certain CAN bus devices), or two-way, such as to other computer devices via a local or wide area digital network. Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0249] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core (1940) of the computer device (1900).

[0250] The core (1940) may include at least one central processing unit (CPU) (1941), a graphics processing unit (GPU) (1942), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1943), a hardware accelerator for a specific task (1944), a graphics adapter (1950), etc. These devices, as well as a read-only memory (ROM) (1945), a random access memory (1946), an internal mass storage (e.g., an internal non-user accessible hard disk drive, a solid state drive, etc.) (1947), etc., may be connected via a system bus (1948). In some computer devices, the system bus (1948) may be accessed in the form of at least one physical plug so that it may be expanded by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus (1948) of the core, or connected via a peripheral bus (1949). In one example, a screen (1910) may be connected to a graphics adapter (1950). The architecture of the peripheral bus includes a peripheral controller interface PCI, a universal serial bus USB, etc.

[0251] The CPU (1941), GPU (1942), FPGA (1943) and accelerator (1944) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1945) or RAM (1946). Transition data can also be stored in RAM (1946), while permanent data can be stored in, for example, internal mass storage (1947). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one CPU (1941), GPU (1942), mass storage (1947), ROM (1945), RAM (1946), etc.

[0252] The computer readable storage medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purpose of this application, or may be medium and code well known and available to those skilled in the art of computer software.

[0253] As an example and not a limitation, a computer device having the architecture (1900), in particular the core (1940), can be provided as a processor (including a CPU, a GPU, an FPGA, an accelerator, etc.) to execute the software contained in at least one tangible computer-readable storage medium. Such a computer-readable storage medium can be a medium associated with the above-mentioned user-accessible mass storage, and a specific memory of the core (1940) having non-volatility, such as a core internal mass storage (1947) or a ROM (1945). Software implementing various embodiments of the present application can be stored in such a device and executed by the core (1940). Depending on specific needs, the computer-readable storage medium may include one or more storage devices or chips. The software can enable the core (1940), in particular the processor therein (including a CPU, a GPU, an FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM (1946) and modifying such a data structure according to a software-defined process. Additionally or alternatively, a computer device may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., an accelerator (1944)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to a computer-readable storage medium may include circuitry (e.g., an integrated circuit (IC)) storing execution software, circuitry containing execution logic, or both. The present application includes any suitable combination of hardware and software.

[0254] Although the present application has described a number of exemplary embodiments, various changes, arrangements and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and are therefore within the spirit and scope of the present application.

[0255] Appendix A: Acronyms

[0256] JEM: Joint Development Model

[0257] VVC: Versatile Video Coding

[0258] BMS: Benchmark Collection

[0259] MV: Motion Vector

[0260] HEVC: High Efficiency Video Coding

[0261] SEI: Supplemental Enhancement Information

[0262] VUI: Video Availability Information

[0263] GOP: Group of Pictures

[0264] TU: Transform Unit

[0265] PU: prediction unit

[0266] CTU: Coding Tree Unit

[0267] CTB: Coding Tree Block

[0268] PB: prediction block

[0269] HRD: Hypothesized Reference Decoder

[0270] SNR: Signal to Noise Ratio

[0271] CPU: Central Processing Unit

[0272] GPU: Graphics Processing Unit

[0273] CRT: cathode ray tube

[0274] LCD: Liquid Crystal Display

[0275] OLED: Organic Light Emitting Diode

[0276] CD: compact disc

[0277] DVD: Digital Video Disc

[0278] ROM: Read Only Memory

[0279] RAM: Random Access Memory

[0280] ASIC: Application-Specific Integrated Circuit

[0281] PLD: Programmable Logic Device

[0282] LAN: Local Area Network

[0283] GSM: Global System for Mobile Communications

[0284] LTE: Long Term Evolution

[0285] CANBus: Controller Area Network Bus

[0286] USB: Universal Serial Bus

[0287] PCI: Peripheral Component Interconnect

[0288] FPGA: Field Programmable Gate Array

[0289] SSD: Solid State Drive

[0290] IC: Integrated Circuit

[0291] HDR: High Dynamic Range

[0292] SDR: Standard Dynamic Range

[0293] JVET: Joint Video Development Team

[0294] MPM: Most Probable Mode

[0295] WAIP: Wide Angle Intra Prediction

[0296] CU: Coding Unit

[0297] PU: prediction unit

[0298] TU: Transformation Unit

[0299] CTU: Coding Tree Unit

[0300] PDPC: Position Dependent Prediction Combination

[0301] ISP: Intra-frame sub-partitioning

[0302] SPS: Sequence Parameter Set

[0303] PPS: Picture Parameter Set

[0304] APS: Adaptive Parameter Set

[0305] VPS: Video Parameter Set

[0306] DPS: Decoding Parameter Set

[0307] ALF: Adaptive Loop Filter

[0308] SAO: Sample Adaptive Alignment CC-ALF: Cross Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter

[0309] CCSO: Cross-Component Sample Bias

[0310] LSO: Local Sample Bias

[0311] LR: Loop Restoration Filter

[0312] AV1: AOM Video 1AV2: AOM Video 2

[0313] MVD: Motion Vector Difference

[0314] CfL: From Luminance to Chroma

[0315] SDT: Semi-Decoupled Tree

[0316] SDP: Semi-Decoupled Partitioning

[0317] SST: Semi-Separating Tree

[0318] SB:Super Block

[0319] IBC (or Intra BC): Intra Block Copy

[0320] CDF: Cumulative Distribution Function

[0321] SCC: Screen Content Codec

[0322] GBI: Generalized Bi-Index

[0323] BCW: Bi-prediction with CU-level weights

[0324] CIIP: Combined Intra-Inter Prediction

[0325] POC: Image Sequential Counting

[0326] RPS: Reference Picture Set

[0327] DPB: Decoded Picture Buffer

[0328] MMVD: Merge mode with motion vector difference

Claims

1. A video decoding method, characterized in that: The method comprises: Receive video stream; Extracting an inter-frame prediction syntax element from the video bitstream, wherein the inter-frame prediction syntax element indicates which single-reference intra-frame prediction mode in a set of single-reference inter-frame prediction modes is enabled for at least one video block of the video bitstream, the set of single-reference inter-frame prediction modes comprising an ADAPTMV mode, a NEWMV mode, a NEARMV mode, and a GLOBALMV mode, the ADAPTMV mode being a single-reference inter-frame prediction mode with an adaptive motion vector difference MVD pixel resolution; determining, using an inter prediction syntax element, whether an ADAPTMV mode is enabled for the at least one video block; When ADAPTMV mode is enabled for the at least one video block, selecting a current MVD pixel resolution from a set of available MVD pixel resolutions; and The one or more video blocks are decoded based on the current MVD pixel resolution.

2. The method according to claim 1, characterized in that The inter prediction syntax element is signaled in the video code stream after signaling an inter prediction reference frame index associated with the at least one video block.

3. The method according to claim 1, characterized in that The NEWMV mode refers to a single reference inter-frame prediction mode with non-adaptive pixel resolution; The NEARMV mode refers to a single reference inter-frame prediction mode that directly predicts motion vectors without any MVD; The GLOBALMV mode refers to a single reference inter-frame prediction mode that uses a global motion parameter set to predict a motion vector.

4. The method according to claim 3, characterized in that: Further including: The information of the ADAPTMV mode, the NEWMV mode, the NEARMV mode, and the GLOBALMV mode is decoded using a shared context.

5. The method according to claim 1, characterized in that Further including: At least one context is derived based on whether the ADAPTMV mode is signaled in the inter-prediction syntax element, the at least one context being used to decode at least one MVD-related syntax element associated with the at least one video block.

6. The method according to claim 5, characterized in that The at least one MVD-related syntax element comprises at least one of the following: A first MVD syntax element, used to indicate which MVD components are non-zero; The second MVD syntax element is used to specify the MVD symbol; The third MVD syntax element is used to specify the MVD value range; a fourth MVD syntax element for specifying an integer MVD magnitude offset within the MVD magnitude range; or, The fifth MVD syntax element is used to specify the MVD pixel resolution.

7. The method according to claim 6, characterized in that The deriving at least one context based on whether the ADAPTMV mode is signaled in the inter-frame prediction syntax element comprises: When the at least one video block is encoded using the ADAPTMV mode, deriving a first context, the first context being used to decode the first MVD syntax element or the third MVD syntax element; When the at least one video block is encoded using an inter-prediction mode other than the ADAPTMV mode, a second context different from the first context is derived, the second context being used to decode the first MVD syntax element or the third MVD syntax element.

8. The method according to any one of claims 1 to 7, further comprising: An MVD magnitude range associated with the at least one video block is determined, wherein fractional MVD pixel resolution is allowed only when the MVD magnitude is equal to or less than a predetermined threshold MVD magnitude.

9. The method according to claim 8, characterized in that The allowed MVD pixel resolutions correspond to different MVD magnitudes in non-increasing order.

10. The method according to any one of claims 1 to 7, further comprising: An MVD category index is obtained from the video code stream, where the MVD category index is used to specify an MVD value range associated with the at least one video block.

11. The method according to claim 10, characterized in that The allowed MVD pixel resolutions correspond to different MVD class indices in non-increasing order.

12. The method according to claim 10, characterized in that Fractional MVD pixel resolution is allowed only when the MVD class index is equal to or less than a predetermined threshold MVD class index.

13. The method according to claim 12, characterized in that Each MVD class index that is equal to or above the predetermined threshold MVD class index is associated with a single allowed integer MVD pixel resolution value.

14. The method according to claim 10, characterized in that The MVD pixel resolutions associated with different MVD class indices are different.

15. A video encoding method, characterized in that: include: Receive video stream; Determine an inter-frame prediction syntax element of the video bitstream, wherein the inter-frame prediction syntax element indicates which single-reference intra-frame prediction mode in a set of single-reference inter-frame prediction modes is enabled for at least one video block of the video bitstream, the set of single-reference inter-frame prediction modes includes an ADAPTMV mode, a NEWMV mode, a NEARMV mode, and a GLOBALMV mode, and the ADAPTMV mode refers to a single-reference inter-frame prediction mode with an adaptive motion vector difference MVD pixel resolution; determining, using an inter prediction syntax element, whether an ADAPTMV mode is enabled for the at least one video block; selecting a current MVD pixel resolution from a set of available MVD pixel resolutions when an ADAPTMV mode is enabled for the at least one video block; and The one or more video blocks are encoded based on the current MVD pixel resolution.

16. A video decoding device, characterized in that: include: A receiving module, used for receiving a video code stream; an extraction module, configured to extract an inter-frame prediction syntax element from the video bitstream, wherein the inter-frame prediction syntax element indicates which single-reference intra-frame prediction mode in a set of single-reference inter-frame prediction modes is enabled for at least one video block of the video bitstream, the set of single-reference inter-frame prediction modes comprising an ADAPTMV mode, a NEWMV mode, a NEARMV mode, and a GLOBALMV mode, the ADAPTMV mode being a single-reference inter-frame prediction mode with an adaptive motion vector difference MVD pixel resolution; A determination module for determining whether an ADAPTMV mode is enabled for the at least one video block using an inter-frame prediction syntax element; A selection module for selecting a current MVD pixel resolution from a set of available MVD pixel resolutions when the ADAPTMV mode is enabled for the at least one video block; and A decoding module is configured to decode the one or more video blocks based on a current MVD pixel resolution.

17. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method according to any one of claims 1 to 15.

18. A non-transitory computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor, the processor is caused to implement the method according to any one of claims 1 to 15.

19. A method for storing a video code stream, characterized in that: The video code stream is decoded according to the video decoding method according to any one of claims 1 to 14, or is generated based on the video encoding method according to claim 15.