Video decoding method and video decoder
The video decoding method addresses the challenge of efficient context model storage and processing by sharing a common context model for syntax elements 1 and 2, thereby improving decoding efficiency and reducing storage needs.
Patent Information
- Application Number
- JP2024070950
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-10
- Filing Date
- 2024-04-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-09-10
AI Technical Summary
Existing video decoding technologies face challenges in efficiently storing and processing context models, which are essential for entropy decoding in video coding, leading to increased storage requirements and decreased decoding efficiency.
The proposed video decoding method analyzes a received bitstream to identify syntax elements for entropy decoding, where syntax elements 1 and 2 share a common context model, reducing the need for multiple context models and minimizing storage requirements.
This approach enhances decoding efficiency by eliminating the need to check multiple context models and reduces storage space requirements for video decoders, allowing for more efficient processing of video data.
Smart Images

Figure 0007693895000026 
Figure 0007693895000027 
Figure 0007693895000028
Abstract
Description
Background Art
[0001] This application claims priority to Chinese Patent Application No. 201811053068.0, titled "Video Decoding Method and Video Decoder", filed with the China National Intellectual Property Administration on September 10, 2018, the entire disclosure of which is incorporated herein by reference in its entirety.
[0002] Technical Field Embodiments of the present application generally relate to the field of video coding, and more specifically to video decoding methods and video decoders.
[0003] Background Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVDs and Blu-ray discs, video content capture and editing systems, and security applications of camcorders.
[0004] The development of block-based hybrid video coding modes in the 1990 H.261 standard has led to the development of new video coding technologies and tools, which form the basis of new video coding standards. Other video coding standards include MPEG-1 video, MPEG-2 video, ITU-T H.262 / MPEG-2, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10: Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards, including, for example, scalability and / or 3D (three-dimensional) extensions of such standards. As video production and consumption become increasingly widespread, video traffic has become the largest burden on communication networks and data storage. Therefore, one of the objectives of many video coding standards is to reduce the bit rate without sacrificing image quality compared to previous standards. The latest High Efficiency Video Coding (HEVC) can compress video at about twice the rate of AVC without sacrificing image quality, but there is still an urgent need for new technologies to further compress video compared to HEVC. SUMMARY OF THE INVENTION
[0005] Embodiments of the present application provide a video decoding method and a video decoder that reduce the space required by an encoder or decoder to store context.
[0006] The above and other objectives are achieved by the subject matter of the independent claims. Other implementations will be apparent from the dependent claims, the specification, and the accompanying drawings.
[0007] According to a first aspect, a video decoding method is provided. The video decoding method includes steps of analyzing a received bitstream to obtain a syntax element to be entropy decoded in a current block, where the syntax element to be entropy decoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block; performing entropy decoding on the syntax element to be entropy decoded in the current block, where entropy decoding for syntax element 1 in the current block is completed by using a pre-set context model, or entropy decoding for syntax element 2 in the current block is completed by using a context model; performing prediction processing on the current block based on the syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block; and obtaining a reconstructed image of the current block based on the predicted block of the current block.
[0008] Since syntax element 1 and syntax element 2 in the current block share one context model, the decoder does not need to check the context model when performing entropy decoding, improving the decoding efficiency of video decoding performed by the decoder. Further, since the video decoder only needs to store one context model for syntax element 1 and syntax element 2, it is possible to occupy less storage space of the video decoder.
[0009] According to a second aspect, a video decoding method is provided. The video decoding method includes steps of analyzing a received bitstream to obtain a syntax element to be entropy decoded in a current block, where the syntax element to be entropy decoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block; obtaining a context model corresponding to the syntax element to be entropy decoded, where the context model corresponding to syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 2 in the current block is determined from the pre-set context model set; performing entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block; performing prediction processing on the current block based on the syntax element in the current block and what is obtained by the entropy decoding to obtain a predicted block of the current block; and obtaining a reconstructed image of the current block based on the predicted block of the current block.
[0010] Since syntax element 1 and syntax element 2 in the current block share one context model, the video decoder needs to store only one context model for syntax element 1 and syntax element 2, occupying less storage space of the video decoder.
[0011] Regarding the second aspect, in a possible implementation, the amount of context models in the pre-set context model set is 2 or 3.
[0012] Regarding the second aspect, in a possible implementation, that the context model corresponding to syntax element 1 in the current block is determined from a pre-set context model set includes determining the context index of syntax element 1 in the current block based on syntax element 1 and syntax element 2 in the left adjacent block of the current block and syntax element 1 and syntax element 2 in the upper adjacent block of the current block. The context index of syntax element 1 in the current block is used to indicate the context model corresponding to syntax element 1 in the current block, or that the context model corresponding to syntax element 2 in the current block is determined from a pre-set context model set includes determining the context index of syntax element 2 in the current block based on syntax element 1 and syntax element 2 in the left adjacent block of the current block and syntax element 1 and syntax element 2 in the upper adjacent block of the current block. The context index of syntax element 2 in the current block is used to indicate the context model corresponding to syntax element 2 in the current block.
[0013] Regarding the second aspect, in a possible implementation, when the number of context models in a pre-set context model set is 3, the value of the context index of syntax element 1 in the current block is the sum of the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or The value of the context index of syntax element 2 in the current block is the sum of the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0014] Regarding the second aspect, in a possible implementation, when the amount of context models in a pre-set context model set is 2, the value of the context index of syntax element 1 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or The value of the context index of syntax element 2 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0015] Regarding the first aspect or the second aspect, in a possible implementation, syntax element 1 in the current block is affine_merge_flag and is used to indicate whether an affine motion model-based merge mode is used in the current block, or syntax element 2 in the current block is affine_inter_flag and is used to indicate whether an affine motion model-based AMVP mode is used in the current block when the slice where the current block is located is a P-type slice or a B-type slice, or The syntax element 1 in the current block is the subblock_merge_flag, which is used to indicate whether the subblock-based merge mode is used in the current block, or the syntax element 2 in the current block is the affine_inter_flag, and when the slice in which the current block is located is a P-type slice or a B-type slice, it is used to indicate whether the affine motion model-based AMVP mode is used in the current block.
[0016] According to a third aspect, a video decoding method is provided. The video decoding method includes steps of analyzing a received bitstream to obtain a syntax element to be entropy decoded in a current block, where the syntax element to be entropy decoded in the current block includes a syntax element 3 or a syntax element 4 in the current block; obtaining a context model corresponding to the syntax element to be entropy decoded, where the context model corresponding to the syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to the syntax element 4 in the current block is determined from a pre-set context model set; performing entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block; performing prediction processing on the current block based on the syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block; and obtaining a reconstructed image of the current block based on the predicted block of the current block.
[0017] Since the syntax elements 3 and 4 in the current block share one context model, the video decoder needs to store only one context model for the syntax elements 3 and 4, occupying less storage space of the video decoder.
[0018] Regarding the third aspect, in a possible implementation, the pre-set context model set includes five context models.
[0019] Regarding the third aspect, in a possible implementation, the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is affine_merge_idx and is used to indicate the index value of the affine merge candidate list of the current block, or the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is subblock_merge_idx and is used to indicate the index value of the sub-block merge candidate list.
[0020] According to a fourth aspect, a video decoding method is provided. The video decoding method includes steps of analyzing a received bitstream to obtain a syntax element to be entropy decoded in a current block, where the syntax element to be entropy decoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block; determining a value of a context index of the syntax element to be entropy decoded in the current block based on values of syntax element 1 and syntax element 2 in a left adjacent block of the current block and values of syntax element 1 and syntax element 2 in an upper adjacent block of the current block; performing entropy decoding on the syntax element to be entropy decoded based on the value of the context index of the syntax element to be entropy decoded in the current block; performing prediction processing on the current block based on the syntax element in the current block obtained by the entropy decoding to obtain a predicted block of the current block; and obtaining a reconstructed image of the current block based on the predicted block of the current block.
[0021] Regarding the fourth aspect, in a possible implementation, syntax element 1 in the current block is affine_merge_flag and is used to indicate whether an affine motion model-based merge mode is used in the current block, or syntax element 2 in the current block is affine_inter_flag and is used to indicate whether an affine motion model-based AMVP mode is used in the current block when the slice in which the current block is located is a P-type slice or a B-type slice, or The syntax element 1 in the current block is the subblock_merge_flag, which is used to indicate whether the subblock-based merge mode is used in the current block, or the syntax element 2 in the current block is the affine_inter_flag, and is used to indicate whether the affine motion model-based AMVP mode is used in the current block when the slice where the current block is located is a P-type slice or a B-type slice.
[0022] Regarding the fourth aspect, in a possible implementation, based on the values of the syntax element 1 and the syntax element 2 in the left adjacent block of the current block and the values of the syntax element 1 and the syntax element 2 in the upper adjacent block of the current block, the step of determining the value of the context index of the syntax element to be entropy decoded in the current block is The value of the context index of the syntax element to be entropy decoded in the current block is given by the following logical expression: Context index = (condL && availableL) + (condA && availableA) including the step of determining according to condL = syntax element 1 [x0-1][y0] | syntax element 2 [x0-1][y0] where syntax element 1 [x0-1][y0] indicates the value of the syntax element 1 in the left adjacent block, and syntax element 2 [x0-1][y0] indicates the said value of the syntax element 2 in the left adjacent block, condA = syntax element 1 [x0][y0-1] | syntax element 2 [x0][y0-1] where syntax element 1 [x0][y0-1] indicates the value of syntax element 1 in the upper adjacent block, and syntax element 2 [x0][y0-1] indicates the value of syntax element 2 in the upper adjacent block, availableL indicates whether the left adjacent block is available, and availableA indicates whether the upper adjacent block is available.
[0023] According to a fifth aspect, a video decoder is provided, the video decoder being an entropy decoding unit configured to analyze a received bitstream to obtain a syntax element to be entropy decoded in a current block, where the syntax element to be entropy decoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block, and the entropy decoding unit determines a value of a context index of the syntax element to be entropy decoded in the current block based on values of syntax element 1 and syntax element 2 in a left adjacent block of the current block and values of syntax element 1 and syntax element 2 in an upper adjacent block of the current block, and performs entropy decoding on the syntax element to be entropy decoded based on the value of the context index of the syntax element to be entropy decoded in the current block; a prediction processing unit configured to perform a prediction process on the current block based on a syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block; and a reconstruction unit configured to obtain a reconstructed image of the current block based on the predicted block of the current block.
[0024] Regarding the fifth aspect, in a possible implementation, in the current block, syntax element 1 is the affine_merge_flag, which is used to indicate whether an affine motion model-based merge mode is used in the current block, or syntax element 2 in the current block is the affine_inter_flag, and when the slice where the current block is located is a P-type slice or a B-type slice, it is used to indicate whether an affine motion model-based AMVP mode is used in the current block, or in the current block, syntax element 1 is the subblock_merge_flag, which is used to indicate whether a subblock-based merge mode is used in the current block, or syntax element 2 in the current block is the affine_inter_flag, and when the slice where the current block is located is a P-type slice or a B-type slice, it is used to indicate whether an affine motion model-based AMVP mode is used in the current block.
[0025] Regarding the fifth aspect, in a possible implementation, in a possible implementation, the entropy decoding unit specifically is configured to determine the value of the context index of the syntax element to be entropy decoded in the current block according to the following logical expression: Context index = (condL && availableL) + (condA && availableA) where condL = syntax element 1 [x0-1][y0] | syntax element 2 [x0-1][y0] and syntax element 1 [x0-1][y0] indicates the value of syntax element 1 in the left adjacent block, and syntax element 2 [x0-1][y0] indicates the value of syntax element 2 in the left adjacent block. condA = syntax element 1 [x0][y0-1] | syntax element 2 [x0][y0-1] where syntax element 1 [x0][y0-1] indicates the value of syntax element 1 in the upper adjacent block, and syntax element 2 [x0][y0-1] indicates the said value of syntax element 2 in the upper adjacent block, availableL indicates whether the left adjacent block is available, and availableA indicates whether the upper adjacent block is available.
[0026] According to a sixth aspect, a video decoder is provided, the video decoder being an entropy decoding unit configured to analyze a received bitstream to obtain a syntax element to be entropy decoded in a current block, the syntax element to be entropy decoded in the current block including syntax element 1 in the current block or syntax element 2 in the current block, the entropy decoding unit being configured to perform entropy decoding on the syntax element to be entropy decoded in the current block, the entropy decoding regarding syntax element 1 in the current block being completed by using a pre-set context model, or the entropy decoding regarding syntax element 2 in the current block being completed by using a context model, an entropy decoding unit, a prediction processing unit configured to perform a prediction process on the current block based on the syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block, and a reconstruction unit configured to obtain a reconstructed image of the current block based on the predicted block of the current block.
[0027] According to a seventh aspect, a video decoder is provided. The video decoder includes an entropy decoding unit configured to analyze a received bitstream to obtain a syntax element to be entropy decoded in a current block. The syntax element to be entropy decoded in the current block includes a syntax element 1 in the current block or a syntax element 2 in the current block. The entropy decoding unit is configured to obtain a context model corresponding to the syntax element to be entropy decoded. The context model corresponding to the syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to the syntax element 2 in the current block is determined from the pre-set context model set. The entropy decoding unit is configured to perform entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block. The video decoder further includes a prediction processing unit configured to perform a prediction process on the current block based on a syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block, and a reconstruction unit configured to obtain a reconstructed image of the current block based on the predicted block of the current block.
[0028] Regarding the seventh aspect, in a possible implementation, the number of context models in the pre-set context model set is 2 or 3.
[0029] Regarding the seventh aspect, in a possible implementation, the entropy decoding unit is specifically configured to determine the context index of the syntax element 1 in the current block based on the syntax element 1 and the syntax element 2 in the left adjacent block of the current block and the syntax element 1 and the syntax element 2 in the upper adjacent block of the current block. The context index of the syntax element 1 in the current block is used to indicate the context model corresponding to the syntax element 1 in the current block, or The entropy decoding unit is specifically configured to determine the context index of the syntax element 2 in the current block based on the syntax element 1 and the syntax element 2 in the left adjacent block of the current block and the syntax element 1 and the syntax element 2 in the upper adjacent block of the current block. The context index of the syntax element 2 in the current block is used to indicate the context model corresponding to the syntax element 2 in the current block.
[0030] Regarding the seventh aspect, in a possible implementation, when the number of context models in the pre-set context model set is 3, the value of the context index of the syntax element 1 in the current block is the sum of the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the left adjacent block, or the value of the context index of the syntax element 2 in the current block is the sum of the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the left adjacent block.
[0031] Regarding the seventh aspect, in a possible implementation, when the amount of context models in a pre-set context model set is 2, the value of the context index of syntax element 1 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or the value of the context index of syntax element 2 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0032] Regarding the sixth or seventh aspect, in a possible implementation, syntax element 1 in the current block is affine_merge_flag and is used to indicate whether an affine motion model-based merge mode is used in the current block, or syntax element 2 in the current block is affine_inter_flag and is used to indicate whether an affine motion model-based AMVP mode is used in the current block when the slice where the current block is located is a P-type slice or a B-type slice, or syntax element 1 in the current block is subblock_merge_flag and is used to indicate whether a subblock-based merge mode is used in the current block, or syntax element 2 in the current block is affine_inter_flag and is used to indicate whether an affine motion model-based AMVP mode is used in the current block when the slice where the current block is located is a P-type slice or a B-type slice.
[0033] According to the eighth aspect, a video decoder is provided. The video decoder is an entropy decoding unit configured to analyze a received bitstream to obtain a syntax element to be entropy decoded in a current block. The syntax element to be entropy decoded in the current block includes syntax element 3 in the current block or syntax element 4 in the current block. The entropy decoding unit is configured to obtain a context model corresponding to the syntax element to be entropy decoded. The context model corresponding to syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 4 in the current block is determined from a pre-set context model set. The entropy decoding unit is configured to perform entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block, an entropy decoding unit; a prediction processing unit configured to perform prediction processing on the current block based on a syntax element in the current block and obtained by entropy decoding to obtain a predicted block of the current block; and a reconstruction unit configured to obtain a reconstructed image of the current block based on the predicted block of the current block.
[0034] Regarding the eighth aspect, in a possible implementation, the pre-set context model set includes five context models.
[0035] Regarding the eighth aspect, in a possible implementation, the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is affine_merge_idx and is used to indicate the index value of the affine merge candidate list of the current block, or the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is subblock_merge_idx and is used to indicate the index value of the subblock merge candidate list.
[0036] According to the ninth aspect, an encoding method is provided. The encoding method includes steps of obtaining a syntax element to be entropy encoded in a current block, where the syntax element to be entropy encoded in the current block includes syntax element 1 or syntax element 2 in the current block; performing entropy encoding on the syntax element to be entropy encoded in the current block, where when entropy encoding is performed on the syntax element to be entropy encoded in the current block, entropy encoding for syntax element 1 in the current block is completed by using a pre-set context model, or entropy encoding for syntax element 2 in the current block is completed by using a context model; and outputting a bitstream including the syntax element in the current block and obtained by entropy encoding.
[0037] For specific syntax elements and specific context models, refer to the first aspect.
[0038] According to the tenth aspect, an encoding method is provided. The encoding method includes a step of obtaining a syntax element to be entropy-encoded in a current block, where the syntax element to be entropy-encoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block, and a step of Symbol obtaining a context model corresponding to the syntax element to be entropy-encoded, where the context model corresponding to syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 2 in the current block is determined from a pre-set context model set, and a step of, based on the context model corresponding to the syntax element to be entropy-encoded in the current block, Symbol performing entropy encoding on the syntax element to be entropy-encoded, and a step of outputting a bitstream including the syntax element in the current block and obtained by entropy encoding.
[0039] For specific syntax elements and specific context models, refer to the second aspect.
[0040] According to the 11th aspect, an encoding method is provided. The encoding method includes a step of obtaining a syntax element to be entropy-encoded in a current block, where the syntax element to be entropy-encoded in the current block includes syntax element 3 in the current block or syntax element 4 in the current block; a step of obtaining a context model corresponding to the syntax element to be entropy-decoded, where the context model corresponding to syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 4 in the current block is determined from a pre-set context model set; a step of performing entropy encoding on the syntax element to be entropy-encoded with respect to the context model corresponding to the syntax element to be entropy-encoded in the current block; and a step of outputting a bitstream including the syntax element in the current block and obtained by entropy encoding. Symbol For specific syntax elements and specific context models, refer to the 3rd aspect.
[0041] For specific syntax elements and specific context models, refer to the 3rd aspect.
[0042] According to the 12th aspect, a video encoder is provided. The video encoder is an entropy encoding unit configured to obtain a syntax element to be entropy encoded in a current block. The syntax element to be entropy encoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block. The entropy encoding unit is configured to perform entropy encoding on the syntax element to be entropy encoded in the current block. When entropy encoding is performed on the syntax element to be entropy encoded in the current block, the entropy encoding for syntax element 1 in the current block is completed by using a pre-set context model, or the entropy encoding for syntax element 2 in the current block is completed by using a context model. It includes the entropy encoding unit and an output configured to output a bitstream including the syntax element in the current block and what is obtained by entropy encoding.
[0043] For specific syntax elements and specific context models, refer to the 4th aspect.
[0044] According to the 13th aspect, a video encoder is provided. The video encoder is an entropy encoding unit configured to obtain a syntax element to be entropy encoded in a current block. The syntax element to be entropy encoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block. The entropy encoding unit is entropy Symbolconfigured to obtain a context model corresponding to the syntax element to be coded, the context model corresponding to the syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to the syntax element 2 in the current block is determined from a pre-set context model set, and the entropy coding unit is based on the context model corresponding to the syntax element to be entropy-coded in the current block, entropy Symbol including an entropy coding unit configured to perform entropy coding on the syntax element to be coded, and an output configured to output a bitstream including the syntax element in the current block and obtained by entropy coding.
[0045] For specific syntax elements and specific context models, refer to the fifth aspect.
[0046] According to the fourteenth aspect, a video encoder is provided, the video encoder is an entropy coding unit configured to obtain the syntax element to be entropy-coded in the current block, the syntax element to be entropy-coded in the current block includes the syntax element 3 in the current block or the syntax element 4 in the current block, the entropy coding unit is configured to obtain a context model corresponding to the syntax element to be entropy-decoded, the context model corresponding to the syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to the syntax element 4 in the current block is determined from a pre-set context model set, and the entropy coding unit is based on the context model corresponding to the syntax element to be entropy-coded in the current block, entropy SymbolAn entropy encoding unit configured to perform entropy encoding on syntax elements to be coded, and an output configured to output a bitstream including syntax elements in a current block and obtained by entropy encoding.
[0047] For specific syntax elements and specific context models, refer to the sixth aspect.
[0048] According to a fifteenth aspect, the present invention relates to an apparatus for decoding a video stream, including a processor and a memory. The memory stores instructions that enable the processor to execute the method in the first aspect, the second aspect, the third aspect, or the fourth aspect, or any possible implementation thereof.
[0049] According to a sixteenth aspect, the present invention relates to an apparatus for decoding a video stream, including a processor and a memory. The memory stores instructions that enable the processor to execute the method in the seventh aspect, the eighth aspect, or the ninth aspect, or any possible implementation thereof.
[0050] According to a seventeenth aspect, a computer-readable storage medium is proposed. The computer-readable storage medium stores instructions that, when executed, enable one or more processors to encode video data. The instructions enable one or more processors to execute the method in the first aspect, the second aspect, the third aspect, the fourth aspect, the seventh aspect, the eighth aspect, or the ninth aspect, or any possible implementation thereof.
[0051] According to an eighteenth aspect, the present invention relates to a computer program including program code. When the program code is executed on a computer, the method in the first aspect, the second aspect, the third aspect, the fourth aspect, the seventh aspect, the eighth aspect, or the ninth aspect, or any possible implementation thereof is executed.
[0052] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments or background of the present application, the accompanying drawings required for explaining the embodiments or background of the present application will be briefly described below.
[0054]
Figure 1
[0055]
Figure 2
[0056]
Figure 3
[0057]
Figure 4
[0058]
Figure 5
[0059]
Figure 6
[0060]
Figure 7
[0061]
Figure 8A
[0062]
Figure 8B
[0063]
Figure 9A
[0064]
Figure 9B
[0065]
Figure 9C
[0066]
Figure 10
[0067]
Figure 11
[0068]
Figure 12
[0069] Hereinafter, unless otherwise specified, the same reference symbols represent the same or at least functionally equivalent features.
DETAILED DESCRIPTION OF THE INVENTION
[0070] In the following description, reference is made to the accompanying drawings, which form a part hereof and which illustrate specific aspects of embodiments of the invention or specific aspects in which embodiments of the invention may be used. It is to be understood that the embodiments of the invention may be used in other aspects and may include structural or logical changes not shown in the accompanying drawings. Accordingly, the following detailed description is not to be construed in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0071] For example, it should be understood that what is disclosed with respect to the described method may also be applicable to a corresponding device or system configured to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., one unit for performing one or more of the described steps, or a plurality of units each for performing one or more of the plurality of steps) such as functional units for performing the one or more described method steps, even if such one or more units are not explicitly described or illustrated in the accompanying drawings. Further, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include one step (e.g., one step for performing the functions of one or more units, or a plurality of steps each for performing one or more of the functions of a plurality of units) for performing the functions of the one or more units, even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. Additionally, it should be understood that the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.
[0072] Video coding typically processes a series of pictures that form a video or video sequence. In the field of video coding, the terms "picture", "frame", and "image" may be used interchangeably. The video coding used in this application (or this disclosure) refers to video encoding or video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves performing the reverse process associated with the encoder to reconstruct the video picture. The "coding" of a video picture (or generally referred to as a picture as described below) in an embodiment should be understood as "encoding" or "decoding" with respect to the video sequence. A combination of encoding and decoding is also referred to as coding (encoding and decoding).
[0073] In the case of lossless video coding, it is possible to reconstruct the original video picture, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss occurs during storage and transmission). In the case of non-lossless video coding, further compression is performed, such as quantization, to reduce the amount of data required to represent the video picture, and the video picture cannot be completely reconstructed on the decoder side, i.e., the quality of the reconstructed video picture is inferior to that of the original video picture.
[0074] Several H.261 video coding standards are related to "lossy hybrid video coding" (i.e., spatial and temporal prediction in the sample domain is combined with 2D transform coding to apply quantization in the transform domain). Each picture of a video sequence is usually divided into a set of non-overlapping blocks, and coding is usually performed at the block level. Specifically, on the encoder side, the video is usually processed, i.e., encoded, at the block (video block) level. For example, prediction blocks are generated by spatial (intra-picture) prediction and temporal (inter-picture) prediction, and the prediction blocks are subtracted from the current block (the block being processed or to be processed) to obtain a residual block, and the residual block is transformed and quantized in the transform domain to reduce (compress) the amount of data to be transmitted. On the decoder side, the inverse process for the encoder is applied to the encoded or compressed block to reconstruct the current block for presentation. Further, the encoder repeats the decoder's processing loop, and as a result, the encoder and decoder generate the same prediction (e.g., intra prediction and inter prediction) and / or reconstruction for processing, i.e., encoding, subsequent blocks.
[0075] As used herein, the term "block" may be part of a picture or frame. For ease of explanation, embodiments of the present invention will be described with reference to Versatile Video Coding (VVC) or High-Efficiency Video Coding (HEVC) developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC, and the block may be a Coding Unit (CU), a Prediction Unit (PU), or a Transform Unit (TU). In HEVC, a Coding Tree Unit (CTU) is divided into a plurality of CUs by using a quadtree structure shown as a coding tree. It is determined whether the picture area is coded by inter-picture (temporal) or intra-picture (spatial) prediction at the CU level. Each CU may further be divided into one, two, or four PUs based on the PU split type. The same prediction process is applied within one PU, and the relevant information is sent to the decoder based on the PU. By applying the prediction process based on the PU split type, after obtaining the residual block, the CU may be divided into a transform unit (TU) based on another quadtree structure similar to the coding tree used for the CU. In the development of the latest video compression technology, the frame is a quadtree plus binary tree (Quad-tree plusIt is divided by a quadtree binary tree (QTBT). In the QTBT block structure, the CU may be square or rectangular. In VVC, a coding tree unit (CTU) is first divided by using a quadtree structure, and a quadtree leaf node is further divided by using a binary tree structure. The binary tree leaf node is referred to as a coding unit (CU), and the division is used for prediction and transformation processing without any other division. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. Further, multiple divisions, such as triple tree division, are used together with the QTBT block structure.
[0076] Hereinafter, embodiments of the encoder 20, decoder 30, encoding system 10, and decoding system 40 will be described based on FIGS. 1-4 (before describing embodiments of the present invention in more detail based on FIG. 10).
[0077] FIG. 1 is a conceptual or schematic block diagram showing an exemplary encoding system 10, for example, a video encoding system 10 capable of using the technology of the present application (this disclosure). The encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) within the video encoding system 10 represent example devices that may be configured to perform techniques for... (division / intra prediction / ...) according to various examples described in the present application. As shown in FIG. 1, the encoding system 10 includes a source device 12 configured to provide encoded data 13, such as an encoded picture 13, to a destination device that decodes the encoded data 13 and the like.
[0078] The source device 12 includes an encoder 20 and may additionally or optionally include a picture source 16, a preprocessing unit 18 such as a picture preprocessing unit 18, and a communication interface or communication unit 22.
[0079] The video source 16 may include any type of picture capture device configured to capture real-world pictures and the like, and / or any type of device that generates pictures or comments (in the case of screen content encoding, any text on the screen is also considered part of the picture or image to be encoded), for example, a computer graphics processing unit configured to generate computer animation pictures, or any type of device configured to acquire and / or provide real-world pictures or computer animation pictures (for example, screen content or virtual reality (VR) pictures), and / or any combination thereof (for example, augmented reality (AR) pictures), or it may be these devices themselves.
[0080] (Digital) pictures are, or may be considered to be, two-dimensional arrays or matrices of samples having luminance values. The samples in the array may be called pixels (short for picture elements) or pels. The amount of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, usually three color components are used, i.e., the picture may be represented as, or may include, three sample arrays. In the RGB format or color space, the picture includes corresponding red, green, and blue sample arrays. However, in video coding, each sample is usually represented in a luminance / chrominance format or color space. For example, a picture in the YCbCr format includes a luminance component (sometimes denoted by L) represented by Y and two chrominance components represented by Cb and Cr. The luminance (abbreviation, luma) component Y indicates luminance or gray-level intensity (e.g., they are the same in a grayscale picture), and the two chrominance components (abbreviation, chroma) Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format can be converted or transformed into a picture in the YCbCr format, and vice versa. This process is also called color conversion or transformation. If the picture is monochrome, the picture may only include a luminance sample array.
[0081] Picture source 16 (e.g., video source 16) may be, for example, a camera configured to capture pictures, a memory such as a picture memory that includes or stores pre-captured or generated pictures, and / or any type of (internal or external) interface for acquiring or receiving pictures. The camera may be, for example, a local camera or an integrated camera integrated into the source device, and the memory may be a local memory or an integrated memory integrated into the source device. The interface may be, for example, an external interface for receiving pictures from an external video source. The external video source is an external picture capture device such as, for example, a camera, an external memory, or an external picture generation device. The external picture generation device is, for example, an external computer graphics processing unit, a computer, or a server. The interface can be assumed to be any type of interface that conforms to any proprietary or standardized interface protocol, such as a wired or wireless interface or an optical interface. The interface for acquiring picture data 17 may be the same interface as communication interface 22 or may be part of communication interface 22.
[0082] Unlike the preprocessing unit 18 and the processing performed by the preprocessing unit 18, pictures 17 and picture data 17 (e.g., video data 16) may also be referred to as original pictures 17 or original picture data 17.
[0083] The preprocessing unit 18 is configured to receive the (original) picture data 17, perform preprocessing on the picture data 17, and obtain the preprocessed picture 19 or preprocessed picture data 19. For example, the preprocessing executed by the preprocessing unit 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It will be understood that the preprocessing unit 18 may be an optional component.
[0084] The encoder 20 (e.g., video encoder 20) is configured to receive the preprocessed picture data 19 and provide the encoded picture data 21 (details will be further described, for example, based on FIGS. 2 or 4 below). In one example, the encoder 20 can be configured to encode pictures.
[0085] The communication interface 22 of the source device 12 is capable of receiving the encoded picture data 21 and transmitting the encoded picture data 21 to other devices, such as the destination device 14 or any other device, for storage or direct reconstruction, or can be configured to correspondingly store the encoded picture data 13 and / or process the encoded picture data 21 before transmitting the encoded data 13 to other devices. The other device is, for example, the destination device 14 or another device used for decoding or storage.
[0086] The destination device 14 includes a decoder 30 (e.g., video decoder 30) and may additionally or optionally include a communication interface or communication unit 28, a post-processing unit 32, and a display device 34.
[0087] For example, the communication interface 28 of the destination device 14 is configured to directly receive the encoded picture data 21 or the encoded data 13 from the source device 12 or any other arbitrary source. Any other arbitrary source is, for example, a storage device, and the storage device is, for example, a storage device for the encoded picture data.
[0088] The communication interface 22 and the communication interface 28 can be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14 or any type of network. The direct communication link is, for example, a direct wired or wireless connection, and any type of network is, for example, a wired or wireless network or any combination thereof, or any type of private network or public network, or any combination thereof.
[0089] The communication interface 22 can be configured to encapsulate, for example, the encoded picture data 21 into an appropriate format such as a packet for transmission on a communication link or a communication network.
[0090] The communication interface 28 as the corresponding part of the communication interface 22 can be configured to decapsulate the encoded data 13 in order to obtain the encoded picture data 21 and the like.
[0091] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface, for example, as an arrow from the source device 12 to the destination device 14 used for the encoded picture data 13 in FIG. 1, or can be configured as a bidirectional communication interface, for example, to send and receive messages to establish a connection, and can be configured to verify and exchange any other information related to data transmission such as a communication link and / or encoded picture data transmission.
[0092] The decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture 31 (Details will be further described below, for example, based on FIG. 3 or FIG. 5).
[0093] The post - processing processor 32 of the destination device 14 post - processes the decoded picture data 31 such as the decoded picture data 31 (also called reconstructed picture data) to obtain post - processed picture data 33 such as the post - processed picture 33. The post - processing executed by the post - processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, resampling, or any other processing for preparing the decoded picture data 31 for display by the display device 34. タ3 1
[0094] The display device 34 of the destination device 14 is configured to receive the post - processed picture data 33 and display the picture to a user, viewer, etc. The display device 34 may be any type of display configured to present the reconstructed picture, such as an integrated or external display or monitor, or may include them. For example, the display may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display.
[0095] FIG. 1 depicts the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both the source device 12 and the destination device 14, or both the functions of the source device 12 and the functions of the destination device 14, i.e., the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented by using the same hardware and / or software, separate hardware and / or software, or any combination thereof.
[0096] Those skilled in the art can easily understand from the specification that the functions of the source device 12 and / or the destination device 14 shown in FIG. 1, the functions of various units, or the presence and (exact) division of functions may vary depending on the actual device and application.
[0097] Encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) can each be implemented as any one of various suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented partially in software, the device can store software instructions in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware by using one or more processors to perform the technology of the present disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered as one or more processors. Video encoder 20 and video decoder 30 can each be included in one or more encoders or decoders, and any one of the encoders or decoders can be integrated as part of a combined encoder / decoder (codec) within the corresponding device.
[0098] Source device 12 may be referred to as a video encoding device or a video encoding apparatus. Destination device 14 may be referred to as a video decoding device or a video decoding apparatus. Source device 12 and destination device 14 can each be an example of a video encoding device or a Restore video encoding apparatus.
[0099] The source device 12 and the destination device 14 can each be any one of a variety of devices including any type of handheld or stationary device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a video camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiving device, or a broadcast transmitting device, and may or may not use any type of operating system.
[0100] In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0101] In some cases, the video encoding system 10 shown in FIG. 1 is merely an example, and the technology of the present application may be applicable to video coding settings (such as video encoding or video decoding) that do not necessarily involve any data communication between the encoding device and the decoding device. In other examples, the data may be retrieved from local memory or streamed via a network. The video encoding device can encode the data and store the data in memory, and / or the video decoding device can retrieve the data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to memory and / or retrieve data from memory to decode the data.
[0102] For each of the foregoing examples described with reference to video encoder 20, it should be understood that video decoder 30 can be configured to perform the reverse process. In the case of signaling syntax elements, video decoder 30 can be configured to receive and parse the syntax elements and, in response, decode the associated video data. In some examples, video encoder 20 can entropy encode one or more syntax elements that define... into an encoded video bitstream. In such examples, video decoder 30 can parse such syntax elements and, in response, decode the associated video data.
[0103] Encoder & Encoding Method
[0104] FIG. 2 is a schematic / conceptual block diagram of an example of a video encoder 20 configured to implement the technology in the present application (disclosure). In the example of FIG. 2, video encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy encoding unit 270. Prediction processing unit 260 can include an inter prediction unit 244, an intra prediction unit 254, and a mode selection unit 262. Inter prediction unit 244 can include a motion estimation unit and a motion compensation unit (not shown in the drawing). Video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.
[0105] For example, the residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 20, and the inverse quantization unit 210, the inverse transformation processing unit section 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the prediction processing unit 260, etc. form the reverse signal path of the encoder. The reverse signal path of the encoder corresponds to the signal path of the decoder (see the decoder 30 in FIG. 3).
[0106] The encoder 20 receives a picture 201 or a block 203 of the picture 201, for example, a picture in a series of pictures forming a video or a video sequence, by using the input 202, etc. The picture block 203 may also be referred to as the current picture block or the picture block to be encoded, and the picture 201 may be referred to as the current picture or the picture to be encoded (in particular, when the current picture is distinguished from other pictures in video coding, for example, other pictures in the same video sequence also include pictures that have been previously encoded and / or decoded in the video sequence of the current picture).
[0107] Partition
[0108] An embodiment of the encoder 20 can include a partitioning unit (not shown in FIG. 2) configured to partition the picture 201 into a plurality of non-overlapping blocks such as block 203. The partitioning unit can be configured to use the same block size and a corresponding raster defining the block size for all pictures in the video sequence, or it can be configured to change the block size between pictures, subsets, or groups of pictures and partition each picture into corresponding blocks.
[0109] In one example, the prediction processing unit 260 of the video encoder 20 may be configured to perform any combination of the foregoing splitting techniques.
[0110] For example, in picture 201, block 203 may also be or be considered as a two-dimensional array or matrix having luminance values (sample values), but the size of block 203 is smaller than that of picture 201. In other words, block 203 may include, for example, one sample array (e.g., the luminance array in the case of a monochrome picture 201), three sample arrays (e.g., one luminance array and two chrominance arrays in the case of a color picture), or any other arbitrary quantity and / or type of array based on the color format used. The quantity of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203.
[0111] The encoder 20 shown in FIG. 2 is configured to perform encoding and prediction, for example, in each block 203, so as to encode picture 201 block by block.
[0112] Residual calculation
[0113] The residual calculation unit 204 is configured to obtain the residual block 205 in the sample domain, for example, by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 sample by sample (pixel by pixel) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 will be provided below) to calculate the residual block 205.
[0114] Transformation
[0115] The transformation processing unit 206 is configured to apply a transformation such as a discrete cosine transform (DCT) or a discrete sine transform (DST) to the sample values of the residual block 205 and obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207 may also be referred to as residual transformation coefficients and represent the residual block 205 in the transform domain.
[0116] The transformation processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the transformation specified in HEVC / H.265. This integer approximation is typically scaled proportionally by a factor comparable to the orthogonal DCT transform. To maintain the norm of the residual block obtained through the forward and inverse transforms, an additional scale factor is applied as part of the transformation process. The scale factor is typically selected based on several constraints, such as a power of 2, the bit depth of the transformation coefficients, or a trade-off between the accuracy used in the shift operation and the implementation cost. For example, by using the inverse transformation processing unit 212, a specific scale factor may be specified for the inverse transformation on the decoder 30 side (correspondingly, by using the inverse transformation processing unit 212, etc., for the inverse transformation on the encoder 20 side), and correspondingly, by using the transformation processing unit 206, a corresponding scale factor may be specified for the forward transformation on the encoder 20 side.
[0117] Quantization
[0118] The quantization unit 208 is configured to quantize the transform coefficient 207 by applying scale quantization, vector quantization, etc., and obtain the quantized transform coefficient 209. The quantized transform coefficient 209 is also referred to as the quantized residual coefficient 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization can be adjusted by adjusting the quantization parameter (QP). For example, for scale quantization, different scales may be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, and larger quantization steps correspond to coarser quantization. An appropriate quantization step may be indicated by using the quantization parameter (QP). For example, the quantization parameter may be an index of a predetermined set of appropriate quantization steps. For example, a smaller quantization parameter corresponds to finer quantization (smaller quantization steps), and a larger quantization parameter corresponds to coarser quantization (larger quantization steps), and vice versa is also possible. Quantization may include division by the quantization step and corresponding quantization or inverse quantization performed by the inverse quantization unit 210, etc., or may include multiplication by the quantization step. In embodiments according to some standards such as HEVC, the quantization parameter may be used to determine the quantization step. Generally, the quantization step may be calculated based on the quantization parameter by fixed-point approximation of an expression including division. Additional scale factors are introduced for quantization and inverse quantization, and it is possible to restore the norm of the residual block, which may be modified due to the scale and quantization parameter used in the fixed-point approximation of the expression for the quantization step. In an exemplary implementation, the scale of the inverse transform may be combined with the scale of the inverse quantization.Alternatively, a customized quantization table may be used and signaled, e.g., in a bitstream, from the encoder to the decoder. Quantization is a lossy operation, and larger quantization steps indicate larger losses.
[0119] The inverse quantization unit 210 is configured to apply an inverse quantization of the quantization unit 208 to the quantization coefficients to obtain inverse quantization coefficients 211, e.g., based on or using the same quantization step as the quantization unit 208, and to apply an inverse quantization method of the quantization method applied by the quantization unit 208. The inverse quantization coefficients 211 are also referred to as inverse quantization residual coefficients 211 and correspond to the transform coefficients 207, but the losses caused by quantization are usually different from the transform coefficients.
[0120] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, e.g., an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as an inverse transform inverse quantization block 213 or an inverse transform residual block 213.
[0121] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, e.g., by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265, to obtain a reconstructed block 215 in the sample domain.
[0122] As an option, a buffer unit 216, such as line buffer 216 (or simply referred to as "buffer" 216), is configured to buffer or store the reconstructed block 215 and corresponding sample values for intra prediction and the like. In other embodiments, the encoder may be configured to use the unfiltered reconstructed blocks and / or corresponding sample values stored in buffer unit 216 for any type of estimation and / or prediction, such as intra prediction.
[0123] For example, an embodiment of encoder 20 is such that buffer unit 216 is configured not only to store reconstructed block 215 for intra prediction Measurement but also to store the filtered blocks 221 of loop filter unit 220 (not shown in FIG. 2), and / or buffer unit 216 and decoded picture buffer unit 230 can be configured to form one buffer. Other embodiments may be used to use blocks or samples from filtered block 221 and / or decoded picture buffer 230 (not shown in FIG. 2) as input or basis for intra prediction 254.
[0124] The loop filter unit 220 (or, abbreviated as "loop filter" 220) is configured to perform filtering on the reconstructed block 215 to obtain the filtered block 221, perform sample conversion smoothly, or improve video quality. The loop filter unit 220 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a collaborative filter. The loop filter unit 220 is shown in FIG. 2 as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221. The decoded picture buffer 230 is capable of storing the reconstructed coding block after the loop filter unit 220 performs a filtering process on the reconstructed coding block.
[0125] An embodiment of the encoder 20 (correspondingly, the loop filter unit 220) can be used to output loop filter parameters (e.g., sample adaptation offset information), for example, to directly output the loop filter parameters, or for example, after the entropy coding unit 270 or some other entropy coding unit performs entropy coding, to output the loop filter parameters, so that the decoder 30 can receive and apply the same loop filter parameters for decoding and the like.
[0126] The decoded picture buffer (DPB) 230 may be a reference picture memory for storing reference picture data for the video encoder 20 to encode video data. The DPB 230 may be any of a plurality of memories, such as a dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), or resistive RAM (RRAM)), or another type of memory. The DPB 230 and the buffer 216 may be provided by the same memory or separate memories. In one example, the decoded picture buffer (DRB) 230 is configured to store the filtered block 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks, such as previously reconstructed and filtered blocks 221 of different pictures, such as the same current picture or a previously reconstructed picture, and it is possible to provide a completely previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and corresponding samples) and / or a partially reconstructed current picture (and corresponding reference blocks and corresponding samples) for intra prediction, etc. In one example, when the reconstruction block 215 is reconstructed without loop filtering, the decoded picture buffer (DPB) 230 is configured to store the reconstruction block 215.
[0127] The prediction processing unit 260, also referred to as the block prediction processing unit 260, is configured to receive or obtain the block 203 (the current block 203 of the current picture 201) and the reconstructed picture data, such as reference samples from the same (current) picture in the buffer 216, and reference picture data 231 from one or more previous decoded pictures in the decoded picture buffer 230, and process that data for prediction, i.e., to provide a prediction block 265 that may be an inter prediction block 245 or an intra prediction block 255.
[0128] The mode selection unit 262 is configured to select a prediction mode (e.g., an intra or inter prediction mode) and / or a corresponding prediction block 245 or 255 as the prediction block 265, calculate the residual block 205, and reconstruct the reconstruction block 215.
[0129] Embodiments of the mode selection unit 262 may be used to select a prediction mode (e.g., from the prediction modes supported by the prediction processing unit 260). The prediction mode provides the best match or the minimum residual (the minimum residual means better compression in transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression in transmission or storage), or considers both or balances them. The mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion optimization, or to select a prediction mode whose associated rate distortion at least meets the prediction mode selection criteria.
[0130] The prediction processing (e.g., by using the prediction processing unit 260) and the mode selection (e.g., by using the mode selection unit 262) performed by an example of the encoder 20 are described in detail below.
[0131] As described above, the encoder 20 is configured to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes can include, for example, an intra prediction mode and / or an inter prediction mode.
[0132] The set of intra prediction modes can include 35 different intra prediction modes, such as non-directional modes like the DC (or average) mode and the planar mode, or directional modes defined in H.265, or can include 67 intra prediction modes, such as non-directional modes like the DC (or average) mode and the planar mode, or advanced directional modes defined in H.266.
[0133] The (possible) set of inter prediction modes depends on the available reference pictures (e.g., at least a part of the decoded pictures stored in the DBP 230) and other inter prediction parameters, for example, whether the entire reference picture is used or only a part of the reference picture is used, for example, whether a search window area surrounding the area of the current block is searched for the best matching reference block, and / or whether sample interpolation such as half sample and / or quarter sample interpolation is applied.
[0134] In addition to the aforementioned prediction modes, a skip mode and / or a direct mode can also be applied.
[0135] The prediction processing unit 260 may further be configured to divide block 203 into smaller block partitions or sub-blocks, for example, by repeatedly using a quad-tree (QT) partition, a binary-tree (BT) partition, a triple-tree or ternary-tree (TT) partition, or a combination thereof, and perform prediction and the like for each of the block partitions or sub-blocks. The mode selection includes selecting the tree structure of the divided block 203 and selecting the prediction mode applied to each of the block partitions or sub-blocks.
[0136] The inter prediction unit 244 can include a motion estimation (ME) unit (not shown in FIG. 2) and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit is configured to receive or obtain, for performing motion estimation, a picture block 203 (the current picture block 203 of the current picture 201) and the decoded picture ャ3 1, or at least one or more previously reconstructed blocks, for example, a previously decoded picture ャ3 1 and one or more other reconstructed blocks different from it. For example, the video sequence may include the current picture and a previously decoded picture 31. In other words, the current picture and the previously decoded picture 31 may be part of a sequence of pictures forming the video sequence or a picture sequence.
[0137] For example, the encoder 20 may be configured to select reference blocks from multiple reference blocks of the same picture or different pictures in multiple other pictures, and provide a reference picture (or reference picture index), and / or an offset (spatial offset) between the position (X-Y coordinates) of the reference block and the position of the current block to a motion estimation unit (not shown in FIG. 2) as an inter-prediction parameter. This offset is also referred to as a motion vector (MV).
[0138] The motion compensation unit is configured to, for example, obtain, for example, receive, inter-prediction parameters, and perform inter-prediction based on or by using the inter-prediction parameters to obtain an inter-prediction block 245. Motion compensation performed by a motion compensation unit (not shown in FIG. 2) may include fetching or generating a prediction block based on a motion / block vector determined by motion estimation (possibly including performing interpolation with sub-sample accuracy). During interpolation filtering, additional samples are generated from known samples, thereby potentially increasing the amount of candidate prediction blocks that may be used to encode a picture block. Once the motion vector used for the PU of the current picture block is received, the motion compensation unit 246 can position the prediction block indicated by the motion vector within the reference picture list. The motion compensation unit 246 is further capable of generating syntax elements related to the block and video slice, such that as a result, the video decoder 30 uses the syntax elements when decoding the picture block of the video slice.
[0139] The intra prediction unit 254 is configured to obtain, for example, receive, a picture block 203 (current picture block) of the same picture and one or more previous reconstructed blocks such as reconstructed adjacent blocks, and perform an intra estimation. For example, the encoder 20 can be configured to select an intra prediction mode from a plurality of (predetermined) intra prediction modes.
[0140] An embodiment of the encoder 20 may be configured to select an intra prediction mode based on an optimization criterion, for example, based on a minimum residual (e.g., an intra prediction mode that provides a prediction block 255 most similar to the current picture block 203) or a minimum rate distortion.
[0141] The intra prediction unit 254 is further configured to determine an intra prediction block 255 based on the intra prediction parameters of the selected intra prediction mode. In any case, after selecting the intra prediction mode to be used for a block, the intra prediction unit 254 is further configured to provide the intra prediction parameters to the entropy encoding unit 270, that is, to provide information indicating the selected intra prediction mode to be used for the block. In one example, the intra prediction unit 254 may be configured to perform any combination of the following intra prediction techniques.
[0142] Entropy encoding unit 270 applies an entropy encoding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, arithmetic coding scheme, context adaptive binary arithmetic coding (CABAC) scheme, syntax-based context-adaptive binary arithmetic coding (SBAC) scheme, probability interval partitioning entropy (PIPE) coding scheme, or other entropy encoding method or technique) to one or more of the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters (or does not apply to any of them), and obtains encoded picture data 21 that can be output in, for example, an encoded bitstream ムの form. The encoded bitstream may be transmitted to video decoder 30 or stored for later transmission or retrieval by video decoder 30. Entropy encoding unit 270 may be further configured to perform entropy encoding with respect to another syntax element of the currently encoded video slice.
[0143] Another structural variation of video encoder 20 may be configured to encode a video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal without using the transform processing unit 206 for some blocks or frames. In another implementation, encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 coupled to one unit.
[0144] Figure 3 shows an example of a video decoder 30 configured to implement the technology in this application. The video decoder 30 receives the encoded picture data (e.g., an encoded bitstream) 21 encoded by the encoder 20 or the like, and is configured to obtain the decoded picture ャ3 1. In the decoding process, the video decoder 30 receives video data, e.g., an encoded video bitstream indicating the picture blocks of the encoded video slice and the related syntax elements, from the video encoder 20.
[0145] In the example of Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded picture buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 30 is capable of performing a decoding traversal that is generally inverse to the encoding traversal described with reference to the video encoder 20 of Figure 2.
[0146] The entropy decoding unit 304 performs entropy decoding on the encoded picture data 21, and is configured to obtain any one or all of the quantization coefficients 309, the decoded coding parameters (not shown in Figure 3), etc., and / or, for example, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements (being decoded). The entropy decoding unit 304 is further configured to transfer the inter prediction parameters, intra prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 30 is capable of receiving syntax elements at the video slice level and / or syntax elements at the video block level.
[0147] The inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transformation processing unit 312 may have the same function as the inverse transformation processing unit 212, the reconstruction unit 314 may have the same function as the reconstruction unit 214, the buffer 316 may have the same function as the buffer 216, the loop filter 320 may have the same function as the loop filter 220, and the decoded picture buffer 330 may have the same function as the decoded picture buffer 230.
[0148] The prediction processing unit 360 can include an inter prediction unit 344 and an intra prediction unit 354. The inter prediction unit 344 may have the same function as the inter prediction unit 244, and the intra prediction unit 354 may have the same function as the intra prediction unit 254. The prediction processing unit 360 generally performs block prediction and / or obtains a prediction block 365 from the encoded data 21, and is configured to receive or obtain prediction-related parameters and / or information regarding the selected prediction mode, for example, from the entropy decoding unit 304 (explicitly or implicitly).
[0149] When a video slice is encoded as intra-coded (I), the intra prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 to be used for a picture block of the current video slice based on data from previously decoded blocks of the current frame or picture and the signaled intra prediction mode. When a video frame is encoded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 to be used for a video block of the current video slice based on the motion vectors received from the entropy decoding unit 304 and other syntax elements. For inter prediction, the prediction block may be generated from one of the reference pictures within one reference picture list. The video decoder 30 may configure the reference frame lists of list 0 and list 1 by using a default configuration technique based on the reference pictures stored in the DPB 330.
[0150] The prediction processing unit 360 is configured to determine prediction information to be used for a video block of the current video slice by analyzing the motion vectors and other syntax elements, and use the prediction information to generate a prediction block to be used for the current video block being decoded. For example, the prediction processing unit 360 may use some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to encode a video block of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information of one or more pictures within the reference picture list used for the slice, the motion vector of each inter-coded video block used for the slice, the inter prediction state of each inter-coded video block used for the slice, and other information, to decode the video block of the current video slice.
[0151] The inverse quantization unit 310 may be configured to perform inverse quantization (i.e., dequantization) on the quantized transform coefficients decoded by the entropy decoding unit 304 provided in the bit stream. The inverse quantization process may include determining the degree of quantization to be applied and the degree of inverse quantization to be applied using the quantization parameters calculated by the video encoder 20 for each video block within the video slice.
[0152] The inverse transform processing unit 312 is configured to apply an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to generate a residual block in the sample domain.
[0153] The reconstruction unit 314 (e.g., adder 314) is configured to add the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365 to obtain a reconstructed block 315 in the sample domain, for example, by adding the sample values of the reconstructed residual block 313 to the sample values of the prediction block 365.
[0154] The loop filter unit 320 (within or after the encoding loop) is configured to filter the reconstructed block 315 to obtain a filtered block 321, perform sample conversion smoothly, or improve video quality. In one example, the loop filter unit 320 can be configured to perform any combination of the following filtering techniques. The loop filter unit 320 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a collaborative filter. The loop filter unit 320 is shown in FIG. 3 as an in-loop filter, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0155] Within a given frame or picture Filtering The obtained block 321 is stored in the decoded picture buffer 330 that stores the reference pictures used for subsequent motion compensation.
[0156] The decoder 30 is configured to output the decoded picture 31 by using the output 332, etc., and present the decoded picture 31 to the user or provide the decoded picture 31 for the user to view.
[0157] Another variation of the video decoder 30 can be configured to decode a compressed bitstream. For example, the decoder 30 can generate an output video stream without the loop filter unit 320. For example, the non-transform-based decoder 30 can directly dequantize the residual signal for some blocks or frames without the inverse transform processing unit 312. In another implementation, the video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 coupled to one unit.
[0158] FIG. 4 is a diagram showing an example of a video coding system 40 including the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3 according to an exemplary embodiment. The system 40 can implement combinations of various techniques of the present application. In the illustrated implementation, the video coding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video デ coder implemented by the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0159] As shown in the drawings, the imaging device 41, the antenna 42, the processing device 46, the logic circuit 47, the video encoder 20, the video decoder 30, the processor 43, the memory 44, and / or the display device 45 can communicate with each other. As described, the video coding system 40 is shown with both the video encoder 20 and the video decoder 30, but in different examples, the video coding system 40 may include only the video encoder 20 or only the video decoder 30.
[0160] In some examples, as shown in the drawings, the video coding system 40 may include an antenna 42. For example, the antenna 42 may be configured to transmit or receive an encoded bitstream of video data. Further, in some examples, the video coding system 40 may include a display device 45. The display device 45 may be configured to present video data. In some examples, as shown in the drawings, the logic circuit 47 may be implemented by a processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, etc. The video coding system 40 may also include an optional processor 43. The optional processor 43 may similarly include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, etc. In some examples, the logic circuit 47 may be implemented by hardware such as video encoding dedicated hardware, and the processor 43 may be implemented by universal software, an operating system, etc. Further, the memory 44 may be any type of memory, such as volatile memory (e.g., Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM)) or non-volatile memory (e.g., flash memory). In a non-limiting example, the memory 44 may be implemented by cache memory. In some examples, the logic circuit 47 may be able to access the memory 44 (e.g., to implement an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 may include a memory (e.g., a cache) for implementing an image buffer or the like.
[0161] In some examples, a video encoder 20 implemented by a logic circuit can include an image buffer (e.g., implemented by a processing unit 46 or a memory 44) and a graphics processing unit (e.g., implemented by a processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit may include a video encoder 20 implemented by a logic circuit 47 to implement various modules described with reference to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit can be configured to perform various operations described herein.
[0162] The video decoder 30 may be similarly implemented by a logic circuit 47 to implement various modules described with reference to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, a video decoder 30 implemented by a logic circuit can include an image buffer (implemented by a processing unit 46 or a memory 44) and a graphics processing unit (e.g., implemented by a processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit may include a video decoder 30 implemented by a logic circuit 47 to implement various modules described with reference to FIG. 3 and / or any other decoder system or subsystem described herein.
[0163] In some examples, the antenna 42 of the video coding system 40 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data, indicators, index values, mode selection data, etc. related to the video frame coding described herein, e.g., data related to coded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as described), and / or data defining the coded partitions). The video coding system 40 may further include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present video frames.
[0164] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as any one or two of the source device 12 and the destination device 14 of FIG. 1 according to an exemplary embodiment. The apparatus 500 can implement the technology in this application. The apparatus 500 can use a form of a computing system including a plurality of computing devices, or can use a form of a single computing device such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, or a desktop computer.
[0165] The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be some other type of existing or future device, or a device capable of controlling or processing information. As shown in the drawings, the disclosed implementation can be implemented by using a single processor such as the processor 502, but by using more than one processor, advantages in speed and efficiency can be achieved.
[0166] In an implementation, the memory 504 within the device 500 may be a read only memory (ROM) ) or or a random access memory (RAM). ) Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that is accessed by the processor 502 by using a bus 512. The memory 504 may further include an operating system 508 and an application program 510. The application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 may include applications 1 through N, and applications 1 through N may further include video encoding applications for executing the methods described herein. The device 500 may further include additional memory in the form of a secondary memory 514. The secondary memory 514 may be, for example, a memory card used with a mobile computing device. Since a video communication session may include a large amount of information, the information may be stored completely or partially in the secondary memory 514 and loaded into the memory 504 for processing as needed.
[0167] Device 500 may further include one or more output devices, such as display 518. In one example, display 518 may be a touch-sensitive display that combines a touch-sensing element that is operable to sense touch input with a display. Display 518 can be coupled to processor 502 by using bus 512. In addition to display 518, another output device that enables a user to program device 500 or otherwise use device 500 may be further provided, or another output device may be provided in place of display 518. If the output device is a display or includes a display, the display may be implemented differently, for example, by using a light-emitting diode (LED) display such as a liquid crystal display (LCD), a cathode-ray tube (CRT) display, a plasma display, or an organic light-emitting diode (OLED) display.
[0168] Device 500 may further include, or be connected to, an image sensing device 520. Image sensing device 520 is, for example, a camera or some other existing or future image sensing device 520 capable of sensing an image. The image is, for example, an image of the user operating device 500. Image sensing device 520 can be positioned facing directly the user operating device 500. In one example, the position and optical axis of image sensing device 520 can be configured such that the field of view of image sensing device 520 includes an area adjacent to display 518 and display 518 is visible from that area.
[0169] Device 500 may further include, or be connected to, a sound sensing device 522. The sound sensing device 522 is, for example, a microphone or any other existing or future sound sensing device capable of sensing sound in the vicinity of the device 500. The sound sensing device 522 can be arranged to face directly the user who moves the device 500 and may be configured to receive sounds such as voices or other sounds generated by the user when the user moves the device 500.
[0170] Although the processor 502 and the memory 504 of the device 500 are integrated into one unit as shown in FIG. 5, other configurations can be used. The operation of the processor 502 may be distributed among a plurality of machines (each machine having one or more processors) that can be directly coupled, or may be distributed in a local area or another network. The memory 504 may also be distributed among a plurality of machines such as network-based memory and memory in a plurality of machines that operate the device 500. Although a single bus is depicted here, a plurality of buses 512 of the device 500 may exist. Further, the secondary memory 514 may be directly coupled to other components of the device 500, or may be accessed via a network, and may include a single integrated unit such as a memory card, or a plurality of units such as a plurality of memory cards. Therefore, the device 500 can be realized in a plurality of configurations.
[0171] Hereinafter, the concept of the present application will be described.
[0172] 1. Inter prediction mode
[0173] In HEVC, two inter prediction modes: advanced motion vector prediction (AMVP) mode and merge mode are used.
[0174] In the AMVP mode, first, the encoded blocks that are spatially or temporally adjacent to the current block (referred to as adjacent blocks) are traversed, and based on the motion information of the adjacent blocks, a candidate motion vector list (also referred to as a motion information candidate list) is constructed. Next, based on the rate-distortion cost, an optimal motion vector is determined from the candidate motion vector list, and the candidate motion information having the minimum rate-distortion cost is used as the motion vector predictor (MVP) of the current block. Both the position of the adjacent block and its traversal order are predefined. The rate-distortion cost is calculated according to Equation (1), where J represents the rate-distortion cost (RD cost), SAD is the sum of absolute differences (SAD) between the original sample values and the predicted sample values obtained by motion estimation using the candidate motion vector predictor, R represents the bit rate, and λ represents the Lagrange multiplier. The encoder side transfers the index of the selected motion vector predictor in the candidate motion vector list and the index value of the reference frame to the decoder side. Further, in order to obtain the actual motion vector of the current block, a motion search is performed in the vicinity centered on the MVP. The encoder side transfers the difference between the MVP and the actual motion vector (motion vector difference) to the decoder side. J = SAD + λR (1)
[0175] In the merge mode, first, a candidate motion vector list is constructed based on the motion information of the encoded blocks that are spatially or temporally adjacent to the current block. Then, based on the rate-distortion cost, optimal motion information is determined from the candidate motion vector list so as to serve as the motion information of the current block. Next, the index value of the position of the optimal motion information in the candidate motion vector list (hereinafter referred to as the merge index) is transferred to the decoder side. The spatial candidate motion information and the temporal candidate motion information of the current block are shown in FIG. 6. The spatial candidate motion information is from five spatially adjacent blocks (A0, A1, B0, B1, B2). When the adjacent blocks are not available (when there are no adjacent blocks, when the adjacent blocks are not encoded, or when the prediction mode used for the adjacent blocks is not the inter prediction mode), the motion information of the adjacent blocks is not added to the candidate motion vector list. The temporal candidate motion information of the current block is obtained by scaling the MV of the block at the corresponding position in the reference frame based on the reference frame and the picture order count (POC) of the current frame. First, it is determined whether the block at the position of T in the reference frame is available. If the block is not available, the block at the position of C is selected.
[0176] Similar to the AMVP mode, in the merge mode, both the positions of the adjacent blocks and their traversal order are determined in advance. Furthermore, the positions of the adjacent blocks and their traversal order may vary depending on the mode.
[0177] It can be seen that the candidate motion vector list must be maintained in both the AMVP mode and the merge mode. Each time before new motion information is added to the candidate list, it is first checked whether the same motion information exists in the list. If the same motion information exists, the motion information is not added to the list. This inspection process is referred to as trimming the candidate motion vector list. Trimming the list is to avoid the same motion information in the list, thereby avoiding redundant rate-distortion cost calculations.
[0178] In HEVC inter prediction, all samples within a coding block use the same motion information. Therefore, in order to obtain the predictor of the samples in the coding block, motion compensation is performed based on the motion information. However, in a coding block, not all samples have the same motion characteristics. Using the same motion information may result in inaccurate motion compensation prediction and more residual information.
[0179] In existing video coding standards, block matching motion estimation based on a translational motion model is used, and it is assumed that the motion of all samples within a block is consistent. However, in the real world, various motions exist. Many objects are in non-translational motion, such as rotating objects, roller coasters rotating in various directions, fireworks displays, and some stunts in movies, especially moving objects in UGC scenarios. For these moving objects, when the block motion compensation technology based on the translational motion model in the existing coding standard is used for coding, the coding efficiency may be greatly affected. Therefore, in order to further improve the coding efficiency, a non-translational motion model such as an affine motion model is introduced.
[0180] Therefore, for various motion models, the AMVP mode can be classified into a translational model-based AMVP mode and a non-translational model-based AMVP mode, and the merge mode can be classified into a translational model-based merge mode and a non-translational model-based merge mode.
[0181] 2. Non-translational motion model
[0182] In non-parallel motion model-based prediction, on the codec side, one motion model is used to derive the motion information of each child motion compensation unit within the current block, and motion compensation is performed based on the motion information of the child motion compensation units to obtain the prediction block, thereby improving the prediction efficiency. A general non-parallel motion model is a 4-parameter affine motion model or a 6-parameter affine motion model.
[0183] The child motion compensation unit in the embodiment of the present application may be a sample or an N1×N2 sample block obtained by splitting according to a specific method, where both N1 and N2 are positive integers, and N1 may or may not be equal to N2.
[0184] The 4-parameter affine motion model is expressed as in Equation (2):
Equation
[0185] The 4-parameter affine motion model can be represented by the motion vectors of two samples and the coordinates of the two samples with respect to the top-left sample of the current block. The samples used to represent the motion model parameters are called control points. When the top-left sample (0, 0) and the top-right sample (W, 0) are used as control points, the motion vectors (vx0, vy0) and (vx1, vy1) of the top-left control point and the top-right control point of the current block are first determined. Then, the motion information of each child motion compensation unit in the current block is obtained according to Equation (3), where (x, y) is the coordinate of the child motion compensation unit with respect to the top-left sample of the current block, and W represents the width of the current block.
Equation
[0186] The 6-parameter affine motion model is expressed as in Equation (4):
Number
[0187] The 6-parameter affine motion model can be represented by the motion vectors of three samples and the coordinates of the three samples with respect to the top-left sample of the current block. When the top-left sample (0, 0), the top-right sample (W, 0), and the bottom-left sample (0, H) are used as control points, the motion vectors (vx0, vy0), (vx1, vy1), and (vx2, vy2) of the top-left control point, the top-right control point, and the bottom-left control point of the current block are first determined. Then, the motion information of each child motion compensation unit in the current block is obtained according to Equation (5), where (x, y) is the coordinate of the child motion compensation unit with respect to the top-left sample of the current block, and W and H represent the width and height of the current block, respectively.
Number
[0188] The coding block predicted by using the affine motion model is referred to as an affine coding block.
[0189] Generally, the motion information of the control points of an affine coding block can be obtained by using the affine motion model-based Advanced Motion Vector Prediction (AMVP) mode or the affine motion model-based Merge mode.
[0190] The motion information of the control points of the current coding block can be obtained by using the inheritance control point motion vector prediction method or the construction control point motion vector prediction method.
[0191] 3. Inheritance Control Point Motion Vector Prediction Method
[0192] The inheritance control point motion vector prediction method determines the candidate control point motion vector of the current block by using the motion models of adjacent encoded affine coding blocks.
[0193] The current block shown in FIG. 7 is used as an example. The adjacent blocks around the current block are traversed in a specified order, for example, A1→B1→B0→A0→B2, to find the affine coding blocks where the adjacent blocks of the current block are located, and obtain the control point motion information of the affine coding blocks. Further, the control point motion vector (in the case of the merge mode) or the control point motion vector predictor (in the case of the AMVP mode) of the current block is derived using the motion model constructed by using the control point motion information of the affine coding blocks. The order of A1→B1→B0→A0→B2 is used only as an example. Other combinations of orders are also applicable to this application. Further, the adjacent blocks are not limited to A1, B1, B0, A0, and B2.
[0194] The adjacent blocks may be pre-set sized samples or sample blocks obtained based on a specific splitting method, for example, 4×4 sample blocks, 4×2 sample blocks, or sample blocks of another size. This is not limited.
[0195] Hereinafter, the determination process by using A1 as an example will be described. The same applies to other cases.
[0196] As shown in FIG. 7, when the coding block where A1 is located is a 4-parameter affine coding block, the motion vectors (vx4, vy4) of the upper left sample (x4, y4) and the motion vectors (vx5, vy5) of the upper right sample (x5, y5) of the affine coding block are obtained. The motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current affine coding block is calculated according to Equation (6), and the motion vector (vx1, vy1) of the upper right sample (x1, y1) of the current affine coding block is calculated according to Equation (7). [Number]
[0197] The combination of the motion vector (vx0, vy0) of the upper left sample (x0, y0) and the motion vector (vx1, vy1) of the upper right sample (x1, y1) of the current block, obtained based on the affine coding block where A1 is located, is the candidate control point motion vector of the current block.
[0198] When the coding block where A1 is located is a 6-parameter affine coding block, the motion vectors (vx4, vy4) of the upper left sample (x4, y4), the motion vectors (vx5, vy5) of the upper right sample (x5, y5), and the motion vectors (vx6, vy6) of the lower left sample (x6, y6) of the affine coding block are obtained. The motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current block is calculated according to Equation (8), the motion vector (vx1, vy1) of the upper right sample (x1, y1) of the current block is calculated according to Equation (9), and the motion vector (vx2, vy2) of the lower left sample (x2, y2) of the current block is calculated according to Equation (10). [Number]
[0199] Based on the affine coding block where A1 is located, the combination of the motion vectors (vx0, vy0) of the upper left sample (x0, y0), the motion vectors (vx1, vy1) of the upper right sample (x1, y1), and the motion vectors (vx2, vy2) of the lower left sample (x2, y2) of the current block is the candidate control point motion vector of the current block.
[0200] It should be noted that other motion models, candidate positions, search and traversal orders are also applicable to this application. Details are not described in the embodiments of this application.
[0201] It should also be noted that the method of using other control points to represent the motion models of adjacent and current coding blocks is also applicable to this application. Details are not described here.
[0202] 4. Constructed control point motion vectors Prediction method 1
[0203] The constructed control point motion vector prediction method combines the motion vectors of adjacent coded blocks around the control point of the current block without considering whether the adjacent coded blocks are affine coding blocks, and uses the combined motion vectors as the control point motion vectors of the current affine coding block.
[0204] The motion vectors of the upper left sample and the upper right sample of the current block are determined by using the motion information of adjacent coded blocks around the current coding block. Figure 8A is used as an example to illustrate the constructed control point motion vector prediction method. It should be noted that Figure 8A is just an example.
[0205] As shown in FIG. 8A, the motion vectors of the adjacent coded blocks A2, B2, and B3 of the upper left sample are used as candidate motion vectors for the motion vector of the upper left sample of the current block, and the motion vectors of the adjacent coded blocks B1 and B0 of the upper right sample are used as candidate motion vectors for the motion vector of the upper right sample of the current block. The candidate motion vectors of the upper left sample and the upper right sample are combined to form a plurality of 2-tuples. The motion vectors of the two coded blocks included in the 2-tuple can be used as the candidate control point motion vector of the current block as shown in the following formula (11A):
Number
[0206] v A2 indicates the motion vector of A2, and v B1 indicates the motion vector of B1, and v B0 indicates the motion vector of B0, and v B2 indicates the motion vector of B2, and v B3 indicates the motion vector of B3.
[0207] As shown in FIG. 8A, the motion vectors of the adjacent coded blocks A2, B2, and B3 of the upper left sample are used as candidate motion vectors for the motion vector of the upper left sample of the current block, the motion vectors of the adjacent coded blocks B1 and B0 of the upper right sample are used as candidate motion vectors for the motion vector of the upper right sample of the current block, and the motion vectors of the adjacent coded blocks A0 and A1 of the lower left sample are used as candidate motion vectors for the motion vector of the lower left sample of the current block. The candidate motion vectors of the upper left sample, the upper right sample, and the lower left sample are combined to form a 3-tuple. The motion vectors of the three coded blocks included in the 3-tuple can be used as the candidate control point motion vector of the current block as shown in the following formulas (11B) and (11C):
Number
[0208] v A2 indicates the motion vector of A2, and v B1 indicates the motion vector of B1, and v B0 indicates the motion vector of B0, and v B2 indicates the motion vector of B2, and v B3 indicates the motion vector of B3, and v A0 indicates the motion vector of A0, and v A1 indicates the motion vector of A1.
[0209] It should be noted that other methods of combining the motion vectors of the control points can also be applied to the present application. Details are not described here.
[0210] It should also be noted that other methods of using other control points to represent the motion models of adjacent and current coding blocks can also be applied to the present application. Details are not described here.
[0211] 5. Method 2 for predicting constructed control point motion vectors: For this, refer to FIG. 8.
[0212] Step 801: Obtain the motion information of the control points of the current block.
[0213] For example, in FIG. 8A, CPk (k = 1, 2, 3, or 4) indicates the k-th control point, A0, A1, A2, B0, B1, B2, and B3 are spatially adjacent positions of the current block, and are used to predict CP1, CP2, or CP3, T is a temporally adjacent position of the current block, and is used to predict CP4.
[0214] Assume that the coordinates of CP1, CP2, CP3, and CP4 are (0, 0), (W, 0), (H, 0), (W, H) respectively, where W and H indicate the width and height of the current block.
[0215] The movement information of each control point is obtained in the following order:
[0216] (1) For CP1, the inspection order is B2 → A2 → B3. If B2 is available, the movement information of B2 is used. If B2 is not available, A2 and B3 are inspected. If the movement information of all three positions is unavailable, the movement information of CP1 cannot be obtained.
[0217] (2) For CP2, the inspection order is B0 → B1. If B0 is available, the movement information of B0 is used for CP2. If B0 is not available, B1 is inspected. If the movement information of both positions is unavailable, the movement information of CP2 cannot be obtained.
[0218] (3) For CP3, the inspection order is A0 → A1.
[0219] (4) For CP4, the movement information of T is used.
[0220] Here, the fact that X is available means that the block at the position of X (where X is A0, A1, A2, B0, B1, B2, B3, or T) is encoded and the inter prediction mode is used. Otherwise, the position of X is not available.
[0221] It should be noted that other methods for obtaining the movement information of the control point can also be applied to this application. Details are not described here.
[0222] Step 802: Combine the movement information of the control points to obtain the constructed control point movement information.
[0223] The movement information of two control points is combined to form a 2-tuple, and a 4-parameter affine movement model is constructed. The ways to combine two control points may be {CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, or {CP3, CP4}. For example, the 4-parameter affine movement model constructed by using a 2-tuple including control points CP1 and CP2 may be denoted as Affine(CP1, CP2).
[0224] The movement information of three control points is combined to form a 3-tuple, and a 6-parameter affine movement model is constructed. The ways to combine three control points may be {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, or {CP1, CP3, CP4}. For example, the 6-parameter affine movement model constructed by using a 3-tuple including control points CP1, CP2, and CP3 may be denoted as Affine(CP1, CP2, CP3).
[0225] The movement information of four control points is combined to form a 4-tuple, and an 8-parameter bilinear model is constructed. The 8-parameter bilinear model constructed by using a 4-tuple including control points CP1, CP2, CP3, and CP4 may be denoted as “Bilinear”(CP1, CP2, CP3, CP4).
[0226] In this embodiment of the present application, for the sake of simplicity of explanation, the combination of the movement information of two control points (or two coding blocks) is simply referred to as a 2-tuple, the combination of the movement information of three control points (or 3 three coding blocks) is simply referred to as a 3-tuple, and the combination of the movement information of four control points (or four coding blocks) is simply referred to as a 4-tuple.
[0227] These models are traversed in a pre-set order. If the movement information of the control points corresponding to the combined model is not available, the model is considered unavailable. Otherwise, the reference frame index of the model is determined and the movement vector of the control point is scaled. If the scaled movement information of all control points is consistent, the model is invalid. If the movement information of all control points that control the model is available and the model is valid, the movement information of the control points that make up the model is added to the movement information candidate list.
[0228] The control point movement vector scaling method is shown in Equation (12):
Equation
[0229] CurPoc indicates the POC number of the current frame, DesPoc indicates the POC number of the reference frame of the current block, SrcPoc indicates the POC number of the reference frame of the control point, and MV s indicates the movement vector obtained by scaling, and MV indicates the movement vector of the control point.
[0230] It should be noted that combinations of different control points may be converted into control points at the same position.
[0231] For example, a 4-parameter affine motion model obtained by combinations such as {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3} or {CP3, CP4} is converted into an expression by {CP1, CP2} or {CP1, CP2, CP3}. The conversion method is to substitute the movement vector and coordinate information of the control point into Equation (2) to obtain the model parameters, and then substitute the coordinate information of {CP1, CP2} into Equation (3) to obtain the movement vector.
[0232] More directly, the conversion can be performed according to the following equations (13)-(21), where W represents the width of the current block and H represents the height of the current block. In equations (13)-(21), (vx0, vy0) represents the motion vector of CP1, (vx1, vy1) represents the motion vector of CP2, (vx2, vy2) represents the motion vector of CP3, and (vx3, vy3) represents the motion vector of CP4.
[0233] {CP1, CP2} can be converted to {CP1, CP2, CP3} according to the following equation (13). In other words, the motion vector of CP3 in {CP1, CP2, CP3} can be determined by equation (13):
Number
[0234] {CP1, CP3} can be converted to {CP1, CP2} or {CP1, CP2, CP3} according to the following equation (14):
Number
[0235] {CP2, CP3} can be converted to {CP1, CP2} or {CP1, CP2, CP3} according to the following equation (15):
Number
[0236] {CP1, CP4} can be converted to {CP1, CP2} or {CP1, CP2, CP3} according to the following equation (16) or (17):
Number
[0237] {CP2, CP4} can be converted to {CP1, CP2} by the following formula (18), and {CP2, CP4} can be converted to {CP1, CP2, CP3} by the following formulas (18) and (19):
Number
[0238] {CP3, CP4} can be converted to {CP1, CP2} by the following formula (20), and {CP3, CP4} can be converted to {CP1, CP2, CP3} by the following formulas (20) and (21):
Number
[0239] For example, the 6-parameter affine motion model obtained by a combination of {CP1, CP2, CP4}, {CP2, CP3, CP4}, or {CP1, CP3, CP4} is converted to the representation by {CP1, CP2, CP3}. The conversion method is to substitute the motion vector and coordinate information of the control points into formula (4) to obtain the model parameters, and then substitute the coordinate information of {CP1, CP2, CP3} into formula (5) to obtain the motion vector.
[0240] More directly, the conversion can be performed according to the following formulas (22)-(24), where W represents the width of the current block and H represents the height of the current block. In formulas (13)-(21), (vx0, vy0) represents the motion vector of CP1, (vx1, vy1) represents the motion vector of CP2, (vx2, vy2) represents the motion vector of CP3, and (vx3, vy3) represents the motion vector of CP4.
[0241] {CP1, CP2, CP4} can be converted to {CP1, CP2, CP3} by the following formula (22):
Number
[0242] {CP2, CP3, CP4} can be converted to {CP1, CP2, CP3} by the following formula (23):
Number
[0243] {CP1, CP3, CP4} can be converted to {CP1, CP2, CP3} by the following formula (24):
Number
[0244] 6. Affine Motion Model - Based Advanced Motion Vector Prediction Mode (Affine AMVP mode)
[0245] (1) Construct a candidate motion vector list
[0246] The candidate motion vector list for the Affine Motion Model - based AMVP mode is constructed by using the inheritance control point motion vector prediction method and / or the construction control point motion vector prediction method. In this embodiment of the present application, the candidate motion vector list for the Affine Motion Model - based AMVP mode may be referred to as a control point motion vectors predictor candidate list. The motion vector predictor for each control point includes the motion vectors of two control points (of the 4 - parameter affine motion model) or the motion vectors of three control points (of the 6 - parameter affine motion model).
[0247] Optionally, the control point motion vectors predictor candidate list can be pruned, sorted according to specific rules, and truncated or padded by a specific amount.
[0248] (2) Determine the optimal control point motion vector predictor
[0249] On the encoder side, for each child motion compensation unit within the current coding block, the motion vector is obtained based on each control point motion vector predictor in the control point motion vector predictor list by using Equation (3) / (5). The sample value at the corresponding position in the reference frame indicated by the motion vector of each child motion compensation unit is obtained, and the sample value is used as a predictor for performing motion compensation by using an affine motion model. The average difference between the original value and the predictor of each sample in the current coding block is calculated. The control point motion vector predictor corresponding to the minimum average difference is selected as the optimal control point motion vector predictor and used as the motion vector predictor for two / three control points of the current coding block. The index number representing the position of the control point motion vector predictor in the control point motion vector predictor candidate list is encoded in the bitstream and transmitted to the decoder.
[0250] On the decoder side, the index number is analyzed, and based on the index number, the control point motion vector predictor (CPMVP) is determined from the control point motion vector predictor candidate list.
[0251] (3) Determine the control point motion vector
[0252] On the encoder side, to obtain the control point motion vectors (CPMV), the control point motion vector predictor is used as the search start point for motion search within a specific search range. The difference between the control point motion vector and the control point motion vector predictor (CPMVD) is transferred to the decoder side.
[0253] On the decoder side, the control point motion vector difference is analyzed and added to the control point motion vector predictor to obtain the control point motion vector.
[0254] 7. Affine Merge mode
[0255] By using the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method, a control point motion vectors merge candidate list is constructed.
[0256] Optionally, the control point motion vector predictor candidate list can be pruned and sorted according to specific rules and truncated or padded by a specific amount.
[0257] On the encoder side, the motion vectors of each child motion compensation unit (N1×N2 sample block or sample obtained by splitting according to a specific method) within the current coding block are obtained based on each control point motion vector in the merge candidate list by using Equation (3) / (5). The sample values at the positions within the reference frame indicated by the motion vectors of each child motion compensation unit are obtained, and the sample values are used as predictors for performing affine motion compensation. The average difference between the original value and the predictor of each sample in the current coding block is calculated. The control point motion vector corresponding to the minimum average difference is selected as the motion vector of two / three control points of the current coding block. The index number representing the position of the control point motion vector in the candidate list is encoded in the bitstream and sent to the decoder.
[0258] On the decoder side, the index number is analyzed, and based on the index number, control point motion vectors (CPMV) are determined from the control point motion vector merge candidate list.
[0259] Note that in this application, "at least one" means one or more, and "a plurality of" means two or more. It should be noted that "and / or" describes the associative relationship for explaining related objects, indicating that there may be three relationships. For example, A and / or B may represent the following cases: only A exists, both A and B exist, or only B exists, where A and B may be singular or plural. The character " / " generally represents the "or" relationship between related objects. "At least one of the following items (pieces)" or similar expressions indicate any combination of these items, including a single item (piece) or any combination of multiple items (pieces). For example, at least one of a, b, or c may indicate a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c may be singular or plural.
[0260] In this application, when the inter prediction mode is used to decode the current block, syntax elements may be used to signal the inter prediction mode.
[0261] For a part of the currently used syntax structure of the inter prediction mode used for analyzing the current block, refer to Table 1. It should be noted that the syntax elements in the syntax structure may alternatively be represented by other identifiers. This is not specifically limited in this application. Table 1
Table 1
[0262] The syntax element merge_flag[x0][y0] can be used to indicate whether the merge mode is currently used for the current block. For example, when merge_flag[x0][y0]=1, it indicates that the merge mode is currently used for the current block, and when merge_flag[x0][y0]=0, it indicates that the merge mode is not currently used for the current block, where x0 and y0 indicate the coordinates of the current block in the video picture.
[0263] The variable allowAffineMerge can be used to indicate whether the current block meets the conditions for using the affine motion model-based merge mode. For example, when allowAffineInter=0, it indicates that the conditions for using the affine motion model-based merge mode are not met, and when allowAffineInter=1, it indicates that the conditions for using the affine motion model-based merge mode are met. The conditions for using the affine motion model-based merge mode may be that both the width and height of the current block are 8 or more, where cbWidth indicates the width of the current block and cbHeight indicates the height of the current block. In other words, when cbWidth<8 or cbHeight<8, alloweAffineMerge=0, and when cbWidth≧8 and cbHeight≧8, allowAffineMerge=1.
[0264] The variable allowAffineInter can be used to indicate whether the current block meets the conditions for using the affine motion model - based AMVP mode. For example, when allowAffineInter = 0, it indicates that the conditions for using the affine motion model - based AMVP mode are not met, and when allowAffineInter = 1, it indicates that the conditions for using the affine motion model - based AMVP mode are met. The conditions for using the affine motion model - based AMVP mode may be that both the width and height of the current block are 16 or more. In other words, when cbWidth < 16 or cbHeight < 16, allowAffineMerge = 0, and when cbWidth ≥ 16 and cbHeight ≥ 16, allowAffineMerge = 1.
[0265] The syntax element affine_merge_flag[x0][y0] may be used to indicate whether the affine motion model - based merge mode is used for the current block. The type of the slice_type of the slice where the current block is located is of type P or B. For example, when affine_merge_flag[x0][y0]=1, it indicates that the affine motion model - based merge mode is used for the current block, and when affine_merge_flag[x0][y0]=0, it indicates that the affine motion model - based merge mode is not used for the current block, but the translational motion model - based merge mode may be used.
[0266] The syntax element merge_idx[x0][y0] may be used to indicate the index value of the merge candidate list.
[0267] The syntax element affine_merge_idx[x0][y0] may be used to indicate the index value of the affine merge candidate list.
[0268] The syntax element affine_inter_flag[x0][y0] may be used to indicate whether the AMVP mode based on the affine motion model is used for the current block when the slice in which the current block is located is a P-type slice or a B-type slice. For example, when allowAffineInter = 0, it indicates that the AMVP mode based on the affine motion model is used for the current block, and when allowAffineInter = 1, it indicates that the AMVP mode based on the affine motion model is not used for the current block, but the AMVP mode based on the translational motion model may be used.
[0269] When the slice in which the current block is located is a P-type slice or a B-type slice, the syntax element affine_type_flag[x0][y0] may be used to indicate whether the 6-parameter affine motion model is used to perform motion compensation for the current block. When affine_type_flag[x0][y0] = 0, it indicates that the 6-parameter affine motion model is not used to perform motion compensation for the current block, and only the 4-parameter affine motion model may be used to perform motion compensation, and when affine_type_flag[x0][y0] = 1, it indicates that the 6-parameter affine motion model is used to perform motion compensation for the current block.
[0270] As shown in Table 2, when MotionModelIdc[x0][y0] = 1, it indicates that the 4-parameter affine motion model is used, when MotionModelIdc[x0][y0] = 2, it indicates that the 6-parameter affine motion model is used, and when MotionModelIdc[x0][y0] = 0, it indicates that the translational motion model is used. Table 2
Table 2
[0271] The variables MaxNumMergeCand and MaxAffineNumMrgCand indicate the maximum list length and are used to indicate the maximum length of the constructed candidate motion vector list. inter_pred_idc[x0][y0] is used to indicate the prediction direction, PRED_L1 is used to indicate backward prediction, num_ref_idx_l0_active_minus1 indicates the amount of reference frames in the forward reference frame list, ref_idx_l0[x0][y0] indicates the forward reference frame index value of the current block, mvd_coding(x0,y0,0,0) indicates the first motion vector difference, mvp_l0_flag[x0][y0] indicates the forward MVP candidate list index value, PRED_L0 indicates forward prediction, num_ref_idx_l1_active_minus1 indicates the amount of reference frames in the backward reference frame list, ref_idx_l1[x0][y0] indicates the backward reference frame index value of the current block, and mvp_l1_flag[x0][y0] indicates the backward MVP candidate list index value.
[0272] In Table 1, ae(v) indicates the syntax element encoded by context-based adaptive binary arithmetic coding ( CABAC ).
[0273] The inter prediction process will be described in detail below. For this, refer to FIG. 9A.
[0274] Step 601: Analyze the bitstream based on the syntax structure shown in Table 1 to determine the inter prediction mode of the current block.
[0275] If it is determined that the inter prediction mode of the current block is the AMVP mode based on the affine motion model, step 602a is executed.
[0276] Specifically, when the syntax element merge_flag = 0 and the syntax element affine_inter_flag = 1, it indicates that the inter prediction mode of the current block is the AMVP mode based on the affine motion model.
[0277] If it is determined that the inter prediction mode of the current block is the merge mode based on the affine motion model, step 602b is executed.
[0278] Specifically, when the syntax element merge_flag = 1 and the syntax element affine_merge_flag = 1, it indicates that the inter prediction mode of the current block is based on the affine motion model Merge· mode.
[0279] Step 602a: Create a candidate motion vector list corresponding to the AMVP mode based on the affine motion model, and execute step 603a.
[0280] The candidate control point motion vector of the current block is derived using the inheritance control point motion vector prediction method and / or the construction control point motion vector prediction method, and added to the candidate motion vector list.
[0281] The candidate motion vector list may include a 2-tuple list (when the 4-parameter affine motion model is used for the current coding block) or a 3-tuple list. The 2-tuple list includes one or more 2-tuples used to construct the 4-parameter affine motion model. The 3-tuple list includes one or more 3-tuples used to construct the 6-parameter affine motion model.
[0282] Optionally, the 2-tuple / 3-tuple list of candidate motion vectors can be pruned, sorted according to specific rules, and truncated or padded by a specific amount.
[0283] A1: A process of constructing a candidate motion vector list by using an inheritance control point motion vector prediction method is described.
[0284] Figure 7 is used as an example. The adjacent blocks around the current block are traversed in the order of A1→B1→B0→A0→B2 in Figure 7, an affine coding block where the adjacent blocks are located is found, and the control point motion information of the affine coding block is obtained. Further, the candidate control point motion information of the current block is derived by using a motion model constructed based on the control point motion information of the affine coding block. For details, refer to the related description of the inheritance control point motion vector prediction method in 3. Details are not described here.
[0285] For example, when the affine motion model used for the current block is a four-parameter affine motion model (i.e., MotionModelIdc = 1), and the four-parameter affine motion model is used for an adjacent affine decoding block, the motion vectors of two control points of the affine decoding block: the motion vector (vx4, vy4) of the upper left control point (x4, y4) and the motion vector (vx5, vy5) of the upper right control point (x5, y5) are obtained. The affine decoding block is an affine coding block predicted in the encoding phase by using the affine motion model.
[0286] The motion vectors of the upper left control point and the upper right control point of the current block are respectively derived according to equations (6) and (7) corresponding to the four-parameter affine motion model by using a four-parameter affine motion model including two control points of the adjacent affine decoding block.
[0287] When the 6-parameter affine motion model is used for an adjacent affine decoding block, the motion vectors of the three control points of the adjacent affine decoding block, for example, the motion vector (vx4, vy4) of the upper-left control point (x4, y4) in FIG. 7, the motion vector (vx5, vy5) of the upper-right control point (x5, y5), and the motion vector (vx6, vy6) of the lower-left control point (x6, y6) are obtained.
[0288] The motion vectors of the upper-left control point and the upper-right control point of the current block are respectively derived according to Expressions (8) and (9) corresponding to the 6-parameter affine motion model by using the 6-parameter affine motion model including the three control points of the adjacent affine decoding block.
[0289] For example, the affine motion model used for the current decoding block is a 6-parameter affine motion model (i.e., MotionModelIdc = 2).
[0290] When the affine motion model used for an adjacent affine decoding block is a 6-parameter affine motion model, the motion vectors of the three control points of the adjacent affine decoding block, for example, the motion vector (vx4, vy4) of the upper-left control point (x4, y4) in FIG. 7, the motion vector (vx5, vy5) of the upper-right control point, and the motion vector (vx6, vy6) of the lower-left control point (x6, y6) are obtained.
[0291] The motion vectors of the upper-left control point, the upper-right control point, and the lower-left control point of the current block are respectively derived according to Expressions (8), (9), and (10) corresponding to the 6-parameter affine motion model by using the 6-parameter affine motion model including the three control points of the adjacent affine decoding block.
[0292] When the affine motion model used for adjacent affine decoding blocks is a four-parameter affine motion model, the motion vectors of two control points of the affine decoding block, for example, the motion vector (vx4, vy4) of the upper left control point (x4, y4) and the motion vector (vx5, vy5) of the upper right control point (x5, y5), are obtained.
[0293] The motion vectors of the upper left control point, the upper right control point, and the lower left control point of the current block are respectively derived according to formulas (6) and (7) corresponding to the four-parameter affine motion model by using a four-parameter affine motion model that includes two control points of an adjacent affine decoding block.
[0294] It should be noted that other motion models, candidate positions, and search orders are also applicable to the present application. Details are not described here. It should also be noted that the method of using other control points to represent the motion models of adjacent and current coding blocks is also applicable to the present application. Details are not described here.
[0295] A2: A process of constructing a candidate motion vector list by using a construction control point motion vector prediction method is described.
[0296] For example, when the affine motion model used for the current decoding block is a four-parameter affine motion model (that is, MotionModelIdc = 1), the motion vectors of the upper left sample and the upper right sample of the current coding block are determined by using the motion information of adjacent coding blocks of the current coding block. Specifically, it is possible to construct a candidate motion vector list by using Construction Control Point Motion Vector Prediction Method 1 or Construction Control Point Motion Vector Prediction Method 2. For specific methods, refer to the descriptions in 4 and 5. Details are not described here.
[0297] For example, when the affine motion model used for the current decoding block is a 6-parameter affine motion model (i.e., MotionModelIdc = 2), the motion vectors of the top-left sample, top-right sample, and bottom-left sample of the current coding block are determined by using the motion information of the adjacent coded blocks of the current coding block. Specifically, the candidate motion vector list can be configured by using Construction Control Point Motion Vector Prediction Method 1 or Construction Control Point Motion Vector Prediction Method 2. For specific methods, refer to the descriptions in 4 and 5. Details are not described here.
[0298] It should be noted that another method of combining control point motion information is also applicable to the present application. Details are not described here.
[0299] Step 603a: Analyze the bitstream to determine the optimal control point motion vector predictor and execute Step 604a.
[0300] B1: When the affine motion model used for the current decoding block is a 4-parameter affine motion model (MotionModelIdc = 1), the index number is analyzed, and based on the index number, the optimal motion vector predictors at two control points are determined from the candidate motion vector list.
[0301] For example, the index number is mvp_l0_flag or mvp_l1_flag.
[0302] B2: When the affine motion model used for the current decoding block is a 6-parameter affine motion model (MotionModelIdc = 2), the index number is analyzed, and based on the index number, the optimal motion vector predictors at three control points are determined from the candidate motion vector list.
[0303] Step 604a: Analyze the bitstream to determine the control point motion vectors.
[0304] C1: When the affine motion model used for the current decoded block is a 4-parameter affine motion model (MotionModelIdc = 1), the motion vector differences of the two control points of the current block are obtained from the bitstream by decoding, and the motion vectors of the control points are obtained based on the motion vector differences of the control points and the motion vector predictors. Taking forward prediction as an example, the motion vector differences of the two control points are mvd_coding(x0,y0,0,0) and mvd_coding(x0,y0,0,1), respectively.
[0305] For example, the motion vector differences of the upper left control point and the upper right control point are obtained from the bitstream by decoding, and are respectively added to the motion vector predictors to obtain the motion vectors of the upper left control point and the upper right control point of the current block.
[0306] C2: The affine motion model used for the current decoded block is a 6-parameter affine motion model (MotionModelIdc = 2).
[0307] The motion vector differences of the three control points of the current block are obtained from the bitstream by decoding, and the control point motion vectors are obtained based on the motion vector differences of the control points and the motion vector predictors. Taking forward prediction as an example, the motion vector differences of the three control points are mvd_coding(x0,y0,0,0), mvd_coding(x0,y0,0,1) and mvd_coding(x0,y0,0,2), respectively.
[0308] For example, the motion vector differences of the upper left control point, the upper right control point, and the lower left control point are obtained from the bitstream by decoding, and are respectively added to the motion vector predictors to obtain the motion vectors of the upper left control point, the upper right control point, and the lower left control point of the current block.
[0309] Step 605a: Based on the motion information of the control points used for the current decoding block and the affine motion model, obtain the motion vectors of each sub-block in the current block.
[0310] For each sub-block within the current affine decoding block (one sub-block is equivalent to one motion compensation unit, and the width and height of the sub-block are smaller than the width and height of the current block), the motion information of the samples at the preset positions within the motion compensation unit can be used to represent the motion information of all the samples within the motion compensation unit. Assuming that the size of the motion compensation unit is MxN, the samples at the preset positions may be the central sample (M / 2, N / 2), the upper left sample (0, 0), the upper right sample (M - 1, 0), or samples at other positions within the motion compensation unit. Hereinafter, as an example for explanation, the central sample of the motion compensation unit is used. Referring to FIG. 9C, V0 indicates the motion vector of the upper left control point, and V1 indicates the motion vector of the upper right control point. Each of the small boxes indicates one compensation unit.
[0311] The coordinates of the central sample of the motion compensation unit with respect to the upper left sample of the current affine decoding block are calculated by using Equation (25), where i indicates the i-th motion compensation unit in the horizontal direction (from left to right), j indicates the j-th motion compensation unit in the vertical direction (from top to bottom), and (x (i,j) , y (i,j) ) indicates the coordinates of the central sample of the (i, j)-th motion compensation unit with respect to the upper left sample of the current affine decoding block.
[0312] When the affine motion model used in the current affine decoding block is a 6-parameter affine motion model, (x (i,j) , y (i,j) ) is substituted into Equation (26) corresponding to the 6-parameter affine motion model to obtain the motion vector of the central sample of each motion compensation unit. The motion vector of the central sample is used as the motion vectors (vx (i,j) , vy (i,j) ) of all samples in the motion compensation unit.
[0313] When the affine motion model used in the current affine decoding block is a 4-parameter affine motion model, (x (i,j) , y (i,j) ) is substituted into Equation (27) corresponding to the 4-parameter affine motion model to obtain the motion vector of the central sample of each motion compensation unit. The motion vector of the central sample is used as the motion vectors (vx (i,j) , vy (i,j) ) of all samples in the motion compensation unit.
Number
[0314] Step 606a: Perform motion compensation for each sub-block based on the determined motion vectors of the sub-blocks, and obtain the sample predictors of the sub-blocks.
[0315] Step 602b: Construct a motion information candidate list corresponding to the affine motion model-based merge mode.
[0316] Specifically, the motion information candidate list corresponding to the affine motion model-based merge mode can be constructed by using the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method.
[0317] As an option, the motion information candidate list can be pruned, sorted according to specific rules, and can be truncated or padded to a specific amount.
[0318] D1: A process of constructing a candidate motion vector list by using an inheritance control point motion vector prediction method is described.
[0319] The candidate control point motion information of the current block is derived by using the inheritance control point motion vector prediction method and added to the motion information candidate list.
[0320] The adjacent blocks around the current block are traversed in the order of A1→B1→B0→A0→B2 in FIG. 8A to find the affine coding blocks whose positions are determined, and the control point motion information of the affine coding blocks is obtained. Further, the candidate control point motion information of the current block is derived by using the motion model of the affine coding block.
[0321] If the candidate motion vector list is empty, the candidate control point motion information is added to the candidate list. Otherwise, the motion information in the candidate motion vector list is sequentially traversed to check whether the same motion information as the candidate control point motion information exists in the candidate motion vector list. If the same motion information as the candidate control point motion information does not exist in the candidate motion vector list, the candidate control point motion information is added to the candidate motion vector list.
[0322] To determine whether two pieces of candidate motion information are the same, it is necessary to sequentially determine whether the horizontal and vertical components of each of the forward reference frame, backward reference frame, forward motion vector, and backward motion vector are the same in the two pieces of candidate motion information. The two pieces of candidate motion information are considered different only when all of the above elements are different.
[0323] When the number of motion information in the candidate motion vector list reaches the maximum list length MaxAffineNumMrgCand (MaxAffineNumMrgCand is a positive integer such as 1, 2, 3, 4, or 5. Hereinafter, for the sake of explanation, a length of 5 is used, and the details are not described here), the candidate list is complete. Otherwise, the next adjacent block is traversed.
[0324] D2: The control point motion information of the current block is derived by using the construction control point motion vector prediction method and added to the motion information candidate list. For this, refer to FIG. 9B.
[0325] Step 601c: Obtain the motion information of the control point of the current block. For this, refer to step 801 in the construction control point motion vector prediction method 2 in 5. The details are not described again here.
[0326] Step 602c: Combine the motion information of the control points to obtain the construction control point motion information. For this, refer to step 801 in FIG. 8B. The details are not described again here.
[0327] Step 603c: Add the construction control point motion information to the candidate motion vector list.
[0328] If the length of the candidate list is shorter than the maximum list length MaxAffineNumMrgCand, the combinations are traversed in a preset order to obtain valid combinations as candidate control point movement information. In this case, if the candidate motion vector list is empty, the candidate control point movement information is added to the candidate motion vector list. Otherwise, the motion information in the candidate motion vector list is sequentially traversed to check whether motion information identical to the candidate control point movement information exists in the candidate motion vector list. If motion information identical to the candidate control point movement information does not exist in the candidate motion vector list, the candidate control point movement information is added to the candidate motion vector list.
[0329] For example, the preset order is as follows: Affine(CP1,CP2,CP3) → Affine (CP1,CP2,CP4) → Affine(CP1,CP3,CP4) → Affine(CP2,CP3,CP4) → Affine(CP1,CP2) → Affine(CP1,CP3) → Affine(CP2,CP3) → Affine(CP1,CP4) → Affine(CP2,CP4) → Affine(CP3,CP4). There are a total of 10 combinations.
[0330] If the control point movement information corresponding to the combination is not available, the combination is considered unavailable. If the combination is available, the reference frame index of the combination is determined (in the case of two control points, the smaller reference frame index is selected as the reference frame index of the combination; in the case of more than two control points, the reference frame index that appears most frequently is selected; if the number of occurrences of multiple reference frame indexes is the same, the smallest reference frame index is selected as the reference frame index of the combination), and the motion vectors of the control points are scaled. If the scaled motion information of all control points is consistent, the combination is invalid.
[0331] As an option, in this embodiment of the present application, the candidate motion vector list may be padded. For example, after the above-described traversal process, if the length of the candidate motion vector list is shorter than the maximum list length MaxAffineNumMrgCand, the candidate motion vector list may be padded until the list length becomes equal to MaxAffineNumMrgCand.
[0332] Padding can be performed by using the zero motion vector padding method or by combining or weighted averaging existing candidate motion information in the existing list. It should be noted that another method for padding the candidate motion vector list is also applicable to the present application. Details are not described here.
[0333] Step S603b: Analyze the bitstream to determine the optimal control point motion information.
[0334] The index number is analyzed, and based on the index number, the optimal control point motion information is determined from the candidate motion vector list.
[0335] Step 604b: Based on the optimal control point motion information and the affine motion model used for the current decoded block, obtain the motion vector of each sub-block of the current block.
[0336] This step is the same as step 605a.
[0337] Step 605b: Based on the determined motion vectors of the sub-blocks, perform motion compensation for each sub-block and obtain the sample predictor of the sub-block.
[0338] The technology in the present invention is related to a context adaptive binary arithmetic coding (CABAC) entropy decoder, or another entropy decoder such as a probability interval partitioning entropy (PIPE) decoder or a related decoder. Arithmetic decoding is a form of entropy decoding used in many compression algorithms with high decoding efficiency because symbols can be mapped to non-integer length codes in arithmetic decoding. Generally, decoding data symbols by CABAC includes one or more of the following steps:
[0339] (1) Binary conversion: If the symbol to be decoded is not binary, the symbol is mapped to a "binary" sequence, and the value of each binary bit can be either "0" or "1".
[0340] (2) Context assignment: (In the normal mode) one context is assigned to each binary bit. The context model is used to determine a method for calculating the context for a given binary bit based on the information available for the binary bit. The information is, for example, the value of a previously decoded symbol or a binary number.
[0341] (3) Binary coding: The arithmetic encoder codes the binary bit. To code the binary bit, the arithmetic encoder requires the probability of the value of the binary bit as input, which is the probability that the value of the binary bit is equal to "0" and the probability that the value of the binary bit is equal to "1". The (estimated) probability of each context is represented by an integer value called the "context state". Each context has a state, and thus the state (i.e., the estimated probability) is the same for the binary bit to which one context is assigned and different between contexts.
[0342] (4) State update: The probability (state) of selecting a context is updated based on the actual decoded value of the binary bit (e.g., if the value of the binary bit is "1", the probability of "1" is increased).
[0343] In the prior art, when analyzing the parameter information of the affine motion model, such as affine_merge_flag, affine_merge_idx, affine_inter_flag, and affine_type_flag in Table 1, by CABAC, it is necessary that different contexts are used for different syntax elements in the CABAC analysis. In the present invention, the amount of context used in CABAC is reduced. Therefore, less space required by the encoder and decoder to store the context is occupied without affecting the coding efficiency.
[0344] Regarding affine_merge_flag and affine_inter_flag, two different context sets (each context set includes three contexts) are used in the CABAC in the prior art. The actual context index used in each set is equal to the sum of the value of the same syntax element in the left adjacent block of the current decoded block and the value of the same syntax element in the upper adjacent block of the current decoded block, as shown in Table 3. Here, avallableL indicates the availability of the left adjacent block of the current decoded block (whether the left adjacent block exists and has been decoded), and avallableA indicates the availability of the upper adjacent block of the current decoded block (whether the upper adjacent block exists and has been decoded). In the prior art, the amount of context for affine_merge_flag and affine_inter_flag is 6. Table 3 Context Index [Table 3]
[0345] Figure 10 describes the procedure of a video decoding method according to an embodiment of the present invention. This embodiment can be executed by the video decoder shown in Figure 3. As shown in Figure 10, the method includes the following steps.
[0346] 1001. Analyze the received bitstream to obtain the syntax elements to be entropy decoded in the current block. Here, the syntax elements to be entropy decoded in the current block include syntax element 1 in the current block or syntax element 2 in the current block.
[0347] In implementation, syntax element 1 in the current block is affine_merge_flag, or syntax element 2 in the current block is affine_inter_flag.
[0348] In implementation, syntax element 1 in the current block is subblock_merge_flag, or syntax element 2 in the current block is affine_inter_flag.
[0349] This step may specifically be executed by the entropy decoding unit 304 in Figure 3.
[0350] The current block in this embodiment of the present invention may be a CU.
[0351] 1002. Perform entropy decoding on the syntax elements to be entropy decoded in the current block. Entropy decoding for syntax element 1 in the current block is completed by using a pre-set context model, or entropy decoding for syntax element 2 in the current block is completed by using a context model.
[0352] This step may specifically be executed by the entropy decoding unit 304 in Figure 3.
[0353] 1003. Based on the syntax elements in the current block and those obtained by entropy decoding, perform prediction processing on the current block to obtain the predicted block of the current block.
[0354] Specifically, this step may be executed by the prediction processing unit 360 in FIG. 3.
[0355] 1004. Based on the predicted block of the current block, obtain the reconstructed image of the current block.
[0356] Specifically, this step may be executed by the reconstruction unit 314 in FIG. 3.
[0357] In this embodiment, since the syntax element 1 and the syntax element 2 in the current block share one context model, the decoder does not need to check the context model when performing entropy decoding, improving the decoding efficiency of video decoding executed by the decoder. Furthermore, since the video decoder only needs to store one context model for the syntax element 1 and the syntax element 2, less storage space of the video decoder is occupied.
[0358] Corresponding to the video decoding method described in FIG. 10, the embodiment of the present invention further provides an encoding method, and the encoding method is: A step of obtaining a syntax element to be entropy-coded in a current block, where the syntax element to be entropy-coded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block; a step of performing entropy coding on the syntax element to be entropy-coded in the current block, where when entropy coding is performed on the syntax element to be entropy-coded in the current block, entropy coding for syntax element 1 in the current block is completed by using a pre-set context model, or entropy coding for syntax element 2 in the current block is completed by using a context model; and a step of outputting a bitstream including the syntax element in the current block and obtained by entropy coding. The context model used when entropy coding is performed on the current block is the same as the context model in the video decoding method described in FIG. 10. In this embodiment, since syntax element 1 and syntax element 2 in the current block share one context model, the encoder does not need to check the context model when performing entropy Symbol coding, and the video coding performed by the encoder Symbol improves the coding efficiency. Further, since the video encoder needs to store only one context model for syntax element 1 and syntax element 2, less storage space of the video encoder is occupied.
[0359] FIG. 11 describes the procedure of a video decoding method according to another embodiment of the present invention. This embodiment can be executed by the video decoder shown in FIG. 3. As shown in FIG. 11, the method includes the following steps.
[0360] Analyze the received bitstream to obtain the syntax elements to be entropy decoded in the current block. The syntax elements to be entropy decoded in the current block include syntax element 1 in the current block or syntax element 2 in the current block.
[0361] In implementation, syntax element 1 in the current block is affine_merge_flag, and syntax element 2 in the current block is affine_inter_flag.
[0362] In implementation, syntax element 1 in the current block is subblock_merge_flag, and syntax element 2 in the current block is affine_inter_flag.
[0363] Specifically, this step may be executed by the entropy decoding unit 304 in FIG. 3.
[0364] 1102. Obtain the context model corresponding to the syntax element to be entropy decoded. The context model corresponding to syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 2 in the current block is determined from a pre-set context model set.
[0365] The video decoder needs to store only one context model set for syntax element 1 and syntax element 2.
[0366] In some implementations, the pre - set context - model set includes only two context models. In some other implementations, the pre - set context - model set includes only three context models. It can be understood that the pre - set context - model set may alternatively include four, five, or six context models. The amount of context models included in the pre - set context - model set is not limited in this embodiment of the present invention.
[0367] In an implementation, determining the context model corresponding to syntax element 1 in the current block from the pre - set context - model set includes determining the context index of syntax element 1 in the current block based on syntax element 1 and syntax element 2 in the left - adjacent block of the current block and syntax element 1 and syntax element 2 in the upper - adjacent block of the current block, and the context index of syntax element 1 in the current block is used to indicate the context model corresponding to syntax element 1 in the current block.
[0368] In another implementation, determining the context model corresponding to syntax element 2 in the current block from the pre - set context - model set includes determining the context index of syntax element 2 in the current block based on syntax element 1 and syntax element 2 in the left - adjacent block of the current block and syntax element 1 and syntax element 2 in the upper - adjacent block of the current block, and the context index of syntax element 2 in the current block is used to indicate the context model corresponding to syntax element 2 in the current block.
[0369] For example, when the amount of context models in a pre-set context model set is 3, the value of the context index of syntax element 1 in the current block is the sum of the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or the value of the context index of syntax element 2 in the current block is the sum of the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0370] Specifically, syntax element 1 affine_merge_flag and syntax element 2 affine_inter_flag can share one context model set (the set contains three context models). As shown in Table 4, the actual context index used in each set is equal to the result obtained by adding the value obtained by performing an OR operation on two syntax elements in the left adjacent block of the current decoding block and the value obtained by performing an OR operation on two syntax elements in the upper adjacent block of the current decoding block. Here, "|" indicates an OR operation. Table 4 Context Index in the Present Invention
Table 4
[0371] For example, when the amount of context models in a pre-set context model set is 2, the value of the context index of syntax element 1 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or the value of the context index of syntax element 2 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0372] Specifically, syntax element 1 affine_merge_flag and syntax element 2 affine_inter_flag share one context model set (the set contains two context models). As shown in Table 5, the actual context index used in each set is equal to the result obtained by performing an OR operation on the value obtained by performing an OR operation on two syntax elements in the left adjacent block of the current decoding block and the value obtained by performing an OR operation on two syntax elements in the upper adjacent block of the current decoding block. Here, "|" indicates an OR operation. In this embodiment of the present invention, the amount of context regarding affine_merge_flag and affine_inter_flag is reduced to 2. Table 5 Context Index in the Present Invention
Table 5
[0373] 1103. Based on the context model corresponding to the syntax element to be entropy decoded in the current block, perform entropy decoding on the syntax element to be entropy decoded.
[0374] Specifically, this step may be executed by the entropy decoding unit 304 in FIG. 3.
[0375] 1104. Based on the syntax element in the current block and obtained by entropy decoding, perform prediction processing on the current block to obtain a predicted block of the current block.
[0376] Specifically, this step may be executed by the prediction processing unit 360 in FIG. 3.
[0377] 1105. Based on the predicted block of the current block, obtain a reconstructed image of the current block.
[0378] Specifically, this step may be executed by the reconstruction unit 314 in FIG. 3.
[0379] In this embodiment, since the syntax element 1 and the syntax element 2 in the current block share one context model, the video decoder needs to store only one context model for the syntax element 1 and the syntax element 2, occupying less storage space of the video decoder.
[0380] Corresponding to the video decoding method described in FIG. 11, an embodiment of the present invention further provides an encoding method, and the encoding method includes: A step of obtaining a syntax element to be entropy encoded in the current block, where the syntax element to be entropy encoded in the current block includes the syntax element 1 in the current block or the syntax element 2 in the current block, and entropy SymbolObtaining a context model corresponding to the syntax element to be coded, wherein the context model corresponding to syntax element 1 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 2 in the current block is determined from a pre-set context model set; and based on the context model corresponding to the syntax element to be entropy-coded in the current block, entropy Symbol Performing entropy coding on the syntax element to be coded; and outputting a bitstream including the syntax element in the current block and obtained by entropy coding. The context model set used when entropy coding is performed on the current block is the same as the context model set in the video decoding method described in FIG. 11. In this embodiment, since syntax element 1 and syntax element 2 in the current block share one context model, the video encoder needs to store only one context model for syntax element 1 and syntax element 2, occupying less storage space of the video encoder.
[0381] FIG. 12 describes the procedure of a video decoding method according to an embodiment of the present invention. This embodiment may be executed by the video decoder shown in FIG. 3. As shown in FIG. 12, the method includes the following steps.
[0382] 1201. Analyzing the received bitstream to obtain the syntax element to be entropy-decoded in the current block. The syntax element to be entropy-decoded in the current block includes syntax element 3 or syntax element 4 in the current block.
[0383] In the implementation, the syntax element 3 in the current block is merge_idx, and the syntax element 4 in the current block is affine_merge_idx.
[0384] In the implementation, the syntax element 3 in the current block is merge_idx, and the syntax element 4 in the current block is subblock_merge_idx.
[0385] This step may specifically be executed by the entropy decoding unit 304 in FIG. 3.
[0386] 1202. Obtain a context model corresponding to the syntax element to be entropy decoded. The context model corresponding to the syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to the syntax element 4 in the current block is determined from a pre-set context model.
[0387] In the implementation, the number of context models included in the pre-set context model set is 5. It can be understood that the number of context models included in the pre-set context model set may alternatively be another value such as 1, 2, 3, or 4. When the number of context models included in the pre-set context model set is 1, the pre-set context model set is one context model. The number of context models included in the pre-set context model set is not limited in this embodiment of the present invention.
[0388] This step may specifically be executed by the entropy decoding unit 304 in FIG. 3.
[0389] Based on a context model corresponding to a syntax element to be entropy decoded in the current block, perform entropy decoding on the syntax element to be entropy decoded.
[0390] Specifically, this step may be performed by the entropy decoding unit 304 in FIG. 3.
[0391] Based on a syntax element in the current block and obtained by entropy decoding, perform prediction processing on the current block to obtain a predicted block of the current block.
[0392] Specifically, this step may be performed by the prediction processing unit 360 in FIG. 3.
[0393] Based on the predicted block of the current block, obtain a reconstructed image of the current block.
[0394] Specifically, this step may be performed by the reconstruction unit 314 in FIG. 3.
[0395] In this embodiment, since the syntax element 3 and the syntax element 4 in the current block share one context model, the video decoder needs to store only one context model for the syntax element 3 and the syntax element 4, occupying less storage space of the video decoder.
[0396] Corresponding to the video decoding method described in FIG. 12, an embodiment of the present invention further provides an encoding method. The encoding method includes a step of obtaining a syntax element to be entropy encoded in the current block, where the syntax element to be entropy encoded in the current block includes the syntax element 3 in the current block or the syntax element 4 in the current block, and entropy SymbolObtaining a context model corresponding to the syntax element to be coded, wherein the context model corresponding to syntax element 3 in the current block is determined from a pre-set context model set, or the context model corresponding to syntax element 4 in the current block is determined from a pre-set context model set; and based on the context model corresponding to the syntax element to be entropy-coded in the current block, entropy Symbol Executing entropy coding on the syntax element to be coded; and outputting a bitstream including the syntax element in the current block and obtained by entropy coding. The context model set used when entropy coding is executed on the current block is the same as the context model set in the video decoding method described in FIG. 12. In this embodiment, since syntax element 3 and syntax element 4 in the current block share one context model, the video encoder only needs to store only one context model for syntax element 3 and syntax element 4, occupying less storage space of the video encoder.
[0397] An embodiment of the present invention provides a video decoder 30 including an entropy decoding unit 304, a prediction processing unit 360, and a reconstruction unit 314.
[0398] The entropy decoding unit 304 is configured to analyze the received bit stream to obtain the syntax elements to be entropy decoded in the current block. The syntax elements to be entropy decoded in the current block include syntax element 1 in the current block or syntax element 2 in the current block. The entropy decoding unit 304 is configured to perform entropy decoding on the syntax elements to be entropy decoded in the current block. The entropy decoding for syntax element 1 in the current block is completed by using a pre-set context model, or the entropy decoding for syntax element 2 in the current block is completed by using a context model.
[0399] In an implementation, syntax element 1 in the current block is affine_merge_flag, and syntax element 2 in the current block is affine_inter_flag.
[0400] In an implementation, syntax element 1 in the current block is subblock_merge_flag, and syntax element 2 in the current block is affine_inter_flag.
[0401] The prediction processing unit 360 is configured to perform prediction processing on the current block based on the syntax elements in the current block and the syntax elements obtained by entropy decoding, and obtain the predicted block of the current block.
[0402] The reconstruction unit 314 is configured to obtain the reconstructed image of the current block based on the predicted block of the current block.
[0403] In this embodiment, since the syntax element 1 and the syntax element 2 in the current block share one context model, the decoder does not need to check the context model when performing entropy decoding, improving the decoding efficiency of video decoding performed by the decoder. Further, since the video decoder needs to store only one context model for the syntax element 1 and the syntax element 2, less storage space of the video decoder is occupied.
[0404] Correspondingly, an embodiment of the present invention provides a video encoder 20, the video encoder comprising an entropy encoding unit 270 configured to obtain a syntax element to be entropy encoded in the current block, the syntax element to be entropy encoded in the current block including the syntax element 1 in the current block or the syntax element 2 in the current block, the entropy encoding unit being configured to perform entropy encoding on the syntax element to be entropy encoded in the current block, and when entropy encoding is performed on the syntax element to be entropy encoded in the current block, the entropy encoding for the syntax element 1 in the current block is completed by using a pre-set context model, or the entropy encoding for the syntax element 2 in the current block is completed by using a context model; and an output 272 configured to output a bitstream including the syntax element in the current block and obtained by entropy encoding. The context model used when entropy encoding is performed on the current block is the same as the context model in the method described in FIG. 10. In this embodiment, since the syntax element 1 and the syntax element 2 in the current block share one context model, the encoder performs entropy SymbolIt is not necessary to check the context model when performing quantization, and video encoding is performed by the encoder. Symbol The encoding efficiency is improved. Further, since the video encoder needs to store only one context model for syntax element 1 and syntax element 2, less storage space of the video encoder is occupied.
[0405] Another embodiment of the present invention provides a video decoder 30 including an entropy decoding unit 304, a prediction processing unit 360, and a reconstruction unit 314.
[0406] The entropy decoding unit 304 is configured to analyze the received bitstream to obtain the syntax element to be entropy decoded in the current block. The syntax element to be entropy decoded in the current block includes syntax element 1 in the current block or syntax element 2 in the current block. The entropy decoding unit 304 is configured to obtain the context model corresponding to the syntax element to be entropy decoded. The context model corresponding to syntax element 1 in the current block is determined from a preset context model set, or the context model corresponding to syntax element 2 in the current block is determined from a preset context model set. The entropy decoding unit is configured to perform entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block.
[0407] In an implementation, syntax element 1 in the current block is affine_merge_flag, and syntax element 2 in the current block is affine_inter_flag.
[0408] In implementation, the syntax element 1 in the current block is subblock_merge_flag, and the syntax element 2 in the current block is affine_inter_flag.
[0409] In implementation, the entropy decoding unit 304 can be specifically configured to determine the context index of the syntax element 1 in the current block based on the syntax element 1 and the syntax element 2 in the left adjacent block of the current block and the syntax element 1 and the syntax element 2 in the upper adjacent block of the current block. The context index of the syntax element 1 in the current block is used to indicate the context model corresponding to the syntax element 1 in the current block, or The entropy decoding unit can be configured to determine the context index of the syntax element 2 in the current block based on the syntax element 1 and the syntax element 2 in the left adjacent block of the current block and the syntax element 1 and the syntax element 2 in the upper adjacent block of the current block. The context index of the syntax element 2 in the current block is used to indicate the context model corresponding to the syntax element 2 in the current block.
[0410] For example, the value of the context index of the syntax element 1 in the current block is the sum of the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on the syntax element 1 and the syntax element 2 in the left adjacent block, or The value of the context index of syntax element 2 in the current block is the sum of the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0411] For example, the value of the context index of syntax element 1 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block, or The value of the context index of syntax element 2 in the current block is the result obtained by performing an OR operation on the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the upper adjacent block and the value obtained by performing an OR operation on syntax element 1 and syntax element 2 in the left adjacent block.
[0412] The prediction processing unit 360 is configured to perform prediction processing on the current block based on the syntax element in the current block and obtained by entropy decoding, and obtain the predicted block of the current block.
[0413] The reconstruction unit 314 is configured to obtain the reconstructed image of the current block based on the predicted block of the current block.
[0414] In this embodiment, since the syntax elements 1 and 2 in the current block share one context model, the video decoder only needs to store only one context model for the syntax elements 1 and 2, occupying less storage space of the video decoder.
[0415] Correspondingly, an embodiment of the present invention further provides a video encoder 20, and the video encoder is an entropy encoding unit 270 configured to obtain a syntax element to be entropy encoded in the current block, where the syntax element to be entropy encoded in the current block includes the syntax element 1 in the current block or the syntax element 2 in the current block, and the entropy encoding unit is configured to obtain a context model corresponding to the syntax element to be entropy encoded. The context model corresponding to the syntax element 1 in the current block is determined from a preset context model set, or the context model corresponding to the syntax element 2 in the current block is determined from a preset context model set. The entropy encoding unit entropy Symbol encodes based on the context model corresponding to the syntax element to be encoded in the current block. SymbolAn entropy encoding unit configured to perform entropy encoding on a syntax element to be coded, and an output 272 configured to output a bitstream including a syntax element in the current block and obtained by entropy encoding. The context model set used when entropy encoding is performed on the current block is the same as the context model set in the video decoding method described in FIG. 11. In this embodiment, since the syntax element 1 and the syntax element 2 in the current block share one context model, the video encoder needs to store only one context model for the syntax element 1 and the syntax element 2, occupying less storage space of the video encoder.
[0416] Another embodiment of the present invention provides a video decoder 30 including an entropy decoding unit 304, a prediction processing unit 360, and a reconstruction unit 314.
[0417] The entropy decoding unit 304 is configured to analyze the received bitstream to obtain a syntax element to be entropy decoded in the current block. The syntax element to be entropy decoded in the current block includes the syntax element 3 in the current block or the syntax element 4 in the current block. The entropy decoding unit is configured to obtain a context model corresponding to the syntax element to be entropy decoded. The context model corresponding to the syntax element 3 in the current block is determined from a preset context model set, or the context model corresponding to the syntax element 4 in the current block is determined from a preset context model set. The entropy decoding unit is configured to perform entropy decoding on the syntax element to be entropy decoded based on the context model corresponding to the syntax element to be entropy decoded in the current block.
[0418] The pre-set context model set includes one, two, three, four, or five context models. If the pre-set context model set includes only one context model, it can be understood that the pre-set context model set is one context model.
[0419] In the implementation, the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is affine_merge_idx and is used to indicate the index value of the affine merge candidate list of the current block, or the syntax element 3 in the current block is merge_idx and is used to indicate the index value of the merge candidate list of the current block, or the syntax element 4 in the current block is subblock_merge_idx and is used to indicate the index value of the sub-block merge candidate list.
[0420] The prediction processing unit 360 is configured to execute prediction processing for the current block based on the syntax element in the current block and obtained by entropy decoding, so as to obtain the predicted block of the current block.
[0421] The reconstruction unit 314 is configured to obtain the reconstructed image of the current block based on the predicted block of the current block.
[0422] In this embodiment, since the syntax element 3 and the syntax element 4 in the current block share one context model, the video decoder needs to store only one context model for the syntax element 3 and the syntax element 4, occupying less storage space of the video decoder.
[0423] Correspondingly, an embodiment of the present invention further provides a video encoder, the video encoder including an entropy encoding unit 270 configured to obtain a syntax element to be entropy encoded in a current block, where the syntax element to be entropy encoded in the current block includes a syntax element 3 in the current block or a syntax element 4 in the current block, and the entropy encoding unit is configured to obtain a context model corresponding to the syntax element to be entropy encoded, the context model corresponding to the syntax element 3 in the current block being determined from a pre-set context model set, or the context model corresponding to the syntax element 4 in the current block being determined from a pre-set context model set, and the entropy encoding unit is configured to perform entropy encoding on the syntax element to be entropy encoded based on the context model corresponding to the syntax element to be entropy encoded in the current block. Symbol The video encoder further includes an output 272 configured to output a bitstream including a syntax element in the current block and obtained by entropy encoding. The context model set used when entropy encoding is performed on the current block is the same as the context model set in the video decoding method described in FIG. 12. In this embodiment, since the syntax element 3 and the syntax element 4 in the current block share one context model, the video encoder only needs to store a single context model for the syntax element 3 and the syntax element 4, occupying less storage space of the video encoder. Symbol
[0424] Embodiment 1 of the present invention proposes that the affine_merge_flag and the affine_inter_flag share one context set (the set includes three contexts), and the actual context index used in each set is, as shown in Table 4, equal to the result obtained by adding the value obtained by performing an OR operation on two syntax elements in the left adjacent block of the current decoded block and the value obtained by performing an OR operation on two syntax elements in the upper adjacent block of the current decoded block. Here, "|" indicates an OR operation. In Embodiment 1 of the present invention, the amount of context regarding the affine_merge_flag and the affine_inter_flag is reduced to 3.
[0425] Embodiment 2 of the present invention proposes that the affine_merge_flag and the affine_inter_flag share one context set (the set includes two contexts), and the actual context index used in each set is, as shown in Table 5, equal to the result obtained by performing an OR operation on the value obtained by performing an OR operation on two syntax elements in the left adjacent block of the current decoded block and the value obtained by performing an OR operation on two syntax elements in the upper adjacent block of the current decoded block. Here, "|" indicates an OR operation. In Embodiment 2 of the present invention, the amount of context regarding the affine_merge_flag and the affine_inter_flag is reduced to 2.
[0426] Embodiment 3 of the present invention proposes that the affine_merge_flag and the affine_inter_flag share one context. In Embodiment 3 of the present invention, the number of affine_merge_flag contexts and the number of affine_inter_flag contexts are reduced to 1.
[0427] In the prior art, by using truncated single - progression codes, binarization is performed with respect to merge_idx and affine_merge_idx, two different context sets (each context set contains five contexts) are used in CABAC, and different contexts are used for each binary bit after binarization. In the prior art, the amount of contexts for merge_idx and affine_merge_idx is 10.
[0428] Embodiment 4 of the present invention proposes that merge_idx and affine_merge_idx share one context set (each context set contains five contexts). In Embodiment 4 of the present invention, the amount of contexts for merge_idx and affine_merge_idx is reduced to 5.
[0429] In some other technologies, the syntax element affine_merge_flag[x0][y0] in Table 1 may be replaced by subblock_merge_flag[x0][y0] and is used to indicate whether the sub - block - based merge mode is currently used for the block, and the syntax element affine_merge_idx[x0][y0] in Table 1 may be replaced by subblock_merge_idx[x0][y0] and is used to indicate the index value of the sub - block merge candidate list.
[0430] In this case, Embodiments 1 to 4 of the present invention are still applicable, that is, subblock_merge_flag and affine_inter_flag share one context set (or context) and one index acquisition method, and merge_idx and subblock_merge_idx share one context set (or context).
[0431] Embodiments of the present invention further provide a video decoder including an execution circuit configured to execute any of the above - mentioned methods.
[0432] Embodiments of the present invention further provide a video decoder including at least one processor and a non-volatile computer-readable storage medium coupled to the at least one processor. The non-volatile computer-readable storage medium stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the video decoder is operable to perform any of the above methods.
[0433] Embodiments of the present invention further provide a computer-readable storage medium configured to store a computer program executable by a processor. When the computer program is executed by the at least one processor, any of the above methods is performed.
[0434] Embodiments of the present invention further provide a computer program. When the computer program is executed, any of the above methods is performed.
[0435] In one or more of the foregoing examples, the described functionality may be implemented by hardware, software, firmware, or any combination thereof. When implemented in software, the functionality may be stored on or transmitted via a computer-readable medium and executed by a hardware-based processing unit as one or more instructions or code. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium. The communication medium may include, for example, any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. Thus, the computer-readable medium may generally correspond to either (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0436] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the required program code in the form of instructions or data structures and that is accessible by a computer. Further, any connection may suitably be referred to as a computer-readable medium. For example, coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, when the instructions are transmitted from a website, server, or other remote source, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or any other transient medium, and are actually directed to non-transient tangible storage media. As used herein, disks and optical disks include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc. Disks generally reproduce data magnetically, and optical disks reproduce data optically by using a laser. Any of the above combinations should fall within the scope of computer-readable media.
[0437] The commands can be executed by one or more processors, which can be, for example, one or more digital signal processors (DSPs), one or more general-purpose microprocessors, one or more application specific integrated circuits (ASICs), one or more field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein can represent any one of the foregoing structures or other structures applicable to the implementation of the technologies described herein. Further, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured to perform encoding and decoding, or may be incorporated into a combined codec. Further, the technology may be fully realized by one or more circuits or logic elements.
[0438] The techniques of the present disclosure may be implemented in a plurality of devices or apparatuses including a wireless handset, an integrated circuit (IC), or an IC set (e.g., a chipset). To emphasize the functions of the devices configured to execute the disclosed techniques, various components, modules, or units are described in this disclosure, which do not necessarily have to be realized by different hardware units. In fact, as described above, the various units may be combined within a codec hardware unit with appropriate software and / or firmware, or may be provided by a set of interoperable hardware units. The hardware units include one or more of the processors described above.
Claims
1. 1. A video decoding method comprising: receiving a bitstream including syntax elements to be entropy decoded for a current block, the syntax elements to be entropy decoded including a first syntax element and a second syntax element, the first syntax element including a first flag indicating whether a sub-block based merging mode is used for the current block, or the second syntax element including a second flag indicating whether an affine motion model based motion compensation is used for the current block, when a slice including the current block is a predictive (P) type slice or a bi-predictive (B) type slice; parsing the bitstream to obtain syntax elements to be entropy decoded; obtaining a first value of a first parameter indicating whether a left neighboring block of the current block is available; obtaining a second value of a second parameter indicating whether an upper neighboring block of the current block is available; obtaining a third value of a third flag or a fourth value of a fourth flag, where the third flag indicates whether a sub-block based merging mode is used for the left neighboring block, or the fourth flag indicates whether an affine motion model based motion compensation is used for the left neighboring block, when a slice including the left neighboring block is a P-type slice or a B-type slice; obtaining a context index for the syntax element to be entropy decoded based on the first value, the second value, and at least one of the third value or the fourth value; performing entropy decoding on the syntax element to be entropy decoded based on the context index to obtain a decoded syntax element corresponding to the syntax element to be entropy decoded; performing a prediction process on the current block based on the decoded syntax elements to obtain a prediction block of the current block; and obtaining a reconstructed image of the current block based on the predicted block; 2. A video decoding method comprising:
2. 10. The video decoding method of claim 1, further comprising: obtaining a fifth value of a fifth flag or a sixth value of a sixth flag, where the fifth flag indicates whether the sub-block based merging mode is used for the above neighboring block, or the sixth flag indicates whether an affine motion model based motion compensation is used for the above neighboring block when a slice including the left neighboring block is a P-type slice or a B-type slice; and further obtaining a context index of the syntax element to be entropy decoded based on the fifth value or the sixth value; 2. A video decoding method comprising:
3. 3. The video decoding method of claim 2, further comprising: performing a first OR operation on the third value and the fourth value to obtain a seventh value; performing a second OR operation on the fifth value and the sixth value to obtain an eighth value; and further obtaining a context index of the syntax element to be entropy decoded based on the seventh value and the eighth value; 2. A video decoding method comprising:
4. 4. The video decoding method of claim 3, further comprising: performing a first AND operation on the first value and the seventh value to obtain a ninth value; performing a second AND operation on the second value and the eighth value to obtain a tenth value; further obtaining a context index of the syntax element to be entropy decoded based on the ninth value and the tenth value; 2. A video decoding method comprising:
5. 5. The video decoding method of claim 4, further comprising obtaining a sum of the ninth value and the tenth value, the sum being the context index.
6. 1. A video decoding device comprising: A non-transitory storage medium configured to store video data in the form of a bitstream, the bitstream including syntax elements to be entropy decoded for a current block, the syntax elements to be entropy decoded including a first syntax element and a second syntax element, the first syntax element including a first flag indicating whether a sub-block based merging mode is used for the current block, or the second syntax element including a second flag indicating whether an affine motion model based motion compensation is used for the current block, when a slice including the current block is a predictive (P) type slice or a bi-predictive (B) type slice; and a video decoder coupled to the non-transitory storage medium; said video decoder comprising: parsing the bitstream to obtain syntax elements to be entropy decoded; obtaining a first value of a first parameter indicating whether a left neighboring block of the current block is available; obtaining a second value of a second parameter indicating whether an upper neighboring block of the current block is available; obtaining a third value of a third flag or a fourth value of a fourth flag, where the third flag indicates whether a sub-block based merging mode is used for the left neighboring block, or the fourth flag indicates whether an affine motion model based motion compensation is used for the left neighboring block, when a slice including the left neighboring block is a P-type slice or a B-type slice; obtaining a context index for the syntax element to be entropy decoded based on the first value, the second value, and at least one of the third value or the fourth value; performing entropy decoding on the syntax element to be entropy decoded based on the context index to obtain a decoded syntax element corresponding to the syntax element to be entropy decoded; performing a prediction process on the current block based on the decoded syntax elements to obtain a prediction block of the current block; and obtaining a reconstructed image of the current block based on the predicted block; a video decoding device configured to:
7. 7. The video decoding device of claim 6, further comprising: obtaining a fifth value of a fifth flag or a sixth value of a sixth flag, where the fifth flag indicates whether the sub-block based merging mode is used for the above neighboring block, or the sixth flag indicates whether an affine motion model based motion compensation is used for the above neighboring block when a slice including the left neighboring block is a P-type slice or a B-type slice; and further obtaining a context index of the syntax element to be entropy decoded based on the fifth value or the sixth value; a video decoding device configured to:
8. 8. The video decoding device of claim 7, further comprising: performing a first OR operation on the third value and the fourth value to obtain a seventh value; performing a second OR operation on the fifth value and the sixth value to obtain an eighth value; and further obtaining a context index of the syntax element to be entropy decoded based on the seventh value and the eighth value; a video decoding device configured to:
9. 9. The video decoding device of claim 8, further comprising: performing a first AND operation on the first value and the seventh value to obtain a ninth value; performing a second AND operation on the second value and the eighth value to obtain a tenth value; further obtaining a context index of the syntax element to be entropy decoded based on the ninth value and the tenth value; a video decoding device configured to:
10. 10. The video decoding device of claim 9, further configured to obtain a sum of the ninth value and the tenth value, the sum being the context index.
11. 7. The video decoding device of claim 6, wherein the first syntax element and the second syntax element share a context model of the context index.
12. One or more memories configured to store programming instructions; and one or more processors coupled to the one or more memories and configured to execute the programming instructions; wherein the programming instructions cause the video encoder to: generating a bitstream, the bitstream including a syntax element of a first entropy decoding target of a current block and a syntax element of a second entropy decoding target of a left neighboring block of the current block; The first entropy decoding target syntax element includes a first syntax element or a second syntax element, and the first syntax element includes a first flag indicating whether a sub-block based merging mode is used for the current block, or the second syntax element includes a second flag indicating whether an affine motion model based motion compensation is used for the current block, when a slice including the current block is a predictive (P) type slice or a bi-predictive (B) type slice; The second entropy decoding target syntax element includes a third syntax element or a fourth syntax element, and the third syntax element includes a third flag indicating whether a sub-block based merging mode is used for the left neighboring block, or the fourth syntax element includes a fourth flag indicating whether an affine motion model based motion compensation is used for the left neighboring block, when a slice including the left neighboring block is a P-type slice or a B-type slice; at least one of the third value of the third flag or the fourth value of the fourth flag is used to obtain a context index of the syntax element of the first entropy decoding target, the context index is used to perform entropy decoding on the syntax element of the first entropy decoding target to obtain a decoded syntax element corresponding to the syntax element of the first entropy decoding target, and the decoded syntax element is used to perform a prediction process on the current block to obtain a predicted block of the current block, the predicted block being used to obtain a reconstructed image of the current block; and transmitting the bitstream to a video decoding device; A video encoder that performs
13. 13. The video encoder of claim 12, wherein the bitstream further includes a syntax element for a third entropy decoding target of an upper neighboring block of the current block, the third entropy decoding target syntax element including a fifth syntax element or a sixth syntax element, the fifth syntax element including a fifth flag indicating whether a subblock-based merging mode is used for the upper neighboring block, or the sixth syntax element including a sixth flag indicating whether an affine motion model based motion compensation is used for the upper neighboring block, when a slice including the left neighboring block is a P type slice or a B type slice, and at least one of a fifth value of the fifth flag or a sixth value of the sixth flag is used to obtain a context index of the first entropy decoding target syntax element.
14. 13. The video encoder of claim 12, wherein a context model corresponding to the first syntax element is based on a preset context model set, or a context model corresponding to the second syntax element is based on the preset context model set.
15. 15. The video encoder of claim 14, wherein the first syntax element and the second syntax element share the context model.
16. 13. The video encoder of claim 12, wherein the one or more processors are further configured to execute the programming instructions to cause the video encoder to further transmit the bitstream to a video decoder.
17. A video decoder comprising an implementation circuit configured to implement the method according to any one of claims 1-5.
18. At least one processor; and a non-volatile computer readable storage medium coupled to the at least one processor, comprising: A video decoder, wherein the non-volatile computer-readable storage medium stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the video decoder is operable to perform the method of any one of claims 1-5.
19. A computer readable storage medium having a program recorded thereon, the program causing a computer to execute the method according to any one of claims 1 to 5.
20. A computer program arranged to cause a computer to carry out the method according to any one of claims 1 to 5.
21. A method of video encoding, comprising the steps of a video encoder generating a bitstream: The bitstream includes a syntax element of a first entropy decoding target of a current block and a syntax element of a second entropy decoding target of a left neighboring block of the current block; The first entropy decoding target syntax element includes a first syntax element or a second syntax element, and the first syntax element includes a first flag indicating whether a sub-block based merging mode is used for the current block, or the second syntax element includes a second flag indicating whether an affine motion model based motion compensation is used for the current block, when a slice including the current block is a predictive (P) type slice or a bi-predictive (B) type slice; The second entropy decoding target syntax element includes a third syntax element or a fourth syntax element, and the third syntax element includes a third flag indicating whether a sub-block based merging mode is used for the left neighboring block, or the fourth syntax element includes a fourth flag indicating whether an affine motion model based motion compensation is used for the left neighboring block, when a slice including the left neighboring block is a P-type slice or a B-type slice; 11. The method of claim 10, wherein at least one of a third value of the third flag or a fourth value of the fourth flag is used to obtain a context index of the first syntax element to be entropy decoded, the context index is used to perform entropy decoding on the syntax element of the first entropy decoding target to obtain a decoded syntax element corresponding to the syntax element of the first entropy decoding target, and the decoded syntax element is used to perform a prediction process on the current block to obtain a predicted block of the current block, the predicted block being used to obtain a reconstructed image of the current block.
22. 22. A video encoder comprising circuitry configured to perform the method of claim 21.
23. At least one processor; a non-volatile computer readable storage medium coupled to the at least one processor, 22. A video encoder, wherein the non-volatile computer-readable storage medium stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, enabling the video encoder to operate to perform the method of claim 21.
24. A computer-readable storage medium having a program recorded thereon, the program causing a computer to execute the method of claim 21.
25. 22. A computer program product arranged to cause a computer to carry out the method of claim 21.
26. A method of storage comprising the steps of generating a bitstream by performing the method of claim 21, and storing the bitstream in a storage device.