Encoding device, decoding device, and non-transitory storage medium

By limiting the number of context-adaptive coding processing times in video coding and combining Columbine coding with context-adaptive coding, the video coding device and method are optimized, solving the shortcomings of coding efficiency and circuit scale in the existing technology and improving image quality and processing speed.

CN120825580APending Publication Date: 2025-10-21PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511250197.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-24
Filing Date
2020-04-24
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing video coding technologies require improvement in coding efficiency, image quality, processing capacity, circuit size, and processing speed. In particular, there are deficiencies in the appropriate selection of elements or actions such as filters, block sizes, motion vectors, and reference images.

Method used

A coding device and method are adopted, which limit the number of context adaptive coding processing times in the case of orthogonal transformation and non-orthogonal transformation, use Columbine coding and context adaptive coding to encode or decode coefficient information flags, appropriately select the coding and decoding methods, simplify the processing flow and optimize the circuit structure.

Benefits of technology

The coding efficiency is improved, the image quality is improved, the processing amount and circuit scale are reduced, the processing speed is increased, and the key elements and actions in encoding and decoding are appropriately selected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825580A_ABST
    Figure CN120825580A_ABST
Patent Text Reader

Abstract

The invention provides an encoding device, a decoding device, and a non-transitory storage medium. The encoding device encodes a flag by context-adaptive encoding if the number of processing times of context-adaptive encoding is limited to allow context-adaptive encoding of a plurality of coefficient information flags together in both cases of applying a different syntax and skipping an orthogonal transform in residual encoding of a current block, and then performs encoding of the flag by the context-adaptive encoding if the number of processing times of the context-adaptive encoding is limited to allow context-adaptive encoding of a plurality of coefficient information flags. And encoding the residual value of the coefficient by Golomb encoding, if not allowed, skipping the encoding of the flag, encoding the value of the coefficient by Golomb encoding, and if the orthogonal transformation is skipped, encoding the coefficient information flag other than the plurality of absolute value flags after encoding and before encoding the residual value of the coefficient. A plurality of absolute value flags are encoded by context adaptive encoding, the plurality of absolute value flags being flags related to whether the absolute value of the coefficient is greater than a prescribed value, the prescribed value being an integer greater than 1, the flags including flags indicating whether the value of the coefficient is zero or non-zero.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional case of the invention patent application with the application date of April 24, 2020, application number 202080030228.9, and invention name “Encoding device, decoding device, encoding method and decoding method”. Technical Field

[0002] The present invention relates to video coding, and for example, to systems, components, and methods for encoding and decoding moving images. Background Art

[0003] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). This advancement has led to a continuous need for improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data used in various applications.

[0004] Furthermore, Non-Patent Document 1 relates to an example of existing standards related to the above-mentioned video encoding technology.

[0005] Prior art literature

[0006] Non-patent literature

[0007] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0008] Problems to be solved by the invention

[0009] Regarding the coding methods described above, it is expected that new methods will be proposed to improve coding efficiency, improve image quality, reduce processing volume, reduce circuit scale, or appropriately select elements or actions such as filters, blocks, sizes, motion vectors, reference pictures or reference blocks.

[0010] The present invention provides a structure or method that can contribute to one or more of the following: improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, the present invention may include structures or methods that can contribute to benefits other than those described above.

[0011] Means for solving problems

[0012] For example, an encoding device according to one aspect of the present invention comprises: a circuit; and a memory connected to the circuit, wherein, in residual encoding of a current block, in both cases of applying orthogonal transforms using different syntaxes and skipping the orthogonal transform, when a processing limit of context adaptive encoding allows context adaptive encoding of a plurality of coefficient information flags related to coefficients included in the current block, the circuit encodes the plurality of coefficient information flags by the context adaptive encoding and encodes residual values ​​of the coefficients by Columbus encoding, the residual values ​​of the coefficients being values ​​used to reconstruct the values ​​of the coefficients using the plurality of coefficient information flags, when the processing limit does not allow context adaptive encoding of the plurality of coefficient information flags. When the context-adaptive encoding is performed on the plurality of coefficient information flags together, the circuit skips encoding the plurality of coefficient information flags and encodes the values ​​of the coefficients by the Columbine coding. When the orthogonal transform is skipped, the circuit encodes the plurality of absolute value flags by the context-adaptive coding after encoding the coefficient information flags other than the plurality of absolute value flags among the plurality of coefficient information flags and before encoding the remaining values ​​of the coefficients. The plurality of absolute value flags are flags related to whether the absolute values ​​of the coefficients are greater than a prescribed value, and the prescribed value is an integer greater than 1. The plurality of coefficient information flags include flags indicating whether the values ​​of the coefficients are zero or non-zero.

[0013] The implementation of several embodiments of the present invention can improve coding efficiency, simplify coding / decoding processing, speed up coding / decoding processing, and efficiently select appropriate components / actions used in coding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.

[0014] The present invention provides further advantages and effects according to the present invention. These advantages and / or effects are achieved through several embodiments and features described in the present invention and the accompanying drawings, but all advantages and / or effects do not necessarily need to be provided in order to achieve one or more advantages and / or effects.

[0015] Furthermore, these general or specific aspects may also be implemented as a system, a method, an integrated circuit, a computer program, a recording medium, or any combination thereof.

[0016] Effects of the Invention

[0017] The structure or method of one aspect of the present invention can contribute to one or more of, for example, improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, the structure or method of one aspect of the present invention can also contribute to benefits other than those listed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a block diagram showing the functional structure of the encoding device according to the embodiment.

[0019] Figure 2 This is a flowchart showing an example of the overall encoding process performed by the encoding device.

[0020] Figure 3 This is a conceptual diagram showing an example of block division.

[0021] Figure 4A This is a conceptual diagram showing an example of the structure of a slice.

[0022] Figure 4B This is a conceptual diagram showing an example of a tile structure.

[0023] Figure 5A This is a table showing the transformation basis functions corresponding to various transformation types.

[0024] Figure 5B This is a conceptual diagram showing an example of SVT (Spatially Varying Transform).

[0025] Figure 6A This is a conceptual diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).

[0026] Figure 6B This is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0027] Figure 6C This is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0028] Figure 7 This is a block diagram showing an example of a detailed configuration of a loop filter unit that functions as a DBF (deblocking filter).

[0029] Figure 8 This is a conceptual diagram showing an example of deblocking filtering having filter characteristics that are symmetric with respect to block boundaries.

[0030] Figure 9This is a conceptual diagram for explaining the block boundary on which the deblocking filtering process is performed.

[0031] Figure 10 This is a conceptual diagram showing an example of the Bs value.

[0032] Figure 11 This is a flowchart showing an example of processing performed by the prediction processing unit of the encoding device.

[0033] Figure 12 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device.

[0034] Figure 13 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device.

[0035] Figure 14 This is a conceptual diagram showing an example of 67 intra prediction modes in the intra prediction according to the embodiment.

[0036] Figure 15 This is a flowchart showing an example of the flow of basic processing of inter-frame prediction.

[0037] Figure 16 This is a flowchart showing an example of motion vector derivation.

[0038] Figure 17 This is a flowchart showing another example of motion vector derivation.

[0039] Figure 18 This is a flowchart showing another example of motion vector derivation.

[0040] Figure 19 This is a flowchart showing an example of inter prediction based on the normal inter mode.

[0041] Figure 20 is a flowchart illustrating an example of inter-frame prediction based on merge mode.

[0042] Figure 21 This is a conceptual diagram for explaining an example of motion vector derivation processing based on the merge mode.

[0043] Figure 22 This is a flowchart showing an example of FRUC (frame rate up conversion) processing.

[0044] Figure 23 This is a conceptual diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0045] Figure 24This is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.

[0046] Figure 25A This is a conceptual diagram for explaining an example of derivation of a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks.

[0047] Figure 25B This is a conceptual diagram for explaining an example of derivation of a motion vector in sub-block units in an affine mode having three control points.

[0048] Figure 26A This is a conceptual diagram used to illustrate the affine merge mode.

[0049] Figure 26B This is a conceptual diagram for explaining the affine merge mode with two control points.

[0050] Figure 26C This is a conceptual diagram for explaining the affine merge mode with three control points.

[0051] Figure 27 This is a flowchart showing an example of processing in the affine merge mode.

[0052] Figure 28A This is a conceptual diagram for explaining the affine inter mode with two control points.

[0053] Figure 28B This is a conceptual diagram for explaining the affine inter-frame mode with three control points.

[0054] Figure 29 This is a flowchart showing an example of processing in the affine inter mode.

[0055] Figure 30A This is a conceptual diagram used to illustrate the affine inter-frame mode in which the current block has 3 control points and the adjacent block has 2 control points.

[0056] Figure 30B This is a conceptual diagram used to illustrate the affine inter-frame mode in which the current block has 2 control points and the adjacent block has 3 control points.

[0057] Figure 31A This is a flowchart showing the merge mode including DMVR (decoder motion vector refinement).

[0058] Figure 31B This is a conceptual diagram for explaining an example of DMVR processing.

[0059] Figure 32 This is a flowchart showing an example of generating a predicted image.

[0060] Figure 33 This is a flowchart showing another example of generating a predicted image.

[0061] Figure 34 This is a flowchart showing another example of generating a predicted image.

[0062] Figure 35 This is a flowchart for explaining an example of a predicted image correction process based on an OBMC (overlapped block motion compensation) process.

[0063] Figure 36 This is a conceptual diagram for explaining an example of predicted image correction processing based on OBMC processing.

[0064] Figure 37 This is a conceptual diagram for explaining the generation of predicted images of two triangles.

[0065] Figure 38 This is a conceptual diagram for explaining a model assuming constant velocity linear motion.

[0066] Figure 39 This is a conceptual diagram for explaining an example of a method for generating a predicted image using a brightness correction process based on LIC (local illumination compensation) processing.

[0067] Figure 40 This is a block diagram showing an example of installing an encoding device.

[0068] Figure 41 This is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0069] Figure 42 This is a flowchart showing an example of the overall decoding process performed by the decoding device.

[0070] Figure 43 This is a flowchart showing an example of processing performed by the prediction processing unit of the decoding device.

[0071] Figure 44 This is a flowchart showing another example of processing performed by the prediction processing unit of the decoding device.

[0072] Figure 45 This is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in a decoding device.

[0073] Figure 46 This is a block diagram showing an implementation example of a decoding device.

[0074] Figure 47 This is a flowchart showing the basic coefficient encoding method of the first aspect.

[0075] Figure 48 This is a flowchart showing the basic first encoding method of the first aspect.

[0076] Figure 49 This is a flowchart showing the basic second encoding method of the first aspect.

[0077] Figure 50 This is a flowchart showing the coefficient encoding method of the first example of the first aspect.

[0078] Figure 51 This is a flowchart showing the coefficient encoding method of the second example of the first aspect.

[0079] Figure 52 This is a flowchart showing the coefficient encoding method of the first example of the second aspect.

[0080] Figure 53 This is a flowchart showing the coefficient encoding method of the second example of the second aspect.

[0081] Figure 54 This is a syntax diagram showing the basic first encoding method of the third form.

[0082] Figure 55 This is a syntax diagram showing the basic second encoding method of the third form.

[0083] Figure 56 This is a syntax diagram showing the second encoding method of the first example of the third aspect.

[0084] Figure 57 This is a syntax diagram showing the second encoding method of the second example of the third aspect.

[0085] Figure 58 This is a flowchart showing the operation of the encoding device according to the embodiment.

[0086] Figure 59 This is a flowchart showing the operation of the decoding device according to the embodiment.

[0087] Figure 60 This is a block diagram showing the overall structure of a content provision system that implements content distribution services.

[0088] Figure 61 This is a conceptual diagram showing an example of a coding structure in the case of scalable coding.

[0089] Figure 62 This is a conceptual diagram showing an example of a coding structure in the case of scalable coding.

[0090] Figure 63This is a conceptual diagram showing an example of a display screen of a web page.

[0091] Figure 64 This is a conceptual diagram showing an example of a display screen of a web page.

[0092] Figure 65 This is a block diagram showing an example of a smart phone.

[0093] Figure 66 This is a block diagram showing a configuration example of a smartphone. DETAILED DESCRIPTION

[0094] For example, when encoding image blocks, the encoding device may be able to transform the blocks into easily compressible data by applying an orthogonal transform to the blocks. Alternatively, when encoding image blocks, the encoding device may be able to reduce processing delay by not applying an orthogonal transform to the blocks.

[0095] Furthermore, the characteristics of a block to which orthogonal transformation is applied are different from the characteristics of a block to which orthogonal transformation is not applied. The encoding method used for the block to which orthogonal transformation is applied and the encoding method used for the block to which orthogonal transformation is not applied may be different.

[0096] However, if an inappropriate coding method is used for a block to which an orthogonal transform is applied, or if an inappropriate coding method is used for a block not to which an orthogonal transform is applied, there is a possibility of an increase in the amount of code or an increase in processing delay. In addition, if there is a significant difference between the coding method used for a block to which an orthogonal transform is applied and the coding method used for a block not to which an orthogonal transform is applied, there is a possibility of increased processing complexity and an increase in circuit scale.

[0097] Therefore, for example, an encoding device according to one aspect of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit limits the number of context adaptive coding processing times during operation and encodes a block of an image, and in the encoding of the block, in both cases of applying an orthogonal transform to the block and not applying an orthogonal transform to the block, a sub-block flag encoding process is performed without including it in the number of processing times, wherein the sub-block flag encoding process encodes a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context adaptive coding.

[0098] This makes it possible to encode sub-block flags using context-adaptive coding, regardless of whether or not an orthogonal transform is applied and whether the number of context-adaptive coding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0099] In addition, for example, the circuit further performs position parameter encoding processing without including it in the processing times when an orthogonal transform is applied to the block, and the position parameter encoding processing encodes the parameters representing the position of the initial non-zero coefficient in the scanning order in the block through context adaptive coding.

[0100] Therefore, when orthogonal transform is applied, the parameter indicating the position of the first non-zero coefficient can be encoded by context adaptive coding regardless of whether the number of context adaptive coding processes is limited, thereby reducing the amount of code.

[0101] Furthermore, for example, when an orthogonal transform is applied to the block, the circuit further determines a limit range of the number of processing times according to a position of the first non-zero coefficient.

[0102] Therefore, when orthogonal transformation is applied, it is possible to appropriately determine the limit number of processing times, and thus it is possible to appropriately adjust the balance between reduction in the amount of code and reduction in processing delay.

[0103] Furthermore, for example, a decoding device according to one aspect of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit decodes a block of an image while limiting the number of context-adaptive decoding processing times during operation, and in decoding the block, a sub-block flag decoding process is performed without including it in the number of processing times, in both cases of applying an inverse orthogonal transform to the block and not applying an inverse orthogonal transform to the block, the sub-block flag decoding process decoding a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context-adaptive decoding.

[0104] This makes it possible to decode sub-block flags using context-adaptive decoding, regardless of whether an inverse orthogonal transform is applied or whether the number of context-adaptive decoding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be minimized, leading to a reduction in circuit size.

[0105] In addition, for example, the circuit further performs position parameter decoding processing without including it in the processing times when applying an inverse orthogonal transform to the block, and the position parameter decoding processing encodes the parameters representing the position of the initial non-zero coefficient in the scanning order in the block through context adaptive decoding.

[0106] Therefore, when an inverse orthogonal transform is applied, the parameter indicating the position of the first non-zero coefficient can be decoded by context adaptive decoding regardless of whether the number of context adaptive decoding processes is limited, thereby reducing the amount of code.

[0107] Furthermore, for example, when applying an inverse orthogonal transform to the block, the circuit further determines a limit range of the number of processing times according to a position of the first non-zero coefficient.

[0108] This makes it possible to appropriately determine the limit on the number of processing times when inverse orthogonal transform is applied, thereby making it possible to appropriately adjust the balance between reducing the amount of code and reducing processing delay.

[0109] In addition, for example, one aspect of the encoding method of the present invention is to limit the number of context adaptive coding processing times and encode a block of an image. In the encoding of the block, in both cases of applying an orthogonal transform to the block and not applying an orthogonal transform to the block, a sub-block flag encoding process is performed without including it in the number of processing times. The sub-block flag encoding process encodes a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context adaptive coding.

[0110] This makes it possible to encode sub-block flags using context-adaptive coding, regardless of whether or not an orthogonal transform is applied and whether the number of context-adaptive coding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0111] Furthermore, for example, a decoding method according to one aspect of the present invention is to limit the number of context-adaptive decoding processing times while decoding a block of an image. In the decoding of the block, in both cases of applying an inverse orthogonal transform to the block and not applying an inverse orthogonal transform to the block, a sub-block flag decoding process is performed without including it in the number of processing times. The sub-block flag decoding process decodes a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context-adaptive decoding.

[0112] This makes it possible to decode sub-block flags using context-adaptive decoding, regardless of whether an inverse orthogonal transform is applied or whether the number of context-adaptive decoding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be minimized, leading to a reduction in circuit size.

[0113] In addition, for example, an encoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit, when in operation, encodes a coefficient information flag indicating the attributes of a coefficient contained in the block by context adaptive coding when the number of processing times of context adaptive coding is within a limit of the number of processing times, skips encoding of the coefficient information flag when the number of processing times is not within the limit of the number of processing times, encodes residual value information for reconstructing the value of the coefficient using the coefficient information flag by Golombirai coding when the coefficient information flag is encoded, and encodes the value of the coefficient by Golombirai coding when the encoding of the coefficient information flag is skipped.

[0114] This makes it possible to skip encoding of coefficient information flags, regardless of whether or not an orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive coding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0115] Furthermore, for example, the coefficient information flag is a flag indicating whether the value of the coefficient is greater than 1.

[0116] Thus, regardless of whether orthogonal transform is applied, it is possible to skip encoding the coefficient information flag indicating whether the coefficient value is greater than 1, in accordance with the limit on the number of context-adaptive coding processes. This makes it possible to suppress increases in processing delay and the amount of code.

[0117] In addition, for example, a decoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit, when in operation, decodes a coefficient information flag indicating the attributes of a coefficient contained in the block by context adaptive decoding in both cases of applying an inverse orthogonal transform to a block of a decoding target image and not applying an inverse orthogonal transform to the block, if the number of processing times of context adaptive decoding is within a limit of the number of processing times; skips decoding of the coefficient information flag when the number of processing times is not within the limit of the number of processing times; decodes residual value information for reconstructing the value of the coefficient using the coefficient information flag by Golombirai decoding when the coefficient information flag is decoded; and decodes the value of the coefficient by Golombirai decoding when the decoding of the coefficient information flag is skipped.

[0118] This makes it possible to skip decoding of coefficient information flags, regardless of whether an inverse orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive decoding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0119] Furthermore, for example, the coefficient information flag is a flag indicating whether the value of the coefficient is greater than 1.

[0120] Thus, regardless of whether an inverse orthogonal transform is applied, it is possible to skip decoding the coefficient information flag indicating whether the coefficient value is greater than 1, in accordance with the limit on the number of times context-adaptive decoding is processed. This makes it possible to suppress increases in processing delay and the amount of code.

[0121] In addition, for example, one form of the encoding method of the present invention is, in both cases of applying an orthogonal transform to a block of an encoding object image and not applying an orthogonal transform to the block, when the number of processing times of context adaptive coding is within a limit range of the processing times, encoding of a coefficient information flag indicating the attributes of a coefficient contained in the block is performed by context adaptive coding, when the number of processing times is not within the limit range of the processing times, skipping encoding of the coefficient information flag, when the coefficient information flag is encoded, encoding of residual value information for reconstructing the value of the coefficient using the coefficient information flag by Golombirai coding, and when encoding of the coefficient information flag is skipped, encoding the value of the coefficient by Golombirai coding.

[0122] This makes it possible to skip encoding of coefficient information flags, regardless of whether or not an orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive coding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0123] In addition, for example, a decoding method of one form of the present invention, in both cases of applying an inverse orthogonal transform to a block of a decoding object image and not applying an inverse orthogonal transform to the block, decodes a coefficient information flag indicating the attributes of a coefficient contained in the block by context adaptive decoding when the number of processing times of context adaptive decoding is within a limit range of the processing times; skips decoding of the coefficient information flag when the coefficient information flag is decoded, decodes residual value information used to reconstruct the value of the coefficient using the coefficient information flag by Columbus decoding, and decodes the value of the coefficient by Columbus decoding when the decoding of the coefficient information flag is skipped.

[0124] This makes it possible to skip decoding of coefficient information flags, regardless of whether an inverse orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive decoding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0125] In addition, for example, an encoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit, when in operation, limits the number of context adaptive coding processing times and encodes a block of an image, and in the encoding of the block, when no orthogonal transform is applied to the block, it is determined whether a plurality of coefficient information flags respectively representing a plurality of attributes of the coefficients contained in the block satisfy a processing condition, and when it is determined that the processing condition is satisfied, the plurality of coefficient information flags are encoded by context adaptive coding, and the processing condition is the following condition: the number of processing times when the number of the plurality of coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0126] Thus, when an orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive coding can be used for multiple coefficient information flags. This simplifies processing and reduces processing delay. Furthermore, when similar processing is performed on blocks to which an orthogonal transform is applied, the difference between the coding method used in blocks to which an orthogonal transform is applied and the coding method used in blocks not to which an orthogonal transform is applied can be reduced, and the circuit scale can be reduced.

[0127] In addition, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0128] This makes it possible to collectively determine a plurality of coefficient information flags, including a coefficient information flag indicating whether the coefficient value is greater than 3 and a coefficient information flag indicating whether the coefficient value is greater than 5. Consequently, processing can be simplified and processing delay can be reduced.

[0129] In addition, for example, the plurality of coefficient information flags further include a coefficient information flag indicating whether the value of the coefficient is greater than 7 and a coefficient information flag indicating whether the value of the coefficient is greater than 9.

[0130] This makes it possible to collectively determine multiple coefficient information flags, including four coefficient information flags: whether the coefficient value is greater than 3, whether the coefficient value is greater than 5, whether the coefficient value is greater than 7, and whether the coefficient value is greater than 9. This simplifies processing and reduces processing delay.

[0131] In addition, for example, a decoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit, wherein the circuit, when in operation, limits the number of processing times of context adaptive decoding and decodes a block of an image, and in decoding the block, when no inverse orthogonal transform is applied to the block, determines whether a plurality of coefficient information flags respectively representing a plurality of attributes of coefficients contained in the block satisfy a processing condition, and when it is determined that the processing condition is satisfied, decodes the plurality of coefficient information flags through context adaptive decoding, and the processing condition is the following condition: the number of processing times when the number of the plurality of coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0132] Thus, even when an inverse orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive decoding can be used for multiple coefficient information flags. Consequently, processing can be simplified and processing delay can be reduced. Furthermore, when similar processing is performed on blocks to which an inverse orthogonal transform is applied, the difference between the decoding method used in blocks to which an inverse orthogonal transform is applied and the decoding method used in blocks not to which an inverse orthogonal transform is applied can be minimized, and the circuit scale can be reduced.

[0133] In addition, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0134] This makes it possible to collectively determine a plurality of coefficient information flags, including a coefficient information flag indicating whether the coefficient value is greater than 3 and a coefficient information flag indicating whether the coefficient value is greater than 5. Consequently, processing can be simplified and processing delay can be reduced.

[0135] In addition, for example, the plurality of coefficient information flags further include a coefficient information flag indicating whether the value of the coefficient is greater than 7 and a coefficient information flag indicating whether the value of the coefficient is greater than 9.

[0136] This makes it possible to collectively determine multiple coefficient information flags, including four coefficient information flags: whether the coefficient value is greater than 3, whether the coefficient value is greater than 5, whether the coefficient value is greater than 7, and whether the coefficient value is greater than 9. This simplifies processing and reduces processing delay.

[0137] In addition, for example, one form of the encoding method of the present invention is to limit the number of context adaptive coding processing times and encode blocks of an image. In the encoding of the block, when no orthogonal transform is applied to the block, it is determined whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the block meet the processing conditions. When it is determined that the processing conditions are met, the multiple coefficient information flags are encoded by context adaptive coding. The processing condition is the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0138] Thus, when an orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive coding can be used for multiple coefficient information flags. This simplifies processing and reduces processing delay. Furthermore, when similar processing is performed on blocks to which an orthogonal transform is applied, the difference between the coding method used in blocks to which an orthogonal transform is applied and the coding method used in blocks not to which an orthogonal transform is applied can be reduced, and the circuit scale can be reduced.

[0139] In addition, for example, one form of a decoding method of the present invention is to decode a block of an image while limiting the number of processing times of context adaptive decoding. In the decoding of the block, without applying an inverse orthogonal transform to the block, it is determined whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the block meet a processing condition. If it is determined that the processing condition is met, the multiple coefficient information flags are decoded by context adaptive decoding, and the processing condition is the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0140] Thus, even when an inverse orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive decoding can be used for multiple coefficient information flags. Consequently, processing can be simplified and processing delay can be reduced. Furthermore, when similar processing is performed on blocks to which an inverse orthogonal transform is applied, the difference between the decoding method used in blocks to which an inverse orthogonal transform is applied and the decoding method used in blocks not to which an inverse orthogonal transform is applied can be minimized, and the circuit scale can be reduced.

[0141] Furthermore, for example, an encoding device according to one aspect of the present invention includes a partitioning unit, an intra-frame prediction unit, an inter-frame prediction unit, a prediction control unit, a transform unit, a quantization unit, an entropy encoding unit, and a loop filter unit.

[0142] The segmentation unit segments a current picture constituting the moving image into a plurality of blocks. The intra-frame prediction unit performs intra-frame prediction to generate the predicted image of the current block in the current picture using a reference image in the current picture. The inter-frame prediction unit performs inter-frame prediction to generate the predicted image of the current block using a reference image in a reference picture different from the current picture.

[0143] The prediction control unit controls intra-frame prediction performed by the intra-frame prediction unit and inter-frame prediction performed by the inter-frame prediction unit. The transformation unit transforms a prediction residual signal between the predicted image generated by the intra-frame prediction unit or the inter-frame prediction unit and the image of the encoding target block to generate a transform coefficient signal for the encoding target block. The quantization unit quantizes the transform coefficient signal. The entropy coding unit encodes the quantized transform coefficient signal. The loop filter unit applies a filter to the encoding target block.

[0144] In addition, for example, the entropy coding unit encodes a block of the image while limiting the number of context adaptive coding processing times in operation. In the encoding of the block, in both cases of applying an orthogonal transform to the block and not applying an orthogonal transform to the block, a sub-block flag encoding process is performed without including it in the number of processing times. The sub-block flag encoding process encodes a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context adaptive coding.

[0145] In addition, for example, when the entropy coding unit is in operation, in both cases of applying an orthogonal transform to a block of the encoding object image and not applying an orthogonal transform to the block, when the number of processing times of context adaptive coding is within the limit of the processing times, the coefficient information flag indicating the attributes of the coefficient contained in the block is encoded by context adaptive coding; when the number of processing times is not within the limit of the processing times, the encoding of the coefficient information flag is skipped; when the coefficient information flag is encoded, the residual value information used to reconstruct the value of the coefficient using the coefficient information flag is encoded by Golombirai coding; when the encoding of the coefficient information flag is skipped, the value of the coefficient is encoded by Golombirai coding.

[0146] In addition, for example, the entropy coding unit limits the number of context adaptive coding processing times during operation and encodes blocks of the image. In the encoding of the blocks, when no orthogonal transform is applied to the blocks, it is determined whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the blocks meet processing conditions. When it is determined that the processing conditions are met, the multiple coefficient information flags are encoded by context adaptive coding. The processing conditions are as follows: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0147] In addition, for example, a decoding device of one form of the present invention uses a predicted image to decode a moving image, and the decoding device has an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra-frame prediction unit, an inter-frame prediction unit, a prediction control unit, an addition unit (reconstruction unit) and a loop filter unit.

[0148] The entropy decoding unit decodes a quantized transform coefficient signal of a decoding target block in a decoding target picture constituting the moving image. The inverse quantization unit inversely quantizes the quantized transform coefficient signal. The inverse transform unit inversely transforms the transform coefficient signal to obtain a prediction residual signal of the decoding target block.

[0149] The intra prediction unit performs intra prediction for generating the predicted image of the current block using a reference image in the current picture. The inter prediction unit performs inter prediction for generating the predicted image of the current block using a reference image in a reference picture different from the current picture. The prediction control unit controls the intra prediction performed by the intra prediction unit and the inter prediction performed by the inter prediction unit.

[0150] The adding unit reconstructs an image of the decoding target block by adding the predicted image generated by the intra prediction unit or the inter prediction unit to the prediction residual signal. The loop filtering unit applies a filter to the decoding target block.

[0151] In addition, for example, the entropy decoding unit limits the number of context adaptive decoding processing times during operation and decodes a block of the image. In decoding the block, in both cases of applying an inverse orthogonal transform to the block and not applying an inverse orthogonal transform to the block, a sub-block flag decoding process is performed without including it in the number of processing times. The sub-block flag decoding process decodes a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context adaptive decoding.

[0152] In addition, for example, when the entropy decoding unit is in operation, in both cases of applying an inverse orthogonal transform to a block of the decoding object image and not applying an inverse orthogonal transform to the block, when the number of processing times of context adaptive decoding is within the limit of the processing times, the coefficient information flag indicating the attributes of the coefficient contained in the block is decoded by context adaptive decoding; when the number of processing times is not within the limit of the processing times, the decoding of the coefficient information flag is skipped; when the coefficient information flag is decoded, the residual value information used to reconstruct the value of the coefficient using the coefficient information flag is decoded by Golombirai decoding; when the decoding of the coefficient information flag is skipped, the value of the coefficient is decoded by Golombirai decoding.

[0153] In addition, for example, the entropy decoding unit limits the number of context adaptive decoding processing times during operation and decodes a block of the image. In the decoding of the block, without applying an inverse orthogonal transform to the block, it is determined whether a plurality of coefficient information flags respectively representing a plurality of attributes of the coefficients contained in the block satisfy a processing condition. If it is determined that the processing condition is satisfied, the plurality of coefficient information flags are decoded by context adaptive decoding. The processing condition is the following condition: the number of processing times when the number of the plurality of coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0154] Moreover, these inclusive or specific forms may be implemented by systems, devices, methods, integrated circuits, computer programs, or non-transitory recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0155] The following embodiments are described in detail with reference to the accompanying drawings. The embodiments described below are intended to be inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connections of components, steps, and the relationships and order of steps shown in the following embodiments are merely examples and are not intended to limit the claims.

[0156] The following describes embodiments of encoding and decoding devices. The embodiments are examples of encoding and decoding devices to which the processing and / or structures described in the various aspects of the present invention can be applied. The processing and / or structures can also be implemented in encoding and decoding devices that differ from the embodiments. For example, the processing and / or structures applied to the embodiments may include any of the following.

[0157] (1) Any of the multiple components of the encoding device or decoding device described in the embodiments of the present invention may be replaced by another component described in any of the embodiments of the present invention, or these components may be combined.

[0158] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be modified by adding, replacing, deleting, or other arbitrary changes. For example, any function or process may be replaced by another function or process described in any of the various aspects of the present invention, or these functions or processes may be combined.

[0159] (3) In the method implemented by the encoding device or decoding device of the embodiment, any changes such as addition, replacement, or deletion may be made to a portion of the multiple processes included in the method. For example, any process in the method may be replaced with another process described in one of the various aspects of the present invention, or these processes may be combined.

[0160] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with a component described in any of the aspects of the present invention, may be combined with a component having a portion of the functions described in any of the aspects of the present invention, or may be combined with a component that performs a portion of the processing performed by a component described in any of the aspects of the present invention.

[0161] (5) A component having a portion of the functions of the encoding device or decoding device of the embodiment, or a component implementing a portion of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced with a component described in any of the aspects of the present invention, a component having a portion of the functions described in any of the aspects of the present invention, or a component implementing a portion of the processing described in any of the aspects of the present invention;

[0162] (6) In a method implemented by an encoding device or decoding device according to an embodiment, one of the multiple processes included in the method is replaced by one of the processes described in each aspect of the present invention or a similar process, or a combination of these processes;

[0163] (7) Some of the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the process described in any of the aspects of the present invention.

[0164] (8) The implementation of the processes and / or structures described in the various aspects of the present invention is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or structures may be implemented in a device used for a purpose different from that of the video encoding or decoding disclosed in the embodiments.

[0165] [Encoding device]

[0166] First, the encoding device according to the embodiment will be described. Figure 1 1 is a block diagram showing the functional structure of the encoding device 100 according to the embodiment. The encoding device 100 is a moving picture encoding device that encodes a moving picture in units of blocks.

[0167] like Figure 1 As shown, the encoding device 100 is a device that encodes an image in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126 and a prediction control unit 128.

[0168] The encoding device 100 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes a software program stored in the memory, the processor functions as the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.

[0169] Below, after describing the overall processing flow of the encoding device 100 , each component included in the encoding device 100 will be described.

[0170] [Overall flow of encoding processing]

[0171] Figure 2 3 is a flowchart showing an example of the overall encoding process performed by the encoding device 100 .

[0172] First, the segmentation unit 102 of the encoding device 100 segments each picture included in an input image, which is a moving image, into a plurality of fixed-size blocks (e.g., 128×128 pixels) (step Sa_1). The segmentation unit 102 then selects a segmentation pattern (also called a block shape) for each of the fixed-size blocks (step Sa_2). Specifically, the segmentation unit 102 further segments the fixed-size blocks into a plurality of blocks that conform to the selected segmentation pattern. The encoding device 100 then performs steps Sa_3 through Sa_9 on each of the plurality of blocks (i.e., the encoding target block).

[0173] That is, the prediction processing unit composed of all or part of the intra-frame prediction unit 124, the inter-frame prediction unit 126 and the prediction control unit 128 generates a prediction signal (also called a prediction block) of the encoding target block (also called the current block) (step Sa_3).

[0174] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).

[0175] Next, the transform unit 106 and the quantization unit 108 transform and quantize the difference block to generate a plurality of quantized coefficients (step Sa_5). In addition, a block composed of a plurality of quantized coefficients is also called a coefficient block.

[0176] Next, the entropy coding unit 110 encodes the coefficient block and the prediction parameters related to the generation of the prediction signal (specifically, entropy coding) to generate a coded signal (step Sa_6). In addition, the coded signal is also called a coded bit stream, a compressed bit stream, or a stream.

[0177] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore a plurality of prediction residuals (ie, difference blocks) by performing inverse quantization and inverse transformation on the coefficient block (step Sa_7).

[0178] Next, the adding unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the prediction block to the restored difference block (step Sa_8).

[0179] When the reconstructed image is generated, the loop filter unit 120 filters the reconstructed image as necessary (step Sa_9).

[0180] Then, the encoding device 100 determines whether encoding of the entire picture is completed (step Sa_10 ), and if it is determined that encoding is not completed (No in step Sa_10 ), it repeats the process from step Sa_2 .

[0181] Furthermore, in the above example, encoding device 100 selects a single partitioning pattern for fixed-size blocks and encodes each block according to that partitioning pattern. However, encoding device 100 may also encode each block according to each of a plurality of partitioning patterns. In this case, encoding device 100 may evaluate the cost for each of the plurality of partitioning patterns and, for example, select as the output encoded signal the encoded signal obtained by encoding according to the partitioning pattern with the lowest cost.

[0182] As shown in the figure, the processes of steps Sa_1 to Sa_10 are sequentially performed by the encoding device 100. Alternatively, a part of a plurality of these processes may be performed in parallel, or the order of these processes may be reversed.

[0183] [Division]

[0184] The segmentation unit 102 segments each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the picture into blocks of a fixed size (e.g., 128×128). Other fixed block sizes may also be used. The fixed-size blocks are sometimes called coding tree units (CTUs). Furthermore, the segmentation unit 102 segments each fixed-size block into blocks of a variable size (e.g., 64×64 or less), for example, based on recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. The variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be used as processing units of CUs, PUs, or TUs.

[0185] Figure 3 This is a conceptual diagram showing an example of block division in an implementation manner. Figure 3 In the figure, the solid line represents the block boundary based on quadtree block partitioning, and the dotted line represents the block boundary based on binary tree block partitioning.

[0186] Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into four square 64×64 blocks (quadtree block division).

[0187] The 64×64 block on the upper left is further vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is further vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the 64×64 block on the upper left is divided into two 16×64 blocks 11 and 12 and a 32×64 block 13.

[0188] The 64×64 block in the upper right corner is horizontally partitioned into two rectangular 64×32 blocks 14 and 15 (binary tree block partitioning).

[0189] The 64×64 block on the lower left is divided into four square 32×32 blocks (quadtree block partitioning). The upper left and lower right blocks of the four 32×32 blocks are further partitioned. The upper left 32×32 block is vertically partitioned into two rectangular 16×32 blocks, and the right 16×32 block is further partitioned horizontally into two 16×16 blocks (binary tree block partitioning). The lower right 32×32 block is horizontally partitioned into two 32×16 blocks (binary tree block partitioning). As a result, the lower left 64×64 block is partitioned into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.

[0190] The lower right 64×64 block 23 is not split.

[0191] As above, in Figure 3 In FIG, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. This type of partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.

[0192] In addition, Figure 3 In the example above, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). Division including this ternary tree division is called MBT (multi type tree) division.

[0193] [Structural slices / tiles of the image]

[0194] In order to decode pictures in parallel, pictures may be constructed in slice units or tile units. The picture constructed in slice units or tile units can be constructed by the partitioning unit 102.

[0195] A slice is a basic coding unit that constitutes a picture. A picture is composed of one or more slices. In addition, a slice is composed of one or more consecutive CTUs (Coding Tree Units).

[0196] Figure 4A This is a conceptual diagram showing an example of the structure of a slice. For example, a picture includes 11×8 CTUs and is divided into 4 slices (slices 1 to 4). Slice 1 consists of 16 CTUs, slice 2 consists of 21 CTUs, slice 3 consists of 29 CTUs, and slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any slice. The shape of the slice becomes the shape that divides the picture in the horizontal direction. The boundary of the slice does not need to be the end of the picture, and can be any position in the boundary of the CTU in the picture. The processing order (encoding order or decoding order) of the CTU in the slice is, for example, a raster scan order. In addition, the slice includes header information and encoded data. The header information may also record the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.

[0197] A tile is a unit of rectangular area that constitutes a picture. Each tile may be assigned a number called a TileId in raster scan order.

[0198] Figure 4BThis is a conceptual diagram showing an example of a tile structure. For example, a picture includes 11×8 CTUs and is divided into four rectangular area tiles (tiles 1 to 4). When tiles are used, the processing order of CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed in raster scan order. When tiles are used, at least one CTU in each of the multiple tiles is processed in raster scan order. For example, Figure 4B As shown, the processing order of multiple CTUs included in tile 1 is from the left end of the 1st row of tile 1 to the right end of the 1st row of tile 1, and then from the left end of the 2nd row of tile 1 to the right end of the 2nd row of tile 1.

[0199] In addition, one tile may include more than one slice, and one slice may include more than one tile.

[0200] [Subtraction Department]

[0201] The subtraction unit 104 subtracts the prediction signal (prediction samples input from the prediction control unit 128, described below) from the original signal (original samples) in units of blocks input from the segmentation unit 102 and segmented by the segmentation unit 102. Specifically, the subtraction unit 104 calculates a prediction error (also referred to as a residual) for the current block to be coded (hereinafter referred to as the current block). The subtraction unit 104 then outputs the calculated prediction error (residual) to the transformation unit 106.

[0202] The original signal is an input signal to the encoding device 100 and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0203] [Conversion Unit]

[0204] The transform unit 106 transforms the spatial domain prediction error into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction error. The predetermined DCT or DST may also be predetermined.

[0205] Alternatively, the transform unit 106 may adaptively select a transform type from a plurality of transform types and transform the prediction error into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes referred to as an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0206] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transformation basis functions corresponding to the transformation type examples. Figure 5A Where N represents the number of input pixels. The selection of a transform type from among these multiple transform types may depend on, for example, the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode.

[0207] Information indicating whether such EMT or AMT is applied (e.g., an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. However, signaling of this information is not limited to the CU level and may be performed at other levels (e.g., bitstream level, picture level, slice level, tile level, or CTU level).

[0208] In addition, the transform unit 106 may also re-transform the transform coefficients (transform results). Such re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 re-transforms each sub-block (for example, a 4×4 sub-block) contained in the block of transform coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but can also be other levels (for example, sequence level, picture level, slice level, tile level or CTU level).

[0209] Separable transformation and non-separable transformation can also be applied in the transformation unit 106. Separable transformation refers to a method of performing multiple transformations in each direction corresponding to the number of dimensions of the input. Non-separable transformation refers to a method of treating two or more dimensions as one dimension when the input is multi-dimensional and transforming them together.

[0210] For example, as an example of non-separable transformation, when a 4×4 block is input, it is considered as an array of 16 elements, and the array is transformed using a 16×16 transformation matrix.

[0211] In a further example of the non-separable transformation, a 4×4 input block may be regarded as an array of 16 elements, and then a transformation (Hypercube Givens Transform) may be performed by performing multiple Givens rotations on the array.

[0212] In the transformation in the transformation unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, Figure 5B As shown, the CU is divided into two equal parts in the horizontal or vertical direction, and only the area on one side is transformed into the frequency area. The type of transformation base can be set for each area, for example, DST7 and DCT8. In this example, only one of the two areas in the CU is transformed, and the other is not transformed, but both areas can also be transformed. In addition, the division method is not limited to bisection, but can be more flexible, such as quartering or encoding the information indicating the division separately, and signaling it in the same way as the CU division. In addition, SVT is sometimes also called SBT (Sub-block Transform).

[0213] [Quantitative Department]

[0214] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112. The predetermined scanning order may also be predetermined.

[0215] The predetermined scanning order is the order used for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order may be defined in ascending order (from low frequency to high frequency) or descending order (from high frequency to low frequency) of frequency.

[0216] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0217] Quantization also sometimes uses a quantization matrix. For example, multiple quantization matrices are sometimes used to correspond to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. Furthermore, quantization refers to digitizing values ​​sampled at specified intervals and assigning them to specified levels. In this technical field, other representations such as rounding, integer scaling, and scaling may also be used. The specified intervals and levels may also be predetermined.

[0218] Methods for using quantization matrices include using a quantization matrix directly set on the encoding device side and using a default quantization matrix (default matrix). By directly setting the quantization matrix on the encoding device side, a quantization matrix suitable for image characteristics can be set. However, this method has the disadvantage of increasing the amount of code due to encoding the quantization matrix.

[0219] On the other hand, there is also a method of performing quantization so that the coefficients of high-frequency components and low-frequency components are the same without using a quantization matrix. This method is equivalent to using a quantization matrix (flat matrix) in which all coefficients have the same value.

[0220] The quantization matrix can be specified by, for example, an SPS (Sequence Parameter Set) or a PPS (Picture Parameter Set). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and PPS are sometimes referred to simply as parameter sets.

[0221] [Entropy coding unit]

[0222] The entropy coding unit 110 generates a coded signal (coded bit stream) based on the quantization coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110 binarizes the quantization coefficients, performs arithmetic coding on the binary signal, and outputs a compressed bit stream or sequence.

[0223] [Inverse quantization unit]

[0224] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. Furthermore, the inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114. The predetermined scanning order may also be predetermined.

[0225] [Inverse transformation unit]

[0226] The inverse transform unit 114 restores the prediction error (residual) by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform performed by the transform unit 106 on the transform coefficients. The inverse transform unit 114 then outputs the restored prediction error to the addition unit 116.

[0227] Furthermore, the restored prediction error generally loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error generally includes a quantization error.

[0228] [Addition Department]

[0229] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a locally decoded block.

[0230] [Block Memory]

[0231] The block memory 118 is a storage unit for storing blocks in a current picture to be coded (referred to as a current picture) to be referenced in intra prediction, for example. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116 .

[0232] [Frame Memory]

[0233] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120 .

[0234] [Loop filter unit]

[0235] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the encoding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0236] In ALF, a least squares error filter is used to remove coding distortion. For example, for each 2×2 sub-block in the current block, one filter is selected from multiple filters based on the direction and activity of the local gradient.

[0237] Specifically, sub-blocks (e.g., 2×2 sub-blocks) are first classified into multiple classes (e.g., 15 or 25 classes). Sub-block classification is performed based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4), a classification value C is calculated (e.g., C = 5D + A). Based on the classification value C, the sub-blocks are then classified into multiple classes.

[0238] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.

[0239] Based on the result of such classification, a filter to be used for the sub-block is determined from among a plurality of filters.

[0240] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 6A to 6C FIG. 1 is a diagram showing a plurality of examples of filter shapes used in ALF. Figure 6A represents a 5×5 diamond-shaped filter, Figure 6B represents a 7×7 diamond-shaped filter, Figure 6C Represents a 9×9 diamond-shaped filter. Information representing the filter shape is typically signaled at the picture level. However, signaling of the filter shape information need not be limited to the picture level and can also be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0241] The on / off of ALF can also be determined at the picture level or the CU level. For example, for luminance, whether to use ALF can be determined at the CU level, and for chrominance, whether to use ALF can be determined at the picture level. The information indicating whether ALF is on / off is usually signaled at the picture level or the CU level. In addition, the signaling of the information indicating whether ALF is on / off does not need to be limited to the picture level or the CU level, and can also be at other levels (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0242] The coefficient sets for multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level. However, the signaling of coefficient sets is not limited to the picture level and can also be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0243] [Loop Filter Section > Deblocking Filter]

[0244] In the deblocking filter, the loop filter unit 120 performs filtering processing on block boundaries of the reconstructed image to reduce distortion generated at the block boundaries.

[0245] Figure 7 This is a block diagram showing an example of a detailed configuration of the loop filter unit 120 functioning as a deblocking filter.

[0246] The loop filter unit 120 includes a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 , and 1206 .

[0247] The boundary determination unit 1201 determines whether or not there is a pixel (ie, a target pixel) to be subjected to deblocking filtering near a block boundary, and outputs the determination result to the switch 1202 and the processing determination unit 1208 .

[0248] When the boundary determination unit 1201 determines that the target pixel exists near a block boundary, the switch 1202 outputs the image before filtering to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel does not exist near a block boundary, the switch 1202 outputs the image before filtering to the switch 1206.

[0249] The filter determination unit 1203 determines whether to perform deblocking filtering on the target pixel based on the pixel value of at least one surrounding pixel located around the target pixel, and then outputs the determination result to the switch 1204 and the processing determination unit 1208 .

[0250] When the filter determination unit 1203 determines that the deblocking filtering process is to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that the deblocking filtering process is not to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the switch 1206.

[0251] When the pre-filtered image is obtained via switches 1202 and 1204 , the filter processing unit 1205 performs deblocking filtering on the target pixel using the filter characteristics determined by the filter characteristic determination unit 1207 . The filter processing unit 1205 then outputs the filtered pixel to the switch 1206 .

[0252] According to the control of the processing determination unit 1208 , the switch 1206 selectively outputs pixels that have not been processed by the deblocking filter and pixels that have been processed by the deblocking filter by the filter processing unit 1205 .

[0253] The processing determination unit 1208 controls the switch 1206 based on the determination results of the boundary determination unit 1201 and the filter determination unit 1203. Specifically, when the boundary determination unit 1201 determines that the target pixel is located near a block boundary and the filter determination unit 1203 determines that the target pixel is to be subjected to deblocking filtering, the processing determination unit 1208 outputs the pixel subjected to deblocking filtering from the switch 1206. Otherwise, in other cases, the processing determination unit 1208 outputs the pixel that has not been subjected to deblocking / filtering from the switch 1206. By repeating this pixel outputting process, the filtered image is output from the switch 1206.

[0254] Figure 8 This is a conceptual diagram showing an example of a deblocking filter having a filter characteristic that is symmetric with respect to a block boundary.

[0255] In the deblocking filter process, for example, using pixel values ​​and quantization parameters, one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter, is selected. Figure 8 As shown, when pixels p0 to p2 and pixels q0 to q2 exist across a block boundary, the pixel values ​​of pixels q0 to q2 are changed to pixel values ​​q'0 to q'2 by performing the calculation shown in the following equation, for example.

[0256] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8

[0257] q'1=(p0+q0+q1+q2+2) / 4

[0258] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0259] In the above equations, p0 through p2 and q0 through q2 are the pixel values ​​of pixels p0 through p2 and q0 through q2, respectively. Furthermore, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the side opposite the block boundary. On the right side of each of the above equations, the coefficients multiplied by the pixel values ​​of each pixel used in the deblocking filter process are the filter coefficients.

[0260] Furthermore, during deblocking filtering, clipping can be performed to ensure that the calculated pixel value does not exceed a threshold. In this clipping process, the pixel value calculated based on the above equation is clipped to "the calculated pixel value ± 2 × the threshold" using a threshold determined by the quantization parameter. This prevents excessive smoothing.

[0261] Figure 9 This is a conceptual diagram for explaining the block boundary on which the deblocking filtering process is performed. Figure 10 This is a conceptual diagram showing an example of the Bs value.

[0262] The block boundary for deblocking filtering is, for example, Figure 9 The deblocking filter can be performed in units of 4 rows or 4 columns. First, for Figure 9 The block P and block Q shown are as follows: Figure 10 That determines the Bs (Boundary Strength) value.

[0263] according to Figure 10 The Bs value determines whether to perform deblocking filtering with different intensities even at block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering is performed on the color difference signal. When the Bs value is 1 or more and the specified conditions are met, deblocking filtering is performed on the luminance signal. The specified conditions can also be predetermined. In addition, the determination conditions of the Bs value are not limited to Figure 10 The conditions shown can also be determined based on other parameters.

[0264] [Prediction Processing Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]

[0265] Figure 11 12 is a flowchart showing an example of processing performed by the prediction processing unit of the encoding device 100. The prediction processing unit is composed of all or part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0266] The prediction processing unit generates a predicted image for the current block (step Sb_1). This predicted image is also referred to as a prediction signal or a prediction block. Prediction signals include, for example, intra-frame prediction signals and inter-frame prediction signals. Specifically, the prediction processing unit generates a predicted image for the current block using the reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0267] The reconstructed image may be, for example, an image of a reference picture, or an image including the current block, that is, an image of an already coded block within the current picture. The already coded block within the current picture may be, for example, an adjacent block of the current block.

[0268] Figure 12 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device 100.

[0269] The prediction processing unit generates a predicted image using the first method (step Sc_1a), generates a predicted image using the second method (step Sc_1b), and generates a predicted image using the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating predicted images, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may also be used in these prediction methods.

[0270] Next, the prediction processing unit selects any one of the multiple prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). The selection of the prediction image, that is, the selection of the method or mode for obtaining the final prediction image can also be performed by calculating the cost for each prediction image generated and based on the cost. In addition, the selection of the prediction image can be performed based on the parameters used for the encoding process. The encoding device 100 can signal information for determining the selected prediction image, method or mode into a coded signal (also called a coded bit stream). The information can be, for example, a flag, etc. Thus, the decoding device can generate a prediction image in accordance with the method or mode selected in the encoding device 100 based on the information. In addition, in Figure 12 In the example shown, the prediction processing unit selects one of the predicted images after generating them using various methods. However, before generating these predicted images, the prediction processing unit may select a method or mode based on the parameters used in the encoding process described above and generate the predicted images based on that method or mode.

[0271] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction processing unit may select a final prediction image for the current block from prediction images generated according to these prediction methods.

[0272] Figure 13 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device 100.

[0273] First, the prediction processing unit generates a predicted image through intra-frame prediction (step Sd_1a), and generates a predicted image through inter-frame prediction (step Sd_1b). In addition, the predicted image generated through intra-frame prediction is also called an intra-frame predicted image, and the predicted image generated through inter-frame prediction is also called an inter-frame predicted image.

[0274] Next, the prediction processing unit evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). Cost can also be used in this evaluation. That is, the prediction processing unit calculates the cost C of each of the intra-frame prediction image and the inter-frame prediction image. The cost C can be calculated by an equation of the RD optimization model, such as C=D+λ×R. In this equation, D is the coding distortion of the predicted image, and is represented by, for example, the sum of the absolute values ​​of the differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. In addition, R is the amount of generated coding of the predicted image, specifically, the amount of coding required for encoding motion information, etc. for generating the predicted image. In addition, λ is, for example, an undetermined multiplier of Lagrange.

[0275] Then, the prediction processing unit selects the prediction image with the minimum cost C from the intra-frame prediction image and the inter-frame prediction image as the final prediction image of the current block (step Sd_3). In other words, the prediction method or mode for generating the prediction image of the current block is selected.

[0276] [Intra-frame prediction unit]

[0277] The intra prediction unit 124 performs intra prediction (also called intra-screen prediction) on the current block by referring to blocks in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra prediction unit 124 generates the intra-prediction signal by performing intra prediction with reference to samples (e.g., luminance values ​​and chrominance values) of blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 128.

[0278] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes. The plurality of predetermined modes may also be predetermined.

[0279] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.

[0280] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. Furthermore, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 14 This is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent the 33 directions specified by the H.265 / HEVC specification, and the dotted arrows represent the additional 32 directions (2 non-directional prediction modes in Figure 14 (not shown in the figure).

[0281] In various processing examples, intra prediction of chrominance blocks can also refer to luma blocks. That is, the chrominance component of the current block can be predicted based on the luma component of the current block. This type of intra prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra prediction mode for chrominance blocks that refers to luma blocks (e.g., CCLM mode) can also be added as one of the intra prediction modes for chrominance blocks.

[0282] The intra-frame prediction unit 124 may also modify the intra-predicted pixel values ​​based on the gradients of reference pixels in the horizontal and vertical directions. Intra-frame prediction with such modification is sometimes called PDPC (position-dependent intraprediction combination). Information indicating whether PDPC is used (e.g., a PDPC flag) is typically signaled at the CU level. However, this information is not necessarily signaled at the CU level and may be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0283] [Inter-frame prediction unit]

[0284] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) on the current block by referring to a reference picture stored in the frame memory 122, which is different from the current picture, to generate a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed on the current block or current sub-block within the current block (e.g., a 4×4 block). For example, the inter-frame prediction unit 126 performs motion estimation on the current block or current sub-block within the reference picture to find the reference block or sub-block that most closely matches the current block or sub-block. Furthermore, the inter-frame prediction unit 126 obtains motion information (e.g., a motion vector) to compensate for motion or changes from the reference block or sub-block to the current block or sub-block. Based on this motion information, the inter-frame prediction unit 126 performs motion compensation (or motion prediction) to generate an inter-frame prediction signal for the current block or sub-block. The inter-frame prediction unit 126 then outputs the generated inter-frame prediction signal to the prediction control unit 128.

[0285] The motion information used in motion compensation can be signaled as an inter-frame prediction signal in various forms. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.

[0286] [Basic process of inter-frame prediction]

[0287] Figure 15 This is a flowchart showing an example of a basic flow of inter-frame prediction.

[0288] The inter prediction unit 126 first generates a predicted image (steps Se_1 to Se_3 ). Next, the subtraction unit 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4 ).

[0289] Here, in generating a predicted image, the inter-frame prediction unit 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Furthermore, in determining the MV, the inter-frame prediction unit 126 determines the MV by selecting candidate motion vectors (candidate MVs) (step Se_1) and deriving the MV (step Se_2). The candidate MV is selected, for example, by selecting at least one candidate MV from a candidate MV list. Furthermore, in deriving the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from the at least one candidate MV and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by each of the at least one selected candidate MVs. The act of searching the region of the reference picture may also be referred to as motion estimation.

[0290] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126 . However, for example, the processing of step Se_1 or step Se_2 may be performed by other components included in the encoding device 100 .

[0291] [Flow of Motion Vector Derivation]

[0292] Figure 16 This is a flowchart showing an example of motion vector derivation.

[0293] The inter-frame prediction unit 126 derives the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information is encoded as prediction parameters and signaled. That is, the encoded motion information is included in the encoded signal (also called the encoded bitstream).

[0294] Alternatively, the inter prediction unit 126 derives the MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the encoded signal.

[0295] Here, the MV derivation modes may include the normal inter mode, merge mode, FRUC mode, and affine mode described later. Among these modes, the modes that encode motion information include the normal inter mode, merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Furthermore, motion information may include not only the MV but also the predicted motion vector selection information described later. Furthermore, modes that do not encode motion information include the FRUC mode. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.

[0296] Figure 17 This is a flowchart showing another example of motion vector derivation.

[0297] The inter-frame prediction unit 126 derives the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. In other words, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0298] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode in which the difference MV is not encoded. In this case, the encoded difference MV is not included in the encoded signal.

[0299] As described above, the MV derivation modes include the normal inter mode, merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the normal inter mode and affine mode (specifically, affine inter mode) encode the difference MV. Furthermore, the FRUC mode, merge mode, and affine mode (specifically, affine merge mode) do not encode the difference MV. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.

[0300] [Flow of Motion Vector Derivation]

[0301] Figure 18This is a flowchart showing another example of motion vector derivation. There are multiple modes for MV derivation, that is, inter-frame prediction modes, which are roughly divided into modes that encode differential MVs and modes that do not encode differential motion vectors. Modes that do not encode differential MVs include merge mode, FRUC mode, and affine mode (specifically, affine merge mode). The details of these modes will be described later. Simply put, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from the surrounding coded blocks, and the FRUC mode is a mode in which the MV of the current block is derived by searching between coded areas. In addition, the affine mode is a mode in which the motion vectors of each of the multiple sub-blocks constituting the current block are derived as the MV of the current block, assuming an affine transformation.

[0302] Specifically, as shown in the figure, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_2) based on the merge mode. In addition, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_3) based on the FRUC mode. In addition, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_4) based on the affine mode (specifically, the affine merge mode). In addition, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_5) based on the mode for encoding the differential MV (for example, the normal inter-frame mode).

[0303] [MV Export > Normal Interframe Mode]

[0304] Normal inter mode is an inter prediction mode that derives the MV of the current block from a block similar to the image of the current block in an area of ​​a reference picture indicated by a candidate MV. In this normal inter mode, a differential MV is encoded.

[0305] Figure 19 This is a flowchart showing an example of inter prediction based on the normal inter mode.

[0306] First, the inter prediction unit 126 obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter prediction unit 126 creates a candidate MV list.

[0307] Next, the inter-frame prediction unit 126 extracts N (N is an integer greater than or equal to 2) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as motion vector predictor candidates (also referred to as predicted MV candidates) in a predetermined order of priority (step Sg_2). Alternatively, the order of priority may be predetermined for each of the N candidate MVs.

[0308] Next, the inter-frame prediction unit 126 selects one motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor (also called predicted MV) for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor into the stream. The stream is the coded signal or coded bitstream described above.

[0309] Next, the inter-frame prediction unit 126 refers to the coded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference between the derived MV and the predicted motion vector as a differential MV into the stream. The coded reference picture is a picture composed of multiple blocks reconstructed after encoding.

[0310] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). Note that the predicted image is the inter-frame prediction signal described above.

[0311] Furthermore, information indicating the inter prediction mode (in the above example, the normal inter mode) used in generating the predicted image, which is included in the encoded signal, is encoded as, for example, a prediction parameter.

[0312] The candidate MV list can also be used in conjunction with lists used in other modes. Furthermore, processing related to the candidate MV list can be applied to processing related to lists used in other modes. Processing related to the candidate MV list includes, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.

[0313] [MV Export > Merge Mode]

[0314] The merge mode is an inter-frame prediction mode that derives the MV of the current block by selecting a candidate MV from a candidate MV list as the MV of the current block.

[0315] Figure 20 is a flowchart illustrating an example of inter-frame prediction based on merge mode.

[0316] First, the inter prediction unit 126 obtains multiple candidate MVs for the current block based on information about multiple coded block MVs located temporally or spatially around the current block (step Sh_1). In other words, the inter prediction unit 126 creates a candidate MV list.

[0317] Next, the inter prediction unit 126 selects one candidate MV from the plurality of candidate MVs acquired in step Sh_1 to derive the MV of the current block (step Sh_2). At this time, the inter prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream.

[0318] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).

[0319] Furthermore, information indicating the inter prediction mode (in the above example, the merge mode) used for generating the predicted image, which is included in the encoded signal, is encoded as, for example, a prediction parameter.

[0320] Figure 21 This is a conceptual diagram for explaining an example of motion vector derivation processing of the current picture based on the merge mode.

[0321] First, a prediction MV list is generated, which registers candidate prediction MVs. These include: spatially neighboring prediction MVs, which are MVs of multiple coded blocks located in the spatial vicinity of the target block; temporally neighboring prediction MVs, which are MVs of blocks near the target block's position in the coded reference picture; combined prediction MVs, which are generated by combining the MV values ​​of spatially neighboring prediction MVs and temporally neighboring prediction MVs; and zero prediction MVs, which have a value of zero.

[0322] Next, one predicted MV is selected from a plurality of predicted MVs registered in the predicted MV list to determine the MV of the target block.

[0323] Then, the variable length coding unit encodes merge_idx, which is a signal indicating which predicted MV is selected, by describing it in the stream.

[0324] In addition, Figure 21 The predicted MVs registered in the predicted MV list described in the figure are just examples, and may be a number different from the number in the figure, or a structure that does not include some of the types of predicted MVs in the figure, or a structure that adds predicted MVs other than the types of predicted MVs in the figure.

[0325] The final MV may be determined by performing a DMVR (decoder motion vector refinement) process described later using the MV of the target block derived in the merge mode.

[0326] In addition, the candidate for the predicted MV is the candidate MV mentioned above, and the predicted MV list is the candidate MV list mentioned above. In addition, the candidate MV list can also be called the candidate list. In addition, merge_idx is MV selection information.

[0327] [MV Export > FRUC Mode]

[0328] Motion information may be derived on the decoder side instead of being signaled on the encoder side. Furthermore, as described above, the merge mode specified in the H.265 / HEVC specification may be used. Furthermore, motion information may be derived, for example, by performing a motion search on the decoder side. In one embodiment, the decoder side performs a motion search without using the pixel values ​​of the current block.

[0329] Here, a mode for performing motion estimation on the decoding device side is described. This mode for performing motion estimation on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0330] In the form of a flow chart Figure 22An example of FRUC processing is shown in Figure 1. First, a list of multiple candidates (i.e., a candidate MV list, which may also be shared with a merge list) each including a predicted motion vector (MV) is generated by referring to the motion vectors of previously coded blocks that are spatially or temporally adjacent to the current block (step Si_1). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and one candidate is selected based on the evaluation value. Furthermore, a motion vector for the current block is derived based on the motion vector of the selected candidate (step Si_4). Specifically, for example, the selected candidate motion vector (the best candidate MV) is derived as is as the motion vector for the current block. Alternatively, the motion vector for the current block can be derived by performing pattern matching in the surrounding area of ​​the position in the reference picture corresponding to the selected candidate motion vector. Specifically, the surrounding area of ​​the best candidate MV can be searched using pattern matching and evaluation values ​​in the reference picture. If an MV with a better evaluation value is found, the best candidate MV is updated to the above MV and used as the final MV for the current block. A configuration may be adopted in which the process of updating to an MV having a better evaluation value is not performed.

[0331] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).

[0332] Exactly the same processing can be performed even when processing is performed in sub-block units.

[0333] The evaluation value can also be calculated using various methods. For example, the reconstructed image of the region within the reference picture corresponding to the motion vector is compared with the reconstructed image of a predetermined region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). The predetermined region may also be predetermined.

[0334] Then, the difference between the pixel values ​​of the two reconstructed images may be calculated and used for the motion vector evaluation value. Alternatively, the evaluation value may be calculated using other information in addition to the difference value.

[0335] Next, we'll explain a pattern matching example in detail. First, a candidate MV from a candidate MV list (e.g., a merge list) is selected as the starting point for a pattern matching search. For example, pattern matching can involve either a first pattern match or a second pattern match. First and second pattern matching are also known as bilateral matching and template matching, respectively.

[0336] [MV Export > FRUC > Bidirectional Matching]

[0337] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, the region within the other reference pictures along the motion trajectory of the current block is used as the predetermined region for calculating the candidate evaluation value. The predetermined region may also be predetermined.

[0338] Figure 23 This is a conceptual diagram for explaining an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along the motion trajectory. Figure 23 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best-matching pair among pairs of blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by a symmetric MV scaled by the display time interval is derived, and the evaluation value is calculated using the resulting difference. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs, achieving excellent results.

[0339] Assuming a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, mirror-symmetric bidirectional motion vectors are derived in the first pattern matching.

[0340] [MV Export > FRUC > Template Matching]

[0341] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the block above and / or to the left)) and a block in the reference picture. Therefore, in the second pattern matching, blocks adjacent to the current block in the current picture are used as the predetermined area for calculating the candidate evaluation value described above.

[0342] Figure 24This is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 24 As shown, in the second pattern matching, the motion vector of the current block is derived by searching the reference picture (Ref0) for a block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded area adjacent to the left and above, or one of the two, and the reconstructed image at the same position in the coded reference picture (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value can be selected as the best candidate MV from among multiple candidate MVs.

[0343] Such information indicating whether FRUC mode is adopted (e.g., called a FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is adopted (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (first pattern matching or second pattern matching) is signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and may be performed at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0344] [MV Export > Affine Mode]

[0345] Next, the affine mode for deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0346] Figure 25A This is a conceptual diagram for explaining an example of deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks. Figure 25A In the example, the current block consists of 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, according to the following formula (1A), the two motion vectors v0 and v1 can be projected, and the motion vectors (v x , v y ).

[0347]

Formula 1

[0348]

[0349] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predetermined weight coefficient. The predetermined weight coefficient may also be predetermined.

[0350] Information indicating this affine mode (e.g., called an affine flag) can be signaled as a CU-level signal. Furthermore, the signaling of information indicating this affine mode need not be limited to the CU-level, and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0351] Furthermore, such affine modes may include several modes that differ in how the motion vectors for the upper left and upper right control points are derived. For example, affine modes include affine inter (also called affine normal inter) mode and affine merge mode.

[0352] [MV Export > Affine Mode]

[0353] Figure 25B This is a conceptual diagram for explaining an example of derivation of a motion vector for a sub-block unit in an affine mode having three control points. Figure 25B In the example, the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vector of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vector of the adjacent block, and the motion vector v2 of the lower left corner control point of the current block is derived based on the motion vector of the adjacent block. Then, according to the following formula (1B), the three motion vectors v0, v1 and v2 can be projected, and the motion vectors (v x , v y ).

[0354]

Formula 2

[0355]

[0356] Here, x and y represent the horizontal and vertical positions of the sub-block center, respectively, w represents the width of the current block, and h represents the height of the current block.

[0357] Affine modes with different numbers of control points (e.g., 2 and 3) can also be switched and signaled at the CU level. Furthermore, information indicating the number of control points in the affine mode used at the CU level can also be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0358] Furthermore, in such an affine mode with three control points, several modes may be included, each with different methods for deriving motion vectors for the upper left, upper right, and lower left control points. For example, the affine mode includes the affine inter (also called affine normal inter) mode and the affine merge mode.

[0359] [MV Export > Affine Merge Mode]

[0360] Figure 26A 、 Figure 26B and Figure 26C This is a conceptual diagram used to illustrate the affine merge mode.

[0361] In affine merge mode, such as Figure 26A As shown, for example, based on multiple motion vectors corresponding to blocks coded in affine mode among the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, a predicted motion vector for each of the control points of the current block is calculated. Specifically, the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in this order to determine the first valid block coded in affine mode. The predicted motion vectors for the control points of the current block are calculated based on the multiple motion vectors corresponding to the determined blocks.

[0362] For example, Figure 26B As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 are derived that are projected onto the positions of the upper left corner and upper right corner of the encoded block including block A. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.

[0363] For example, Figure 26C As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, a predicted motion vector v0 of the control point at the upper left corner, a predicted motion vector v1 of the control point at the upper right corner, and a predicted motion vector v2 of the control point at the lower left corner of the current block are calculated.

[0364] In addition, in the following Figure 29 This method of deriving a predicted motion vector may also be used in deriving the predicted motion vectors of the control points of the current block in step Sj_1.

[0365] Figure 27 This is a flowchart showing an example of the affine merge mode.

[0366] In the affine merge mode, as shown in the figure, first, the inter-frame prediction unit 126 derives the predicted MV of each control point of the current block (step Sk_1). Figure 25A As shown, it is the top left and top right corner points of the current block, or as Figure 25B As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block.

[0367] That is to say, if Figure 26A As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.

[0368] Then, in the case where block A is determined and block A has 2 control points, as Figure 26B As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner and the motion vector v1 of the control point at the upper right corner of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block including the block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner and the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0369] Alternatively, in the case where block A is determined and block A has 3 control points, as Figure 26C As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner, the motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner, the predicted motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block.

[0370] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. Specifically, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two predicted motion vectors v0 and v1 and the above-mentioned equation (1A), or three predicted motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the motion vector for that sub-block as an affine MV (step Sk_2). The inter-frame prediction unit 126 then performs motion compensation on that sub-block using these affine MVs and the coded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image for the current block is generated.

[0371] [MV Export > Affine Inter-frame Mode]

[0372] Figure 28A This is a conceptual diagram for explaining the affine inter mode with two control points.

[0373] In this affine inter-frame mode, if Figure 28A As shown in FIG. 1 , a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0374] Figure 28B This is a conceptual diagram for explaining the affine inter-frame mode with three control points.

[0375] In this affine inter-frame mode, if Figure 28B As shown in FIG. 1 , a motion vector selected from the motion vectors of blocks A, B, and C that have been coded adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of blocks D and E that have been coded adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block. Furthermore, a motion vector selected from the motion vectors of blocks F and G that have been coded adjacent to the current block is used as the predicted motion vector v2 of the control point at the lower left corner of the current block.

[0376] Figure 29 This is a flowchart showing an example of the affine inter mode.

[0377] As shown in the figure, in the affine inter mode, first, the inter prediction unit 126 derives the predicted MV (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). Figure 25A or Figure 25B As shown, the control point is the point at the upper left corner, upper right corner or lower left corner of the current block.

[0378] That is, the inter-frame prediction unit 126 selects Figure 28A or Figure 28B The inter prediction unit 126 derives the predicted motion vector (v0, v1) or (v0, v1, v2) of the control point of the current block based on the motion vector of a block in the coded blocks near each control point of the current block. At this time, the inter prediction unit 126 encodes predicted motion vector selection information for identifying the two selected motion vectors into the stream.

[0379] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select as the predicted motion vector of the control point from the encoded blocks adjacent to the current block by using cost evaluation, etc., and can record a flag indicating which predicted motion vector is selected in the bitstream.

[0380] Next, the inter-frame prediction unit 126 performs a motion search (steps Sj_3 and Sj_4) while updating the predicted motion vector selected or derived in step Sj_1 (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vector of each sub-block corresponding to the predicted motion vector to be updated as an affine MV and calculates it using the above-mentioned equation (1A) or equation (1B) (step Sj_3). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on each sub-block (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference between the determined MV and the predicted motion vector as a differential MV into the stream.

[0381] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).

[0382] [MV Export > Affine Inter-frame Mode]

[0383] When affine modes with different numbers of control points (for example, 2 and 3) are switched at the CU level for signaling, the number of control points may differ between the coded block and the current block. Figure 30A as well as Figure 30B This is a conceptual diagram for explaining a method for deriving a prediction vector for control points when the number of control points in an already coded block and a current block is different.

[0384] For example, Figure 30AAs shown in FIG. 1 , when the current block has three control points, namely, the upper left corner, the upper right corner, and the lower left corner, and the block A adjacent to the left of the current block is coded in an affine mode having two control points, motion vectors v3 and v4 are derived, projected onto the positions of the upper left corner and the upper right corner of the coded block including block A. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated. Furthermore, the predicted motion vector v2 of the control point at the lower left corner is calculated based on the derived motion vectors v0 and v1.

[0385] For example, Figure 30B As shown in FIG, when the current block has two control points, namely, the upper left corner and the upper right corner, and the block A adjacent to the left of the current block is coded in an affine mode having three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the coded block including block A. Then, based on the derived motion vectors v3, v4, and v5, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.

[0386] exist Figure 29 This method of deriving a predicted motion vector may also be used in deriving each predicted motion vector of the control point of the current block in step Sj_1.

[0387] [MV Export>DMVR]

[0388] Figure 31A This is a flowchart showing the relationship between the merge mode and DMVR.

[0389] The inter-frame prediction unit 126 derives a motion vector for the current block in merge mode (step S1_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, i.e., a motion search (step S1_2). If it is determined not to perform a motion search (No in step S1_2), the inter-frame prediction unit 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). In other words, in this case, the motion vector for the current block is determined in merge mode.

[0390] On the other hand, if it is determined in step S1_1 that a motion search is to be performed (Yes in step S1_2), the inter-frame prediction unit 126 searches the surrounding area of ​​the reference picture represented by the motion vector derived in step S1_1 to derive a final motion vector for the current block (step S1_3). In other words, in this case, the motion vector of the current block is determined by DMVR.

[0391] Figure 31B This is a conceptual diagram for explaining an example of DMVR processing for determining MV.

[0392] First, the best MVP set for the current block (for example, in merge mode) is set as a candidate MV. Next, based on the candidate MV (L0), reference pixels are determined from the first reference picture (L0), which is the coded picture in the L0 direction. Similarly, based on the candidate MV (L1), reference pixels are determined from the second reference picture (L1), which is the coded picture in the L1 direction. A template is generated by averaging these reference pixels.

[0393] Next, using the template, the surrounding areas of the candidate MVs in the first reference image (L0) and the second reference image (L1) are searched, and the MV with the lowest cost is determined as the final MV. Alternatively, the cost value can be calculated using, for example, the difference between the pixel values ​​of the template and the pixel values ​​of the search area, as well as the candidate MV values.

[0394] Typically, the configuration and operation of the processing described here are basically common in the encoding device and the decoding device described later.

[0395] Even if it is not the processing example described here, any processing can be used as long as it is a processing that can search the vicinity of the candidate MV and derive the final MV.

[0396] [Motion Compensation > BIO / OBMC]

[0397] In motion compensation, there are modes for generating a predicted image and then correcting the predicted image. Examples of such modes include BIO and OBMC, which will be described later.

[0398] Figure 32 This is a flowchart showing an example of generating a predicted image.

[0399] The inter prediction unit 126 generates a predicted image (step Sm_1 ), and corrects the predicted image using, for example, any of the above-described modes (step Sm_2 ).

[0400] Figure 33 This is a flowchart showing another example of generating a predicted image.

[0401] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, if it is determined that a correction process is to be performed (yes in step Sn_3), the inter-frame prediction unit 126 corrects the predicted image to generate a final predicted image (step Sn_4). On the other hand, if it is determined that a correction process is not to be performed (no in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correction (step Sn_5).

[0402] Furthermore, in motion compensation, there is a mode for correcting brightness when generating a predicted image. This mode is, for example, LIC, which will be described later.

[0403] Figure 34 This is a flowchart showing another example of generating a predicted image.

[0404] The inter-frame prediction unit 126 derives the motion vector for the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform brightness correction processing (step So_2). If it is determined that brightness correction processing is to be performed (yes in step So_2), the inter-frame prediction unit 126 generates a predicted image while performing brightness correction (step So_3). In other words, the predicted image is generated using LIC. On the other hand, if it is determined that brightness correction processing is not to be performed (no in step So_2), the inter-frame prediction unit 126 generates a predicted image using standard motion compensation without performing brightness correction (step So_4).

[0405] [Motion Compensation > OBMC]

[0406] Inter-frame prediction signals can be generated using not only the motion information of the current block obtained through motion search, but also the motion information of neighboring blocks. Specifically, inter-frame prediction signals can be generated in sub-block units within the current block by weighted addition of prediction signals based on motion information obtained through motion search (within the reference picture) and prediction signals based on motion information of neighboring blocks (within the current picture). This type of inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0407] In OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) can also be signaled at the sequence level. Furthermore, information indicating whether OBMC mode is applied (e.g., OBMC flag) can also be signaled at the CU level. Furthermore, the signaling level for this information need not be limited to the sequence and CU levels and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0408] An example of the OBMC mode will be described in more detail. Figure 35 and Figure 36 It is a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.

[0409] First, if Figure 36 As shown in FIG, the motion vector (MV) assigned to the processing target (current) block is used to obtain the predicted image (Pred) based on the usual motion compensation. Figure 36 In FIG, the arrow “MV” points to the reference picture and indicates which block the current block of the current picture refers to to obtain the predicted image.

[0410] Next, the motion vector (MV_L) derived for the already coded left-adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The motion vector (MV_L) is represented by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by overlaying the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0411] Similarly, the motion vector (MV_U) derived for the previously coded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The motion vector (MV_U) is represented by an arrow "MV_U" pointing from the current block to the reference picture. The predicted image Pred_U is then overlaid with the predicted image (e.g., Pred and Pred_L) that has undergone the first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image for the current block, with the boundaries with adjacent blocks blended (smoothed).

[0412] In addition, the above example is a two-path correction method using left-adjacent and upper-adjacent blocks, but the correction method can also be a three-path correction method or more than three-path correction method using right-adjacent and / or lower-adjacent blocks.

[0413] Furthermore, the area to be overlapped may not be the entire pixel area of ​​the block, but may be only a partial area near the block boundary.

[0414] In this description, the predicted image correction process using OBMC is described as obtaining a single predicted image Pred by superimposing a single reference picture with the additional predicted images Pred_L and Pred_U. However, when correcting the predicted image based on multiple reference pictures, the same process can be applied to each of the multiple reference pictures. In this case, after performing OBMC image correction based on multiple reference pictures and obtaining a corrected predicted image from each reference picture, the final predicted image is obtained by further superimposing the obtained multiple corrected predicted images.

[0415] In OBMC, the target block unit may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.

[0416] One method for determining whether to apply OBMC processing involves using, for example, obmc_flag, a signal indicating whether OBMC processing should be applied. As a specific example, the encoding device may determine whether the target block belongs to a region with complex motion. If the target block belongs to a region with complex motion, the encoding device sets obmc_flag to a value of 1 and applies OBMC processing to the target block. If the target block does not belong to a region with complex motion, the encoding device sets obmc_flag to a value of 0 and does not apply OBMC processing to the target block. Meanwhile, the decoding device decodes obmc_flag described in a stream (e.g., a compressed sequence) and switches whether to apply OBMC processing based on the value of the flag during decoding.

[0417] In the above example, the inter-frame prediction unit 126 generates a single rectangular predicted image for a rectangular current block. However, the inter-frame prediction unit 126 may generate multiple predicted images having shapes other than a rectangle for the rectangular current block and may combine these multiple predicted images to generate a final rectangular predicted image. A shape other than a rectangle may be, for example, a triangle.

[0418] Figure 37 This is a conceptual diagram for explaining the generation of predicted images of two triangles.

[0419] The inter-frame prediction unit 126 performs motion compensation on the first triangular partition within the current block using the first MV of the first partition to generate a triangular predicted image. Similarly, the inter-frame prediction unit 126 performs motion compensation on the second triangular partition within the current block using the second MV of the second partition to generate a triangular predicted image. The inter-frame prediction unit 126 then combines these predicted images to generate a predicted image with the same rectangular shape as the current block.

[0420] In addition, Figure 37 In the example shown, the first partition and the second partition are each a triangle, but they may also be a trapezoid or have different shapes. Figure 37 In the example shown, the current block is composed of 2 partitions, but it can also be composed of 3 or more partitions.

[0421] Furthermore, the first and second partitions may overlap. That is, the first and second partitions may contain the same pixel region. In this case, the predicted image in the first and second partitions can be used to generate the predicted image for the current block.

[0422] In addition, although this example shows an example in which predicted images are generated by inter-frame prediction for both of the two partitions, a predicted image may be generated by intra-frame prediction for at least one partition.

[0423] [Motion Compensation > BIO]

[0424] Next, the method for deriving motion vectors will be described. First, a mode for deriving motion vectors based on a model assuming constant-speed linear motion will be described. This mode is sometimes called the BIO (bidirectional optical flow) mode.

[0425] Figure 38 This is a conceptual diagram used to explain a model assuming constant velocity linear motion. Figure 38 In (v x , v y ) represents the velocity vector, τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represent the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represent the motion vector corresponding to the reference picture Ref1.

[0426] At this time, it can also be that the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, using the following optical flow equation (2).

[0427]

Formula 3

[0428]

[0429] Here, I(k) represents the luminance value of reference image k (k = 0, 1) after motion compensation. This optical flow equation states that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Block-level motion vectors obtained from a merge list, etc., may be corrected on a pixel-by-pixel basis based on a combination of this optical flow equation and Hermite interpolation.

[0430] Furthermore, the decoding device may derive motion vectors using a method different from the method based on a model assuming constant velocity linear motion. For example, motion vectors may be derived on a sub-block basis based on motion vectors of a plurality of adjacent blocks.

[0431] [Motion Compensation > LIC]

[0432] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.

[0433] Figure 39 This is a conceptual diagram for explaining an example of a method for generating a predicted image using a brightness correction process based on an LIC process.

[0434] First, MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.

[0435] Next, information is extracted for the current block, indicating how the luminance values ​​vary between the reference picture and the current picture. This extraction is based on the luminance pixel values ​​of the coded left-adjacent reference region (peripheral reference region) and the coded upper-adjacent reference region (peripheral reference region) in the current picture, as well as the luminance pixel values ​​at equivalent positions within the reference picture specified by the derived MV. This information, indicating how the luminance values ​​vary, is then used to calculate the luminance correction parameters.

[0436] The brightness correction parameters are applied to the reference image in the reference picture specified by the MV to perform brightness correction processing, thereby generating a predicted image for the current block.

[0437] in addition, Figure 39 The shape of the peripheral reference area in the figure is an example, and shapes other than these may be used.

[0438] In addition, the process of generating a predicted image based on one reference image is described here, but the same is true when generating a predicted image based on multiple reference images. The predicted image can also be generated after brightness correction processing is performed on the reference images obtained from each reference image in the same way as described above.

[0439] One method for determining whether to use the LIC process involves using, for example, a lic_flag, which serves as a signal indicating whether the LIC process is used. As a specific example, the encoder determines whether the current block belongs to an area where luminance changes. If the current block belongs to an area where luminance changes, the lic_flag is set to a value of 1, and encoding is performed using the LIC process. If the current block does not belong to an area where luminance changes, the lic_flag is set to a value of 0, and encoding is performed without using the LIC process. Alternatively, the decoder can decode the lic_flag described in the stream and switch whether to use the LIC process based on its value for decoding.

[0440] Another method for determining whether to use the LIC process is to determine whether the LIC process was used in surrounding blocks. As a specific example, when the current block is in merge mode, a determination is made as to whether the surrounding coded blocks selected during MV derivation in merge mode were coded using the LIC process. Based on the determination, whether or not to use the LIC process is switched for coding. In this example, the same process also applies to the decoding device.

[0441] use Figure 39 The form of the LIC process (luminance correction process) has been described, and its details will be described below.

[0442] First, the inter prediction unit 126 derives a motion vector for acquiring a reference image corresponding to a current block to be encoded from a reference picture that is an already encoded picture.

[0443] Next, the inter-frame prediction unit 126 uses the luminance pixel values ​​of the coded neighboring reference areas to the left and above, as well as the luminance pixel values ​​at the same position in the reference picture specified by the motion vector, to extract information indicating how the luminance values ​​in the reference picture and the current picture change, and calculates a luminance correction parameter. For example, the luminance pixel value of a pixel in the neighboring reference area in the current picture is set to p0, and the luminance pixel value of the pixel in the neighboring reference area at the same position in the reference picture is set to p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1+B=p0 for multiple pixels in the neighboring reference area as luminance correction parameters.

[0444] Next, the inter-frame prediction unit 126 uses the brightness correction parameters to perform brightness correction on the reference image within the reference picture specified by the motion vector to generate a predicted image for the current block. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after brightness correction is set to p3. The inter-frame prediction unit 126 calculates A×p2+B=p3 for each pixel in the reference image to generate the predicted image after brightness correction.

[0445] also, Figure 39 The shape of the peripheral reference area in is an example, and other shapes can also be used. Figure 39 A portion of the surrounding reference area shown is used. For example, an area including a predetermined number of pixels thinned out from the upper and left adjacent pixels may be used as the surrounding reference area. Furthermore, the surrounding reference area is not limited to an area adjacent to the encoding target block and may also be an area not adjacent to the encoding target block. The predetermined number of pixels may also be predetermined.

[0446] In addition, Figure 39 In the example shown, the peripheral reference region within the reference picture is an area specified by the motion vector of the current picture from among the peripheral reference regions within the current picture, but it may also be an area specified by another motion vector. For example, the other motion vector may be the motion vector of the peripheral reference region within the current picture.

[0447] Here, the operation in the encoding device 100 is described, but typically, the operation in the decoding device 200 is also similar.

[0448] Furthermore, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for all of them.

[0449] Furthermore, the LIC process may be applied in sub-block units. For example, the modification parameters may be derived using the surrounding reference region of the current sub-block and the surrounding reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0450] [Prediction Control Department]

[0451] The prediction control unit 128 selects one of the intra-frame prediction signal (the signal output from the intra-frame prediction unit 124) and the inter-frame prediction signal (the signal output from the inter-frame prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0452] like Figure 1 As shown, in various examples of encoding devices, the prediction control unit 128 may also output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate a coded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may also be used in the decoding device. The decoding device may also receive and decode the coded bitstream, performing the same prediction processing as that performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction parameters may include a prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra-frame prediction unit 124 or the inter-frame prediction unit 126), or any index, flag, or value based on the prediction processing performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128 or indicating the prediction processing.

[0453] [Encoding device installation example]

[0454] Figure 40 1 is a block diagram showing an implementation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. For example, Figure 1The multiple components of the encoding device 100 shown are Figure 40 The processor a1 and the memory a2 shown are implemented.

[0455] Processor a1 is a circuit that processes information and can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. Processor a1 can also be a processor such as a CPU. In addition, processor a1 can also be a collection of multiple electronic circuits. In addition, for example, processor a1 can also play a role. Figure 1 The functions of multiple components of the encoding device 100 shown in FIG.

[0456] Memory a2 is a dedicated or general-purpose memory that stores information used by processor a1 to encode moving images. Memory a2 can be an electronic circuit or connected to processor a1. Alternatively, memory a2 can be included in processor a1. Alternatively, memory a2 can be a collection of multiple electronic circuits. Memory a2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory a2 can be either non-volatile or volatile memory.

[0457] For example, the memory a2 may store a coded moving image or a bit string corresponding to the coded moving image. In addition, the memory a2 may store a program for the processor a1 to encode the moving image.

[0458] In addition, for example, memory a2 can also serve as Figure 1 The memory a2 can be used as a component for storing information among the multiple components of the encoding device 100 shown in FIG. Figure 1 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks and reconstructed pictures.

[0459] In addition, in the encoding device 100, it is not necessary to install Figure 1 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 1 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.

[0460] [Decoding device]

[0461] Next, a decoding device that can decode the coded signal (coded bit stream) output from, for example, the above-described coding device 100 will be described. Figure 412 is a block diagram showing the functional structure of a decoding device 200 according to an embodiment. The decoding device 200 is a moving picture decoding device that decodes a moving picture in units of blocks.

[0462] like Figure 41 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.

[0463] Decoding device 200 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes a software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. Alternatively, decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to entropy decoding unit 202, inverse quantization unit 204, inverse transform unit 206, addition unit 208, loop filter unit 212, intra-frame prediction unit 216, inter-frame prediction unit 218, and prediction control unit 220.

[0464] Below, after describing the overall processing flow of the decoding device 200 , each component included in the decoding device 200 will be described.

[0465] [Overall decoding process]

[0466] Figure 42 This is a flowchart showing an example of the overall decoding process performed by the decoding device 200 .

[0467] First, the entropy decoding unit 202 of the decoding device 200 determines a partitioning pattern for a fixed-size block (e.g., 128×128 pixels) (step Sp_1). This partitioning pattern is the same as the partitioning pattern selected by the encoding device 100. The decoding device 200 then performs steps Sp_2 to Sp_6 on each of the multiple blocks that make up this partitioning pattern.

[0468] That is, the entropy decoding unit 202 decodes (specifically, performs entropy decoding) the encoded quantization coefficients and prediction parameters of the decoding target block (also referred to as the current block) (step Sp_2).

[0469] Next, the inverse quantization unit 204 and the inverse transformation unit 206 perform inverse quantization and inverse transformation on the plurality of quantized coefficients to restore a plurality of prediction residuals (ie, difference blocks) (step Sp_3).

[0470] Next, the prediction processing unit composed of all or part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction signal (also referred to as a prediction block) of the current block (step Sp_4).

[0471] Next, the adding unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the differential block (step Sp_5).

[0472] Then, when the reconstructed image is generated, the loop filter unit 212 filters the reconstructed image (step Sp_6).

[0473] Then, the decoding device 200 determines whether decoding of the entire picture is completed (step Sp_7 ). If it is determined that decoding is not completed (No in step Sp_7 ), the processing from step Sp_1 is repeatedly executed.

[0474] As shown in the figure, the processing of steps Sp_1 to Sp_7 is sequentially performed by the decoding device 200, or a plurality of processes of some of these processes may be performed in parallel, or the order may be reversed.

[0475] [Entropy decoding unit]

[0476] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the coded bit stream to the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220 in the embodiment (see Figure 1 The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.

[0477] [Inverse quantization unit]

[0478] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantization unit 204 then outputs the inversely quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0479] [Inverse transformation unit]

[0480] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients input from the inverse quantization unit 204 .

[0481] For example, when the information read from the coded bitstream indicates that EMT or AMT is adopted (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.

[0482] Furthermore, for example, when the information decoded from the coded bit stream indicates that NSST is adopted, the inverse transform unit 206 applies inverse re-transformation to the transform coefficients.

[0483] [Addition Department]

[0484] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction sample input from the prediction control unit 220 . The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212 .

[0485] [Block Memory]

[0486] The block memory 210 is a storage unit for storing blocks in a decoding target picture (hereinafter referred to as a current picture) to be referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the adding unit 208 .

[0487] [Loop filter unit]

[0488] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device.

[0489] When the ALF on / off information read from the coded bitstream indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.

[0490] [Frame Memory]

[0491] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0492] [Prediction Processing Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]

[0493] Figure 432 is a flowchart showing an example of processing performed by the prediction processing unit of the decoding device 200. The prediction processing unit is composed of all or part of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0494] The prediction processing unit generates a predicted image for the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Prediction signals include, for example, intra-frame prediction signals and inter-frame prediction signals. Specifically, the prediction processing unit generates a predicted image for the current block using the reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0495] The reconstructed image may be, for example, an image of a reference picture, or an image including the current block, that is, an image of a decoded block within the current picture. The decoded block within the current picture may be, for example, an adjacent block of the current block.

[0496] Figure 44 This is a flowchart showing another example of processing performed by the prediction processing unit of the decoding device 200 .

[0497] The prediction processing unit determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode can be determined based on prediction parameters or the like.

[0498] If the first mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the first mode (step Sr_2a). Furthermore, if the second mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the second mode (step Sr_2b). Furthermore, if the third mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the third mode (step Sr_2c).

[0499] The first, second, and third methods are different methods for generating predicted images, and may be, for example, inter-frame prediction, intra-frame prediction, or other prediction methods. In such prediction methods, the above-mentioned reconstructed image may also be used.

[0500] [Intra-frame prediction unit]

[0501] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the coded bitstream, referring to blocks in the current picture stored in the block memory 210, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 performs intra prediction by referring to samples (e.g., luminance values ​​and chrominance values) of blocks adjacent to the current block, thereby generating an intra prediction signal and outputting the intra prediction signal to the prediction control unit 220.

[0502] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0503] Furthermore, when the information read from the coded bit stream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.

[0504] [Inter-frame prediction unit]

[0505] The inter-frame prediction unit 218 predicts the current block by referring to the reference picture stored in the frame memory 214. Prediction is performed on the current block or a sub-block (e.g., a 4×4 block) within the current block. For example, the inter-frame prediction unit 218 performs motion compensation using motion information (e.g., motion vectors) decoded from the coded bitstream (e.g., prediction parameters output from the entropy decoding unit 202). This generates an inter-frame prediction signal for the current block or sub-block and outputs the inter-frame prediction signal to the prediction control unit 220.

[0506] When the information read from the encoded bit stream indicates that the OBMC mode is adopted, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks.

[0507] Furthermore, when the information decoded from the coded bitstream indicates that the FRUC mode is used, the inter-frame prediction unit 218 performs motion estimation using the pattern matching method (bidirectional matching or template matching) decoded from the coded stream to derive motion information. Furthermore, the inter-frame prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0508] Furthermore, when the BIO mode is used, the inter-frame prediction unit 218 derives a motion vector based on a model assuming constant-speed linear motion. Furthermore, when the information decoded from the coded bitstream indicates that the affine motion compensation prediction mode is used, the inter-frame prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0509] [MV Export > Normal Interframe Mode]

[0510] When the information read from the encoded bit stream indicates that the normal inter mode is applied, the inter prediction unit 218 derives an MV based on the information read from the encoded bit stream, and performs motion compensation (prediction) using the MV.

[0511] Figure 45 This is a flowchart illustrating an example of inter prediction based on the normal inter mode in the decoding apparatus 200 .

[0512] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation on each block. Based on information such as the MVs of multiple decoded blocks temporally and spatially surrounding the current block, the inter-frame prediction unit 218 obtains multiple candidate MVs for the current block (step Ss_1). In other words, the inter-frame prediction unit 218 creates a candidate MV list.

[0513] Next, the inter-frame prediction unit 218 extracts N (N is an integer greater than or equal to 2) candidate MVs from the multiple candidate MVs obtained in step Ss_1 as motion vector predictor candidates (also referred to as predicted MV candidates) in a predetermined priority order (step Ss_2). Alternatively, the priority order may be predetermined for each of the N predicted MV candidates.

[0514] Next, the inter-frame prediction unit 218 decodes the predicted motion vector selection information from the input stream (i.e., the encoded bit stream), and uses the decoded predicted motion vector selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted motion vector (also called predicted MV) of the current block (step Ss_3).

[0515] Next, the inter prediction section 218 decodes the difference MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded difference MV to the selected predicted motion vector (step Ss_4).

[0516] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).

[0517] [Prediction Control Department]

[0518] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as the prediction signal to the adder 208. Generally speaking, the structure, function, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side can correspond to the structure, function, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.

[0519] [Decoding device installation example]

[0520] Figure 46 2 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 41 The components of the decoding device 200 are shown in FIG. Figure 46The processor b1 and memory b2 are shown as being installed.

[0521] Processor b1 is a circuit that processes information and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes the encoded moving image (i.e., the encoded bit stream). Processor b1 can also be a processor such as a CPU. In addition, processor b1 can also be a collection of multiple electronic circuits. In addition, for example, processor b1 can also play a role in Figure 41 The functions of multiple components of the decoding device 200 shown in FIG.

[0522] Memory b2 is a dedicated or general-purpose memory that stores information used by processor b1 to decode the coded bit stream. Memory b2 can be an electronic circuit or connected to processor b1. Alternatively, memory b2 can be included in processor b1. Alternatively, memory b2 can be a collection of multiple electronic circuits. Alternatively, memory b2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory b2 can be either non-volatile or volatile memory.

[0523] For example, the memory b2 may store moving images or coded bit streams. In addition, the memory b2 may also store a program for the processor b1 to decode the coded bit stream.

[0524] In addition, for example, memory b2 can serve as Figure 41 The memory b2 can be used as a component for storing information among the multiple components of the decoding device 200 shown in FIG. Figure 41 The functions of the block memory 210 and the frame memory 214 are shown. More specifically, the memory b2 can store reconstructed blocks and reconstructed pictures.

[0525] In addition, in the decoding device 200, it is not necessary to install Figure 41 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 41 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.

[0526] [Definition of each term]

[0527] As an example, each term may be defined as follows.

[0528] A picture is an arrangement of multiple luma samples in monochrome format, or an arrangement of multiple luma samples and two corresponding arrangements of multiple color difference samples in 4:2:0, 4:2:2, and 4:4:4 color formats. A picture can be a frame or a field.

[0529] A frame is a composition of a top field that generates a plurality of sample rows 0, 2, 4, ... and a bottom field that generates a plurality of sample rows 1, 3, 5, ... .

[0530] A slice is an integer number of coding tree units contained in an independent slice segment and all subsequent dependent slice segments before the next independent slice segment (if any) in the same access unit.

[0531] A tile is a rectangular area of ​​multiple coding tree blocks within a specific tile column and a specific tile row in a picture. Tiles can still apply loop filters across the edges of the tiles, but can also be rectangular areas of a frame that are intended to be independently decoded and encoded.

[0532] A block is an MxN (N rows and M columns) array of samples, or an MxN array of transform coefficients. A block can also be a square or rectangular area of ​​multiple pixels consisting of multiple matrices of one luma and two chroma.

[0533] A CTU (Coding Tree Unit) can be a coding tree block of multiple luma samples for a picture with a three-sample arrangement, or two corresponding coding tree blocks of multiple chroma samples. Alternatively, a CTU can be a coding tree block of any number of samples in a monochrome picture or a picture encoded using the same syntax used for encoding three separate color planes and multiple samples.

[0534] A super block may be composed of one or two mode information blocks, or may be recursively divided into four 32×32 blocks, or further divided into a 64×64 pixel square block.

[0535] [First form of coefficient coding]

[0536] Figure 47 : is a flowchart showing the basic coefficient coding method of the first aspect. Specifically, Figure 47 This represents a coefficient encoding method for a region where a prediction residual is obtained through intra-frame coding or inter-frame coding. The following description shows the operations performed by encoding device 100. Decoding device 200 can perform operations corresponding to those performed by encoding device 100. For example, decoding device 200 can perform inverse orthogonal transform and decoding corresponding to the orthogonal transform and encoding performed by encoding device 100.

[0537] exist Figure 47In the figure, last_sig_coeff, subblock_flag, thres, and CCB are shown. last_sig_coeff is a parameter that indicates the coordinate position of the first non-zero coefficient (non-zero coefficient) when scanning within a block. subblock_flag is a flag that indicates whether there are non-zero coefficients in a 4×4 subblock (also known as 16 transform coefficient levels). subblock_flag can also be expressed as coded_sub_block_flag or subblock flag.

[0538] Thres is a constant determined per block. Thres can be predetermined. Thres can vary depending on the block size or remain constant regardless of block size. Thres can take different values ​​when orthogonal transformation is applied and when not. Thres can be determined based on the coordinate position within the block determined by last_sig_coeff.

[0539] CCB represents the number of bins encoded in the context mode of CABAC (Context Adaptive Binary Arithmetic Coding). That is, CCB represents the number of times the encoding based on the context mode of CABAC is processed. The context mode is also called the regular mode. Here, the encoding based on the context mode of CABAC is called CABAC encoding or context adaptive coding. In addition, the encoding based on the bypass mode of CABAC is called bypass coding. The processing of bypass coding is simpler than that of CABAC coding.

[0540] CABAC coding is a process that converts the bin string obtained by binarizing the signal to be coded into a coded bit string based on the probability of occurrence of 0 and 1 for each bin. In addition, CCB can count the number of all flags used in residual coefficient coding, or count the number of flags used in residual coefficient coding. Bypass coding is a process that does not use the variable probability of occurrence of 0 and 1 for each bin (in other words, uses a fixed probability), and encodes 1 bin in the bin string as 1 bit of the coded bit string.

[0541] For example, the encoding apparatus 100 compares the CCB value and the thres value to determine the coefficient encoding method.

[0542] Specifically, in Figure 47In the encoding process, first, the CCB is initialized to 0 (S101). Then, it is determined whether or not an orthogonal transform is applied to the block (S102). If an orthogonal transform is applied to the block (Yes in S102), the encoding device 100 encodes last_sig_coeff (S131). Then, the encoding device 100 performs a loop process for each sub-block (S141-S148).

[0543] During the loop processing ( S141 - S148 ) for each subblock, encoding device 100 encodes the subblock_flag associated with that subblock. If subblock_flag is different from 0 ( YES in S146 ), encoding device 100 encodes the 16 coefficients within that subblock using the first encoding scheme described below ( S147 ).

[0544] If orthogonal transform is not applied to the block (No in S102 ), the encoding apparatus 100 performs loop processing on each subblock ( S121 to S128 ).

[0545] During the loop processing (S121-S128) for each sub-block, the encoding device 100 determines whether the CCB is less than or equal to thres (S122). If the CCB is less than or equal to thres (Yes in S122), the encoding device 100 encodes the subblock_flag using CABAC encoding (S123). The encoding device 100 then increments the CCB (S124). Otherwise (No in S122), the encoding device 100 encodes the subblock_flag using bypass encoding (S125).

[0546] Then, when subblock_flag is different from 0 (Yes in S126 ), encoding apparatus 100 encodes the 16 coefficients in the subblock using a second encoding method described later ( S127 ).

[0547] The case where an orthogonal transform is not applied to a block can be, for example, when the orthogonal transform is skipped. CCBs are also used in the first and second coding modes. CCBs can be initialized per sub-block. In this case, thres can be a value that varies for each sub-block, rather than a fixed value within the block.

[0548] In addition, here, the CCB counts up from 0 and determines whether it reaches thres, but the CCB may count down from thres (or a specific value) and determine whether it reaches 0.

[0549] Figure 48 Yes Figure 47Detailed flowchart of the first encoding method shown in FIG. In the first encoding method, multiple coefficients within a sub-block are encoded. In this case, a first loop (S151 to S156) is performed for each coefficient information flag of each coefficient within the sub-block, and a second loop (S161 to S165) is performed for each coefficient within the sub-block.

[0550] In the first loop (S151-S156), one or more coefficient information flags, each indicating one or more attributes of a coefficient, are sequentially encoded. The one or more coefficient information flags may include sig_flag, gt1_flag, parity_flag, and gt3_flag, described below. Then, within the range where the CCB does not exceed thres, the one or more coefficient information flags are sequentially encoded using CABAC encoding, with the CCB incremented by 1 each time the CCB is encoded. After the CCB exceeds thres, the coefficient information flags are not encoded.

[0551] Specifically, in the first loop (S151-S156), the encoding device 100 determines whether the CCB is less than or equal to thres (S152). If the CCB is less than or equal to thres (Yes in S152), the encoding device 100 encodes the coefficient information flag using CABAC coding (S153). The encoding device 100 then increments the CCB (S154). If the CCB is not less than or equal to thres (No in S152), the encoding device 100 terminates the first loop (S151-S156).

[0552] In the second loop (S161-S165), for coefficients whose coefficient information flags are encoded, remainders (remainders) not represented by the coefficient information flags (i.e., remainders used to reconstruct the values ​​of the coefficients using the coefficient information flags) are encoded using Golombre coding. Coefficients whose coefficient information flags are not encoded are directly encoded using Golombre coding. Alternatively, the remainders may be encoded using other coding methods rather than Golombre coding.

[0553] Specifically, in the second loop (S161-S165), the encoding device 100 determines whether the coefficient information flag corresponding to the coefficient to be processed has been encoded (S162). If the coefficient information flag has been encoded (Yes in S162), the encoding device 100 encodes the remainder using Golombirai coding (S163). If the coefficient information flag has not been encoded (No in S162), the encoding device 100 encodes the value of the coefficient using Golombirai coding (S164).

[0554] In addition, although the number of loop processing is 2 here, the number of loop processing may be different from 2.

[0555] The sig_flag flag indicates whether AbsLevel is non-zero. AbsLevel is the value of the coefficient, more specifically, the absolute value of the coefficient. gt1_flag indicates whether AbsLevel is greater than 1. parity_flag is the first bit of AbsLevel, indicating whether AbsLevel is odd or even. gt3_flag indicates whether AbsLevel is greater than 3.

[0556] gt1_flag and gt3_flag may be expressed as abs_gt1_flag and abs_gt3_flag, respectively. In addition, for example, as the above-mentioned remainder, a value of (Abslevel-4) / 2 may be encoded by Columbrax coding.

[0557] One or more coefficient information flags other than the one or more coefficient information flags described above may also be encoded. For example, some coefficient information flags may not be encoded. A coefficient information flag included in the one or more coefficient information flags described above may be replaced with a coefficient information flag or parameter having another meaning.

[0558] Figure 49 Yes Figure 47 Detailed flowchart of the second encoding method shown in FIG. In the second encoding method, multiple coefficients within a sub-block are encoded. In this case, a first loop (S171 to S176) is performed for each coefficient information flag of each coefficient within the sub-block, and a second loop (S181 to S185) is performed for each coefficient within the sub-block.

[0559] In the first loop processing (S171-S176), one or more coefficient information flags each indicating one or more attributes of a coefficient are sequentially encoded. The one or more coefficient information flags may include sig_flag, sign_flag, gt1_flag, parity_flag, gt3_flag, gt5_flag, gt7_flag, and gt9_flag.

[0560] Here, sign_flag is a flag indicating the sign of the coefficient. gt5_flag is a flag indicating whether AbsLevel is greater than 5. gt7_flag is a flag indicating whether AbsLevel is greater than 7. gt9_flag is a flag indicating whether AbsLevel is greater than 9. gt5_flag, gt7_flag, and gt9_flag are sometimes expressed as abs_gt5_flag, abs_gt7_flag, and abs_gt9_flag, respectively. Furthermore, flags indicating whether AbsLevel is greater than x (x is an integer greater than 1) may be combined and expressed as gtx_flag or abs_gtx_flag. AbsLevel is, for example, the absolute value of the transform coefficient level.

[0561] Furthermore, one or more coefficient information flags other than the one or more coefficient information flags described above may be encoded. For example, some coefficient information flags may not be encoded. A coefficient information flag included in the one or more coefficient information flags described above may be replaced with a coefficient information flag or parameter having another meaning.

[0562] The one or more coefficient information flags are sequentially encoded using CABAC coding. The CCB is incremented by 1 each time coding is performed. After the CCB exceeds the thres, the coefficient information flags are encoded using bypass coding.

[0563] Specifically, in the first loop (S171-S176), the encoding device 100 determines whether the CCB is less than or equal to thres (S172). If the CCB is less than or equal to thres (Yes in S172), the encoding device 100 encodes the coefficient information flag using CABAC encoding (S173). The encoding device 100 then increments the CCB (S174). If the CCB is not less than or equal to thres (No in S172), the encoding device 100 encodes the coefficient information flag using bypass encoding (S175).

[0564] Figure 49 The syntax of the second loop processing in remains unchanged before and after the CCB exceeds thres. That is, regardless of whether the coefficient information flag is encoded using CABAC coding or bypass coding, the following identical processing is performed in the second loop processing (S181 to S185).

[0565] Specifically, in the second loop (S181-S185), the encoding device 100 encodes the remainder, which is the residual value that cannot be represented by the coefficient information flag (i.e., the residual value used to reconstruct the coefficient value using the coefficient information flag), using Golombre coding (S183). Alternatively, the remainder may be encoded using another coding method instead of Golombre coding.

[0566] In addition, although the number of loop processing is 2 here, the number of loop processing may be different from 2.

[0567] like Figure 47 、 Figure 48 and Figure 49 As shown, in the basic operation of this aspect, there are different flags for whether or not the CABAC encoding processing times are limited, depending on whether or not an orthogonal transform is applied. Furthermore, the coefficient coding syntax differs between when orthogonal transform is applied and when not. Consequently, separate circuits may need to be prepared, potentially complicating the circuit structure.

[0568] [First Example of the First Form of Coefficient Coding]

[0569] Figure 50 This is a flowchart showing the coefficient encoding method of the first example of the first aspect. Figure 50 In the example, the post-processing of last_sig_coeff (S132) and the processing of subblock_flag (S142 to S145) are the same as Figure 47 The examples are different.

[0570] exist Figure 47 In the case of applying orthogonal transform, CCB counts up sig_flag, parity_flag and gtX_flag (X=1, 3). Figure 50 In the example, CCB also increments last_sig_coeff and subblock_flag. On the other hand, the processing flow is the same as that in the case where orthogonal transformation is not applied. Figure 47 Same as the example.

[0571] That is, in Figure 50 In the example of , the encoding device 100 encodes last_sig_coeff ( S131 ), and then adds the number of CABAC encoding processes in the encoding of last_sig_coeff to the CCB ( S132 ).

[0572] Furthermore, before encoding subblock_flag, the encoding device 100 determines whether the CCB is less than or equal to thres (S142). If the CCB is less than or equal to thres (Yes in S142), the encoding device 100 encodes subblock_flag using CABAC encoding (S143). The encoding device 100 then increments the CCB by 1 (S144). On the other hand, if the CCB is not less than or equal to thres (No in S142), the encoding device 100 encodes subblock_flag using bypass encoding (S145).

[0573] [Effect of the First Example of the First Form of Coefficient Coding]

[0574] according to Figure 50 For example, the encoding process flow for subblock_flag can sometimes be standardized and unified between cases where orthogonal transformation is performed and cases where orthogonal transformation is not performed. Therefore, it is possible to share some circuits between cases where orthogonal transformation is performed and cases where orthogonal transformation is not performed, potentially reducing circuit size. As a result, multiple process flows that are divided based on whether or not orthogonal transformation is performed may become the same, except for the presence or absence of last_sig_coeff.

[0575] For example, even if the number of CABAC encoding passes is limited at the block level, Figure 47 In the example, after CCB reaches thres, subblock_flag will be encoded by CABAC coding. Figure 50 In the case of , after the CCB reaches thres, subblock_flag is not encoded by CABAC encoding. Therefore, the number of CABAC encoding processes can be appropriately limited to thres.

[0576] Furthermore, the number of times CABAC encoding is processed in last_sig_coeff may not be included in the CCB. Furthermore, thres may be determined depending on the coordinate position determined by last_sig_coeff in the block.

[0577] Furthermore, the encoding apparatus 100 may perform encoding by always determining the value of subblock_flag to be 1 after the CCB exceeds thres. If the value of subblock_flag is always determined to be 1 after the CCB exceeds thres, the encoding apparatus 100 may not encode subblock_flag after the CCB exceeds thres.

[0578] Furthermore, even when an orthogonal transform is not applied, the encoding apparatus 100 may always determine the value of subblock_flag to be 1 after the CCB exceeds thres and perform encoding. Furthermore, in this case, if the value of subblock_flag is always determined to be 1 after the CCB exceeds thres, the encoding apparatus 100 may not encode subblock_flag after the CCB exceeds thres.

[0579] [Second example of the first aspect of coefficient encoding]

[0580] Figure 51 This is a flowchart showing the coefficient encoding method of the second example of the first aspect. Figure 51 In the example, the processing of subblock_flag (S123) is the same as Figure 47 The examples are different.

[0581] exist Figure 47 In the case where orthogonal transformation is not applied, the CCB has incremented the sig_flag, parity_flag, gtX_flag (X=1, 3, 5, 7, 9) and subblock_flag. Figure 51 In the example, subblock_flag, CCB are not counted up. On the other hand, the processing flow in the case of applying orthogonal transform is the same as Figure 47 Same as the example.

[0582] That is, in Figure 51 In the example of FIG. 1 , regardless of whether the CCB exceeds the thres, the encoding apparatus 100 always encodes the subblock_flag by CABAC encoding without incrementing the CCB ( S123 ).

[0583] [Effect of the Second Example of the First Form of Coefficient Coding]

[0584] according to Figure 51 For example, when performing orthogonal transformation and when not performing orthogonal transformation, the coding process of subblock_flag can sometimes be made common and unified. Therefore, when performing orthogonal transformation and when not performing orthogonal transformation, some circuits may be shared, potentially reducing circuit size. As a result, multiple processing flows divided according to the presence or absence of orthogonal transformation may become the same, except for the presence or absence of last_sig_coeff.

[0585] In addition, with Figure 50 In comparison, Figure 51This simplifies processing. Therefore, it is possible to reduce circuit size. Furthermore, it is assumed that the frequency of occurrence of 0 or 1 associated with subblock_flag is likely to vary depending on surrounding conditions. Therefore, in CABAC encoding of subblock_flag, it is assumed that the reduction in the amount of code is greater than the increase in processing delay. Therefore, it is useful to limit the number of CABAC encoding operations so that CABAC encoding of subblock_flag is not counted towards the number of CABAC encoding operations.

[0586] In addition, the number of CABAC encoding processes for last_sig_coeff may be included in the CCB. In addition, thres may be determined depending on the coordinate position determined by last_sig_coeff in the block.

[0587] [Second form of coefficient coding]

[0588] [First example of the second aspect of coefficient encoding]

[0589] Figure 52 This is a flowchart showing the coefficient encoding method of the first example of the second aspect. Figure 52 In the example of FIG. 1 , even when no orthogonal transform is applied to the block, the case where the 16 coefficients in the sub-block are encoded by the first encoding method (S127a) is the same as the case where the 16 coefficients in the sub-block are encoded by the first encoding method (S127a). Figure 47 The examples are different.

[0590] That is, in Figure 52 In the example of Figure 48 The first encoding method shown is not Figure 49 The 16 coefficients in the sub-block are encoded using the second encoding method shown in FIG. 127a. That is, whether or not orthogonal transformation is applied, the encoding apparatus 100 does not encode the 16 coefficients in the sub-block using the second encoding method shown in FIG. Figure 49 The second encoding method shown is through Figure 48 The first coding method shown encodes 16 coefficients in a sub-block.

[0591] More specifically, regardless of whether or not there is an orthogonal transform, the encoding apparatus 100 follows Figure 48 In the first encoding method shown, in the first processing loop, bypass coding is not used and the encoding of the coefficient information flag is skipped when the CCB exceeds the thres. Then, in the second processing loop, if the coefficient information flag corresponding to the coefficient to be processed has not been encoded, the encoding device 100 encodes the coefficient value using Golombrite coding without using the coefficient information flag.

[0592] In addition, Figure 48In the first loop of the processing, the syntax for encoding the coefficient information flags may differ between when an orthogonal transform is applied and when an orthogonal transform is not applied. For example, one or more coefficient information flags when an orthogonal transform is applied may differ from one or more coefficient information flags when an orthogonal transform is not applied, and some or all of these coefficient information flags may differ.

[0593] [Effect of the First Example of the Second Aspect of Coefficient Coding]

[0594] according to Figure 52 For example, even if the coding syntax of the coefficient information flag differs depending on whether or not an orthogonal transform is used, the coding syntax of the 16 coefficients in the sub-block after the CCB exceeds the threshold can be made common regardless of whether or not an orthogonal transform is used. Therefore, it is possible to share some circuits between the case where an orthogonal transform is used and the case where an orthogonal transform is not used, thereby reducing the circuit size.

[0595] Furthermore, after the CCB exceeds the thres, the coefficient is encoded without being divided into the coefficient information flag encoded by bypass coding and the residual value information encoded by Golombre coding. Therefore, it is possible to suppress an increase in the amount of information and an increase in the amount of code.

[0596] [Second Example of the Second Aspect of Coefficient Coding]

[0597] Figure 53 This is a flowchart showing the coefficient encoding method of the second example of the second aspect. Figure 53 In the example of FIG. 14 , even when orthogonal transformation is applied to the block, the case where 16 coefficients in the sub-block are encoded using the second encoding method (S147a) is the same as Figure 47 The examples are different.

[0598] That is, in Figure 53 In the example of FIG. 1 , when an orthogonal transform is applied to a block, the encoding apparatus 100 does not use Figure 48 The first encoding method shown is used Figure 49 The second encoding method shown in FIG. 1 encodes 16 coefficients in the sub-block ( S147a ). That is, the encoding apparatus 100 does not use the orthogonal transform in either case. Figure 48 The first encoding method shown is used Figure 49 The second coding method shown encodes 16 coefficients in a sub-block.

[0599] More specifically, the encoding apparatus 100 performs the encoding according to Figure 49In the second coding method shown, in the first loop, if the CCB exceeds the thres, the coefficient information flag is encoded using bypass coding instead of skipping coding. Then, in the second loop, the coding apparatus 100 encodes the remainder based on the coefficient information flag using Golombre coding.

[0600] In addition, Figure 49 The syntax used to encode the coefficient information flags in the first loop of the process may differ between when an orthogonal transform is applied and when an orthogonal transform is not applied. For example, one or more coefficient information flags when an orthogonal transform is applied may differ from one or more coefficient information flags when an orthogonal transform is not applied, and some or all of these coefficient information flags may differ.

[0601] [Effect of the Second Example of the Second Aspect of Coefficient Coding]

[0602] according to Figure 53 For example, even if the syntax for encoding the coefficient information flag differs depending on whether or not an orthogonal transform is used, the syntax for encoding the 16 coefficients in the sub-block after the CCB exceeds the thres value can be made common regardless of whether or not an orthogonal transform is used. Therefore, it is possible to share some circuitry between the case where an orthogonal transform is used and the case where an orthogonal transform is not used, thereby reducing circuit size.

[0603] [The third form of coefficient coding]

[0604] Figure 54 This is a syntax diagram showing the basic first encoding method of the third form. Figure 54 The syntax shown corresponds to Figure 47 An example of the syntax of the first coding method is shown in FIG. Basically, the first coding method is used when orthogonal transformation is applied.

[0605] The coefficient information flags and parameters here are the same as those shown in the first aspect. Furthermore, the multiple coefficient information flags shown here are merely examples, and other multiple coefficient information flags may be encoded. For example, some coefficient information flags may not be encoded. Furthermore, the coefficient information flags shown here may be replaced with coefficient information flags or parameters having other meanings.

[0606] Figure 54 The initial for loop in the example corresponds to Figure 48The first loop in the example is processed. In this initial for loop, if CCBs remain, that is, if the CCBs do not exceed the threshold, coefficient information flags such as sig_flag are encoded using CABAC coding. If no CCBs remain, the coefficient information flags are not encoded. Furthermore, in this example, the second example of the second aspect can also be applied. That is, if no CCBs remain, the coefficient information flags can be encoded using bypass coding.

[0607] The second for loop from the top and the third for loop from the top correspond to Figure 48 The second loop in the example is processed. In the second for loop from the top, the coefficients whose coefficient information flags have been encoded are encoded using Golombirai coding for the remaining values. In the third for loop from the top, the coefficients whose coefficient information flags have not been encoded are encoded using Golombirai coding for the remaining values. Furthermore, by applying the second example of the second aspect to this example, the remaining values ​​can also be always encoded using Golombirai coding.

[0608] In the fourth for loop from the top, sign_flag is encoded by bypass encoding.

[0609] The syntax described in this form can be applied to Figure 47 、 Figure 48 、 Figure 50 、 Figure 51 and Figure 52 Examples of .

[0610] Figure 55 This is a syntax diagram showing the basic second encoding method of the third form. Figure 55 The syntax shown corresponds to Figure 47 An example of the syntax of the second encoding method. Basically, the second encoding method is used when no orthogonal transform is applied.

[0611] The multiple coefficient information flags shown here are examples, and other multiple coefficient information flags may be encoded. For example, some coefficient information flags may not be encoded. Furthermore, the coefficient information flags shown here may be replaced with coefficient information flags or parameters having other meanings.

[0612] Figure 55 The first five for loops in the example correspond to Figure 49The first loop in the example of is processed. If CCBs remain in the first five for loops, that is, if the CCBs do not exceed the threshold, coefficient information flags such as sig_flag are encoded using CABAC coding. If no CCBs remain, the coefficient information flags are encoded using bypass coding. Furthermore, this example can also be applied to the first example of the second aspect. That is, if no CCBs remain, the coefficient information flags do not need to be encoded.

[0613] The sixth for loop from the top corresponds to Figure 49 The second loop in the example is processed. In the sixth for loop from the top, the residual value is encoded using Golombirai coding. Furthermore, in this example, the first example of the second aspect can also be applied. That is, for coefficients whose coefficient information flags are encoded, the residual value can be encoded using Golombirai coding. Then, for coefficients whose coefficient information flags are not encoded, the coefficients can be encoded using Golombirai coding.

[0614] The syntax described in this form can be applied to Figure 47 、 Figure 49 、 Figure 50 、 Figure 51 and Figure 53 Each example in .

[0615] Without applying orthogonal transformation ( Figure 55 ), the number of loop processes for encoding the coefficient information flag is greater than that in the case of applying orthogonal transform ( Figure 54 ). Therefore, when orthogonal transform is not applied, the hardware processing load may increase compared to when orthogonal transform is applied. In addition, since the syntax of coefficient coding differs depending on whether or not an orthogonal transform is performed, circuits may need to be prepared separately. As a result, the circuitry may become complex.

[0616] [First example of the third aspect of coefficient encoding]

[0617] Figure 56 This is a syntax diagram showing the second encoding method of the first example of the third aspect. Figure 56 The syntax shown corresponds to Figure 47 An example of the second encoding method. Figure 47 The first encoding method can also be used Figure 54 In addition, this example can be combined with other examples of the third form, and can also be combined with other forms.

[0618] Figure 56 The initial for loop in the example corresponds to Figure 49The first loop in the example is processed. If, in this initial for loop, 8 or more CCBs remain, that is, if the CCB after adding 8 does not exceed the threshold, up to 8 coefficient information flags are encoded using CABAC encoding for each coefficient, and the CCB is incremented up to 8 times. If the CCB does not remain at least 8, that is, if the CCB after adding 8 exceeds the threshold, 8 coefficient information flags are encoded using bypass encoding for each coefficient.

[0619] In other words, before encoding the eight coefficient information flags, it is comprehensively determined whether the eight coefficient information flags can be encoded using CABAC encoding. Then, if the eight coefficient information flags can be encoded using CABAC encoding, a maximum of eight coefficient information flags are encoded using CABAC encoding.

[0620] Furthermore, in this example, the first example of the second aspect can also be applied. That is, if there are not more than eight CCBs remaining, the eight coefficient information flags do not need to be encoded. In other words, in this case, rather than encoding the eight coefficient information flags through bypass coding, the encoding of the eight coefficient information flags can be skipped.

[0621] In addition, if Figure 56 As shown, the coding of one or more coefficient information flags among the eight coefficient information flags may be omitted depending on the value of the coefficient. For example, when sig_flag is 0, the coding of the remaining seven coefficient information flags may be omitted.

[0622] The second for loop from the top corresponds to Figure 49 The second loop in the example is processed. In the second for loop from the top, the residual value is encoded using Golombre coding. Furthermore, in this example, the first example of the second aspect can also be applied. That is, for the coefficients whose eight coefficient information flags are encoded, the residual value can be encoded using Golombre coding, and for the coefficients whose eight coefficient information flags are not encoded, the coefficient can be encoded using Golombre coding.

[0623] The multiple coefficient information flags shown here are examples, and other multiple coefficient information flags may be encoded. For example, some coefficient information flags may not be encoded. Furthermore, the coefficient information flags shown here may be replaced with coefficient information flags or parameters having other meanings.

[0624] In addition, you can also combine Figure 55 Examples and Figure 56 For example, in Figure 55In the example, before the four coefficient information flags such as sig_flag and sign_flag are encoded, it can be comprehensively determined whether the four coefficient information flags can be encoded by CABAC encoding.

[0625] [Effect of the first example of the third aspect of coefficient encoding]

[0626] exist Figure 56 In the example of , all coefficient information flags encoded by CABAC encoding are encoded in one loop process. Figure 55 Compared to the example in Figure 56 In the example above, the number of loop processes is small. Therefore, the processing volume may be reduced.

[0627] In addition, Figure 54 Examples and Figure 56 In the example of , the number of loop processes for encoding a plurality of coefficient information flags by CABAC encoding is the same. Figure 54 Examples and Figure 55 Compared to the combination of examples in Figure 54 Examples and Figure 56 In the combination of the examples, the number of circuit changes may be reduced.

[0628] Furthermore, before a plurality of coefficient information flags are encoded, it is collectively determined whether or not the plurality of coefficient information flags can be encoded by CABAC encoding. This simplifies processing and may reduce processing delay.

[0629] Furthermore, whether or not the plurality of coefficient information flags can be encoded using CABAC coding can be determined collectively before the plurality of coefficient information flags are encoded, both in the case where orthogonal transform is applied and in the case where orthogonal transform is not applied. Consequently, the difference between the encoding method used in a block to which orthogonal transform is applied and the encoding method used in a block not to which orthogonal transform is applied is further reduced, and the circuit scale can be further reduced.

[0630] In addition, Figure 56 In the example, sig_flag to abs_gt9_flag are included in one loop, but the encoding method is not limited to this. Multiple loops (for example, two loops) can be used, and it can also be determined in general whether it is possible to encode multiple coefficient information flags of each loop by CABAC encoding. Compared with one loop, the processing increases, but it is the same as Figure 55 Compared with the example, the same processing reduction effect can be obtained.

[0631] [Second example of the third aspect of coefficient encoding]

[0632] Figure 57This is a syntax diagram showing the second encoding method of the second example of the third aspect. Figure 57 The syntax shown corresponds to Figure 47 An example of the second encoding method. Figure 47 The first encoding method can also be used Figure 54 In addition, this example can be combined with other examples of the third form, and can also be combined with other forms.

[0633] Figure 57 The initial for loop in the example corresponds to Figure 49 The first loop process in the example of is performed. If at least 7 remain in the CCB in this initial for loop, that is, if the CCB after adding 7 does not exceed the threshold, up to seven coefficient information flags are encoded using CABAC encoding for each coefficient, and the CCB is incremented up to seven times. If at least 7 remain in the CCB, that is, if the CCB after adding 7 exceeds the threshold, seven coefficient information flags are encoded using bypass encoding for each coefficient.

[0634] In other words, before encoding the seven coefficient information flags, it is comprehensively determined whether the seven coefficient information flags can be encoded using CABAC encoding. Then, if it is possible to encode the seven coefficient information flags using CABAC encoding, the seven coefficient information flags are encoded using CABAC encoding.

[0635] Furthermore, in this example, the first example of the second aspect can also be applied. That is, if there are seven or more CCBs remaining, the seven coefficient information flags do not need to be encoded. In other words, in this case, the encoding of the seven coefficient information flags can be skipped instead of encoding them using bypass coding.

[0636] In addition, if Figure 57 As shown, the coding of one or more coefficient information flags among the seven coefficient information flags may be omitted depending on the value of the coefficient. For example, when sig_flag is 0, the coding of the remaining six coefficient information flags may be omitted.

[0637] The second for loop from the top corresponds to Figure 49 The second loop in the example is processed. In the second for loop from the top, the residual value is encoded using Golombirai coding. Furthermore, in this example, the first example of the second aspect can also be applied. That is, for coefficients whose seven coefficient information flags are encoded, the residual value can be encoded using Golombirai coding, and for coefficients whose seven coefficient information flags are not encoded, the coefficient can be encoded using Golombirai coding.

[0638] In the third for loop from the top, if CCB remains, that is, CCB does not exceed the threshold, sign_flag is encoded by CABAC encoding, and CCB is counted up. If CCB does not remain, sign_flag is encoded by bypass encoding. Figure 54 As in the example of , sign_flag can always be encoded through bypass coding.

[0639] The multiple coefficient information flags shown here are examples, and other multiple coefficient information flags may be encoded. For example, some coefficient information flags may not be encoded. Furthermore, the coefficient information flags shown here may be replaced with coefficient information flags or parameters having other meanings.

[0640] In addition, you can also combine Figure 55 Examples and Figure 57 For example, in Figure 55 In the example of , before the four coefficient information flags such as sig_flag and sign_flag are encoded, it can be comprehensively determined whether the four coefficient information flags can be encoded by CABAC encoding.

[0641] [Effect of the Second Example of the Third Aspect of Coefficient Coding]

[0642] and Figure 56 Same as the example, in Figure 57 In the example of , multiple coefficient information flags (specifically, abs_gt3_flag and abs_gt5_flag, etc.) for comparing the magnitude of the coefficient with the threshold are encoded in one loop process. Figure 55 Compared to the example in Figure 57 In the example, the number of loop processes is small. Therefore, the processing volume may be reduced.

[0643] and Figure 56 Compared to the example in Figure 57 In the example of , the number of loop processes for encoding a plurality of coefficient information flags increases. However, Figure 56 Compared to the example in Figure 57 In the example, there is Figure 54 For example, sign_flag is encoded last. Therefore, Figure 54 Examples and Figure 56 Compared to the combination of examples in Figure 54 Examples and Figure 57 In the combination of the examples, the number of circuit changes may be reduced.

[0644] Furthermore, before a plurality of coefficient information flags are encoded, it is collectively determined whether or not the plurality of coefficient information flags can be encoded by CABAC encoding. This simplifies processing and may reduce processing delay.

[0645] In both cases of applying orthogonal transform and not applying orthogonal transform, it is possible to comprehensively determine whether the plurality of coefficient information flags can be coded using CABAC coding before the plurality of coefficient information flags are coded. Therefore, the difference between the coding method used in blocks to which orthogonal transform is applied and the coding method used in blocks not applied is further reduced, and the circuit scale can be further reduced.

[0646] In addition, Figure 57 In the example, sig_flag to abs_gt9_flag are included in one loop, but the encoding method is not limited to this. Multiple loops (for example, two loops) can be used, and it can be determined in general whether multiple coefficient information flags of each loop can be encoded by CABAC encoding. Compared with one loop, the processing increases, but it is the same as Figure 55 Compared with the example, the same processing reduction effect can be obtained.

[0647] [Modification of Coefficient Coding]

[0648] Any of the above-described multiple forms and examples related to coefficient coding can be combined. Furthermore, any of the above-described multiple forms and examples related to coefficient coding, and any combination thereof, can be applied to both luma blocks and chroma blocks. In this case, different thres can be used for the luma blocks and the chroma blocks.

[0649] Furthermore, any one of the aforementioned multiple forms, multiple examples, and any combination thereof related to coefficient coding can be used for a block to which an orthogonal transform is not applied and to which BDPCM (Block-based Delta Pulse Code Modulation) is applied. In a block to which BDPCM is applied, the amount of information in each residual signal within the block is reduced by subtracting the residual signal vertically or horizontally adjacent to the residual signal from the residual signal.

[0650] Furthermore, any one of the above-described multiple forms, multiple examples, and any combination thereof related to coefficient encoding may be used for a block to which BDPCM is applied and which is a chrominance block.

[0651] Furthermore, for blocks to which ISP (Intra Sub-Partitions) is applied, any of the above-described multiple forms, multiple examples, and any combination thereof related to coefficient coding may be used. In ISP, an intra block is divided vertically or horizontally, and the pixel values ​​of adjacent sub-blocks are used to perform intra prediction for each sub-block.

[0652] Furthermore, any one of the plurality of forms, plurality of examples, and any combination thereof related to the above-mentioned coefficient encoding may be used for a block to which ISP is applied and which is a chrominance block.

[0653] Furthermore, when chroma joint coding is used as the coding mode for chroma blocks, any of the multiple forms, multiple examples, and any combination thereof related to the above-mentioned coefficient coding can be used. Here, chroma joint coding is a coding method that derives the value of Cr from the value of Cb.

[0654] Furthermore, the thres value when orthogonal transformation is applied may be twice the thres value when orthogonal transformation is not applied. Alternatively, the thres value when orthogonal transformation is not applied may be twice the thres value when orthogonal transformation is applied.

[0655] Furthermore, when only chrominance combined coding is used, the thres value of the CCB when orthogonal transform is applied may be twice the thres value of the CCB when orthogonal transform is not applied. Alternatively, when only chrominance combined coding is used, the thres value of the CCB when orthogonal transform is not applied may be twice the thres value of the CCB when orthogonal transform is applied.

[0656] In addition, in the above-mentioned multiple forms and multiple examples related to coefficient encoding, the scanning order of multiple coefficients in a block to which orthogonal transform is not applied can be the same as the scanning order of multiple coefficients in a block to which orthogonal transform is applied.

[0657] Furthermore, although several examples of syntax are shown in the third aspect and its multiple examples, the syntax to be applied is not limited to these examples. For example, in multiple aspects different from the third aspect and its multiple examples, syntax different from any of the multiple syntaxes shown in the third aspect and its multiple examples can be used. Various syntaxes for encoding 16 coefficients can be applied.

[0658] In addition, the various forms and examples of coefficient coding illustrate the encoding process flow, but the decoding process flow is essentially the same as the encoding process flow, except for the difference in whether the bit stream is transmitted or received. For example, decoding apparatus 200 can perform inverse orthogonal transform and decoding corresponding to the orthogonal transform and encoding performed by encoding apparatus 100.

[0659] Furthermore, each flowchart related to a plurality of forms and a plurality of examples of coefficient coding is merely an example, and conditions or processes may be newly added to each flowchart, deleted, or modified.

[0660] Here, coefficients are values ​​that constitute an image, such as a block or sub-block. Specifically, the multiple coefficients that constitute an image can be obtained from multiple pixel values ​​in the image via an orthogonal transform. Alternatively, the multiple coefficients that constitute an image can be obtained from multiple pixel values ​​in the image without undergoing an orthogonal transform. In other words, the multiple coefficients that constitute an image can be the multiple pixel values ​​themselves. Furthermore, each pixel value can be a pixel value of the original image or a prediction residual value. Furthermore, the coefficients can be quantized.

[0661] [Representative examples of structure and processing]

[0662] Representative examples of the configuration and processing of the encoding apparatus 100 and decoding apparatus 200 described above are described below.

[0663] Figure 58 is a flowchart showing the operation of the encoding device 100. For example, the encoding device 100 includes a circuit and a memory connected to the circuit. The circuit and memory included in the encoding device 100 may also correspond to Figure 40 The circuit of the encoding device 100 performs Figure 58 Specifically, the circuit of the encoding device 100 encodes a block of an image during operation ( S211 ).

[0664] In one example, the circuitry of encoding apparatus 100 may encode a block of an image by limiting the number of context-adaptive encoding processes. Furthermore, the sub-block flag encoding process may be performed without including it in the number of processes subject to the limitation, both when an orthogonal transform is applied to the block and when an orthogonal transform is not applied to the block.

[0665] Here, the subblock flag encoding process is a process of encoding a subblock flag indicating whether a subblock included in a block includes a non-zero coefficient through context adaptive encoding.

[0666] This makes it possible to encode sub-block flags using context-adaptive coding, regardless of whether or not an orthogonal transform is applied and whether the number of context-adaptive coding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0667] Furthermore, when applying an orthogonal transform to a block, the circuitry of encoding apparatus 100 may perform position parameter encoding without including the position parameter encoding in the number of processing times subject to the restriction. Here, position parameter encoding is a process of encoding a parameter indicating the position of the first non-zero coefficient in the block in scanning order using context-adaptive coding.

[0668] Therefore, when orthogonal transform is applied, the parameter indicating the position of the first non-zero coefficient can be encoded by context adaptive coding regardless of whether the number of context adaptive coding processes is limited, thereby reducing the amount of code.

[0669] Furthermore, when applying an orthogonal transform to a block, the circuitry of encoding device 100 can also determine a limit on the number of processing operations based on the position of the first non-zero coefficient. Therefore, when applying an orthogonal transform, it is possible to appropriately determine the limit on the number of processing operations. This allows for an appropriate balance between reducing the amount of code and reducing processing delay.

[0670] In another example, the circuit of the encoding device 100 may operate as follows in both cases of applying an orthogonal transform to a block of an encoding target image and not applying an orthogonal transform to the block.

[0671] Specifically, in both cases, if the number of context-adaptive coding processes is within the processing number limit, the circuitry of encoding device 100 may encode the coefficient information flag using context-adaptive coding. Furthermore, in both cases, if the number of context-adaptive coding processes is not within the processing number limit, the circuitry of encoding device 100 may skip encoding the coefficient information flag. Here, the coefficient information flag indicates the attributes of the coefficients included in the block.

[0672] Then, if the coefficient information flag is encoded, the circuitry of encoding device 100 may encode the residual value information using Golombirai coding. Here, the residual value information is information used to reconstruct the value of the coefficient using the coefficient information flag. Furthermore, if encoding of the coefficient information flag is skipped, the circuitry of encoding device 100 may encode the value of the coefficient using Golombirai coding.

[0673] This makes it possible to skip encoding of coefficient information flags, regardless of whether or not an orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive coding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the encoding method used in blocks where an orthogonal transform is applied and the encoding method used in blocks where an orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0674] Alternatively, the coefficient information flag may be a flag indicating whether the coefficient value is greater than 1. Thus, regardless of whether orthogonal transform is applied, encoding of the coefficient information flag indicating whether the coefficient value is greater than 1 may be skipped in accordance with the limit on the number of processing times for context-adaptive coding. This makes it possible to suppress increases in processing delay and the amount of code.

[0675] In another example, the circuitry of encoding device 100 may limit the number of context-adaptive coding processes and encode a block of an image. Furthermore, when no orthogonal transform is applied to the block, a determination may be made as to whether multiple coefficient information flags, each representing multiple attributes of a coefficient included in the block, satisfy processing conditions. If the processing conditions are determined to be satisfied, the multiple coefficient information flags may be encoded using context-adaptive coding.

[0676] Here, the processing condition is a condition that the number of processing times obtained by adding a plurality of coefficient information flags to the number of processing times is within a limited range of the number of processing times.

[0677] Thus, when an orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive coding can be used for multiple coefficient information flags. This simplifies processing and reduces processing delay. Furthermore, when similar processing is performed on blocks to which an orthogonal transform is applied, the difference between the coding method used in blocks to which an orthogonal transform is applied and the coding method used in blocks not to which an orthogonal transform is applied can be reduced, and the circuit scale can be reduced.

[0678] Furthermore, the plurality of coefficient information flags may include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Thus, it is possible to collectively determine the plurality of coefficient information flags including the coefficient information flag indicating whether the value of the coefficient is greater than 3 and the coefficient information flag indicating whether the value of the coefficient is greater than 5. Consequently, processing can be simplified and processing delay can be reduced.

[0679] Furthermore, the plurality of coefficient information flags may include a coefficient information flag indicating whether the coefficient value is greater than 7 and a coefficient information flag indicating whether the coefficient value is greater than 9. This makes it possible to collectively determine the plurality of coefficient information flags including four coefficient information flags: whether the coefficient value is greater than 3, whether the coefficient value is greater than 5, whether the coefficient value is greater than 7, and whether the coefficient value is greater than 9. Consequently, processing can be simplified and processing delay can be reduced.

[0680] Note that the above-described operations performed by the circuits of the encoding device 100 may also be performed by the entropy encoding unit 110 of the encoding device 100 .

[0681] Figure 59 is a flowchart showing the operation of the decoding device 200. For example, the decoding device 200 includes a circuit and a memory connected to the circuit. The circuit and memory included in the decoding device 200 may correspond to Figure 46 The circuit of the decoding device 200 performs Figure 59 Specifically, the circuit of the decoding device 200 decodes a block of an image during operation ( S221 ).

[0682] In one example, the circuitry of decoding apparatus 200 may limit the number of context-adaptive decoding processes and decode the image block. Furthermore, the circuitry may also perform sub-block flag decoding without including it in the target number of processes being limited, both in cases where an inverse orthogonal transform is applied to the block and in cases where an inverse orthogonal transform is not applied to the block.

[0683] Here, the sub-block flag decoding process is a process of decoding a sub-block flag indicating whether a sub-block included in a block includes a non-zero coefficient by context-adaptive decoding.

[0684] This makes it possible to decode sub-block flags using context-adaptive decoding, regardless of whether an inverse orthogonal transform is applied or whether the number of context-adaptive decoding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be minimized, leading to a reduction in circuit size.

[0685] Furthermore, when applying an inverse orthogonal transform to a block, the circuitry of decoding apparatus 200 may perform position parameter decoding processing without including it in the number of processing times subject to the restriction. Here, position parameter decoding processing is processing for decoding parameters indicating the position of the first non-zero coefficient in the block in scanning order using context-adaptive decoding.

[0686] Therefore, when an inverse orthogonal transform is applied, the parameter indicating the position of the first non-zero coefficient can be decoded by context adaptive decoding regardless of whether the number of context adaptive decoding processes is limited, thereby reducing the amount of code.

[0687] Furthermore, when applying an inverse orthogonal transform to a block, the circuitry of decoding apparatus 200 can determine a limit on the number of processing operations based on the position of the first non-zero coefficient. This makes it possible to appropriately determine the limit on the number of processing operations when applying an inverse orthogonal transform. Consequently, it is possible to appropriately balance reducing the amount of code and reducing processing delay.

[0688] In another example, the circuits of the decoding apparatus 200 may operate as follows in both cases of applying an inverse orthogonal transform to a block of a decoding target image and not applying an inverse orthogonal transform to the block.

[0689] Specifically, in both cases, if the number of context-adaptive decoding processes is within the processing number limit, the circuitry of decoding apparatus 200 may also decode the coefficient information flag through context-adaptive decoding. Furthermore, in both cases, if the number of context-adaptive decoding processes is not within the processing number limit, the circuitry of decoding apparatus 200 may also skip decoding the coefficient information flag. Here, the coefficient information flag indicates the attributes of the coefficients included in the block.

[0690] Then, if the coefficient information flag is decoded, the circuitry of decoding apparatus 200 can decode the residual value information through Golombirai decoding. Here, the residual value information is information used to reconstruct the value of the coefficient using the coefficient information flag. Furthermore, if decoding of the coefficient information flag is skipped, the circuitry of decoding apparatus 200 can decode the value of the coefficient through Golombirai decoding.

[0691] This makes it possible to skip decoding of coefficient information flags, regardless of whether an inverse orthogonal transform is applied, in accordance with the limit on the number of times context-adaptive decoding is performed. This makes it possible to suppress increases in processing delay and the amount of code. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be reduced, and the circuit scale can be reduced.

[0692] Furthermore, the coefficient information flag may be a flag indicating whether the coefficient value is greater than 1. Thus, regardless of whether an inverse orthogonal transform is applied, it is possible to skip decoding the coefficient information flag indicating whether the coefficient value is greater than 1, in accordance with the limit on the number of context-adaptive decoding processes. This makes it possible to suppress increases in processing delay and the amount of code.

[0693] In another example, the circuitry of decoding apparatus 200 may limit the number of context-adaptive decoding processes and decode a block of an image. Then, without applying an inverse orthogonal transform to the block, it may determine whether multiple coefficient information flags, each representing multiple attributes of a coefficient included in the block, satisfy processing conditions. If the processing conditions are determined to be satisfied, the multiple coefficient information flags may be decoded using context-adaptive decoding.

[0694] Here, the processing condition is a condition that the number of processing times obtained by adding a plurality of coefficient information flags to the number of processing times is within a limited range of the number of processing times.

[0695] Thus, even when an inverse orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive decoding can be used for multiple coefficient information flags. Consequently, processing can be simplified and processing delay can be reduced. Furthermore, when similar processing is performed on blocks to which an inverse orthogonal transform is applied, the difference between the decoding method used in blocks to which an inverse orthogonal transform is applied and the decoding method used in blocks not to which an inverse orthogonal transform is applied can be minimized, and the circuit scale can be reduced.

[0696] Furthermore, the plurality of coefficient information flags may include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Thus, it is possible to collectively determine the plurality of coefficient information flags including the coefficient information flag indicating whether the value of the coefficient is greater than 3 and the coefficient information flag indicating whether the value of the coefficient is greater than 5. Consequently, processing can be simplified and processing delay can be reduced.

[0697] Furthermore, the plurality of coefficient information flags may include a coefficient information flag indicating whether the coefficient value is greater than 7 and a coefficient information flag indicating whether the coefficient value is greater than 9. This makes it possible to collectively determine the plurality of coefficient information flags including four coefficient information flags: whether the coefficient value is greater than 3, whether the coefficient value is greater than 5, whether the coefficient value is greater than 7, and whether the coefficient value is greater than 9. Consequently, processing can be simplified and processing delay can be reduced.

[0698] Note that the above-described operations performed by the circuits of the decoding device 200 may also be performed by the entropy decoding unit 202 of the decoding device 200 .

[0699] [Other examples]

[0700] The encoding device 100 and the decoding device 200 in each of the above examples can be used as an image encoding device and an image decoding device, respectively, or can be used as a moving picture encoding device and a moving picture decoding device, respectively.

[0701] Furthermore, the encoding device 100 and the decoding device 200 may only perform some of the above-mentioned operations, while other devices may perform other operations. Furthermore, the encoding device 100 and the decoding device 200 may only include some of the above-mentioned components, while other devices may include other components.

[0702] Furthermore, at least a portion of each of the above examples may be utilized as an encoding method or a decoding method, or may be utilized as other methods.

[0703] In addition, each component may be formed by dedicated hardware, or implemented by executing a software program suitable for each component. Each component may be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0704] Specifically, the encoding device 100 and the decoding device 200 may each include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to processor a1 or b1, and the storage device corresponds to memory a2 or b2.

[0705] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using the storage device. In addition, when the processing circuit includes the program execution unit, the storage device stores a software program executed by the program execution unit.

[0706] Here, the software that realizes the above-mentioned encoding device 100 or decoding device 200 is the following program.

[0707] For example, the program can cause a computer to execute a coding method, wherein the coding method is to encode a block of an image with a limited number of context adaptive coding processing times, and in the coding of the block, in both cases of a case where an orthogonal transform is applied to the block and a case where an orthogonal transform is not applied to the block, a sub-block flag coding process is performed without including it in the number of processing times, the sub-block flag coding process encoding a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context adaptive coding.

[0708] Furthermore, for example, the program can cause a computer to execute a decoding method, the decoding method comprising: decoding a block of an image while limiting the number of context-adaptive decoding processing times; in decoding the block, in both cases of a case where an inverse orthogonal transform is applied to the block and a case where an inverse orthogonal transform is not applied to the block, performing a sub-block flag decoding process without including it in the number of processing times; the sub-block flag decoding process decoding a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient through context-adaptive decoding.

[0709] In addition, for example, the program can cause a computer to execute a coding method, which is: in both cases of applying an orthogonal transform to a block of an encoding object image and not applying an orthogonal transform to the block, when the number of processing times of context adaptive coding is within a limit range of the processing times, encoding a coefficient information flag representing the attributes of a coefficient contained in the block is performed by context adaptive coding; when the number of processing times is not within the limit range of the processing times, skipping the encoding of the coefficient information flag; when the coefficient information flag is encoded, encoding residual value information for reconstructing the value of the coefficient using the coefficient information flag by Columbus coding; when the encoding of the coefficient information flag is skipped, encoding the value of the coefficient by Columbus coding.

[0710] In addition, for example, the program can cause a computer to execute a decoding method, which is: in both cases of applying an inverse orthogonal transform to a block of a decoding object image and not applying an inverse orthogonal transform to the block, when the number of processing times of context adaptive decoding is within a limit range of the processing times, decoding of a coefficient information flag indicating the attributes of a coefficient contained in the block is performed by context adaptive decoding; when the number of processing times is not within the limit range of the processing times, decoding of the coefficient information flag is skipped; when the coefficient information flag is decoded, residual value information used to reconstruct the value of the coefficient using the coefficient information flag is decoded by Columbus decoding; and when decoding of the coefficient information flag is skipped, decoding of the value of the coefficient by Columbus decoding.

[0711] In addition, for example, the program can cause a computer to execute a coding method, which is: encoding a block of an image while limiting the number of processing times of context adaptive coding, in the coding of the block, when no orthogonal transform is applied to the block, determining whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the block meet a processing condition, and when it is determined that the processing condition is met, encoding the multiple coefficient information flags through context adaptive coding, the processing condition being the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0712] In addition, for example, the program can cause a computer to execute a decoding method, which is: decoding a block of an image while limiting the number of processing times of context adaptive decoding, in the decoding of the block, without applying an inverse orthogonal transform to the block, determining whether multiple coefficient information flags respectively representing multiple attributes of coefficients contained in the block meet a processing condition, and when it is determined that the processing condition is met, decoding the multiple coefficient information flags through context adaptive decoding, the processing condition being the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0713] In addition, as described above, each component may also be a circuit. These circuits may constitute a single circuit as a whole, or they may be different circuits. In addition, each component may be implemented by a general-purpose processor or a dedicated processor.

[0714] Furthermore, the processing performed by a specific component may be performed by another component. Furthermore, the order in which the processing is performed may be changed, and multiple processing may be performed simultaneously. Furthermore, the encoding and decoding apparatus may include the encoding apparatus 100 and the decoding apparatus 200.

[0715] In addition, the ordinal numbers such as 1 and 2 used in the description may be appropriately interchanged. In addition, new ordinal numbers may be assigned to components, or ordinal numbers may be removed.

[0716] While the embodiments of the encoding device 100 and the decoding device 200 have been described above based on a plurality of examples, the embodiments of the encoding device 100 and the decoding device 200 are not limited to these examples. The embodiments of the encoding device 100 and the decoding device 200 may also include various modifications that can be conceived by those skilled in the art to the respective examples, or embodiments constructed by combining components from different examples, without departing from the spirit of the present invention.

[0717] One or more aspects disclosed herein may be combined with at least a portion of other aspects of the present invention for implementation. In addition, a portion of the processing, a portion of the structure of the device, a portion of the syntax, etc. recorded in the flowcharts of one or more aspects disclosed herein may be combined with other aspects for implementation.

[0718] [Implementation and Application]

[0719] In each of the above embodiments, each functional block or active block can generally be implemented by an MPU (microprocessing unit) and a memory. In addition, the processing of each functional block can also be implemented by a program execution unit such as a processor that reads and executes software (programs) recorded in a recording medium such as a ROM. The software can be distributed. The software can also be recorded in various recording media such as semiconductor memories. In addition, each functional block can also be implemented by hardware (dedicated circuit). Various combinations of hardware and software can be used.

[0720] The processing described in each embodiment can be implemented by centralized processing using a single device (system) or by distributed processing using multiple devices. In addition, the processors that execute the above programs can be single or multiple. In other words, centralized processing can be performed or distributed processing can be performed.

[0721] The aspects of the present invention are not limited to the above-described embodiments, and various modifications are possible, which are also included in the scope of the aspects of the present invention.

[0722] Furthermore, here, application examples of the moving picture encoding method (image encoding method) or moving picture decoding method (image decoding method) described in each of the above embodiments and various systems implementing these application examples are described. Such a system may be characterized by including an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding and decoding device including both. Other configurations of such a system may be modified as appropriate depending on the circumstances.

[0723] [Example of use]

[0724] Figure 60 This diagram shows the overall structure of a content supply system ex100 for implementing content distribution services. The communication service provision area is divided into cells of desired sizes, and in the illustrated example, base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.

[0725] In the content delivery system ex100, various devices, such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115, are connected to the Internet ex101 via an Internet service provider ex102, a communication network ex104, and base stations ex106-ex110. The content delivery system ex100 may also connect some of these devices in combination. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network, short-range wireless, or other means, rather than via base stations ex106-ex110. Furthermore, the streaming server ex103 may be connected to various devices, such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115, via the Internet ex101. Furthermore, the streaming server ex103 may be connected to a terminal, such as a hotspot within an airplane ex117, via a satellite ex116.

[0726] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly to the aircraft ex117 without going through the satellite ex116.

[0727] The camera ex113 is a device such as a digital camera capable of capturing both still and moving images. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handy-phone System) compatible with mobile communication systems such as 2G, 3G, 3.9G, 4G, and what will be called 5G in the future.

[0728] The home appliance ex114 is a refrigerator or equipment included in a household fuel cell cogeneration system.

[0729] In the content delivery system ex100, terminals with camera functions are connected to the streaming server ex103 via a base station ex106 or the like, enabling on-site distribution and the like. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) can perform the encoding processing described in the above embodiments on still images or moving image content captured by users using these terminals. Furthermore, the terminals can multiplex the encoded video data with the audio data obtained by encoding the corresponding audio, and transmit the resulting data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present invention.

[0730] Meanwhile, the streaming server ex103 streams content data sent by requesting clients. Clients are computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or terminals inside airplanes ex117, all capable of decoding the encoded data. Each device that receives the distributed data can also decode and reproduce the data. In other words, each device can function as an image decoding device according to one aspect of the present invention.

[0731] [Distributed Processing]

[0732] Alternatively, the streaming server ex103 can consist of multiple servers or computers, distributing data by distributing processing or recording. For example, the streaming server ex103 can be implemented as a CDN (Content Delivery Network), which distributes content through a network connecting numerous edge servers distributed worldwide. In a CDN, physically close edge servers can be dynamically assigned to clients. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Furthermore, in the event of various errors or changes in communication status due to increased traffic, processing can be distributed across multiple edge servers, distribution can be switched to other edge servers, or delivery can be continued by bypassing a faulty portion of the network, thus achieving high-speed and stable delivery.

[0733] In addition, the encoding process of the captured data is not limited to the distributed processing itself, and can be performed by each terminal, on the server side, or shared. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or encoding amount of the image of the frame or scene unit is detected. In addition, in the second cycle, a process is performed to maintain the image quality and improve the encoding efficiency. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, the quality and efficiency of the content can be improved while reducing the processing load in each terminal. In this case, if there is a request for almost real-time reception and decoding, the data completed by the first encoding by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be achieved.

[0734] As another example, cameras ex113 and others extract features (features or feature quantities) from images, compress the feature data as metadata, and transmit it to a server. The server, for example, determines the importance of an object based on the feature data and switches the quantization precision, performing compression appropriate to the meaning (or content) of the image. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during recompression in the server. Alternatively, the terminal can perform simple encoding such as VLC (Variable Length Coding), while the server performs more processing-intensive encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding).

[0735] As another example, in stadiums, shopping malls, factories, and other locations, there may be multiple video data sets generated by capturing roughly the same scene using multiple terminals. In such cases, encoding can be distributed using the multiple terminals that captured the images, as well as other terminals and servers that did not capture the images as needed, for example, by allocating the encoding processing to each GOP (Group of Picture) unit, picture unit, or tile unit obtained by dividing the picture. This reduces latency and achieves better real-time performance.

[0736] Because multiple image data sets represent roughly the same scene, the server can manage and / or instruct the image data captured by each terminal to cross-reference each other. Furthermore, the server can receive encoded data from each terminal and change the reference relationship between the multiple data sets, or modify or replace the image itself before re-encoding it. This allows the generation of a stream with improved quality and efficiency for each data set.

[0737] Furthermore, the server may also perform transcoding to change the encoding method of the video data before distributing the video data. For example, the server may convert the encoding method of the MPEG type to the VP type (such as VP9), or convert H.264 to H.265.

[0738] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions of "server" or "terminal" refer to the entity performing the processing. However, some or all of the processing performed by the server can also be performed by the terminal, and some or all of the processing performed by the terminal can also be performed by the server. Furthermore, the same applies to the decoding process.

[0739] [3D, multi-angle]

[0740] There is an increasing trend for combining and utilizing images or videos captured by multiple devices, such as cameras ex113 and / or smartphones ex115, that are roughly synchronized with each other, to capture different scenes or the same scene from different angles. The images captured by each device can be combined based on the relative positional relationship between the devices, or based on areas with consistent feature points contained in the images.

[0741] The server not only encodes two-dimensional moving images but can also encode still images automatically or at user-specified times based on scene analysis of moving images and transmit them to the receiving terminal. Furthermore, if the server can determine the relative positional relationship between the capturing terminals, it can generate not only two-dimensional moving images but also three-dimensional shapes of the scene based on images of the same scene captured from different angles. The server can also separately encode three-dimensional data generated from point clouds, etc., or, based on the results of identifying or tracking people or objects using three-dimensional data, select or reconstruct images captured by multiple terminals to generate images for transmission to the receiving terminal.

[0742] This allows users to arbitrarily select the images corresponding to each camera terminal to enjoy the scene, or to enjoy content that extracts images from a selected viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, along with the video, audio can be collected from multiple angles. The server multiplexes the audio from a specific angle or space with the corresponding video and transmits the multiplexed video and audio.

[0743] Furthermore, content that connects the real and virtual worlds, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular in recent years. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. These can be encoded using techniques such as Multi-View Coding (MVC) to allow for reference between viewpoint images, or encoded as separate streams without reference to each other. When these separate streams are decoded, they can be played back in sync with the user's viewpoint, recreating a virtual three-dimensional space.

[0744] In the case of AR images, the server can also overlay virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or movement of the user's viewpoint. The decoding device obtains or stores the virtual object information and three-dimensional data, generates a two-dimensional image based on the movement of the user's viewpoint, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can also send the user's viewpoint movement to the server in addition to the request for virtual object information. Alternatively, the server can create overlay data based on the three-dimensional data stored on the server, matching the received viewpoint movement, encode the overlay data, and distribute it to the decoding device. In addition, the overlay data typically has an alpha value indicating transparency in addition to RGB. The server sets the alpha value of the portion other than the target generated based on the three-dimensional data to 0, for example, and encodes the portion in a transparent state. Alternatively, the server can set the RGB value of a specified value as the background, as in a chroma key, and generate data with the portion other than the target as the background color. The specified RGB value can also be predetermined.

[0745] Similarly, the decoding process of the distributed data can be performed by the client (for example, the terminal), can be performed on the server side, or can be shared and performed. As an example, a terminal may first send a reception request to the server, and other terminals may receive the content corresponding to the request and perform decoding processing, and send the decoded signal to a device with a display. By distributing the processing regardless of the performance of the communicative terminal itself and selecting appropriate content, data with better image quality can be reproduced. In addition, as another example, large-size image data can also be received by a TV, etc., and a portion of the image, such as tiles, can be decoded and displayed by the viewer's personal terminal. In this way, while sharing the overall image, it is possible to confirm one's own area of ​​responsibility or the area that one wants to confirm in more detail at hand.

[0746] In situations where multiple short-range, medium-range, or long-range wireless communications can be used indoors and outdoors, it may be possible to seamlessly receive content using distribution system standards such as MPEG-DASH. Users can also freely select their own terminals, decoding devices such as displays installed indoors and outdoors, and switch in real time. In addition, it is possible to switch the decoding terminal and the display terminal and perform decoding using their own location information. As a result, it is also possible to map and display information on a part of the wall or ground of a building next to a display device while the user is moving to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as caching the encoded data in a server that can be accessed from the receiving terminal in a short time, or copying the encoded data in the edge server of the content distribution service.

[0747] [Scalable Coding]

[0748] To switch content, use Figure 61 The example illustrates a scalable stream compressed and encoded using the moving picture coding method described in the above embodiments. For the server, multiple streams with the same content but different qualities can be provided as a single stream. Alternatively, a structure can be employed to switch content by leveraging the temporally and spatially scalable nature of streams achieved through layered coding, as shown in the figure. Specifically, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. This allows the decoding side to freely switch between low-resolution and high-resolution content. For example, if a user wants to watch a video viewed on a smartphone ex115 while on the go, then watch it later on a device such as an internet TV at home, the device can simply decode the same stream into different layers, reducing the burden on the server.

[0749] Furthermore, in addition to the hierarchical structure of encoding pictures per layer and implementing enhancement layers above the base layer as described above, the enhancement layers may also include metadata such as statistical information based on the image. Alternatively, the decoding side may generate high-definition content by super-resolutioning the base layer pictures based on this metadata. Super-resolution can improve the signal-to-noise ratio while maintaining and / or increasing resolution. Meta-information includes information for determining linear or nonlinear filter coefficients used in super-resolution processing, as well as information for determining parameter values ​​in filtering, machine learning, or least-squares operations used in super-resolution processing.

[0750] Alternatively, a structure can be provided that divides a picture into tiles according to the meaning of the object in the image. The decoding side decodes only a part of the area by selecting the tile to be decoded. Moreover, by storing the attributes of the object (people, cars, balls, etc.) and the position in the image (coordinate position in the same image, etc.) as meta information, the decoding side can determine the position of the desired object based on the meta information and decide the tile that includes the object. For example, Figure 62 As shown, SEI (supplemental enhancement information) messages in HEVC, which are different from pixel data, can also be used to store meta-information. This meta-information indicates, for example, the position, size, or color of the main object.

[0751] Meta-information can also be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in the image, and by matching this information with the time information of the picture unit, it can determine the picture in which the target appears and the location of the target within the picture.

[0752] [Web page optimization]

[0753] Figure 63 This is a diagram showing an example of a display screen of a web page on the computer ex111 or the like. Figure 64 1 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. Figure 63 and Figure 64 As shown, a web page may contain multiple link images that serve as links to image content. The display format of these images may vary depending on the viewing device. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture associated with each content as a link image until the user explicitly selects a link image, or until the link image approaches the center of the screen, or until the entire link image enters the screen. Alternatively, the display device (decoding device) may display an image such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer, decode the image, and display it.

[0754] When a linked image is selected by the user, the display device, for example, sets the base layer as the top priority and decodes it. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device can also decode to the enhancement layer. Moreover, in order to ensure real-time performance, before selection or when the communication bandwidth is very tight, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of decoding to the start of display of the content) by decoding and displaying only the forward reference pictures (I pictures, P pictures, and B pictures that are only forward referenced). Furthermore, the display device can also forcibly ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward references, and perform normal decoding as the number of pictures received increases over time.

[0755] [Automatic driving]

[0756] Furthermore, when transmitting and receiving still images or video data, such as two-dimensional or three-dimensional map information, for autonomous driving or driving assistance, the receiving terminal may receive weather or construction information as metadata in addition to image data belonging to one or more layers, and decode these metadata by associating them with each other. Furthermore, the metadata may belong to a layer or be multiplexed solely with the image data.

[0757] In this case, since the receiving terminal, such as a car, drone, or airplane, is moving, the receiving terminal transmits its location information, enabling seamless reception and decoding while switching between base stations ex106-ex110. Furthermore, the receiving terminal can dynamically switch the level of metadata received and the level of map information updated based on user preferences, user status, and / or communication band conditions.

[0758] In the content providing system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0759] [Distribution of Personal Content]

[0760] Furthermore, the content delivery system ex100 can deliver not only high-quality, long-duration content provided by video distributors, but also low-quality, short-duration content provided by individuals, either unicast or multicast. Such personal content is expected to increase in the future. To enhance personal content, the server can also perform encoding after editing. This can be achieved, for example, with the following configuration.

[0761] After taking the photos in real time or accumulating them, the server performs recognition processing such as shooting errors, scene search, meaning analysis, and target detection based on the original image data or encoded data. In addition, based on the recognition results, the server manually or automatically corrects focus deviation or hand shaking, deletes less important scenes such as scenes with lower brightness than other pictures or scenes that are not in focus, emphasizes the edges of the target, or changes the color tone. The server encodes the edited data based on the editing results. In addition, it is known that the viewing rate will decrease if the shooting time is too long. The server can also automatically limit not only the less important scenes as mentioned above, but also scenes with less movement based on the image processing results, so that the content is within a specific time range. Alternatively, the server can also generate a summary based on the results of the scene's meaning analysis and encode it.

[0762] There are cases where personal content in its original state may be written with content that infringes copyright, author's personality rights or portrait rights, etc., or there are cases where the scope of sharing exceeds the desired scope, which is inconvenient for individuals. Therefore, for example, the server can also forcibly change the faces of people in the peripheral part of the screen, or the home, etc. to an out-of-focus image for encoding. In addition, the server can also identify whether the face of a person different from the pre-registered person is captured in the image to be encoded, and if so, perform processing such as applying mosaics to the face part. Alternatively, as pre-processing or post-processing for encoding, the user can also specify the person or background area that he wants to process the image from the perspective of copyright, etc. The server can also replace the specified area with another image, or blur the focus, etc. If it is a person, it can track the person in the moving image and replace the image of the person's face.

[0763] The viewing of personal content with a small amount of data has a strong demand for real-time performance, so although it also depends on the bandwidth, the decoding device can also receive, decode, and reproduce the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the playback is repeated more than twice, such as in the case of looped playback, the enhancement layer is also included in the playback of high-definition images. In this way, if the stream is scalable, it can provide an experience in which the moving image is relatively rough when it is not selected or at the beginning of viewing, but the stream gradually becomes smoother and the image becomes better. In addition to scalable coding, the same experience can be provided when the relatively rough stream played back the first time and the second stream encoded with reference to the moving image of the first time are composed of a single stream.

[0764] [Other application examples]

[0765] In addition, these encoding and decoding processes are usually processed in the LSI ex500 of each terminal. Figure 60 ) can be a single chip or a multi-chip configuration. Alternatively, video encoding or decoding software can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) readable by the computer ex111 or the like, and encoding and decoding can be performed using this software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera can be transmitted. In this case, the video data can be encoded by the LSI ex500 in the smartphone ex115.

[0766] Alternatively, the LSIex500 can be configured to download and activate application software. In this case, the terminal first determines whether it is compatible with the content's encoding scheme or has the capability to perform a specific service. If the terminal is not compatible with the content's encoding scheme or does not have the capability to perform a specific service, it can download the codec or application software to retrieve and play the content.

[0767] Furthermore, the content delivery system ex100 is not limited to the content delivery system ex100 via the Internet ex101; at least one of the video encoding devices (image encoding devices) or video decoding devices (image decoding devices) described in the above-mentioned embodiments can also be incorporated into a digital broadcasting system. Since multiplexed data containing multiplexed video and audio is transmitted and received over broadcast radio waves using satellites, the content delivery system ex100 differs from the unicast-friendly structure of the content delivery system ex100 in that it is suitable for multicast. However, the encoding and decoding processes can be applied in the same manner.

[0768] [Hardware structure]

[0769] Figure 65 Is a further detailed representation Figure 60 FIG. 1 shows a diagram of the smartphone ex115. Figure 66 This diagram shows an example of the configuration of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of capturing both video and still images, and a display unit ex458 that displays images captured by the camera unit ex465 and decoded data such as images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio and sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded or decoded data such as captured videos or still images, recorded audio, received videos or still images, and emails, and a slot unit ex464 that serves as an interface with a SIM card ex468 for identifying users and authenticating access to various data, including the network. Alternatively, an external memory card can be used in place of the memory unit ex467.

[0770] The main control unit ex460, which can perform integrated control of the display unit ex458 and the operation unit ex466, is synchronously connected to the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the sound signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via a bus ex470.

[0771] When the user turns on the power button, the power supply circuit unit ex461 activates the smartphone ex115 to be operational and supplies power to various components from the battery pack.

[0772] The smartphone ex115 performs processes such as calls and data communications under the control of the main control unit ex460, which includes a CPU, ROM, and RAM. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. The signal is then subjected to spread spectrum processing by the modulation / demodulation unit ex452. The transmission / reception unit ex451 then performs digital-to-analog conversion and frequency conversion, and the resulting signal is transmitted via the antenna ex450. Furthermore, received data is amplified, subjected to frequency conversion and analog-to-digital conversion, and then subjected to inverse spread spectrum processing by the modulation / demodulation unit ex452. The audio signal processing unit ex454 converts the signal into an analog audio signal, which is then output from the audio output unit ex457. During data communications, text, still images, or video data can be transmitted via the operation input control unit ex462 under the control of the main control unit ex460, based on operations on the main unit's operation unit ex466. Similar transmission and reception processes are performed. In data communication mode, when transmitting video, still images, or both video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and then sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of capturing video or still images by the camera unit ex465, and then sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and audio data in a predetermined format. The modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451 perform modulation and conversion processing, and then transmit the data via the antenna ex450. The predetermined format can also be predetermined.

[0773] When receiving a video file attached to an email or chat tool, or a video file linked to a webpage, the multiplexing / demultiplexing unit ex453 demultiplexes the multiplexed data received via antenna ex450 into a bitstream of video data and a bitstream of audio data. The multiplexing / demultiplexing unit ex453 then supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronous bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in the above embodiments. The video or still image contained in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs the audio through the audio output unit ex457. As live streaming becomes increasingly common, audio reproduction may become socially inappropriate depending on the user's circumstances. Therefore, it is preferable that as an initial value, a configuration is adopted in which only the video data is reproduced without reproducing the audio signal, and the audio is reproduced in synchronization only when the user performs an operation such as clicking on the video data.

[0774] While the smartphone ex115 is used as an example, other possible terminal configurations include transmitting and receiving terminals with both an encoder and a decoder, as well as transmitting terminals with only an encoder and receiving terminals with only a decoder. In the digital broadcasting system, the description assumes the reception and transmission of multiplexed data containing audio data multiplexed with video data. However, multiplexed data can also contain text data associated with the video in addition to audio data. Furthermore, it is also possible to receive or transmit video data itself, rather than multiplexed data.

[0775] While the main control unit ex460, which includes a CPU, has been described as controlling the encoding and decoding processes, many terminals also include GPUs. Therefore, a configuration can also be implemented where a memory shared by the CPU and GPU, or a memory whose addresses are managed in a mutually usable manner, allows the GPU to process larger areas simultaneously. This can shorten encoding time, ensure real-time performance, and achieve low latency. In particular, it is more efficient if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing are performed simultaneously on a per-picture basis, such as on the GPU, rather than on the CPU.

[0776] Industrial applicability

[0777] The present invention can be utilized in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conferencing system, or an electronic mirror.

[0778] Description of Reference Numerals

[0779] 100 Encoding device

[0780] 102 Division

[0781] 104 Subtraction Department

[0782] 106 Transformation Unit

[0783] 108 Quantitative Department

[0784] 110 Entropy Coding Unit

[0785] 112, 204 Inverse Quantization Unit

[0786] 114, 206 Inverse transformation unit

[0787] 116, 208 Addition Department

[0788] 118, 210 block memory

[0789] 120, 212 loop filter unit

[0790] 122, 214 frame memories

[0791] 124, 216 Intra-frame prediction unit

[0792] 126, 218 Inter-frame prediction unit

[0793] 128, 220 Prediction and Control Department

[0794] 200 decoding device

[0795] 202 Entropy Decoding Unit

[0796] 1201 Boundary Judgment Department

[0797] 1202, 1204, 1206 switches

[0798] 1203 Filtering and Judgment Unit

[0799] 1205 Filter Processing Unit

[0800] 1207 Filter characteristic determination unit

[0801] 1208 Processing and Judgment Unit

[0802] a1 and b1 processors

[0803] a2, b2 memory

Claims

1. An encoding device, wherein: have: circuits; and a memory connected to the circuit, In the residual coding of the current block, in both cases of applying orthogonal transform using mutually different syntaxes and skipping the orthogonal transform, When a limit on the number of processing times of context adaptive coding allows context adaptive coding to be performed on a plurality of coefficient information flags related to the coefficients included in the current block, the circuit encodes the plurality of coefficient information flags by the context adaptive coding and encodes residual values ​​of the coefficients by Columbus coding, the residual values ​​of the coefficients being values ​​used to reconstruct values ​​of the coefficients using the plurality of coefficient information flags. When the processing number limit does not allow the context-adaptive coding of the plurality of coefficient information flags to be performed simultaneously, the circuit skips coding the plurality of coefficient information flags and encodes the values ​​of the coefficients by the Golomblais coding. In a case where the orthogonal transform is skipped, the circuit encodes the plurality of absolute value flags by the context adaptive coding after encoding coefficient information flags other than a plurality of absolute value flags among the plurality of coefficient information flags and before encoding residual values ​​of the coefficients, the plurality of absolute value flags being flags related to whether the absolute values ​​of the coefficients are greater than a prescribed value, and the prescribed value being an integer greater than 1. The plurality of coefficient information flags include a flag indicating whether the value of the coefficient is zero or non-zero.

2. A decoding device, wherein: have: circuits; and a memory connected to the circuit, In residual decoding of the current block, in both cases of applying inverse orthogonal transform using mutually different syntaxes and skipping the inverse orthogonal transform, When a limit on the number of processing times of context adaptive decoding allows context adaptive decoding to be performed on a plurality of coefficient information flags related to the coefficients included in the current block, the circuit decodes the plurality of coefficient information flags by the context adaptive decoding and decodes residual values ​​of the coefficients by Columbus decoding, the residual values ​​of the coefficients being values ​​used to reconstruct values ​​of the coefficients using the plurality of coefficient information flags. When the processing number limit does not allow the context adaptive decoding of the plurality of coefficient information flags at once, the circuit skips decoding the plurality of coefficient information flags and decodes the values ​​of the coefficients by the Columbus decoding. In a case where the inverse orthogonal transform is skipped, the circuit decodes the plurality of absolute value flags by the context adaptive decoding after decoding coefficient information flags other than a plurality of absolute value flags among the plurality of coefficient information flags and before decoding residual values ​​of the coefficients, the plurality of absolute value flags being flags related to whether the absolute values ​​of the coefficients are greater than a prescribed value, and the prescribed value being an integer greater than 1. The plurality of coefficient information flags include a flag indicating whether the value of the coefficient is zero or non-zero.

3. A computer-readable non-transitory storage medium storing a bitstream, wherein: The bit stream includes information for causing a decoding device to perform a residual decoding process. In residual decoding of the current block, in both cases of applying inverse orthogonal transform using mutually different syntaxes and skipping the inverse orthogonal transform, When a limit on the number of processing times of context adaptive decoding allows context adaptive decoding to be performed on a plurality of coefficient information flags related to the coefficients included in the current block, the plurality of coefficient information flags are decoded by the context adaptive decoding, and residual values ​​of the coefficients are decoded by Columbus decoding, the residual values ​​of the coefficients being values ​​used to reconstruct values ​​of the coefficients using the plurality of coefficient information flags. When the processing number limit does not allow the context adaptive decoding of the plurality of coefficient information flags at once, decoding of the plurality of coefficient information flags is skipped and the values ​​of the coefficients are decoded by the Columbus decoding. In a case where the inverse orthogonal transform is skipped, after decoding coefficient information flags other than a plurality of absolute value flags among the plurality of coefficient information flags and before decoding residual values ​​of the coefficients, the plurality of absolute value flags are decoded by the context adaptive decoding, the plurality of absolute value flags being flags related to whether the absolute values ​​of the coefficients are greater than a prescribed value, and the prescribed value being an integer greater than 1, The plurality of coefficient information flags include a flag indicating whether the value of the coefficient is zero or non-zero.