Encoding device, decoding device, and non-transitory storage medium

By combining context-adaptive coding and Columbine coding, the use of orthogonal transform and inverse orthogonal transform is optimized, which solves the problems of coding efficiency and circuit scale in existing video coding technology and achieves more efficient video encoding and decoding processing.

CN120812271APending Publication Date: 2025-10-17PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511250188.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-24
Filing Date
2020-04-24
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing video coding technologies require improvements in coding efficiency, image quality, processing capacity, circuit size, and processing speed. In particular, there is a lack of effective methods for selecting elements or actions such as blocks, sizes, motion vectors, and reference images.

Method used

A combination of context-adaptive coding and Columbine coding is adopted to limit the number of processing times and appropriately select the coding method, optimize the use of orthogonal transform and inverse orthogonal transform, perform sub-block flag coding and position parameter coding, simplify the processing flow and reduce the circuit scale.

Benefits of technology

It improves coding efficiency, improves image quality, reduces processing volume and circuit scale, increases processing speed, and enables appropriate selection of coding methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812271A_ABST
    Figure CN120812271A_ABST
Patent Text Reader

Abstract

The invention provides an encoding device, a decoding device, and a non-transitory storage medium. The encoding device includes: a circuit; and the memory is connected with the circuit, and the circuit is used for carrying out residual encoding on the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of the patent application with application number 202080030228.9, filed on April 24, 2020, and titled “Encoding apparatus, decoding apparatus, encoding method, and decoding method”. TECHNICAL FIELD

[0002] The present application relates to video encoding, for example, systems, components, and methods in encoding and decoding of moving images, and the like. BACKGROUND

[0003] Video encoding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to handle the increasing amount of digital video data in various uses, there has been a constant need to provide improvements and optimizations in video encoding technology.

[0004] Furthermore, Non-Patent Literature 1 relates to an example of an existing standard related to the above-described video encoding technology.

[0005] Prior Art Documents

[0006] Non-Patent Literature

[0007] Non-Patent Literature 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) SUMMARY

[0008] Problems to be Solved by the Invention

[0009] With regard to the above-described encoding method, it is desirable to propose a new method for improvement of encoding efficiency, improvement of picture quality, reduction of processing amount, reduction of circuit size, or appropriate selection of elements or actions such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, and the like.

[0010] The present application provides a structure or method that can contribute to one or more of, for example, improvement of encoding efficiency, improvement of picture quality, reduction of processing amount, reduction of circuit size, improvement of processing speed, and appropriate selection of elements or actions. Furthermore, the present application can include a structure or method that can contribute to benefits other than the above.

[0011] Means for Solving the Problems

[0012] For example, an encoding device of an aspect of the present technology includes a circuit that performs residual coding of a current block, and a memory that is connected to the circuit.

[0013] The installation of the several embodiments of the present technology can improve coding efficiency, can simplify encoding / decoding processing, can speed up the encoding / decoding processing, and can efficiently select appropriate components / operations used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, and reference blocks.

[0014] Further advantages and effects of an aspect of the present technology are apparent from the description and the drawings. The advantages and / or effects are respectively obtained by the features recited in the several embodiments and the description and the drawings, but all of the features are not necessarily required to obtain one or more of the advantages and / or effects.

[0015] Furthermore, these general or specific aspects can also be realized by a system, a method, an integrated circuit, a computer program, a recording medium or any combination of them.

[0016] Effects of the Invention

[0017] An aspect of the present technology can contribute to one or more of improvement of coding efficiency, improvement of image quality, reduction of processing amount, reduction of circuit size, improvement of processing speed, and appropriate selection of components or operations. In addition, an aspect of the present technology can contribute to benefits other than the above. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a block diagram showing a functional configuration of an encoding device of an embodiment.

[0019] Figure 2 is a flowchart showing an example of overall encoding processing performed by an encoding device.

[0020] Figure 3 is a conceptual diagram showing an example of block partitioning.

[0021] Figure 4A is a conceptual diagram showing an example of a structure of a slice.

[0022] Figure 4B is a conceptual diagram showing an example of a structure of a tile.

[0023] Figure 5A is a table showing transform basis functions corresponding to various transform types.

[0024] Figure 5Bis a conceptual diagram showing an example of a concept of SVT (Spatially Varying Transform).

[0025] Figure 6A is a conceptual diagram showing an example of a shape of a filter used in ALF (adaptive loop filter).

[0026] Figure 6B is a conceptual diagram showing another example of a shape of a filter used in ALF.

[0027] Figure 6C is a conceptual diagram showing another example of a shape of a filter used in ALF.

[0028] Figure 7 is a block diagram showing an example of a detailed structure of a loop filter section functioning as a DBF (deblocking filter).

[0029] Figure 8 is a conceptual diagram showing an example of deblocking filtering having a filter characteristic symmetrical with respect to a block boundary.

[0030] Figure 9 is a conceptual diagram for explaining a block boundary on which deblocking filtering is performed.

[0031] Figure 10 is a conceptual diagram showing an example of a Bs value.

[0032] Figure 11 is a flowchart showing an example of a process performed by a prediction processing section of an encoding apparatus.

[0033] Figure 12 is a flowchart showing another example of a process performed by a prediction processing section of an encoding apparatus.

[0034] Figure 13 is a flowchart showing another example of a process performed by a prediction processing section of an encoding apparatus.

[0035] Figure 14 is a conceptual diagram showing an example of 67 intra prediction modes in intra prediction of an embodiment.

[0036] Figure 15 is a flowchart showing an example of a flow of a basic process of inter prediction.

[0037] Figure 16 is a flowchart showing an example of motion vector derivation.

[0038] Figure 17 is a flowchart showing another example of motion vector derivation.

[0039] Figure 18 is a flowchart showing another example of motion vector derivation.

[0040] Figure 19 is a flowchart showing an example of inter prediction based on normal inter mode.

[0041] Figure 20 is a flowchart showing an example of inter prediction based on merge mode.

[0042] Figure 21 is a conceptual diagram for explaining an example of motion vector derivation processing based on merge mode.

[0043] Figure 22 is a flowchart showing an example of FRUC (frame rate up conversion) processing.

[0044] Figure 23 is a conceptual diagram for explaining an example of pattern matching (bi-directional matching) between two blocks along a motion trajectory.

[0045] Figure 24 is a conceptual diagram for explaining an example of pattern matching (template matching) between a template within a current picture and a block within a reference picture.

[0046] Figure 25A is a conceptual diagram for explaining an example of derivation of motion vector of a sub-block unit based on motion vectors of a plurality of neighboring blocks.

[0047] Figure 25B is a conceptual diagram for explaining an example of derivation of motion vector of a sub-block unit in an affine mode with three control points.

[0048] Figure 26A is a conceptual diagram for explaining an affine merge mode.

[0049] Figure 26B is a conceptual diagram for explaining an affine merge mode with two control points.

[0050] Figure 26C is a conceptual diagram for explaining an affine merge mode with three control points.

[0051] Figure 27 is a flowchart showing an example of processing of an affine merge mode.

[0052] Figure 28A is a conceptual diagram for explaining an affine inter mode with two control points.

[0053] Figure 28Bis a conceptual diagram for explaining an affine inter mode with 3 control points.

[0054] Figure 29 is a flowchart representing an example of the process of the affine inter mode.

[0055] Figure 30A is a conceptual diagram for explaining an affine inter mode in which a current block has 3 control points and a neighboring block has 2 control points.

[0056] Figure 30B is a conceptual diagram for explaining an affine inter mode in which a current block has 2 control points and a neighboring block has 3 control points.

[0057] Figure 31A is a flowchart representing a merge mode including DMVR (decoder motion vector refinement).

[0058] Figure 31B is a conceptual diagram for explaining an example of the DMVR process.

[0059] Figure 32 is a flowchart representing an example of the generation of a prediction image.

[0060] Figure 33 is a flowchart representing another example of the generation of a prediction image.

[0061] Figure 34 is a flowchart representing another example of the generation of a prediction image.

[0062] Figure 35 is a flowchart for explaining an example of a prediction image correction process based on an OBMC (overlapped block motion compensation) process.

[0063] Figure 36 is a conceptual diagram for explaining an example of a prediction image correction process based on an OBMC process.

[0064] Figure 37 is a conceptual diagram for explaining the generation of a prediction image of 2 triangles.

[0065] Figure 38 is a conceptual diagram for explaining a model assuming constant velocity straight line motion.

[0066] Figure 39 is a conceptual diagram for explaining an example of a prediction image generation method using a brightness correction process based on LIC (local illumination compensation).

[0067] Figure 40 This is a block diagram showing an example of an implementation of an encoding device.

[0068] Figure 41 This is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0069] Figure 42 This is a flowchart showing an example of the overall decoding process performed by the decoding device.

[0070] Figure 43 This is a flowchart showing an example of processing performed by the prediction processing unit of the decoding device.

[0071] Figure 44 This is a flowchart showing another example of processing performed by the prediction processing unit of the decoding device.

[0072] Figure 45 This is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in a decoding device.

[0073] Figure 46 This is a block diagram showing an implementation example of a decoding device.

[0074] Figure 47 This is a flowchart showing the basic coefficient encoding method of the first aspect.

[0075] Figure 48 This is a flowchart showing the basic first encoding method of the first aspect.

[0076] Figure 49 This is a flowchart showing the basic second encoding method of the first aspect.

[0077] Figure 50 This is a flowchart showing the coefficient encoding method of the first example of the first aspect.

[0078] Figure 51 This is a flowchart showing the coefficient encoding method of the second example of the first aspect.

[0079] Figure 52 This is a flowchart showing the coefficient encoding method of the first example of the second aspect.

[0080] Figure 53 This is a flowchart showing the coefficient encoding method of the second example of the second aspect.

[0081] Figure 54 This is a syntax diagram showing the basic first encoding method of the third form.

[0082] Figure 55 This is a syntax diagram showing the basic second encoding method of the third form.

[0083] Figure 56 is a syntax diagram of the second encoding method representing the first example of the third modality.

[0084] Figure 57 is a syntax diagram of the second encoding method representing the second example of the third modality.

[0085] Figure 58 is a flowchart showing an action of an encoding apparatus of the embodiment.

[0086] Figure 59 is a flowchart showing an action of a decoding apparatus of the embodiment.

[0087] Figure 60 is a block diagram showing an overall structure of a content supply system that implements a content distribution service.

[0088] Figure 61 is a conceptual diagram showing an example of an encoding configuration at the time of scalable encoding.

[0089] Figure 62 is a conceptual diagram showing an example of an encoding configuration at the time of scalable encoding.

[0090] Figure 63 is a conceptual diagram showing an example of a display screen of a web page.

[0091] Figure 64 is a conceptual diagram showing an example of a display screen of a web page.

[0092] Figure 65 is a block diagram showing an example of a smart phone.

[0093] Figure 66 is a block diagram showing an example of a structure of a smart phone. DETAILED DESCRIPTION

[0094] For example, in encoding of a block of an image, an encoding apparatus can sometimes transform the block into data that is easy to compress by applying an orthogonal transform to the block. On the other hand, in encoding of a block of an image, an encoding apparatus can sometimes reduce processing delay by not applying an orthogonal transform to the block.

[0095] Furthermore, a characteristic of a block to which an orthogonal transform is applied and a characteristic of a block to which an orthogonal transform is not applied are different from each other. An encoding method used for a block to which an orthogonal transform is applied and an encoding method used for a block to which an orthogonal transform is not applied can be different from each other.

[0096] However, in a case where an inappropriate coding mode is used for a block to which an orthogonal transform is applied, or in a case where an inappropriate coding mode is used for a block to which an orthogonal transform is not applied, an increase in coding amount or an increase in processing delay, or the like can occur. Further, in a case where there is a large difference between a coding mode used for a block to which an orthogonal transform is applied and a coding mode used for a block to which an orthogonal transform is not applied, processing can be complicated, and the circuit size can increase.

[0097] Therefore, for example, an encoding apparatus of one aspect of the present application includes a circuit that encodes a block of an image while limiting a number of times of processing of context adaptive encoding in operation, and a memory connected to the circuit, in the encoding of the block, in both a case where an orthogonal transform is applied to the block and a case where the orthogonal transform is not applied to the block, a sub-block flag encoding process that encodes, by context adaptive encoding, a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient is performed without being included in the number of times of processing.

[0098] Thus, regardless of whether an orthogonal transform is applied and regardless of whether the number of times of processing of context adaptive encoding is limited, a sub-block flag is likely to be encoded by context adaptive encoding. Therefore, the coding amount is likely to be reduced. Further, a difference between a coding mode used in a block to which an orthogonal transform is applied and a coding mode used in a block to which an orthogonal transform is not applied is likely to be small, and the circuit size is likely to be small.

[0099] Further, for example, the circuit further performs, without being included in the number of times of processing, a position parameter encoding process that encodes, by context adaptive encoding, a parameter indicating a position of a first non-zero coefficient in the block in a case where the orthogonal transform is applied to the block.

[0100] Therefore, in the case where the orthogonal transform is applied, regardless of whether the number of times of processing of context adaptive encoding is limited, the parameter indicating the position of the first non-zero coefficient is likely to be encoded by context adaptive encoding. Therefore, the coding amount is likely to be reduced.

[0101] Further, for example, the circuit further determines a range of limitation of the number of times of processing by the position of the first non-zero coefficient in the case where the orthogonal transform is applied to the block.

[0102] Therefore, in the case where the orthogonal transform is applied, it is possible to appropriately determine the number of times of limitation of processing. Therefore, it is possible to appropriately adjust a balance between reduction of the coding amount and reduction of the processing delay.

[0103] Further, for example, a decoding device of an aspect of the present application includes a circuit that decodes a block of an image while limiting a number of times of processing of context-adaptive decoding in operation, and a memory connected to the circuit, and in the decoding of the block, in both a case where inverse orthogonal transform is applied to the block and a case where inverse orthogonal transform is not applied to the block, sub-block flag decoding processing that decodes, by context-adaptive decoding, a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient is performed without being included in the number of times of processing.

[0104] Thus, it is possible to decode the sub-block flag by context-adaptive decoding regardless of whether inverse orthogonal transform is applied or not and whether the number of times of processing of context-adaptive decoding is limited or not. Therefore, it is possible to reduce the amount of encoding. Further, it is possible to make a difference between a decoding method used in a block to which inverse orthogonal transform is applied and a decoding method used in a block to which inverse orthogonal transform is not applied small, and it is possible to make a circuit scale small.

[0105] Further, for example, the circuit further performs, in a case where inverse orthogonal transform is applied to the block, position parameter decoding processing that encodes, by context-adaptive decoding, a parameter indicating a position of a first non-zero coefficient in the block in a scanning order without being included in the number of times of processing.

[0106] Thus, in the case where inverse orthogonal transform is applied, it is possible to decode the parameter indicating the position of the first non-zero coefficient by context-adaptive decoding regardless of whether the number of times of processing of context-adaptive decoding is limited or not. Therefore, it is possible to reduce the amount of encoding.

[0107] Further, for example, the circuit further determines a range of limitation of the number of times of processing by the position of the first non-zero coefficient in a case where inverse orthogonal transform is applied to the block.

[0108] Thus, in the case where inverse orthogonal transform is applied, it is possible to appropriately determine the number of limitation of the number of times of processing. Therefore, it is possible to appropriately adjust a balance between reduction of the amount of encoding and reduction of processing delay.

[0109] Further, for example, an encoding method of an aspect of the present application encodes a block of an image while limiting a number of times of processing of context-adaptive encoding, and in the encoding of the block, in both a case where orthogonal transform is applied to the block and a case where orthogonal transform is not applied to the block, sub-block flag encoding processing that encodes, by context-adaptive encoding, a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient is performed without being included in the number of times of processing.

[0110] Thus, the sub-block flag can be encoded by the context adaptive encoding regardless of whether or not the orthogonal transform is applied and regardless of whether or not the number of times of the context adaptive encoding is limited. Therefore, the amount of encoding can be reduced. Further, the difference between the encoding method used in the block to which the orthogonal transform is applied and the encoding method used in the block to which the orthogonal transform is not applied can be made small, and the circuit size can be made small.

[0111] Further, for example, a decoding method of an aspect of the present application decodes a block of an image while limiting the number of times of context adaptive decoding, and in the decoding of the block, in both a case where inverse orthogonal transform is applied to the block and a case where the inverse orthogonal transform is not applied to the block, a sub-block flag decoding process that decodes a sub-block flag indicating whether or not a sub-block included in the block includes a non-zero coefficient by context adaptive decoding is performed without being included in the number of times.

[0112] Thus, the sub-block flag can be decoded by the context adaptive decoding regardless of whether or not the inverse orthogonal transform is applied and regardless of whether or not the number of times of the context adaptive decoding is limited. Therefore, the amount of encoding can be reduced. Further, the difference between the decoding method used in the block to which the inverse orthogonal transform is applied and the decoding method used in the block to which the inverse orthogonal transform is not applied can be made small, and the circuit size can be made small.

[0113] Further, for example, an encoding device of an aspect of the present application includes a circuit and a memory connected to the circuit, and in operation, in both a case where orthogonal transform is applied to a block of an encoding target image and a case where the orthogonal transform is not applied to the block, in a case where the number of times of context adaptive encoding is within a limit range of the number of times, encodes coefficient information flags indicating properties of coefficients included in the block by the context adaptive encoding, in a case where the number of times is not within the limit range of the number of times, skips encoding of the coefficient information flags, in a case where the coefficient information flags are encoded, encodes residual value information for reconstructing values of the coefficients using the coefficient information flags by Golomb encoding, and in a case where the encoding of the coefficient information flags is skipped, encodes the values of the coefficients by Golomb encoding.

[0114] Thus, the encoding of the coefficient information flags can be skipped in accordance with the limitation of the number of times of the context adaptive encoding regardless of whether or not the orthogonal transform is applied. Therefore, it is possible to suppress an increase in processing delay and an increase in the amount of encoding. Further, the difference between the encoding method used in the block to which the orthogonal transform is applied and the encoding method used in the block to which the orthogonal transform is not applied can be made small, and the circuit size can be made small.

[0115] Further, for example, the coefficient information flag is a flag indicating whether the value of the coefficient is greater than 1.

[0116] Thus, regardless of whether or not the inverse orthogonal transform is applied, it is possible to skip the encoding of the coefficient information flag indicating whether the value of the coefficient is greater than 1, in accordance with the limit on the number of times of processing of context-adaptive encoding. Therefore, it is possible to suppress an increase in processing delay, and to suppress an increase in the amount of encoding.

[0117] Further, for example, a decoding device of an aspect of the present application includes a circuit and a memory connected to the circuit, and in operation, in a case where the inverse orthogonal transform is applied to a block of a decoded image and in a case where the inverse orthogonal transform is not applied to the block, in a case where the number of times of processing of context-adaptive decoding is within a limit on the number of times of processing, the circuit decodes a coefficient information flag indicating a property of a coefficient included in the block by context-adaptive decoding, in a case where the number of times of processing is not within the limit on the number of times of processing, the circuit skips the decoding of the coefficient information flag, in a case where the coefficient information flag is decoded, the circuit decodes residual value information for reconstructing a value of the coefficient using the coefficient information flag by Golomb-Rice decoding, and in a case where the decoding of the coefficient information flag is skipped, the circuit decodes the value of the coefficient by Golomb-Rice decoding.

[0118] Thus, regardless of whether or not the inverse orthogonal transform is applied, it is possible to skip the decoding of the coefficient information flag, in accordance with the limit on the number of times of processing of context-adaptive decoding. Therefore, it is possible to suppress an increase in processing delay, and to suppress an increase in the amount of encoding. Further, it is possible to make the difference between the decoding method used in a block to which the inverse orthogonal transform is applied and the decoding method used in a block to which the inverse orthogonal transform is not applied small, and to make the circuit size small.

[0119] Further, for example, the coefficient information flag is a flag indicating whether the value of the coefficient is greater than 1.

[0120] Thus, regardless of whether or not the inverse orthogonal transform is applied, it is possible to skip the decoding of the coefficient information flag indicating whether the value of the coefficient is greater than 1, in accordance with the limit on the number of times of processing of context-adaptive decoding. Therefore, it is possible to suppress an increase in processing delay, and to suppress an increase in the amount of encoding.

[0121] Further, for example, the coding method of one aspect of the present application is such that, in both a case where orthogonal transform is applied to a block of a coding target image and a case where orthogonal transform is not applied to the block, in a case where the number of times of processing of context adaptive coding is within a limit range of the number of times of processing, a coefficient information flag indicating a property of a coefficient included in the block is coded by context adaptive coding, in a case where the number of times of processing is not within the limit range of the number of times of processing, coding of the coefficient information flag is skipped, in a case where the coefficient information flag is coded, residual value information for reconstructing a value of the coefficient using the coefficient information flag is coded by Golomb coding, and in a case where coding of the coefficient information flag is skipped, the value of the coefficient is coded by Golomb coding.

[0122] Thus, it is possible to skip coding of the coefficient information flag in accordance with the limit of the number of times of processing of context adaptive coding, regardless of whether orthogonal transform is applied or not. Therefore, it is possible to suppress an increase in processing delay and an increase in the amount of coding. Further, it is possible to make a difference between the coding manner used in a block to which orthogonal transform is applied and the coding manner used in a block to which orthogonal transform is not applied small, and to make the circuit scale small.

[0123] Further, for example, the coding method of one aspect of the present application is such that, in both a case where orthogonal transform is applied to a block of a coding target image and a case where orthogonal transform is not applied to the block, in a case where the number of times of processing of context adaptive coding is within a limit range of the number of times of processing, a coefficient information flag indicating a property of a coefficient included in the block is coded by context adaptive coding, in a case where the number of times of processing is not within the limit range of the number of times of processing, coding of the coefficient information flag is skipped, in a case where the coefficient information flag is coded, residual value information for reconstructing a value of the coefficient using the coefficient information flag is coded by Golomb coding, and in a case where coding of the coefficient information flag is skipped, the value of the coefficient is coded by Golomb coding.

[0124] Thus, it is possible to skip coding of the coefficient information flag in accordance with the limit of the number of times of processing of context adaptive coding, regardless of whether orthogonal transform is applied or not. Therefore, it is possible to suppress an increase in processing delay and an increase in the amount of coding. Further, it is possible to make a difference between the coding manner used in a block to which orthogonal transform is applied and the coding manner used in a block to which orthogonal transform is not applied small, and to make the circuit scale small.

[0125] Further, for example, an encoding device of an aspect of the present application includes a circuit and a memory connected to the circuit, and in operation, the circuit encodes a block of an image while limiting the number of times of processing of context adaptive encoding, and in the encoding of the block, in a case where the block is not subjected to orthogonal transform, it is determined whether a plurality of coefficient information flags respectively indicating a plurality of attributes of a coefficient included in the block satisfy a processing condition, and in a case where it is determined that the processing condition is satisfied, the plurality of coefficient information flags are encoded by context adaptive encoding, the processing condition being a condition in which the number of times of processing, in a case where the number of times of processing is added to the number of the plurality of coefficient information flags, is within a range of limitation of the number of times of processing.

[0126] Thus, in a case where orthogonal transform is not applied, it is possible to comprehensively determine whether it is possible to use context adaptive encoding for a plurality of coefficient information flags. Therefore, it is possible to simplify processing, and reduce processing delay. Further, in a case where similar processing is performed on a block to which orthogonal transform is applied, it is possible to reduce the difference between the encoding method used in the block to which orthogonal transform is applied and the encoding method used in the block to which orthogonal transform is not applied, and it is possible to reduce the circuit size.

[0127] Further, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0128] Thus, it is possible to comprehensively determine a plurality of coefficient information flags including a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Therefore, it is possible to simplify processing, and reduce processing delay.

[0129] Further, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0130] Thus, it is possible to comprehensively determine a plurality of coefficient information flags including a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Therefore, it is possible to simplify processing, and reduce processing delay.

[0131] Further, for example, a decoding device of an aspect of the present application includes a circuit and a memory connected to the circuit, and the circuit, in operation, limits a number of times of processing of context-adaptive decoding to decode a block of an image, and in the decoding of the block, in a case where inverse orthogonal transform is not applied to the block, determines whether a plurality of coefficient information flags each indicating a plurality of attributes of a coefficient included in the block satisfy a processing condition, and in a case where it is determined that the processing condition is satisfied, decodes the plurality of coefficient information flags by context-adaptive decoding, the processing condition being a condition in which the number of times of processing, in a case where the number of times of processing is added to the number of the plurality of coefficient information flags, is within a range of limitation of the number of times of processing.

[0132] Thus, in a case where inverse orthogonal transform is not applied, it is possible to generally determine whether context-adaptive decoding can be used for a plurality of coefficient information flags. Therefore, processing can be simplified, and processing delay can be reduced. Further, in a case where similar processing is performed on a block to which inverse orthogonal transform is applied, a difference between a decoding method used in a block to which inverse orthogonal transform is applied and a decoding method used in a block to which inverse orthogonal transform is not applied can be small, and a circuit size can be small.

[0133] Further, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether a value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0134] Thus, it is possible to generally determine a plurality of coefficient information flags including a coefficient information flag indicating whether a value of a coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Therefore, processing can be simplified, and processing delay can be reduced.

[0135] Further, for example, the plurality of coefficient information flags include a coefficient information flag indicating whether a value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5.

[0136] Thus, it is possible to generally determine a plurality of coefficient information flags including a coefficient information flag indicating whether a value of a coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Therefore, processing can be simplified, and processing delay can be reduced.

[0137] In addition, for example, one form of the encoding method of the present invention is to limit the number of context adaptive coding processing times and encode blocks of an image. In the encoding of the block, when no orthogonal transform is applied to the block, it is determined whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the block meet the processing conditions. When it is determined that the processing conditions are met, the multiple coefficient information flags are encoded by context adaptive coding. The processing condition is the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0138] Thus, when an orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive coding can be used for multiple coefficient information flags. This simplifies processing and reduces processing delay. Furthermore, when similar processing is performed on blocks to which an orthogonal transform is applied, the difference between the coding method used in blocks to which an orthogonal transform is applied and the coding method used in blocks not to which an orthogonal transform is applied can be reduced, and the circuit scale can be reduced.

[0139] In addition, for example, one form of a decoding method of the present invention is to decode a block of an image while limiting the number of processing times of context adaptive decoding. In the decoding of the block, without applying an inverse orthogonal transform to the block, it is determined whether multiple coefficient information flags respectively representing multiple attributes of the coefficients contained in the block meet a processing condition. If it is determined that the processing condition is met, the multiple coefficient information flags are decoded by context adaptive decoding, and the processing condition is the following condition: the number of processing times when the number of the multiple coefficient information flags is added to the number of processing times is within the limited range of the processing times.

[0140] Thus, even when an inverse orthogonal transform is not applied, it is possible to collectively determine whether context-adaptive decoding can be used for multiple coefficient information flags. Consequently, processing can be simplified and processing delay can be reduced. Furthermore, when similar processing is performed on blocks to which an inverse orthogonal transform is applied, the difference between the decoding method used in blocks to which an inverse orthogonal transform is applied and the decoding method used in blocks not to which an inverse orthogonal transform is applied can be minimized, and the circuit scale can be reduced.

[0141] Furthermore, for example, an encoding device according to one aspect of the present invention includes a partitioning unit, an intra-frame prediction unit, an inter-frame prediction unit, a prediction control unit, a transform unit, a quantization unit, an entropy encoding unit, and a loop filter unit.

[0142] The division section divides an encoding target picture constituting the moving picture into a plurality of blocks. The intra prediction section performs intra prediction that generates the prediction picture of the encoding target block in the encoding target picture using a reference picture in the encoding target picture. The inter prediction section performs inter prediction that generates the prediction picture of the encoding target block using a reference picture in a reference picture different from the encoding target picture.

[0143] The prediction control section controls the intra prediction performed by the intra prediction section and the inter prediction performed by the inter prediction section. The transform section transforms a prediction residual signal between the prediction picture generated by the intra prediction section or the inter prediction section and the picture of the encoding target block, and generates a transform coefficient signal of the encoding target block. The quantization section quantizes the transform coefficient signal. The entropy coding section codes the quantized transform coefficient signal. The loop filter section applies a filter to the encoding target block.

[0144] Further, for example, the entropy coding section codes a block of a picture in an operation of limiting the number of times of processing of context adaptive coding, in encoding of the block, in both a case where an orthogonal transform is applied to the block and a case where the orthogonal transform is not applied to the block, performs sub-block flag coding processing of coding, by context adaptive coding, a sub-block flag indicating whether or not a sub-block included in the block includes a non-zero coefficient, without including the sub-block flag coding processing in the number of times of processing.

[0145] Further, for example, the entropy coding section, in an operation, in both a case where an orthogonal transform is applied to a block of an encoding target picture and a case where the orthogonal transform is not applied to the block, in a case where the number of times of processing of context adaptive coding is within a limit range of the number of times of processing, codes, by context adaptive coding, a coefficient information flag indicating a property of a coefficient included in the block, in a case where the number of times of processing is not within the limit range of the number of times of processing, skips coding of the coefficient information flag, in a case where the coefficient information flag is coded, codes, by Golomb coding, residual value information for reconstructing a value of the coefficient using the coefficient information flag, and in a case where coding of the coefficient information flag is skipped, codes, by Golomb coding, the value of the coefficient.

[0146] Further, for example, the entropy decoding section, in operation, decodes a block of an image while limiting the number of times of processing of context-adaptive decoding, and in decoding the block, in a case where inverse orthogonal transform is not applied to the block, determines whether a plurality of coefficient information flags each indicating a plurality of properties of a coefficient included in the block satisfy a processing condition, encodes the plurality of coefficient information flags by context-adaptive decoding in a case where it is determined that the processing condition is satisfied, the processing condition being a condition under which the number of times of processing, to which the number of the plurality of coefficient information flags is added, is within a limit range of the number of times of processing.

[0147] Further, for example, a decoding device of an aspect of the present application decodes a moving image using a prediction image, the decoding device including an entropy decoding section, an inverse quantization section, an inverse transform section, an intra prediction section, an inter prediction section, a prediction control section, an addition section (reconstruction section), and a loop filter section.

[0148] The entropy decoding section decodes a quantized transform coefficient signal of a decoding target block in a decoding target picture constituting the moving image. The inverse quantization section inverse quantizes the quantized transform coefficient signal. The inverse transform section inverse transforms the transform coefficient signal to obtain a prediction residual signal of the decoding target block.

[0149] The intra prediction section performs intra prediction using a reference image in the decoding target picture to generate the prediction image of the decoding target block. The inter prediction section performs inter prediction using a reference image in a reference picture different from the decoding target picture to generate the prediction image of the decoding target block. The prediction control section controls the intra prediction performed by the intra prediction section and the inter prediction performed by the inter prediction section.

[0150] The addition section reconstructs an image of the decoding target block by adding the prediction image generated by the intra prediction section or the inter prediction section to the prediction residual signal. The loop filter section applies a filter to the decoding target block.

[0151] Further, for example, the entropy decoding section, in operation, decodes a block of an image while limiting the number of times of processing of context-adaptive decoding, and in decoding the block, in a case where inverse orthogonal transform is not applied to the block, determines whether a plurality of coefficient information flags each indicating a plurality of properties of a coefficient included in the block satisfy a processing condition, encodes the plurality of coefficient information flags by context-adaptive decoding in a case where it is determined that the processing condition is satisfied, the processing condition being a condition under which the number of times of processing, to which the number of the plurality of coefficient information flags is added, is within a limit range of the number of times of processing.

[0152] Further, for example, the entropy decoding section, in operation, decodes, by context adaptive decoding, a coefficient information flag indicating a property of a coefficient included in a block of a decoded image, in a case where a number of times of processing of the context adaptive decoding is within a limit range of the number of times of processing in both a case where inverse orthogonal transform is applied to the block and a case where the inverse orthogonal transform is not applied to the block, skips decoding of the coefficient information flag, in a case where the number of times of processing is not within the limit range of the number of times of processing, decodes, by Golomb-Rice decoding, residual value information for reconstructing a value of the coefficient using the coefficient information flag, in a case where the coefficient information flag is decoded, and decodes, by Golomb-Rice decoding, the value of the coefficient, in a case where decoding of the coefficient information flag is skipped.

[0153] Further, for example, the entropy decoding section, in operation, decodes, by context adaptive decoding, a coefficient information flag indicating a property of a coefficient included in a block of a decoded image, in a case where a number of times of processing of the context adaptive decoding is within a limit range of the number of times of processing in both a case where inverse orthogonal transform is applied to the block and a case where the inverse orthogonal transform is not applied to the block, skips decoding of the coefficient information flag, in a case where the number of times of processing is not within the limit range of the number of times of processing, decodes, by Golomb-Rice decoding, residual value information for reconstructing a value of the coefficient using the coefficient information flag, in a case where the coefficient information flag is decoded, and decodes, by Golomb-Rice decoding, the value of the coefficient, in a case where decoding of the coefficient information flag is skipped.

[0154] Moreover, these inclusive or specific modes can be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a CD-ROM that is readable by a computer, and can be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0155] Hereinafter, the embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below indicate inclusive or specific examples. The numerical values, shapes, materials, component elements, arrangement positions of component elements, connection modes, steps, relationships and orders of steps, and the like indicated in the embodiments below are examples, and are not intended to limit the claims.

[0156] Hereinafter, embodiments of an encoding apparatus and a decoding apparatus will be described. The embodiments are examples of an encoding apparatus and a decoding apparatus capable of applying the processes and / or structures described in each mode of the present application. The processes and / or structures can also be implemented in an encoding apparatus and a decoding apparatus different from the embodiments. For example, with respect to the processes and / or structures applied to the embodiments, for example, any one of the following can be performed.

[0157] (1) The configuration elements of the encoding apparatus or the decoding apparatus of the embodiments described in each aspect of the present application can be replaced with other configuration elements described in any of the aspects of the present application, or combined with them.

[0158] (2) In the encoding apparatus or the decoding apparatus of the embodiments, the functions or processes performed by some of the configuration elements of the encoding apparatus or the decoding apparatus can be arbitrarily changed by addition, substitution, deletion, or the like of the functions or processes. For example, any of the functions or processes can be substituted with other functions or processes described in any of the aspects of the present application, or combined with them.

[0159] (3) In the method performed by the encoding apparatus or the decoding apparatus of the embodiments, some of the processes included in the method can be arbitrarily changed by addition, substitution, deletion, or the like. For example, any of the processes in the method can be substituted with other processes described in any of the aspects of the present application, or combined with them.

[0160] (4) Some of the configuration elements constituting the encoding apparatus or the decoding apparatus of the embodiments can be combined with the configuration elements described in any of the aspects of the present application, combined with the configuration elements having some of the functions described in any of the aspects of the present application, or combined with the configuration elements performing some of the processes performed by the configuration elements described in any of the aspects of the present application.

[0161] (5) The configuration elements having some of the functions of the encoding apparatus or the decoding apparatus of the embodiments, or the configuration elements performing some of the processes of the encoding apparatus or the decoding apparatus of the embodiments, can be combined with or substituted with the configuration elements described in any of the aspects of the present application, the configuration elements having some of the functions described in any of the aspects of the present application, or the configuration elements performing some of the processes described in any of the aspects of the present application.

[0162] (6) In the method performed by the encoding apparatus or the decoding apparatus of the embodiments, some of the processes included in the method can be substituted with the processes described in any of the aspects of the present application or the same processes, or combined with them.

[0163] (7) Some of the processes included in the method performed by the encoding apparatus or the decoding apparatus of the embodiments can be combined with the processes described in any of the aspects of the present application.

[0164] (8) The implementation of the processes and / or structures described in the aspects of the present application is not limited to the encoding apparatus or the decoding apparatus of the embodiments. For example, the processes and / or structures can also be implemented in an apparatus used for a purpose different from the motion picture encoding or the motion picture decoding disclosed in the embodiments.

[0165] [Encoding apparatus]

[0166] First, an encoding apparatus relating to the embodiments will be described. Figure 1 is a block diagram showing a functional structure of an encoding apparatus 100 relating to the embodiments. The encoding apparatus 100 is a motion picture encoding apparatus that encodes a motion picture in a block unit.

[0167] As shown in Figure 1 , the encoding apparatus 100 is an apparatus that encodes an image in a block unit, and includes a division section 102, a subtraction section 104, a transform section 106, a quantization section 108, an entropy encoding section 110, an inverse quantization section 112, an inverse transform section 114, an addition section 116, a block memory 118, a loop filter section 120, a frame memory 122, an intra prediction section 124, an inter prediction section 126, and a prediction control section 128.

[0168] The encoding apparatus 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program saved in the memory is executed by the processor, the processor functions as the division section 102, the subtraction section 104, the transform section 106, the quantization section 108, the entropy encoding section 110, the inverse quantization section 112, the inverse transform section 114, the addition section 116, the loop filter section 120, the intra prediction section 124, the inter prediction section 126, and the prediction control section 128. Alternatively, the encoding apparatus 100 can be realized as one or more electronic circuits dedicated to the division section 102, the subtraction section 104, the transform section 106, the quantization section 108, the entropy encoding section 110, the inverse quantization section 112, the inverse transform section 114, the addition section 116, the loop filter section 120, the intra prediction section 124, the inter prediction section 126, and the prediction control section 128.

[0169] Hereinafter, after the flow of the overall process of the encoding apparatus 100 is described, each constituent element included in the encoding apparatus 100 will be described.

[0170] [Overall flow of encoding process]

[0171] Figure 2 is a flowchart showing an example of the overall encoding process performed by the encoding apparatus 100.

[0172] First, the division section 102 of the encoding apparatus 100 divides each picture included in an input image, which is a moving image, into a plurality of fixed-size blocks (for example, 128 x 128 pixels) (step Sa_l). Then, the division section 102 selects a division pattern (also referred to as a block shape) for the fixed-size blocks (step Sa_2). That is, the division section 102 further divides the blocks having the fixed size into a plurality of blocks constituting the selected division pattern. Then, the encoding apparatus 100 performs the processing of steps Sa_3 to Sa_9 on each of the plurality of blocks, for the block (i.e., an encoding target block).

[0173] That is, a prediction processing section constituted by all or a part of the intra prediction section 124, the inter prediction section 126, and the prediction control section 128 generates a prediction signal (also referred to as a prediction block) of an encoding target block (also referred to as a current block) (step Sa_3).

[0174] Next, the subtraction section 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).

[0175] Next, the transform section 106 and the quantization section 108 generate a plurality of quantized coefficients by performing transform and quantization on the difference block (step Sa_5). Further, a block constituted by the plurality of quantized coefficients is also referred to as a coefficient block.

[0176] Next, the entropy coding section 110 generates an encoded signal by performing encoding (specifically, entropy coding) on the coefficient block and a prediction parameter related to the generation of the prediction signal (step Sa_6). In addition, the encoded signal is also referred to as an encoded bitstream, a compressed bitstream, or a stream.

[0177] Next, the inverse quantization section 112 and the inverse transform section 114 reproduce a plurality of prediction residuals (i.e., difference blocks) by performing inverse quantization and inverse transform on the coefficient block (step Sa_7).

[0178] Next, the addition section 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the prediction block to the reproduced difference block (step Sa_8). Thereby, a reconstructed image is generated.

[0179] When the reconstructed image is generated, the loop filter section 120 filters the reconstructed image as necessary (step Sa_9).

[0180] Then, the encoding apparatus 100 determines whether or not the encoding of the entire picture has been completed (step Sa_lO), and in the case where it is determined that the encoding has not been completed (NO in step Sa_lO), the processing from step Sa_2 is repeated.

[0181] In addition, in the above example, the encoding device 100 selects one partitioning pattern for the fixed-size blocks and performs encoding of each block according to the partitioning pattern, but can perform encoding of each block according to each of a plurality of partitioning patterns. In this case, the encoding device 100 can evaluate the cost for each of the plurality of partitioning patterns, and can select, for example, an encoded signal obtained by encoding according to the partitioning pattern of the smallest cost as the output encoded signal.

[0182] As illustrated, the processes of these steps Sa_1 to Sa_10 are sequentially performed by the encoding device 100. Alternatively, a part of the plurality of processes among these processes can be performed in parallel, or the order of these processes can be changed.

[0183] [Partitioning Section]

[0184] The partitioning section 102 partitions each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction section 104. For example, the partitioning section 102 first partitions a picture into blocks of a fixed size (for example, 128 x 128). Other fixed block sizes can also be used. The fixed-size blocks are referred to as coding tree units (CTUs). Further, the partitioning section 102 partitions each of the fixed-size blocks into blocks of variable sizes (for example, 64 x 64 or less) based on, for example, a recursive quadtree and / or binary tree block partitioning. That is, the partitioning section 102 selects a partitioning pattern. The variable-size blocks are referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). In various processing examples, the CUs, PUs, and TUs do not need to be distinguished, and a part or all of the blocks within a picture can be used as the processing units of the CUs, PUs, and TUs.

[0185] Figure 3 is a conceptual diagram illustrating an example of block partitioning according to an embodiment. In Figure 3 In, a solid line indicates a block boundary based on quadtree block partitioning, and a dashed line indicates a block boundary based on binary tree block partitioning.

[0186] Here, the block 10 is a square block of 128 x 128 pixels (128 x 128 block). The 128 x 128 block 10 is first partitioned into four square 64 x 64 blocks (quadtree block partitioning).

[0187] The upper left 64 x 64 block is further vertically partitioned into two rectangular 32 x 64 blocks, and the left 32 x 64 block is further vertically partitioned into two rectangular 16 x 64 blocks (binary tree block partitioning). As a result, the upper left 64 x 64 block is partitioned into two 16 x 64 blocks 11 and 12 and a 32 x 64 block 13.

[0188] The 64x64 block on the upper right is horizontally split into 2 rectangular 64x32 blocks 14, 15 (binary tree block split).

[0189] The 64x64 block on the lower left is split into 4 square 32x32 blocks (quad tree block split). The upper left block and the lower right block among the 4 32x32 blocks are further split. The upper left 32x32 block is vertically split into 2 rectangular 16x32 blocks, and the right 16x32 block is horizontally split into 2 16x16 blocks (binary tree block split). The lower right 32x32 block is horizontally split into 2 32x16 blocks (binary tree block split). As a result, the 64x64 block on the lower left is split into the 16x32 block 16, 2 16x16 blocks 17, 18, 2 32x32 blocks 19, 20, and 2 32x16 blocks 21, 22.

[0190] The 64x64 block 23 on the lower right is not split.

[0191] As above, in the case of Figure 3 , the block 10 is split into 13 blocks 11 to 23 of variable sizes based on the recursive quad tree and binary tree block split. Such split is referred to as QTBT (quad-tree plus binary tree) split.

[0192] In addition, in the case of Figure 3 , 1 block is split into 4 or 2 blocks (quad tree or binary tree block split), but the split is not limited to these. For example, 1 block can be split into 3 blocks (ternary tree split). The split including such ternary tree split is referred to as MBT (multi type tree) split.

[0193] [Structure of picture slice / tile]

[0194] In order to decode a picture in parallel, a picture is sometimes configured in a slice unit or in a tile unit. A picture configured in a slice unit or a tile unit can be configured by the split section 102.

[0195] A slice is a basic unit of encoding that configures a picture. A picture is configured by 1 or more slices, for example. In addition, a slice is configured by 1 or more consecutive CTUs (Coding Tree Unit).

[0196] Figure 4Ais a conceptual diagram showing an example of the structure of a slice. For example, a picture includes 11 x 8 CTUs, and is divided into 4 slices (slices 1 to 4). Slice 1 is constituted of 16 CTUs, slice 2 is constituted of 21 CTUs, slice 3 is constituted of 29 CTUs, and slice 4 is constituted of 22 CTUs. Here, each CTU within the picture belongs to any one of the slices. The shape of the slice becomes a shape in which the picture is divided in the horizontal direction. The boundary of the slice need not be the picture end, and can be any position in the boundary of the CTU within the picture. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. Further, the slice contains header information and encoded data. In the header information, the CTU address of the beginning of the slice, the slice type, and the like, which are characteristics of the slice, can also be described.

[0197] A tile is a unit of a rectangular region constituting a picture. Each tile can also be assigned a number called Tileld in the raster scan order.

[0198] Figure 4B is a conceptual diagram showing an example of the structure of a tile. For example, a picture includes 11 x 8 CTUs, and is divided into 4 rectangular region tiles (tiles 1 to 4). In the case of using tiles, the processing order of the CTUs is changed compared to the case of not using tiles. In the case of not using tiles, a plurality of CTUs within the picture are processed in the raster scan order. In the case of using tiles, in each of a plurality of tiles, at least one CTU is processed in the raster scan order. For example, as shown in Figure 4B , the processing order of the plurality of CTUs included in tile 1 is the order from the left end of the 1st row of tile 1 to the right end of the 1st row of tile 1, and then from the left end of the 2nd row of tile 1 to the right end of the 2nd row of tile 1.

[0199] In addition, one tile sometimes contains one or more slices, and one slice sometimes contains one or more tiles.

[0200] [Subtracting section]

[0201] The subtracting section 104 subtracts a prediction signal (a prediction sample input from the prediction control section 128 shown below) from an original signal (an original sample) in the unit of a block input from and divided by the dividing section 102. That is, the subtracting section 104 calculates a prediction error (also called a residual) of an encoding target block (hereinafter referred to as a current block). Further, the subtracting section 104 outputs the calculated prediction error (residual) to the transforming section 106.

[0202] The original signal is an input signal of the encoding apparatus 100, and is a signal (for example, a luma signal and two chroma signals) showing an image of each picture constituting a moving image. Hereinafter, there is also a case where the signal showing the image is called a sample.

[0203] [Transform unit]

[0204] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106, for example, performs a prescribed discrete cosine transform (DCT) or a discrete sine transform (DST) on the prediction error in the spatial domain. The prescribed DCT or DST can also be determined in advance.

[0205] In addition, the transform unit 106 can also adaptively select a transform type from among a plurality of transform types, and transform the prediction error into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is referred to as an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).

[0206] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table indicating transform basis functions corresponding to the transform types. In Figure 5A N indicates the number of input pixels. The selection of the transform type from among these plurality of transform types can depend, for example, on the type of prediction (intra prediction and inter prediction), or on the intra prediction mode.

[0207] Information indicating whether to apply such an EMT or AMT (for example, referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. In addition, the signaling of these pieces of information does not need to be limited to the CU level, and can be another level (for example, a bit sequence level, a picture level, a slice level, a tile level, or a CTU level).

[0208] Further, the transform unit 106 can also perform a re-transformation on the transform coefficients (transform result). Such a re-transformation has a case called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (for example, 4 x 4 sub-block) included in a block of transform coefficients corresponding to the intra prediction error. Information indicating whether or not to apply NSST and information related to a transform matrix used in NSST are generally signaled at the CU level. In addition, the signaling of these pieces of information does not need to be limited to the CU level, and can be another level (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0209] In the transform unit 106, a Separable transform and a Non-Separable transform can be applied. The Separable transform refers to a manner in which a plurality of transforms are performed in each direction by separating the dimensions of the input. The Non-Separable transform refers to a manner in which, when the input is multi-dimensional, two or more dimensions are regarded as one dimension and a transform is performed on the dimensions together.

[0210] For example, as an example of the Non-Separable transform, there is a manner in which, when the input is a 4 x 4 block, the block is regarded as one arrangement having 16 elements, and a transform process is performed on the arrangement with a 16 x 16 transform matrix.

[0211] Further, in a further example of the Non-Separable transform, a Hypercube Givens Transform in which Givens rotation is performed on an arrangement having 16 elements after the 4 x 4 input block is regarded as the arrangement can be performed.

[0212] In the transform in the transform unit 106, the type of basis to be transformed into a frequency region can be switched according to the region within the CU. As an example, there is an SVT (Spatially Varying Transform). In the SVT, as shown in FIG. 6, the type of basis to be transformed into a frequency region is switched according to the region within the CU. Figure 5BAs shown, the CU is bisected in the horizontal or vertical direction, and the transform to the frequency region is performed only on the region of either side. The type of transform basis can be set for each region, such as using DST7 and DCT8. In this example, the transform is performed only on one of the two regions within the CU, and the other is not transformed, but both of the two regions can also be transformed. In addition, the division method is not only bisected, but can also be more flexible, such as quartered or the information indicating the division is additionally encoded, and is signaled in the same manner as the CU division, and the like. In addition, the SVT is sometimes referred to as SBT (Sub-block Transform).

[0213] [Quantization section]

[0214] The quantization section 108 quantizes the transform coefficients output from the transform section 106. Specifically, the quantization section 108 scans the transform coefficients of the current block in a prescribed scan order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Furthermore, the quantization section 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding section 110 and the inverse quantization section 112. The prescribed scan order can also be determined in advance.

[0215] The prescribed scan order is the order for the quantization / inverse quantization of the transform coefficients. For example, the prescribed scan order can be defined in the ascending order of the frequency (order from low frequency to high frequency) or the descending order of the frequency (order from high frequency to low frequency).

[0216] The quantization parameter (QP) refers to a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0217] In addition, in the quantization, a quantization matrix is sometimes used. For example, a plurality of quantization matrices are sometimes used in correspondence with the frequency transform size such as 4 x 4 and 8 x 8, the prediction mode such as intra prediction and inter prediction, the pixel component such as luminance and color difference, and the like. In addition, quantization refers to digitizing the values sampled at a prescribed interval in correspondence with a prescribed level, and in this technical field, other expressions such as rounding, rounding, scaling can also be used for reference, and rounding, rounding, scaling can also be adopted. The prescribed interval and the level can also be determined in advance.

[0218] As a method of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device side and a method of using a default quantization matrix (default matrix). On the encoding device side, by directly setting the quantization matrix, it is possible to set a quantization matrix corresponding to the characteristics of the image. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix.

[0219] On the other hand, there is also a method of quantizing in a manner that the coefficients of high frequency components and the coefficients of low frequency components are all the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix in which all the coefficients are the same value (flat matrix).

[0220] The quantization matrix can be specified by, for example, an SPS (Sequence Parameter Set) or a PPS (Picture Parameter Set). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and the PPS are sometimes referred to simply as parameter sets.

[0221] [Entropy encoding section]

[0222] The entropy encoding section 110 generates an encoded signal (encoded bit stream) based on the quantized coefficients input from the quantization section 108. Specifically, the entropy encoding section 110, for example, binarizes the quantized coefficients, arithmetically encodes the binarized signal, and outputs a compressed bit stream or sequence.

[0223] [Inverse quantization section]

[0224] The inverse quantization section 112 inverse quantizes the quantized coefficients input from the quantization section 108. Specifically, the inverse quantization section 112 inverse quantizes the quantized coefficients of the current block in a prescribed scan order. Also, the inverse quantization section 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform section 114. The prescribed scan order can also be determined in advance.

[0225] [Inverse transform section]

[0226] The inverse transform section 114 restores the prediction error (residual) by inverse transforming the transform coefficients input from the inverse quantization section 112. Specifically, the inverse transform section 114 restores the prediction error of the current block by inverse transforming the transform coefficients corresponding to the transform of the transform section 106. Also, the inverse transform section 114 outputs the restored prediction error to the addition section 116.

[0227] In addition, the restored prediction error is generally not consistent with the prediction error calculated by the subtraction section 104 because information is lost by quantization. That is, the restored prediction error generally includes a quantization error.

[0228] [Addition section]

[0229] The addition section 116 reconstructs the current block by adding the prediction error input from the inverse transform section 114 to the prediction sample input from the prediction control section 128. Also, the addition section 116 outputs the reconstructed block to the block memory 118 and the loop filter section 120. The reconstructed block is sometimes referred to as a local decoded block.

[0230] [Block memory]

[0231] The block memory 118 is, for example, a storage section for storing a block within an encoding target picture (referred to as a current picture) referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition section 116.

[0232] [Frame memory]

[0233] The frame memory 122 is, for example, a storage section for storing a reference picture used in inter prediction, and is also referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter 120.

[0234] [Loop filter]

[0235] The loop filter 120 applies loop filtering to the block reconstructed by the addition section 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within an encoding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), adaptive loop filtering (ALF), and the like.

[0236] In ALF, a least square error filter for removing encoding distortion is adopted, and, for example, one filter selected from a plurality of filters based on a direction and activity of a gradient in each 2x2 sub-block within a current block is adopted.

[0237] Specifically, first, a sub-block (for example, a 2x2 sub-block) is classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-block is performed based on a direction and activity of a gradient. For example, using a direction value D (for example, 0-2 or 0-4) of a gradient and an activity value A (for example, 0-4) of a gradient, a classification value C (for example, C=5D+A) is calculated. And, based on the classification value C, the sub-block is classified into a plurality of classes.

[0238] The direction value D of the gradient is derived, for example, by comparing gradients of a plurality of directions (for example, horizontal, vertical, and 2 diagonal directions). Further, the activity value A of the gradient is derived, for example, by adding gradients of a plurality of directions and quantizing the addition result.

[0239] Based on the result of such classification, a filter for the sub-block is decided from among a plurality of filters.

[0240] As a shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figure 6A-6C is a diagram showing a plurality of examples of a shape of a filter used in ALF. Figure 6Arepresents a 5x5 diamond-shaped filter, Figure 6B represents a 7x7 diamond-shaped filter, Figure 6C represents a 9x9 diamond-shaped filter. The information indicating the shape of the filter is usually signaled at the picture level. In addition, the signaling of the information indicating the shape of the filter need not be limited to the picture level, but can be another level (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0241] The on / off of the ALF can be determined, for example, at the picture level or the CU level. For example, as for the luminance, whether to employ the ALF can be determined at the CU level, and as for the chrominance, whether to employ the ALF can be determined at the picture level. The information indicating the on / off of the ALF is usually signaled at the picture level or the CU level. In addition, the signaling of the information indicating the on / off of the ALF need not be limited to the picture level or the CU level, but can be another level (e.g., sequence level, slice level, tile level, or CTU level).

[0242] The coefficient set of the selectable multiple filters (e.g., up to 15 or 25 filters) is usually signaled at the picture level. In addition, the signaling of the coefficient set need not be limited to the picture level, but can be another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0243] [Loop filter > deblocking filter]

[0244] In the deblocking filter, the loop filter 120 reduces distortion generated at the block boundary by performing a filter process on the block boundary of the reconstructed image.

[0245] Figure 7 is a block diagram indicating an example of the detailed structure of the loop filter 120 functioning as a deblocking filter.

[0246] The loop filter 120 includes a boundary determination section 1201, a filter determination section 1203, a filter processing section 1205, a process determination section 1208, a filter characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0247] The boundary determination section 1201 determines whether there is a pixel (i.e., an object pixel) subjected to the deblocking filter process in the vicinity of the block boundary. Then, the boundary determination section 1201 outputs the determination result to the switch 1202 and the process determination section 1208.

[0248] In a case where the boundary determination section 1201 determines that the object pixel exists in the vicinity of the block boundary, the switch 1202 outputs the image before the filter processing to the switch 1204. In contrast, in a case where the boundary determination section 1201 determines that the object pixel does not exist in the vicinity of the block boundary, the switch 1202 outputs the image before the filter processing to the switch 1206.

[0249] The filter determination section 1203 determines whether or not to perform the deblocking filter processing on the object pixel on the basis of the pixel value of at least one of the surrounding pixels located in the vicinity of the object pixel. Then, the filter determination section 1203 outputs the determination result to the switch 1204 and the processing determination section 1208.

[0250] In a case where the filter determination section 1203 determines to perform the deblocking filter processing on the object pixel, the switch 1204 outputs the image before the filter processing, which is acquired via the switch 1202, to the filter processing section 1205. In contrast, in a case where the filter determination section 1203 determines not to perform the deblocking filter processing on the object pixel, the switch 1204 outputs the image before the filter processing, which is acquired via the switch 1202, to the switch 1206.

[0251] In a case where the image before the filter processing is acquired via the switches 1202 and 1204, the filter processing section 1205 performs the deblocking filter processing on the object pixel with the filter characteristic decided by the filter characteristic decision section 1207. Then, the filter processing section 1205 outputs the pixel after the filter processing to the switch 1206.

[0252] The switch 1206 selectively outputs the pixel which is not subjected to the deblocking filter processing and the pixel which is subjected to the deblocking filter processing by the filter processing section 1205, in accordance with the control of the processing determination section 1208.

[0253] The processing determination section 1208 controls the switch 1206 on the basis of the respective determination results of the boundary determination section 1201 and the filter determination section 1203. That is, the processing determination section 1208 outputs the pixel after the deblocking filter processing from the switch 1206 in a case where the boundary determination section 1201 determines that the object pixel exists in the vicinity of the block boundary and the filter determination section 1203 determines to perform the deblocking filter processing on the object pixel. In addition, the processing determination section 1208 outputs the pixel which is not subjected to the deblocking / filter processing from the switch 1206 in a case other than the above-described case. The image after the filter processing is output from the switch 1206 by repeating the output of the pixel in this way.

[0254] Figure 8 is a conceptual diagram showing an example of the deblocking filter having a filter characteristic symmetrical with respect to the block boundary.

[0255] In the deblocking filter processing, for example, using a pixel value and a quantization parameter, either of two deblocking filters, that is, a strong filter and a weak filter, which are different in characteristics, is selected. In the strong filter, as shown in Figure 8

[0256] q'0 = (p1 + 2 x p0 + 2 x q0 + 2 x q1 + q2 + 4) / 8

[0257] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0258] q'2 = (p0 + q0 + q1 + 3 x q2 + 2 x q3 + 4) / 8

[0259] Further, in the above formulas, p0 to p2 and q0 to q2 are pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is a pixel value of a pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. Further, on the right side of each of the above formulas, a coefficient multiplied by a pixel value of each pixel used in the deblocking filter processing is a filter coefficient.

[0260] Further, in the deblocking filter processing, clipping processing can also be performed in a manner that an operated pixel value does not exceed a threshold value set. In this clipping processing, using a threshold value decided in accordance with the quantization parameter, the operated pixel value based on the above formula is clipped to "operated pixel value ± 2 x threshold value". Thereby, excessive smoothing can be prevented.

[0261] Figure 9 is a conceptual diagram for explaining a block boundary in which the deblocking filter processing is performed. Figure 10 is a conceptual diagram showing an example of a Bs value.

[0262] A block boundary in which the deblocking filter processing is performed is, for example, a boundary of a PU (Prediction Unit) or a TU (Transform Unit) of an 8 x 8 pixel block as shown in Figure 9 The deblocking filter processing can be performed in units of 4 rows or 4 columns. First, for the block P and the block Q as shown in Figure 9 Bs (Boundary Strength) values are decided as shown in Figure 10

[0263] According to Figure 10 ​​Bs value, it is determined whether or not to perform the de-blocking filter processing of different strength even if the block boundary belongs to the same picture. In the case where the Bs value is 2, the de-blocking filter processing for the color difference signal is performed. In the case where the Bs value is 1 or more and a prescribed condition is satisfied, the de-blocking filter processing for the luminance signal is performed. The prescribed condition can be determined in advance. In addition, the determination condition of the Bs value is not limited to Figure 10 the condition shown in the table, it can be determined based on other parameters.

[0264] [Prediction processing section (intra prediction section / inter prediction section / prediction control section)]

[0265] Figure 11 is a flowchart showing an example of the processing performed by the prediction processing section of the encoding apparatus 100. In addition, the prediction processing section is constituted by all or a part of the constituent elements of the intra prediction section 124, the inter prediction section 126, and the prediction control section 128.

[0266] The prediction processing section generates a prediction image of the current block (step Sb_1). The prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there are an intra prediction signal or an inter prediction signal. Specifically, the prediction processing section generates the prediction image of the current block using a reconstructed image that has been obtained by performing the generation of the prediction block, the generation of the difference block, the generation of the coefficient block, the restoration of the difference block, and the generation of the decoded image block.

[0267] The reconstructed image can be, for example, an image of a reference picture, or an image of a picture that contains the coded block in the current picture, i.e., a current picture. The coded block in the current picture is, for example, a neighboring block of the current block.

[0268] Figure 12 is a flowchart showing another example of the processing performed by the prediction processing section of the encoding apparatus 100.

[0269] The prediction processing section generates a prediction image by a first method (step Sc_1a), generates a prediction image by a second method (step Sc_1b), and generates a prediction image by a third method (step Sc_1c). The first method, the second method, and the third method are mutually different methods for generating a prediction image, and each can be, for example, an inter prediction method, an intra prediction method, and another prediction method. In such a prediction method, the reconstructed image described above can be used.

[0270] Next, the prediction processing section selects any one of the plurality of prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 2). The selection of the prediction image, that is, the selection of the manner or mode for obtaining the final prediction image can also be performed based on the cost calculated for each of the generated prediction images. Furthermore, the selection of the prediction image can be performed based on the parameters for the encoding processing. The encoding apparatus 100 can signal information for determining the selected prediction image, manner, or mode as an encoded signal (also referred to as an encoded bit stream). The information can be, for example, a flag or the like. Thereby, a decoding apparatus can generate a prediction image in the manner or mode selected in the encoding apparatus 100 based on the information. Furthermore, in the example shown in FIG. 8, the prediction processing section selects any one of the prediction images after generating the prediction images by each manner. However, the prediction processing section can select the manner or mode based on the parameters for the above-described encoding processing before generating the prediction images, and can generate the prediction images according to the manner or mode. Figure 12 In the example shown in FIG. 8, the prediction processing section selects any one of the prediction images after generating the prediction images by each manner. However, the prediction processing section can select the manner or mode based on the parameters for the above-described encoding processing before generating the prediction images, and can generate the prediction images according to the manner or mode.

[0271] For example, the first manner and the second manner are intra prediction and inter prediction, respectively, and the prediction processing section can select the final prediction image for the current block from the prediction images generated in accordance with these prediction manners.

[0272] Figure 13 is a flowchart of another example of the processing performed by the prediction processing section of the encoding apparatus 100.

[0273] First, the prediction processing section generates a prediction image by intra prediction (step Sd la), and generates a prediction image by inter prediction (step Sd lb). Furthermore, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0274] Next, the prediction processing section evaluates each of the intra prediction image and the inter prediction image (step Sd 2). A cost can also be used in this evaluation. That is, the prediction processing section calculates the cost C of each of the intra prediction image and the inter prediction image. The cost C can be calculated by the equation of the R-D optimization model, for example, C = D + λ x R. In this equation, D is the encoding distortion of the prediction image, and is represented by, for example, the sum of the absolute values of the differences between the pixel values of the current block and the pixel values of the prediction image. In addition, R is the amount of generated encoding of the prediction image, specifically, the amount of encoding required for encoding the motion information or the like used for generating the prediction image, or the like. In addition, λ is, for example, an undetermined multiplier of Lagrange.

[0275] Then, the prediction processing section selects the prediction image in which the minimum cost C is calculated from among the intra prediction image and the inter prediction image as the final prediction image of the current block (step Sd_3). That is, the prediction method or mode used to generate the prediction image of the current block is selected.

[0276] [Intra prediction section]

[0277] The intra prediction section 124 performs intra prediction (also referred to as in-picture prediction) of the current block with reference to the blocks within the current picture stored in the block memory 118, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction section 124 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, color difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control section 128.

[0278] For example, the intra prediction section 124 performs intra prediction using one of a plurality of prescribed intra prediction modes. The plurality of intra prediction modes generally includes one or more non-directional prediction modes and a plurality of directional prediction modes. The prescribed plurality of modes can also be determined in advance.

[0279] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode prescribed by the H.265 / HEVC specification.

[0280] The plurality of directional prediction modes include, for example, 33 directional prediction modes prescribed by the H.265 / HEVC specification. In addition, the plurality of directional prediction modes can also include 32 directional prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 14 is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows indicate the 33 directions prescribed by the H.265 / HEVC specification, and the dashed arrows indicate the additional 32 directions (the 2 non-directional prediction modes are not shown in Figure 14 ).

[0281] In various processing examples, in the intra prediction of the color difference block, the luminance block can also be referred to. That is, the color difference component of the current block can also be predicted based on the luminance component of the current block. Such intra prediction is referred to as CCLM (cross-component linear model) prediction. Such an intra prediction mode of the color difference block that refers to the luminance block (e.g., referred to as a CCLM mode) can also be added as one of the intra prediction modes of the color difference block.

[0282] The intra prediction section 124 can also correct the pixel values after the intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction. Intra prediction with such correction is referred to as PDPC (position dependent intra prediction combination). Information indicating whether or not PDPC is applied (e.g., referred to as a PDPC flag) is generally signaled at the CU level. In addition, the signaling of this information need not be limited to the CU level, and can be another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0283] [Inter prediction section]

[0284] The inter prediction section 126 performs inter prediction (also referred to as inter picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). Inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 block) within the current block. For example, the inter prediction section 126 performs motion search within the reference picture for the current block or the current sub-block, to find a reference block or sub-block most consistent with the current block or the current sub-block. Also, the inter prediction section 126 acquires motion information (e.g., motion vector) for compensating for motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction section 126 performs motion compensation (or motion prediction) based on the motion information, thereby generating an inter prediction signal of the current block or sub-block. Also, the inter prediction section 126 outputs the generated inter prediction signal to the prediction control section 128.

[0285] The motion information used in the motion compensation is signaled as the inter prediction signal in various forms. For example, the motion vector can be signaled. As another example, the difference between the motion vector and the prediction motion vector (motion vector predictor) can be signaled.

[0286] [Basic flow of inter prediction]

[0287] Figure 15 is a flowchart showing an example of the basic flow of inter prediction.

[0288] The inter prediction section 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction section 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0289] Here, in the generation of the prediction image, the inter prediction section 126 generates the prediction image by performing decision of a motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Further, in the decision of the MV, the inter prediction section 126 decides the MV by performing selection of a candidate motion vector (candidate MV) (step Se_1) and derivation of the MV (step Se_2). The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. Further, in the derivation of the MV, the inter prediction section 126 can also decide the MV of the current block by further selecting at least one candidate MV from the at least one candidate MV. Alternatively, the inter prediction section 126 can decide the MV of the current block by searching a region of a reference picture indicated by each of the selected at least one candidate MV. Further, the action of searching the region of the reference picture can be referred to as motion estimation.

[0290] Further, in the above example, the steps Se_1 to Se_3 are performed by the inter prediction section 126, but the processing of, for example, the step Se_1 or the step Se_2 can be performed by another constituent element included in the encoding apparatus 100.

[0291] [Flow of derivation of motion vector]

[0292] Figure 16 is a flowchart showing an example of motion vector derivation.

[0293] The inter prediction section 126 derives the MV of the current block in a mode in which motion information (for example, the MV) is encoded. In this case, for example, the motion information is encoded as a prediction parameter, and is signaled. That is, the encoded motion information is included in an encoded signal (also referred to as an encoded bitstream).

[0294] Alternatively, the inter prediction section 126 derives the MV in a mode in which the motion information is not encoded. In this case, the motion information is not included in the encoded signal.

[0295] Here, the mode of the MV derivation can be the normal inter mode, the merge mode, the FRUC mode, the affine mode, and the like described later. Among these modes, the mode in which the motion information is encoded is the normal inter mode, the merge mode, and the affine mode (specifically, the affine inter mode and the affine merge mode), and the like. Further, the motion information can include not only the MV but also prediction motion vector selection information described later. Further, the mode in which the motion information is not encoded is the FRUC mode, and the like. The inter prediction section 126 selects the mode for deriving the MV of the current block from among these multiple modes, and derives the MV of the current block using the selected mode.

[0296] Figure 17 is a flowchart showing another example of the derivation of a motion vector.

[0297] The inter prediction section 126 derives the MV of the current block in a mode in which the difference MV is encoded. In this case, for example, the difference MV is encoded as a prediction parameter, and is signaled. That is, the encoded difference MV is included in the encoded signal. The difference MV is the difference between the MV of the current block and the predicted MV thereof.

[0298] Alternatively, the inter prediction section 126 derives the MV in a mode in which the difference MV is not encoded. In this case, the encoded difference MV is not included in the encoded signal.

[0299] Here, as described above, the derivation mode of the MV has the normal inter mode, the merge mode, the FRUC mode, the affine mode, and the like, which will be described later. Among these modes, the mode in which the difference MV is encoded has the normal inter mode, the affine mode (specifically, the affine inter mode), and the like. Further, the mode in which the difference MV is not encoded has the FRUC mode, the merge mode, and the affine mode (specifically, the affine merge mode), and the like. The inter prediction section 126 selects the mode for deriving the MV of the current block from among these multiple modes, and derives the MV of the current block using the selected mode.

[0300] [Flow of derivation of motion vector]

[0301] Figure 18 is a flowchart showing another example of the derivation of a motion vector. The mode of the derivation of the MV, that is, the inter prediction mode has multiple modes, and is roughly classified into a mode in which the difference MV is encoded and a mode in which the difference MV is not encoded. The mode in which the difference MV is not encoded has the merge mode, the FRUC mode, and the affine mode (specifically, the affine merge mode). Details of these modes will be described later, but briefly, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from among the periphery of the encoded blocks, the FRUC mode is a mode in which the MV of the current block is derived by performing a search between the encoded regions. Further, the affine mode is a mode in which the MV of the current block is derived by assuming an affine transformation, and taking the motion vector of each of the sub-blocks constituting the current block as the MV of the current block.

[0302] Specifically, as illustrated, in a case where the inter prediction mode information indicates 0 (0 in Sf_1), the inter prediction section 126 derives a motion vector based on the merge mode (Sf_2). Further, in a case where the inter prediction mode information indicates 1 (1 in Sf_1), the inter prediction section 126 derives a motion vector according to the FRUC mode (Sf_3). Further, in a case where the inter prediction mode information indicates 2 (2 in Sf_1), the inter prediction section 126 derives a motion vector according to the affine mode (specifically, affine merge mode) (Sf_4). Further, in a case where the inter prediction mode information indicates 3 (3 in Sf_1), the inter prediction section 126 derives a motion vector according to a mode in which a differential MV is encoded (for example, normal inter mode) (Sf_5).

[0303] [MV derivation > normal inter mode]

[0304] The normal inter mode is an inter prediction mode in which a MV of a current block is derived based on a block similar to the current block from a region of a reference picture represented by a candidate MV. Further, in this normal inter mode, a differential MV is encoded.

[0305] Figure 19 is a flowchart representing an example of inter prediction based on the normal inter mode.

[0306] First, the inter prediction section 126 acquires a plurality of candidate MVs for a current block based on information of MVs and the like of a plurality of coded blocks located around the current block in time or space (step Sg_1). That is, the inter prediction section 126 makes a candidate MV list.

[0307] Next, the inter prediction section 126 extracts N (N is an integer of 2 or more) candidate MVs as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order from the plurality of candidate MVs acquired in step Sg_1 (step Sg_2). Note that the priority order can be determined in advance for each of the N candidate MVs.

[0308] Next, the inter prediction section 126 selects one prediction motion vector candidate from the N prediction motion vector candidates as a prediction motion vector (also referred to as a prediction MV) of the current block (step Sg_3). At this time, the inter prediction section 126 encodes prediction motion vector selection information for identifying the selected prediction motion vector to a stream. Note that the stream is the above-described coded signal or coded bitstream.

[0309] Next, the inter prediction section 126 refers to the coded reference picture to derive the MV of the current block (step Sg_4). At this time, the inter prediction section 126 also codes the difference value between the derived MV and the predicted motion vector as a difference MV to the stream. Note that the coded reference picture is a picture composed of a plurality of blocks reconstructed after coding.

[0310] Finally, the inter prediction section 126 generates the prediction image of the current block by performing motion compensation on the current block using the derived MV and the coded reference picture (step Sg_5). Note that the prediction image is the inter prediction signal described above.

[0311] Further, information indicating the inter prediction mode (in the example described above, the normal inter mode) used in the generation of the prediction image included in the coded signal is coded as, for example, a prediction parameter.

[0312] Further, the candidate MV list can be used in common with the list used in other modes. Further, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to the candidate MV list is, for example, extraction or selection of a candidate MV from the candidate MV list, rearrangement of the candidate MVs, or deletion of a candidate MV, and the like.

[0313] [MV derivation > merge mode]

[0314] The merge mode is an inter prediction mode in which the MV of the current block is derived by selecting a candidate MV from the candidate MV list as the MV of the current block.

[0315] Figure 20 is a flowchart indicating an example of inter prediction based on the merge mode.

[0316] First, the inter prediction section 126 acquires a plurality of candidate MVs for the current block based on information of a plurality of coded block MVs and the like located around the current block in time or space (step Sh_l). That is, the inter prediction section 126 creates a candidate MV list.

[0317] Next, the inter prediction section 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs acquired in step Sh_l (step Sh_2). At this time, the inter prediction section 126 codes MV selection information for identifying the selected candidate MV to the stream.

[0318] Finally, the inter prediction section 126 generates the prediction image of the current block by performing motion compensation on the current block using the derived MV and the coded reference picture (step Sh_3).

[0319] Further, information indicating an inter prediction mode (in the above example, the merge mode) used in generation of the prediction image is coded as, for example, a prediction parameter, in the coded signal.

[0320] Figure 21 is a conceptual diagram for explaining an example of a motion vector derivation process of a current picture based on the merge mode.

[0321] First, a prediction MV list in which candidates of a prediction MV are registered is generated. As the candidates of the prediction MV, there are: a spatial neighboring prediction MV, which is an MV possessed by a plurality of coded blocks located in the periphery of the space of the target block; a temporal neighboring prediction MV, which is an MV possessed by a block in the vicinity of the target block projected in the position of the coded reference picture; a combined prediction MV, which is an MV generated by combining the MV values of the spatial neighboring prediction MV and the temporal neighboring prediction MV; and a zero prediction MV, which is an MV having a value of zero, and the like.

[0322] Next, the MV for the target block is decided by selecting one prediction MV from among the plurality of prediction MVs registered in the prediction MV list.

[0323] Further, in the variable length coding section, a signal indicating which prediction MV is selected, that is, merge_idx, is coded by being described in the stream.

[0324] In addition, the prediction MVs registered in the prediction MV list explained in Figure 21 may be a different number from that in the drawing, or a structure in which a part of the kinds of the prediction MVs in the drawing is not included, or a structure in which a prediction MV other than the kinds of the prediction MVs in the drawing is added.

[0325] The MV of the target block derived by the merge mode can also be used to decide the final MV by performing a DMVR (decoder motion vector refinement) process described later.

[0326] In addition, the candidates of the prediction MV are the above-described candidate MVs, and the prediction MV list is the above-described candidate MV list. Further, the candidate MV list can also be referred to as a candidate list. Further, the merge_idx is the MV selection information.

[0327] [MV derivation > FRUC mode]

[0328] The motion information can also not be signaled from the encoding device side but derived at the decoding device side. In addition, as described above, the merge mode prescribed by the H.265 / HEVC standard can also be used. Further, for example, the motion information can also be derived by performing a motion search at the decoding device side. In the embodiment, the motion search is performed at the decoding device side without using the pixel values of the current block.

[0329] Here, a mode for performing motion estimation on the decoding device side is described. This mode for performing motion estimation on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0330] In the form of a flow chart Figure 22 An example of FRUC processing is shown in Figure 1. First, a list of multiple candidates (i.e., a candidate MV list, which may also be shared with a merge list) each including a predicted motion vector (MV) is generated by referring to the motion vectors of previously coded blocks that are spatially or temporally adjacent to the current block (step Si_1). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and one candidate is selected based on the evaluation value. Furthermore, a motion vector for the current block is derived based on the motion vector of the selected candidate (step Si_4). Specifically, for example, the selected candidate motion vector (the best candidate MV) is derived as is as the motion vector for the current block. Alternatively, the motion vector for the current block can be derived by performing pattern matching in the surrounding area of ​​the position in the reference picture corresponding to the selected candidate motion vector. Specifically, the surrounding area of ​​the best candidate MV can be searched using pattern matching and evaluation values ​​in the reference picture. If an MV with a better evaluation value is found, the best candidate MV is updated to the above MV and used as the final MV for the current block. A configuration may be adopted in which the process of updating to an MV having a better evaluation value is not performed.

[0331] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).

[0332] Exactly the same processing can be performed even when processing is performed in sub-block units.

[0333] The evaluation value can also be calculated using various methods. For example, the reconstructed image of the region within the reference picture corresponding to the motion vector is compared with the reconstructed image of a predetermined region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). The predetermined region may also be predetermined.

[0334] Then, a difference in pixel values of the two reconstructed images can also be calculated for the evaluation value of the motion vector. Alternatively, other information can also be used in addition to the difference value to calculate the evaluation value.

[0335] Next, an example of pattern matching is described in detail. First, one of the candidate MVs included in the candidate MV list (for example, the merge list) is selected as the starting point of the search based on pattern matching. For example, as the pattern matching, the 1st pattern matching or the 2nd pattern matching can be used. The 1st pattern matching and the 2nd pattern matching are respectively referred to as bilateral matching and template matching.

[0336] [MV derivation > FRUC > bilateral matching]

[0337] In the 1st pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in different two reference pictures. Thus, in the 1st pattern matching, as the specified region for the calculation of the evaluation value for the candidate, the region in the other reference picture along the motion trajectory of the current block is used. The specified region can also be determined in advance.

[0338] Figure 23 is a conceptual diagram for explaining an example of the 1st pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As shown in Figure 23 In the 1st pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, a difference in the reconstructed image at a specified position in the 1st encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the 2nd encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the above-mentioned candidate MV by a display time interval is calculated, and the obtained difference value is used to calculate the evaluation value. The candidate MV with the best evaluation value can be selected from among a plurality of candidate MVS as the final MV, and a good result can be obtained.

[0339] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) of the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, in the case where the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the 1st pattern matching, the bidirectional motion vectors that are mirror-symmetrical are derived.

[0340] [MV derivation > FRUC > template matching]

[0341] In the 2nd pattern matching (template matching), pattern matching is performed between a template (a block adjacent to the current block (e.g., the upper and / or left adjacent block) within the current picture) and a block within the reference picture. Thus, in the 2nd pattern matching, as the specified region for the calculation of the evaluation value for the candidate described above, the block adjacent to the current block within the current picture is used.

[0342] Figure 24 is a conceptual diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within the reference picture. As shown in Figure 24 , in the 2nd pattern matching, the motion vector of the current block is derived by searching for the block within the reference picture (Ref0) that is most matched to the block adjacent to the current block (Cur block) within the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of the left adjacent and upper adjacent both or one and the reconstructed image at the equivalent position within the encoded reference picture (Ref0) specified by the candidate MV is calculated, the evaluation value is calculated using the obtained difference value, and the candidate MV whose evaluation value is the best value can be selected as the best candidate MV among the plurality of candidate MVs.

[0343] Such information indicating whether to adopt the FRUC mode (e.g., called FRUC flag) is signaled at the CU level. In addition, in the case where the FRUC mode is adopted (e.g., in the case where the FRUC flag is true), information indicating the method of the pattern matching (the 1st pattern matching or the 2nd pattern matching) that can be adopted is signaled at the CU level. In addition, the signaling of these information does not need to be limited to the CU level, and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0344] [MV derivation > affine mode]

[0345] Next, an affine mode that derives a motion vector in a sub-block unit based on motion vectors of a plurality of neighboring blocks will be described. This mode is sometimes referred to as an affine motion compensation prediction mode.

[0346] Figure 25A is a conceptual diagram for explaining an example of derivation of a motion vector in a sub-block unit based on motion vectors of a plurality of neighboring blocks. In Figure 25A , a current block includes 16 4x4 sub-blocks. Here, a motion vector v0 of a top-left control point of the current block is derived based on motion vectors of neighboring blocks, and similarly, a motion vector v1 of a top-right control point of the current block is derived based on motion vectors of neighboring sub-blocks. Then, 2 motion vectors v0 and v1 can be projected, and motion vectors (v x , v y ) of each sub-block within the current block can be derived according to the following equation (1A).

[0347] [Equation 1]

[0348]

[0349] Here, x and y represent a horizontal position and a vertical position of a sub-block, respectively, and w represents a predetermined weight coefficient. The predetermined weight coefficient can be determined in advance.

[0350] Information indicating such an affine mode (for example, referred to as an affine flag) can be signaled as a signal of a CU level. Furthermore, the signaling of the information indicating the affine mode need not be limited to the CU level, and can be another level (for example, a sequence level, a picture level, a slice level, a tile level, a CTU level, or a sub-block level).

[0351] In addition, in such an affine mode, several modes having different derivation methods of motion vectors of top-left and top-right control points can be included. For example, in the affine mode, there are 2 modes of an affine inter (also referred to as an affine normal inter) mode and an affine merge mode.

[0352] [MV derivation > affine mode]

[0353] Figure 25B is a conceptual diagram for explaining an example of derivation of a motion vector in a sub-block unit in an affine mode having 3 control points. In Figure 25BIn the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the top-left corner control point of the current block is derived based on the motion vectors of the neighboring blocks, likewise, the motion vector v1 of the top-right corner control point of the current block is derived based on the motion vectors of the neighboring blocks, and the motion vector v2 of the bottom-left corner control point of the current block is derived based on the motion vectors of the neighboring blocks. Then, the three motion vectors v0, v1 and v2 can be projected, and the motion vectors (v x , v y ) of the sub-blocks in the current block can be derived according to the following formula (1B).

[0354] [Formula 2]

[0355]

[0356] Here, x and y represent the horizontal position and the vertical position of the center of the sub-block, respectively, w represents the width of the current block, and h represents the height of the current block.

[0357] The affine modes with different numbers of control points (for example, 2 and 3) can also be switched and signaled at the CU level. In addition, information indicating the number of control points of the affine mode used in the CU level can also be signaled at other levels (for example, the sequence level, the picture level, the slice level, the tile level, the CTU level or the sub-block level).

[0358] In addition, in such an affine mode with 3 control points, several modes with different derivation methods of the motion vectors of the top-left, top-right and bottom-left corner control points can also be included. For example, in the affine mode, there are two modes, namely, the affine inter (also referred to as affine normal inter) mode and the affine merge mode.

[0359] [MV derivation > affine merge mode]

[0360] Figure 26A , Figure 26B and Figure 26C are conceptual diagrams for illustrating the affine merge mode.

[0361] In the affine merge mode, as shown in Figure 26A , for example, the prediction motion vector of each of the control points of the current block is calculated based on the plurality of motion vectors corresponding to the blocks coded in the affine mode in the neighboring blocks A (left), B (top), C (top-right), D (bottom-left) and E (top-left) of the current block. Specifically, the coded blocks A (left), B (top), C (top-right), D (bottom-left) and E (top-left) are checked in order, and the first valid block coded in the affine mode is determined. The prediction motion vector of the control point of the current block is calculated based on the plurality of motion vectors corresponding to the determined block.

[0362] For example, as shown inFigure 26B As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 are derived that are projected onto the positions of the upper left corner and upper right corner of the encoded block including block A. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.

[0363] For example, Figure 26C As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, a predicted motion vector v0 of the control point at the upper left corner, a predicted motion vector v1 of the control point at the upper right corner, and a predicted motion vector v2 of the control point at the lower left corner of the current block are calculated.

[0364] In addition, in the following Figure 29 This method of deriving a predicted motion vector may also be used in deriving the predicted motion vectors of the control points of the current block in step Sj_1.

[0365] Figure 27 This is a flowchart showing an example of the affine merge mode.

[0366] In the affine merge mode, as shown in the figure, first, the inter-frame prediction unit 126 derives the predicted MV of each control point of the current block (step Sk_1). Figure 25A As shown, it is the top left and top right corner points of the current block, or as Figure 25B As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block.

[0367] That is to say, if Figure 26A As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.

[0368] Then, in the case where block A is determined and block A has 2 control points, as Figure 26B As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner and the motion vector v1 of the control point at the upper right corner of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block including the block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner and the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0369] Alternatively, in the case where block A is determined and block A has 3 control points, as Figure 26C As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner, the motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner, the predicted motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block.

[0370] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. Specifically, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two predicted motion vectors v0 and v1 and the above-mentioned equation (1A), or three predicted motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the motion vector for that sub-block as an affine MV (step Sk_2). The inter-frame prediction unit 126 then performs motion compensation on that sub-block using these affine MVs and the coded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image for the current block is generated.

[0371] [MV Export > Affine Inter-frame Mode]

[0372] Figure 28A This is a conceptual diagram for explaining the affine inter mode with two control points.

[0373] In this affine inter-frame mode, if Figure 28A As shown, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0374] Figure 28B This is a conceptual diagram for explaining the affine inter-frame mode with three control points.

[0375] In this affine inter-frame mode, if Figure 28BAs shown, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as a prediction motion vector v0 of the top-left corner of the current block. Likewise, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as a prediction motion vector v1 of the top-right corner of the current block. Further, a motion vector selected from the motion vectors of the coded blocks F and G adjacent to the current block is used as a prediction motion vector v2 of the bottom-left corner of the current block.

[0376] Figure 29 is a flowchart showing an example of the affine inter mode.

[0377] As shown, in the affine inter mode, first, the inter prediction section 126 derives the prediction MVs (v0, v1) or (v0, v1, v2) of the 2 or 3 control points of the current block (step Sj_1). As shown in Figure 25A or Figure 25B As shown, the control points are the points of the top-left corner, the top-right corner, or the bottom-left corner of the current block.

[0378] That is, the inter prediction section 126 derives the prediction motion vectors (v0, v1) or (v0, v1, v2) of the control points of the current block by selecting the motion vector of a certain block among the coded blocks in the vicinity of each control point of the current block as shown in Figure 28A or Figure 28B At this time, the inter prediction section 126 encodes prediction motion vector selection information for identifying the selected 2 motion vectors into the stream.

[0379] For example, the inter prediction section 126 can decide which motion vector of the coded blocks adjacent to the current block to select as the prediction motion vector of the control point by using a cost evaluation or the like, and can describe a flag indicating which prediction motion vector is selected in the bitstream.

[0380] Next, the inter prediction section 126 performs a motion search (steps Sj_3 and Sj_4) while updating the prediction motion vectors selected or derived in step Sj_1 (step Sj_2). That is, the inter prediction section 126 calculates the motion vectors of each sub-block corresponding to the prediction motion vectors to be updated as affine MVs using the above-described equation (1A) or equation (1B) (step Sj_3). Then, the inter prediction section 126 performs motion compensation on each sub-block using these affine MVs and the coded reference picture (step Sj_4). As a result, in the motion search loop, the inter prediction section 126 decides, for example, the prediction motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter prediction section 126 also encodes the difference value between this decided MV and the prediction motion vector as a difference MV into the stream.

[0381] Finally, the inter prediction unit 126 generates a prediction image of the current block by performing motion compensation on the current block using the decided MV and the coded reference picture (step Sj-6).

[0382] [MV derivation > affine inter mode]

[0383] In a case where affine modes with different numbers of control points (e.g., 2 and 3) are signaled at the CU level, sometimes the number of control points is different between the coded block and the current block. Figure 30A and Figure 30B is a conceptual diagram for explaining a method of deriving prediction vectors of control points in a case where the number of control points is different between the coded block and the current block.

[0384] For example, as shown in Figure 30A , in a case where the current block has 3 control points of the upper left corner, the upper right corner, and the lower left corner, and a block A adjacent to the left side of the current block is coded in an affine mode with 2 control points, motion vectors v3 and v4 projecting to positions of the upper left corner and the upper right corner of the coded block including the block A are derived. Then, according to the derived motion vectors v3 and v4, a prediction motion vector v0 of the control point of the upper left corner of the current block and a prediction motion vector v1 of the control point of the upper right corner of the current block are calculated. Further, a prediction motion vector v2 of the control point of the lower left corner is calculated according to the derived motion vectors v0 and v1.

[0385] For example, as shown in Figure 30B , in a case where the current block has 2 control points of the upper left corner and the upper right corner, and a block A adjacent to the left side of the current block is coded in an affine mode with 3 control points, motion vectors v3, v4, and v5 projecting to positions of the upper left corner, the upper right corner, and the lower left corner of the coded block including the block A are derived. Then, according to the derived motion vectors v3, v4, and v5, a prediction motion vector v0 of the control point of the upper left corner of the current block and a prediction motion vector v1 of the control point of the upper right corner of the current block are calculated.

[0386] In the derivation of each prediction motion vector of the control points of the current block in step Sj-1 of Figure 29 , the prediction motion vector derivation method can also be used.

[0387] [MV derivation > DMVR]

[0388] Figure 31A is a flowchart showing the relationship of the merge mode and DMVR.

[0389] The inter prediction section 126 derives a motion vector of the current block in the merge mode (step Sl_l). Next, the inter prediction section 126 determines whether or not to perform a motion vector search, that is, a motion search (step Sl_2). Here, when it is determined not to perform the motion search (NO in step Sl_2), the inter prediction section 126 decides the motion vector derived in step Sl_l as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is decided in the merge mode.

[0390] On the other hand, when it is determined to perform the motion search in step Sl_l (YES in step Sl_2), the inter prediction section 126 derives the final motion vector for the current block by searching a peripheral region of a reference picture indicated by the motion vector derived in step Sl_l (step Sl_3). That is, in this case, the motion vector of the current block is decided by the DMVR.

[0391] Figure 31B is a conceptual diagram for explaining an example of the DMVR process for deciding an MV.

[0392] First, the optimal MVP set to the current block (for example, in the merge mode) is set as a candidate MV. Then, in accordance with the candidate MV (L0), a reference pixel is determined from an encoded picture in the L0 direction, that is, a first reference picture (L0). Similarly, in accordance with the candidate MV (Ll), a reference pixel is determined from an encoded picture in the Ll direction, that is, a second reference picture (Ll). A template is generated by taking an average of these reference pixels.

[0393] Next, using the above template, a peripheral region of the candidate MV of the first reference picture (L0) and the second reference picture (Ll) is searched, respectively, and an MV having the minimum cost is decided as the final MV. Further, the cost value can be calculated using, for example, a difference value of each pixel value of the template and each pixel value of the search region, and a candidate MV value, and the like.

[0394] In addition, the structure and the operation of the process explained here are basically common in the encoding apparatus and the decoding apparatus described later.

[0395] Even if it is not the process example explained here itself, as long as it is a process capable of deriving the final MV by searching the peripheral of the candidate MV, an arbitrary process can be used.

[0396] [Motion compensation > BIO / OBMC]

[0397] In the motion compensation, there is a mode of generating a prediction image and correcting the prediction image. This mode is, for example, BIO and OBMC described later.

[0398] Figure 32is a flowchart showing one example of generation of a prediction image.

[0399] The inter prediction section 126 generates a prediction image (step Sm_1), and corrects the prediction image by, for example, any of the above-described modes (step Sm_2).

[0400] Figure 33 is a flowchart showing another example of generation of a prediction image.

[0401] The inter prediction section 126 decides a motion vector of the current block (step Sn_1). Next, the inter prediction section 126 generates a prediction image (step Sn_2), and determines whether or not to perform correction processing (step Sn_3). Here, when it is determined to perform the correction processing (Yes in step Sn_3), the inter prediction section 126 generates a final prediction image by correcting the prediction image (step Sn_4). On the other hand, when it is determined not to perform the correction processing (No in step Sn_3), the inter prediction section 126 outputs the prediction image as the final prediction image without correcting the prediction image (step Sn_5).

[0402] Further, in the motion compensation, there is a mode of correcting the luminance at the time of generation of the prediction image. This mode is, for example, LIC described later.

[0403] Figure 34 is a flowchart showing another example of generation of a prediction image.

[0404] The inter prediction section 126 derives a motion vector of the current block (step So_1). Next, the inter prediction section 126 determines whether or not to perform luminance correction processing (step So_2). Here, when it is determined to perform the luminance correction processing (Yes in step So_2), the inter prediction section 126 generates a prediction image while performing luminance correction (step So_3). That is, the prediction image is generated by LIC. On the other hand, when it is determined not to perform the luminance correction processing (No in step So_2), the inter prediction section 126 generates a prediction image by the usual motion compensation without performing luminance correction (step So_4).

[0405] [Motion compensation > OBMC]

[0406] Not only the motion information of the current block obtained by the motion search, but also the motion information of the neighboring block can be used to generate the inter prediction signal. Specifically, the inter prediction signal can also be generated in the sub-block unit within the current block by weighted adding a prediction signal based on the motion information obtained by the motion search (within the reference picture) and a prediction signal based on the motion information of the neighboring block (within the current picture). Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0407] In the OBMC mode, information indicating the size of the sub-block for OBMC (e.g., referred to as OBMC block size) can also be signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (e.g., referred to as OBMC flag) can also be signaled at the CU level. In addition, the level of the signaling of these information does not need to be limited to the sequence level and the CU level, but can be other levels (e.g., the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0408] An example of the OBMC mode is described in more detail. Figure 35 and Figure 36 are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process based on the OBMC process.

[0409] First, as shown in Figure 36 , a prediction image (Pred) based on the usual motion compensation is obtained using the motion vector (MV) assigned to the processing target (current) block. In Figure 36 , the arrow "MV" points to the reference picture, and indicates which block of the current picture refers to in order to obtain the prediction image.

[0410] Next, the motion vector (MV_L) already derived for the already encoded left neighboring block is applied (reused) to the encoding target block, and a prediction image (Pred_L) is obtained. The motion vector (MV_L) is indicated by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two prediction images Pred and Pred_L, the 1st correction of the prediction image is performed. This has the effect of mixing the boundaries between the neighboring blocks.

[0411] Likewise, the motion vector (MV_U) that has been derived for the coded upper neighboring block is applied (reused) to the block being coded to obtain a prediction image (Pred_U). The motion vector (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, a second correction of the prediction image is performed by superimposing the prediction image Pred_U with the prediction image that has been corrected for the first time (e.g., Pred and Pred_L). This has the effect of blending the boundaries between neighboring blocks. The prediction image obtained by the second correction is the final prediction image for the current block with the boundaries to the neighboring blocks being blended (smoothed).

[0412] Further, the above-described example is a two-path correction method using the left and upper neighboring blocks, but the correction method can also be a three-path or more path correction method using the right and / or lower neighboring blocks as well.

[0413] In addition, the region to which the superimposition is performed can also not be the entire pixel region of the block, but only a partial region near the block boundary.

[0414] In addition, the prediction image correction processing of the OBMC that is described herein is processing for obtaining one prediction image Pred by superimposing one reference picture with the additional prediction images Pred_L and Pred_U. However, in the case where the prediction image is corrected based on multiple reference pictures, the same processing can also be applied to the multiple reference pictures respectively. In this case, by performing the image correction of the OBMC based on multiple reference pictures, after the corrected prediction images are obtained from each of the reference pictures, the final prediction image is obtained by further superimposing the obtained multiple corrected prediction images.

[0415] In addition, in the OBMC, the unit of the block being coded can be a prediction block unit, or a sub-block unit obtained by further dividing the prediction block.

[0416] As a method of determining whether or not to apply the OBMC processing, for example, there is a method of using a signal indicating whether or not to apply the OBMC processing, i.e., obmc_flag. As a specific example, the coding device can determine whether or not the block being coded belongs to a region of complex motion. The coding device sets the value of obmc_flag to 1 and applies the OBMC processing to code in the case where the block belongs to the region of complex motion, and sets the value of obmc_flag to 0 and codes without applying the OBMC processing in the case where the block does not belong to the region of complex motion. On the other hand, in the decoding device, by decoding the obmc_flag described in the stream (e.g., compressed sequence), whether or not to apply the OBMC processing is switched according to the value to decode.

[0417] In the example described above, the inter prediction section 126 generates one rectangular prediction image for the rectangular current block. However, the inter prediction section 126 can generate a plurality of prediction images having shapes different from a rectangle for the rectangular current block, and can generate a final rectangular prediction image by combining the plurality of prediction images. The shapes different from a rectangle can be, for example, triangles.

[0418] Figure 37 is a conceptual diagram for explaining generation of prediction images of two triangles.

[0419] The inter prediction section 126 generates a triangular prediction image by performing motion compensation using the first MV of the first triangular partition within the current block. Likewise, the inter prediction section 126 generates a triangular prediction image by performing motion compensation using the second MV of the second triangular partition within the current block. Then, the inter prediction section 126 generates a rectangular prediction image identical to the current block by combining the prediction images.

[0420] Further, in the example shown in Figure 37 , the first and second partitions are each a triangle, but can be a trapezoid, and can each be a shape different from each other. Also, in the example shown in Figure 37 , the current block is composed of two partitions, but can be composed of three or more partitions.

[0421] In addition, the first and second partitions can be repeated. That is, the first and second partitions can include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition can be used to generate the prediction image of the current block.

[0422] In addition, in this example, an example in which both of the two partitions generate prediction images by inter prediction is shown, but prediction images can be generated by intra prediction for at least one of the partitions.

[0423] [Motion compensation > BIO]

[0424] Next, a method of deriving a motion vector will be described. First, a mode of deriving a motion vector based on a model assuming constant velocity straight line motion will be described. This mode is sometimes referred to as a BIO (bi-directional optical flow) mode.

[0425] Figure 38 is a conceptual diagram for explaining a model assuming constant velocity straight line motion. In Figure 38 , (v x , v y) indicates a velocity vector, τ0, τ1 indicate a temporal distance between a current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) indicates a motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates a motion vector corresponding to the reference picture Ref1.

[0426] At this time, it can also be that, under the assumption of constant velocity straight line motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (−vxτ1, −vyτ1), respectively, using the following optical flow equation (2).

[0427] [Equation 3]

[0428]

[0429] Here, I(k) indicates a luminance value of a reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) a temporal differential of the luminance value, (ii) a product of a horizontal component of a velocity in the horizontal direction and a spatial gradient of the reference image, and (iii) a product of a vertical component of a velocity in the vertical direction and a spatial gradient of the reference image, is equal to zero. It can also be that, based on a combination of this optical flow equation and Hermite interpolation, a block unit motion vector obtained from a merge list or the like is corrected in pixel units.

[0430] In addition, it can also be that a motion vector is derived on the decoding device side by a method different from derivation of a motion vector based on a model assuming constant velocity straight line motion. For example, it can also be that a motion vector is derived in sub-block units based on motion vectors of a plurality of neighboring blocks.

[0431] [Motion compensation > LIC]

[0432] Next, an example of a mode in which a prediction image (prediction) is generated using LIC (local illumination compensation) processing will be described.

[0433] Figure 39 is a conceptual diagram for explaining an example of a prediction image generation method using luminance correction processing based on LIC processing.

[0434] First, a MV is derived from an encoded reference image, and a reference image corresponding to a current block is obtained.

[0435] Next, information indicating how the luminance value changes in the reference picture and the current picture is extracted for the current block. The extraction is made based on the luminance pixel values of the coded left neighboring reference region (the surrounding reference region) and the coded upper neighboring reference region (the surrounding reference region) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture designated by the derived MV. Then, using the information indicating how the luminance value changes, the luminance correction parameter is calculated.

[0436] By applying the above luminance correction parameter to the reference picture within the reference picture designated by the MV, the luminance correction processing is performed on the reference picture, and a prediction picture for the current block is generated.

[0437] In addition, Figure 39 The shape of the above surrounding reference region in the above embodiment is an example, and a shape other than the above can also be used.

[0438] Further, the processing of generating a prediction picture from one reference picture is described here, but the same applies to the case of generating a prediction picture from a plurality of reference pictures, and a prediction picture can also be generated after performing luminance correction processing on the reference pictures obtained from each of the reference pictures in the same manner as described above.

[0439] As a method of determining whether to apply the LIC processing, for example, there is a method of using lic_flag as a signal indicating whether to apply the LIC processing. As a specific example, in the encoding device, it is determined whether the current block belongs to a region where luminance change has occurred, and in the case of belonging to a region where luminance change has occurred, the value 1 is set as lic_flag, and the LIC processing is applied to encode, and in the case of not belonging to a region where luminance change has occurred, the value 0 is set as lic_flag, and the LIC processing is not applied to encode. On the other hand, in the decoding device, lic_flag described in the stream can be decoded, and whether to apply the LIC processing is switched according to the value to decode.

[0440] As another method of determining whether to apply the LIC processing, for example, there is a method of determining whether the LIC processing is applied to the surrounding blocks in the surrounding blocks. As a specific example, in the case of the current block being in the merge mode, it is determined whether the coded blocks of the surrounding selected in the derivation of the MV in the merge mode processing are encoded with the LIC processing, and according to the result, whether to apply the LIC processing is switched to encode. In addition, in the case of the example, the same processing also applies to the decoding device side.

[0441] The above is the processing of the LIC processing (luminance correction processing), and the details thereof will be described below. Figure 39

[0442] ​First, the inter prediction section 126 derives a motion vector for obtaining a reference image corresponding to the encoding target block from a reference picture that is an already encoded picture.

[0443] Next, the inter prediction section 126 extracts information indicating how luminance values change in the reference picture and the encoding target picture using luminance pixel values of the already encoded peripheral reference region of the left neighbor and the upper neighbor and luminance pixel values at equivalent positions in the reference picture specified by the motion vector, and calculates luminance correction parameters for the encoding target block. For example, let the luminance pixel value of a certain pixel in the peripheral reference region in the encoding target picture be pO, and let the luminance pixel value of the pixel in the peripheral reference region in the reference picture at the equivalent position be pi. The inter prediction section 126 calculates coefficients A and B for optimizing A x pi + B = pO as the luminance correction parameters for a plurality of pixels in the peripheral reference region.

[0444] Next, the inter prediction section 126 generates a prediction image for the encoding target block by performing luminance correction processing on the reference image in the reference picture specified by the motion vector using the luminance correction parameters. For example, let the luminance pixel value in the reference image be p2, and let the luminance pixel value of the prediction image after the luminance correction processing be p3. The inter prediction section 126 generates the prediction image after the luminance correction processing by calculating A x p2 + B = p3 for each pixel in the reference image.

[0445] Further, Figure 39 The shape of the peripheral reference region in the above-described example is one example, and shapes other than this can also be used. Also, a part of the peripheral reference region shown in FIG. 6 can be used. For example, a region including a prescribed number of pixels excluding the upper neighbor pixel and the left neighbor pixel can be used as the peripheral reference region. Further, the peripheral reference region is not limited to a region adjacent to the encoding target block, and can be a region that is not adjacent to the encoding target block. The prescribed number related to the pixel can be determined in advance. Figure 39 Further, in the example shown in FIG. 6, the peripheral reference region in the reference picture is a region specified by the motion vector of the encoding target picture from among the peripheral reference regions in the encoding target picture, but can be a region specified by another motion vector. For example, the other motion vector can be a motion vector of the peripheral reference region in the encoding target picture.

[0446] Figure 39 Also, in this embodiment, the operation in the encoding apparatus 100 is described, but typically, the operation in the decoding apparatus 200 is the same.

[0447] Further, in this embodiment, the operation in the encoding apparatus 100 is described, but typically, the operation in the decoding apparatus 200 is the same.

[0448] ​Furthermore, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for all of them.

[0449] Furthermore, the LIC process may be applied in sub-block units. For example, the modification parameters may be derived using the surrounding reference region of the current sub-block and the surrounding reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0450] [Prediction Control Department]

[0451] The prediction control unit 128 selects one of the intra-frame prediction signal (the signal output from the intra-frame prediction unit 124) and the inter-frame prediction signal (the signal output from the inter-frame prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0452] like Figure 1 As shown, in various examples of encoding devices, the prediction control unit 128 may also output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate a coded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may also be used in the decoding device. The decoding device may also receive and decode the coded bitstream, performing the same prediction processing as that performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction parameters may include a prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra-frame prediction unit 124 or the inter-frame prediction unit 126), or any index, flag, or value based on the prediction processing performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128 or indicating the prediction processing.

[0453] [Encoding device installation example]

[0454] Figure 40 1 is a block diagram showing an implementation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. For example, Figure 1 The multiple components of the encoding device 100 shown are Figure 40 The processor a1 and the memory a2 shown are implemented.

[0455] Processor a1 is a circuit that processes information and can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. Processor a1 can also be a processor such as a CPU. In addition, processor a1 can also be a collection of multiple electronic circuits. In addition, for example, processor a1 can also play a role. Figure 1The functions of multiple components of the encoding device 100 shown in FIG.

[0456] Memory a2 is a dedicated or general-purpose memory that stores information used by processor a1 to encode moving images. Memory a2 can be an electronic circuit or connected to processor a1. Alternatively, memory a2 can be included in processor a1. Alternatively, memory a2 can be a collection of multiple electronic circuits. Memory a2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory a2 can be either non-volatile or volatile memory.

[0457] For example, the memory a2 may store a coded moving image or a bit string corresponding to the coded moving image. In addition, the memory a2 may store a program for the processor a1 to encode the moving image.

[0458] In addition, for example, memory a2 can also serve as Figure 1 The memory a2 can be used as a component for storing information among the multiple components of the encoding device 100 shown in FIG. Figure 1 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks and reconstructed pictures.

[0459] In addition, in the encoding device 100, it is not necessary to install Figure 1 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 1 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.

[0460] [Decoding device]

[0461] Next, a decoding device that can decode the coded signal (coded bit stream) output from, for example, the above-described coding device 100 will be described. Figure 41 2 is a block diagram showing the functional structure of a decoding device 200 according to an embodiment. The decoding device 200 is a moving picture decoding device that decodes a moving picture in units of blocks.

[0462] like Figure 41 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.

[0463] The decoding device 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding section 202, the inverse quantization section 204, the inverse transform section 206, the addition section 208, the loop filtering section 212, the intra prediction section 216, the inter prediction section 218, and the prediction control section 220. Alternatively, the decoding device 200 can be realized as one or more dedicated electronic circuits corresponding to the entropy decoding section 202, the inverse quantization section 204, the inverse transform section 206, the addition section 208, the loop filtering section 212, the intra prediction section 216, the inter prediction section 218, and the prediction control section 220.

[0464] Hereinafter, after the flow of the overall processing of the decoding device 200 is described, each constituent element included in the decoding device 200 is described.

[0465] [Overall flow of decoding processing]

[0466] Figure 42 is a flowchart showing an example of the overall decoding processing performed by the decoding device 200.

[0467] First, the entropy decoding section 202 of the decoding device 200 determines a partitioning pattern of a fixed-size block (for example, 128 x 128 pixels) (step Sp_1). The partitioning pattern is the partitioning pattern selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 on each of the plurality of blocks constituting the partitioning pattern.

[0468] That is, the entropy decoding section 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and the prediction parameters of the block to be decoded (also referred to as a current block) (step Sp_2).

[0469] Next, the inverse quantization section 204 and the inverse transform section 206 reproduce the plurality of prediction residuals (that is, the difference blocks) by inverse quantizing and inverse transforming the plurality of quantization coefficients (step Sp_3).

[0470] Next, a prediction processing section constituted by all or a part of the intra prediction section 216, the inter prediction section 218, and the prediction control section 220 generates a prediction signal (also referred to as a prediction block) of the current block (step Sp_4).

[0471] Next, the addition section 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the difference block (step Sp_5).

[0472] Further, when the reconstructed image is generated, the loop filtering section 212 filters the reconstructed image (step Sp_6).

[0473] Then, the decoding device 200 determines whether decoding of the entire picture is completed (step Sp_7 ). If it is determined that decoding is not completed (No in step Sp_7 ), the processing from step Sp_1 is repeatedly executed.

[0474] As shown in the figure, the processing of steps Sp_1 to Sp_7 is sequentially performed by the decoding device 200, or a plurality of processes of some of these processes may be performed in parallel, or the order may be reversed.

[0475] [Entropy decoding unit]

[0476] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the coded bit stream to the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220 in the embodiment (see Figure 1 The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.

[0477] [Inverse quantization unit]

[0478] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantization unit 204 then outputs the inversely quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0479] [Inverse transformation unit]

[0480] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients input from the inverse quantization unit 204 .

[0481] For example, when the information read from the coded bitstream indicates that EMT or AMT is adopted (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.

[0482] Furthermore, for example, when the information decoded from the coded bit stream indicates that NSST is adopted, the inverse transform unit 206 applies inverse re-transformation to the transform coefficients.

[0483] [addition section]

[0484] The addition section 208 reconstructs the current block by adding the prediction error as an input from the inverse transform section 206 to the prediction sample as an input from the prediction control section 220. Also, the addition section 208 outputs the reconstructed block to the block memory 210 and the loop filter section 212.

[0485] [block memory]

[0486] The block memory 210 is a storage section for storing a block within a decoded target picture (hereinafter referred to as a current picture) referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the addition section 208.

[0487] [loop filter section]

[0488] The loop filter section 212 applies loop filtering to the block reconstructed by the addition section 208, and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.

[0489] In a case where the information indicating the on / off of the ALF read out from the coded bitstream indicates the on of the ALF, one filter is selected from among a plurality of filters based on the direction and the activity of the gradient of the locality, and the selected filter is applied to the reconstructed block.

[0490] [frame memory]

[0491] The frame memory 214 is a storage section for storing a reference picture used in inter prediction, and is also referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter section 212.

[0492] [prediction processing section (intra prediction section / inter prediction section / prediction control section)]

[0493] Figure 43 is a flowchart showing an example of the processing performed by the prediction processing section of the decoding device 200. Further, the prediction processing section is constituted by all or a part of the constituent elements of the intra prediction section 216, the inter prediction section 218, and the prediction control section 220.

[0494] The prediction processing section generates a prediction image of the current block (step Sq_1). The prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there are an intra prediction signal or an inter prediction signal. Specifically, the prediction processing section generates the prediction image of the current block using the reconstructed image already obtained by performing the generation of the prediction block, the generation of the difference block, the generation of the coefficient block, the restoration of the difference block, and the generation of the decoded image block.

[0495] The reconstructed image can be, for example, an image of a reference picture or an image of a picture including a decoded block within a current picture, i.e., a current picture, that contains the current block. The decoded block within the current picture is, for example, a neighboring block of the current block.

[0496] Figure 44 is a flowchart of another example of processing performed by the prediction processing section of the decoding apparatus 200.

[0497] The prediction processing section determines a manner or mode for generating a prediction image (step Sr_1). The manner or mode can be determined, for example, based on a prediction parameter or the like.

[0498] In a case where it is determined that the first manner is the mode for generating the prediction image, the prediction processing section generates the prediction image in accordance with the first manner (step Sr_2a). Further, in a case where it is determined that the second manner is the mode for generating the prediction image, the prediction processing section generates the prediction image in accordance with the second manner (step Sr_2b). Further, in a case where it is determined that the third manner is the mode for generating the prediction image, the prediction processing section generates the prediction image in accordance with the third manner (step Sr_2c).

[0499] The first manner, the second manner, and the third manner are mutually different manners for generating a prediction image, and can be, for example, an inter prediction manner, an intra prediction manner, and another prediction manner. In such a prediction manner, the reconstructed image described above can also be used.

[0500] [Intra Prediction Section]

[0501] The intra prediction section 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to blocks within the current picture stored in the block memory 210 based on an intra prediction mode read out from the coded bitstream. Specifically, the intra prediction section 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, color difference values) of blocks neighboring the current block, and outputs the intra prediction signal to the prediction control section 220.

[0502] In addition, in a case where an intra prediction mode that refers to a luminance block is selected in the intra prediction of a color difference block, the intra prediction section 216 can also predict a color difference component of the current block based on a luminance component of the current block.

[0503] Further, in a case where information read out from the coded bitstream indicates that PDPC is adopted, the intra prediction section 216 corrects pixel values after intra prediction based on gradients of reference pixels in horizontal / vertical directions.

[0504] [Inter Prediction Section]

[0505] The inter prediction section 218 refers to the reference picture stored in the frame memory 214 to predict the current block. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction section 218 performs motion compensation using motion information (e.g., motion vectors) read from the coded bitstream (e.g., prediction parameters output from the entropy decoding section 202) to generate an inter prediction signal of the current block or sub-block, and outputs the inter prediction signal to the prediction control section 220.

[0506] In a case where the information read from the coded bitstream indicates that the OBMC mode is employed, the inter prediction section 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of the neighboring blocks.

[0507] Further, in a case where the information read from the coded bitstream indicates that the FRUC mode is employed, the inter prediction section 218 performs motion estimation according to the pattern matching method (bi-directional matching or template matching) read from the coded bitstream to derive the motion information. Then, the inter prediction section 218 performs motion compensation (prediction) using the derived motion information.

[0508] Further, in a case where the BIO mode is employed, the inter prediction section 218 derives a motion vector based on a model assuming constant velocity straight line motion. Further, in a case where the information read from the coded bitstream indicates that the affine motion compensation prediction mode is employed, the inter prediction section 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of neighboring blocks.

[0509] [MV derivation > normal inter mode]

[0510] In a case where the information read from the coded bitstream indicates that the normal inter mode is applied, the inter prediction section 218 derives an MV based on the information read from the coded bitstream, and performs motion compensation (prediction) using the MV.

[0511] Figure 45 is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding apparatus 200.

[0512] The inter prediction section 218 of the decoding apparatus 200 performs motion compensation for each block. The inter prediction section 218 acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded blocks temporally or spatially around the current block (step Ss_1). That is, the inter prediction section 218 makes a candidate MV list.

[0513] Next, the inter prediction section 218 extracts N (N is an integer of 2 or more) candidate MVs as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order from the plurality of candidate MVs acquired in step Ss_1 (step Ss_2). Note that the priority order can be determined in advance for each of the N prediction MV candidates.

[0514] Next, the inter prediction section 218 decodes the prediction motion vector selection information from the input stream (i.e., the encoded bitstream), and uses the decoded prediction motion vector selection information to select one of the N prediction MV candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Ss_3).

[0515] Next, the inter prediction section 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the differential value of the decoded differential MV to the selected prediction motion vector (step Ss_4).

[0516] Finally, the inter prediction section 218 generates the prediction image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).

[0517] [Prediction control section]

[0518] The prediction control section 220 selects one of the intra prediction signal and the inter prediction signal, and outputs the selected signal as the prediction signal to the addition section 208. In general, the structure, functions, and processing of the prediction control section 220, the intra prediction section 216, and the inter prediction section 218 on the decoding device side can correspond to the structure, functions, and processing of the prediction control section 128, the intra prediction section 124, and the inter prediction section 126 on the encoding device side.

[0519] [Installation example of decoding device]

[0520] Figure 46 is a block diagram showing an installation example of the decoding device 200. The decoding device 200 is provided with a processor b1 and a memory b2. For example, Figure 41 The plurality of constituent elements of the decoding device 200 shown in Figure 46 are installed by the processor b1 and the memory b2.

[0521] The processor b1 is a circuit that performs information processing, and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general electronic circuit that decodes an encoded moving image (i.e., an encoded bitstream). The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can be a collection of a plurality of electronic circuits. In addition, for example, the processor b1 can function asFigure 41 The memory b2 functions as a storage unit of the decoding device 200 illustrated in FIG. 1, for example. Specifically, the memory b2 functions as a storage unit of the decoding device 200 illustrated in FIG. 1.

[0522] The memory b2 is a dedicated or general-purpose memory that stores information used for the processor b1 to decode the coded bitstream. The memory b2 can be an electronic circuit, and can be connected to the processor b1. Alternatively, the memory b2 can be included in the processor b1. Alternatively, the memory b2 can be a collection of a plurality of electronic circuits. Alternatively, the memory b2 can be a magnetic disk or an optical disk, and can be a storage or a recording medium. Alternatively, the memory b2 can be a nonvolatile memory, or a volatile memory.

[0523] The memory b2 can store a moving image, or can store a coded bitstream, for example. Alternatively, a program for the processor b1 to decode the coded bitstream can be stored in the memory b2.

[0524] Alternatively, the memory b2 can function as a storage unit of the decoding device 200 illustrated in FIG. 1, for example. Figure 41 The memory b2 functions as a storage unit of the decoding device 200 illustrated in FIG. 1, for example. Specifically, the memory b2 functions as a storage unit of the decoding device 200 illustrated in FIG. 1. Figure 41 The memory b2 functions as the block memory 210 and the frame memory 214 illustrated in FIG. 1, for example. More specifically, the memory b2 can store a reconstructed block and a reconstructed picture, and the like.

[0525] Alternatively, not all of the plurality of components illustrated in FIG. 1 can be installed in the decoding device 200, and not all of the plurality of processes described above can be performed. Figure 41 Alternatively, not all of the plurality of components illustrated in FIG. 1 can be installed in the decoding device 200, and not all of the plurality of processes described above can be performed. Figure 41 Alternatively, part of the plurality of components illustrated in FIG. 1 can be included in another device, and part of the plurality of processes described above can be performed by another device.

[0526] [Definitions of Terms]

[0527] The terms can be defined as follows, for example.

[0528] A picture is an arrangement of a plurality of luma samples in monochrome format, or an arrangement of a plurality of luma samples and two corresponding arrangements of a plurality of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats. The picture can be a frame or a field.

[0529] A frame is a combination of a top field generating a plurality of sample lines 0, 2, 4, … and a bottom field generating a plurality of sample lines 1, 3, 5, ….

[0530] A slice contains an integer number of coding tree units of one independent slice segment, and, if any, all subsequent dependent slice segments up to, but not including, the next independent slice segment, if any, within the same access unit.

[0531] A tile is a rectangular region of a picture that is a specific tile column and a specific tile row within a plurality of coding tree blocks. A tile can still apply a loop filter across the edges of the tile, but can also be a rectangular region of a frame intended to be able to be independently decoded and encoded.

[0532] A block is an MxN (N rows by M columns) arrangement of a plurality of samples, or an MxN arrangement of a plurality of transform coefficients. A block can also be a square or rectangular region of a plurality of pixels consisting of 1 luma and 2 chroma matrices.

[0533] A CTU (coding tree unit) can be a coding tree block of a plurality of luma samples of a picture having 3 sample arrangements, or 2 corresponding coding tree blocks of a plurality of chroma samples. Alternatively, a CTU can be a coding tree block of any plurality of samples in a monochrome picture, and a picture encoded using a syntax structure used in encoding of 3 separate color planes and a plurality of samples.

[0534] A super block constitutes 1 or 2 mode information blocks, or can also be recursively split into 4 32x32 blocks, and further into a square block of 64x64 pixels that can be split.

[0535] [1st mode of coefficient encoding]

[0536] Figure 47 is a flowchart showing a basic coefficient encoding method of the 1st mode. Specifically, Figure 47 is a coefficient encoding method of a region for which a prediction residual is obtained by intra-frame encoding or inter-frame encoding. In the following description, actions performed by the encoding apparatus 100 are shown. The decoding apparatus 200 can perform actions corresponding to the actions performed by the encoding apparatus 100. For example, the decoding apparatus 200 can perform inverse orthogonal transformation and decoding corresponding to orthogonal transformation and encoding performed by the encoding apparatus 100.

[0537] In Figure 47 , last_sig_coeff, subblock_flag, thres, and CCB are shown. last_sig_coeff is a parameter indicating a coordinate position at which a coefficient (non-zero coefficient) not being zero first appears when scanning within a block. subblock_flag is a flag indicating whether there is a coefficient not being zero in a 4x4 subblock (also referred to as a 16 transform coefficient level). subblock_flag can also be expressed as coded_sub_block_flag or subblock flag.

[0538] thres is a constant determined in units of blocks. thres can be determined in advance. thres can be a value that differs depending on the size of a block, or can be the same value regardless of the size of a block. thres can take different values in the case where a orthogonal transform is applied and in the case where a orthogonal transform is not applied. thres can be determined depending on the coordinate position determined in a block by last_sig_coeff.

[0539] CCB represents the number of bins that are coded in a context mode of CABAC (Context-Adaptive Binary Arithmetic Coding). That is, CCB represents the number of times of processing of coding in the context mode of CABAC. The context mode is also referred to as a regular mode. Here, coding in the context mode of CABAC is referred to as CABAC coding or context-adaptive coding. In addition, coding in a bypass mode of CABAC is referred to as bypass coding. The processing of bypass coding is simpler than that of CABAC coding.

[0540] CABAC coding is processing of transforming a string of bins obtained by binarizing a signal to be coded into a string of coding bits, based on the occurrence probability of 0 and 1 for each bin. In addition, CCB can count the number of all flags used in residual coefficient coding, or can count the number of a part of flags used in residual coefficient coding. Bypass coding is processing of coding 1 bin in a string of bins as 1 bit of a string of coding bits without using the variable occurrence probability of 0 and 1 for each bin (in other words, using a fixed probability).

[0541] For example, the coding device 100 compares the CCB value and the thres value to decide the manner of coefficient coding.

[0542] Specifically, in the first embodiment, Figure 47 First, CCB is initialized to 0 (S101). Then, it is determined whether or not a orthogonal transform is applied to a block (S102). In the case where a orthogonal transform is applied to a block (Yes in S102), the coding device 100 codes last_sig_coeff (S131). Then, the coding device 100 performs loop processing for each subblock (S141 to S148).

[0543] In the loop processing for each subblock (S141 to S148), the coding device 100 codes subblock_flag related to the subblock. Then, in the case where subblock_flag is different from 0 (Yes in S146), the coding device 100 codes 16 coefficients in the subblock by the first coding manner described later (S147).

[0544] If orthogonal transform is not applied to the block (No in S102 ), the encoding apparatus 100 performs loop processing on each subblock ( S121 to S128 ).

[0545] During the loop processing (S121-S128) for each sub-block, the encoding device 100 determines whether the CCB is less than or equal to thres (S122). If the CCB is less than or equal to thres (Yes in S122), the encoding device 100 encodes the subblock_flag using CABAC encoding (S123). The encoding device 100 then increments the CCB (S124). Otherwise (No in S122), the encoding device 100 encodes the subblock_flag using bypass encoding (S125).

[0546] Then, when subblock_flag is different from 0 (Yes in S126 ), encoding apparatus 100 encodes the 16 coefficients in the subblock using a second encoding method described later ( S127 ).

[0547] The case where an orthogonal transform is not applied to a block can be, for example, when the orthogonal transform is skipped. CCBs are also used in the first and second coding modes. CCBs can be initialized per sub-block. In this case, thres can be a value that varies for each sub-block, rather than a fixed value within the block.

[0548] In addition, here, the CCB counts up from 0 and determines whether it reaches thres, but the CCB may count down from thres (or a specific value) and determine whether it reaches 0.

[0549] Figure 48 Yes Figure 47 Detailed flowchart of the first encoding method shown in FIG. In the first encoding method, multiple coefficients within a sub-block are encoded. In this case, a first loop (S151 to S156) is performed for each coefficient information flag of each coefficient within the sub-block, and a second loop (S161 to S165) is performed for each coefficient within the sub-block.

[0550] In the first loop process (S151 to S156), one or more coefficient information flags each indicating one or more properties of a coefficient are sequentially encoded. The one or more coefficient information flags can include sig_flag, gt1_flag, parity_flag, and gt3_flag described later. Then, within a range where CCB does not exceed thres, the one or more coefficient information flags are sequentially encoded by CABAC encoding, and CCB is incrementally counted by one each time of encoding. After CCB exceeds thres, the coefficient information flags are not encoded.

[0551] That is, in the first loop process (S151 to S156), the encoding apparatus 100 determines whether CCB is equal to or less than thres (S152). Then, in a case where CCB is equal to or less than thres (Yes in S152), the encoding apparatus 100 encodes the coefficient information flags by CABAC encoding (S153). Then, the encoding apparatus 100 incrementally counts CCB (S154). In a case where CCB is not equal to or less than thres (No in S152), the encoding apparatus 100 ends the first loop process (S151 to S156).

[0552] In the second loop process (S161 to S165), for a coefficient for which the coefficient information flag is encoded, remainder, which is a remaining value of a value for reconstructing the coefficient using the coefficient information flag (i.e., a remaining value of a value for which the coefficient information flag is used), is encoded by Golomb encoding. A coefficient for which the coefficient information flag is not encoded is directly encoded by Golomb encoding. Alternatively, remainder can be encoded using another encoding method instead of Golomb encoding.

[0553] That is, in the second loop process (S161 to S165), the encoding apparatus 100 determines whether the coefficient information flag corresponding to the coefficient of the processing target has been encoded (S162). Then, in a case where the coefficient information flag is encoded (Yes in S162), the encoding apparatus 100 encodes remainder by Golomb encoding (S163). In a case where the coefficient information flag is not encoded (No in S162), the encoding apparatus 100 encodes the value of the coefficient by Golomb encoding (S164).

[0554] Note that the number of loop processes is two in this example, but the number of loop processes can be different from two.

[0555] The sig_flag described above is a flag indicating whether the AbsLevel is nonzero. The AbsLevel is the value of the coefficient, more specifically, the absolute value of the coefficient. The gt1_flag is a flag indicating whether the AbsLevel is greater than 1. The parity_flag is a flag of the 1st bit of the AbsLevel, and is a flag indicating whether the AbsLevel is odd or even. The gt3_flag is a flag indicating whether the AbsLevel is greater than 3.

[0556] The gt1_flag and the gt3_flag are sometimes expressed as abs_gt1_flag and abs_gt3_flag, respectively. Further, for example, as the remainder described above, the value of (Abslevel - 4) / 2 can be encoded by Golomb coding.

[0557] Other one or more coefficient information flags different from the one or more coefficient information flags described above can also be encoded. For example, a part of the coefficient information flags can not be encoded. The coefficient information flags included in the one or more coefficient information flags described above can be replaced with coefficient information flags or parameters having other meanings.

[0558] Figure 49 is a flag indicating Figure 47 a flowchart of details of the second encoding mode shown in FIG. 2. In the second encoding mode, a plurality of coefficients within a subblock are encoded. At this time, a first loop process (S171 to S176) is performed for each coefficient information flag of each coefficient within the subblock, and a second loop process (S181 to S185) is performed for each coefficient within the subblock.

[0559] In the first loop process (S171 to S176), one or more coefficient information flags respectively indicating one or more properties of the coefficient are sequentially encoded. The one or more coefficient information flags can include a sig_flag, a sign_flag, a gt1_flag, a parity_flag, a gt3_flag, a gt5_flag, a gt7_flag, and a gt9_flag.

[0560] Here, the sign_flag is a flag indicating the sign of the coefficient. The gt5_flag is a flag indicating whether the AbsLevel is greater than 5. The gt7_flag is a flag indicating whether the AbsLevel is greater than 7. The gt9_flag is a flag indicating whether the AbsLevel is greater than 9. The gt5_flag, the gt7_flag, and the gt9_flag are sometimes expressed as abs_gt5_flag, abs_gt7_flag, and abs_gt9_flag, respectively. In addition, a flag indicating whether the AbsLevel is greater than x (x is an integer of 1 or more) can be collectively expressed as gtx_flag or abs_gtx_flag. The AbsLevel is, for example, the absolute value of the transform coefficient level.

[0561] In addition, other one or more coefficient information flags different from the above-described one or more coefficient information flags can be encoded. For example, a part of the coefficient information flags can not be encoded. The coefficient information flags included in the above-described one or more coefficient information flags can be replaced with coefficient information flags or parameters having other meanings.

[0562] The above-described one or more coefficient information flags are sequentially encoded by CABAC encoding. Then, the CCB is incrementally counted by 1 each time of encoding. After the CCB exceeds the thres, the coefficient information flags are encoded by bypass encoding.

[0563] That is, in the first loop process (S171 to S176), the encoding apparatus 100 determines whether the CCB is equal to or less than the thres (S172). Then, in a case where the CCB is equal to or less than the thres (Yes in S172), the encoding apparatus 100 encodes the coefficient information flags by CABAC encoding (S173). Then, the encoding apparatus 100 incrementally counts the CCB (S174). In a case where the CCB is not equal to or less than the thres (No in S172), the encoding apparatus 100 encodes the coefficient information flags by bypass encoding (S175).

[0564] Figure 49 The syntax of the second loop process in the above-described (1) is not changed before and after the CCB exceeds the thres. That is, whether the coefficient information flags are encoded by CABAC encoding or the coefficient information flags are encoded by bypass encoding, the following same process is performed in the second loop process (S181 to S185).

[0565] Specifically, in the 2nd loop processing (S181-S185), the encoding device 100 encodes remainder, which is a residual value that cannot be represented by the coefficient information flag (i.e., a residual value for reconstructing the coefficient value using the coefficient information flag), by the Golomb encoding (S183). In addition, remainder can be encoded by another encoding method instead of the Golomb encoding.

[0566] In addition, the number of loop processes is 2 here, but the number of loop processes can be different from 2.

[0567] As shown in Figure 47 , Figure 48 and Figure 49 , in the basic operation of the present aspect, there are different flags included in the number of times of CABAC encoding depending on whether or not the orthogonal transform is applied. In addition, the syntax of the coefficient encoding is different between the case where the orthogonal transform is applied and the case where the orthogonal transform is not applied. Thus, it can be necessary to prepare circuits separately. Therefore, the circuit structure can become complicated.

[0568] [1st example of 1st aspect of coefficient encoding]

[0569] Figure 50 is a flowchart showing the coefficient encoding method of the 1st example of the 1st aspect. In Figure 50 , the post-processing of last_sig_coeff (S132) and the processing of subblock_flag (S142-S145) are different from those of Figure 47 .

[0570] In Figure 47 , in the case where the orthogonal transform is applied, CCB counts sig_flag, parity_flag, and gtX_flag (X=1, 3) incrementally. In Figure 50 , CCB counts last_sig_coeff and subblock_flag incrementally as well. On the other hand, the processing flow in the case where the orthogonal transform is not applied is the same as that of Figure 47 .

[0571] That is, in the example of Figure 50 , the encoding device 100 encodes last_sig_coeff (S131), and then adds the number of times of CABAC encoding in the encoding of last_sig_coeff to CCB (S132).

[0572] Further, the encoding device 100 determines whether the CCB is equal to or lower than the thres before encoding the subblock_flag (S142). Then, in a case where the CCB is equal to or lower than the thres (Yes in S142), the encoding device 100 encodes the subblock_flag by CABAC encoding (S143). Then, the encoding device 100 adds 1 to the CCB (S144). On the other hand, in a case where the CCB is not equal to or lower than the thres (No in S142), the encoding device 100 encodes the subblock_flag by bypass encoding (S145).

[0573] [Effects of the first example of the first mode of coefficient encoding]

[0574] According to Figure 50 , in a case where the orthogonal transform is performed and in a case where the orthogonal transform is not performed, it is sometimes possible to commonize and unify the processing flow of the encoding of the subblock_flag. Therefore, in the case where the orthogonal transform is performed and in the case where the orthogonal transform is not performed, it is possible to share a part of the circuit, and it is possible to reduce the circuit size. As a result, in addition to the last_sig_coeff, the processing flows divided according to the presence or absence of the orthogonal transform can become the same.

[0575] For example, even in a case where the number of times of CABAC encoding is limited at the block level, in the example of Figure 47 , the subblock_flag is encoded by CABAC encoding after the CCB reaches the thres. On the other hand, in the case of Figure 50 , the subblock_flag is not encoded by CABAC encoding after the CCB reaches the thres. Thus, the number of times of CABAC encoding can be appropriately limited to the thres.

[0576] Further, the number of times of CABAC encoding in the last_sig_coeff can not be included in the CCB. Further, the thres can be determined depending on the coordinate position determined in the block by the last_sig_coeff.

[0577] Further, the encoding device 100 can also encode the subblock_flag with the value always determined as 1 after the CCB exceeds the thres. In a case where the value of the subblock_flag is always determined as 1 after the CCB exceeds the thres, the encoding device 100 can not encode the subblock_flag after the CCB exceeds the thres.

[0578] Further, even in the case where the orthogonal transform is not applied, the encoding apparatus 100 can always determine the value of the subblock_flag as 1 and encode it after the CCB exceeds the thres. Further, in this case, in the case where the value of the subblock_flag is always determined as 1 after the CCB exceeds the thres, the encoding apparatus 100 can not encode the subblock_flag after the CCB exceeds the thres.

[0579] [Example 2 of the 1st Mode of the Coefficient Encoding]

[0580] Figure 51 is a flowchart showing the coefficient encoding method of Example 2 of the 1st Mode. In Figure 51 , the processing of the subblock_flag (S123) is different from that of Figure 47 .

[0581] In Figure 47 , in the case where the orthogonal transform is not applied, the CCB has already been incremented for the sig_flag, the parity_flag, the gtX_flag (X = 1, 3, 5, 7, 9), and the subblock_flag. In Figure 51 , the CCB is not incremented for the subblock_flag. On the other hand, the processing flow in the case where the orthogonal transform is applied is the same as that of Figure 47 .

[0582] That is, in the example of Figure 51 , the encoding apparatus 100 always encodes the subblock_flag by the CABAC encoding regardless of whether the CCB exceeds the thres or not, without incrementing the CCB (S123).

[0583] [Effects of Example 2 of the 1st Mode of the Coefficient Encoding]

[0584] According to the example of Figure 51 , in the case where the orthogonal transform is performed and in the case where the orthogonal transform is not performed, it is sometimes possible to commonize and unify the processing of the encoding of the subblock_flag. Therefore, in the case where the orthogonal transform is performed and in the case where the orthogonal transform is not performed, it is possible that a part of the circuit is commonized, and it is possible to reduce the circuit size. As a result, except for the last_sig_coeff, it is possible that the plurality of processing flows divided according to the presence or absence of the orthogonal transform become the same.

[0585] Further, in Figure 50 , compared to Figure 51This simplifies processing. Therefore, it is possible to reduce circuit size. Furthermore, it is assumed that the frequency of occurrence of 0 or 1 associated with subblock_flag is likely to vary depending on surrounding conditions. Therefore, in CABAC encoding of subblock_flag, it is assumed that the reduction in the amount of code is greater than the increase in processing delay. Therefore, it is useful to limit the number of CABAC encoding operations so that CABAC encoding of subblock_flag is not counted towards the number of CABAC encoding operations.

[0586] In addition, the number of CABAC encoding processes for last_sig_coeff may be included in the CCB. In addition, thres may be determined depending on the coordinate position determined by last_sig_coeff in the block.

[0587] [Second form of coefficient coding]

[0588] [First example of the second aspect of coefficient encoding]

[0589] Figure 52 This is a flowchart showing the coefficient encoding method of the first example of the second aspect. Figure 52 In the example of FIG. 1 , even when no orthogonal transform is applied to the block, the case where the 16 coefficients in the sub-block are encoded by the first encoding method (S127a) is the same as the case where the 16 coefficients in the sub-block are encoded by the first encoding method (S127a). Figure 47 The examples are different.

[0590] That is, in Figure 52 In the example of Figure 48 The first encoding method shown is not Figure 49 The 16 coefficients in the sub-block are encoded using the second encoding method shown in FIG. 127a. That is, whether or not orthogonal transformation is applied, the encoding apparatus 100 does not encode the 16 coefficients in the sub-block using the second encoding method shown in FIG. Figure 49 The second encoding method shown is through Figure 48 The first coding method shown encodes 16 coefficients in a sub-block.

[0591] More specifically, regardless of whether or not there is an orthogonal transform, the encoding apparatus 100 follows Figure 48 In the first encoding method shown, in the first processing loop, bypass coding is not used and the encoding of the coefficient information flag is skipped when the CCB exceeds the thres. Then, in the second processing loop, if the coefficient information flag corresponding to the coefficient to be processed has not been encoded, the encoding device 100 encodes the coefficient value using Golombrite coding without using the coefficient information flag.

[0592] In addition, Figure 48In the 1st cycle processing of the 1st embodiment, the syntax for encoding the coefficient information flag can be different between the case where the orthogonal transform is applied and the case where the orthogonal transform is not applied. For example, between the one or more coefficient information flags in the case where the orthogonal transform is applied and the one or more coefficient information flags in the case where the orthogonal transform is not applied, some or all of the coefficient information flags can also be different.

[0593] [Effect of the 1st example of the 2nd form of coefficient encoding]

[0594] According to Figure 52 the example, even in the case where the encoding syntax of the coefficient information flag is different depending on the presence or absence of the orthogonal transform, the encoding syntax of the 16 coefficients within the sub-block after the CCB exceeds the thres can be commonized regardless of the presence or absence of the orthogonal transform. Therefore, in the case where the orthogonal transform is applied and in the case where the orthogonal transform is not applied, it is possible that some of the circuits are commonized, and it is possible to reduce the circuit size.

[0595] Further, after the CCB exceeds the thres, the coefficients are no longer encoded into the coefficient information flag encoded by the bypass encoding and the residual value information encoded by the Golomb encoding. Therefore, it is possible to suppress the increase in the amount of information, and it is possible to suppress the increase in the amount of encoding.

[0596] [2nd example of the 2nd form of coefficient encoding]

[0597] Figure 53 is a flowchart showing the coefficient encoding method of the 2nd example of the 2nd form. In Figure 53 the example, even in the case where the orthogonal transform is applied to the block, the case where the 16 coefficients within the sub-block are encoded by the 2nd encoding method (S147a) is different from Figure 47 the example.

[0598] That is, in Figure 53 the example, in the case where the orthogonal transform is applied to the block, the encoding device 100 does not encode the 16 coefficients within the sub-block using the 1st encoding method shown in Figure 48 but uses the 2nd encoding method shown in Figure 49 (S147a). That is, in the case where the orthogonal transform is applied and in the case where the orthogonal transform is not applied, the encoding device 100 does not encode the 16 coefficients within the sub-block using the 1st encoding method shown in Figure 48 but uses the 2nd encoding method shown in Figure 49 .

[0599] More specifically, the encoding device 100, regardless of the presence or absence of the orthogonal transform, encodes the 16 coefficients within the sub-block using the 2nd encoding method shown in Figure 49In the second coding mode shown, in the first loop process, in a case where the CCB exceeds the thres, the coefficient information flag is not skipped from coding but is coded by the bypass coding. Then, in the second loop process, the coding device 100 codes the remainder dependent on the coefficient information flag by the Golomb coding.

[0600] Further, in Figure 49 The syntax for coding the coefficient information flag in the first loop process of the first coding mode can be different between a case where the orthogonal transform is applied and a case where the orthogonal transform is not applied. For example, one or more of the coefficient information flags in the case where the orthogonal transform is applied and one or more of the coefficient information flags in the case where the orthogonal transform is not applied can also be different.

[0601] [Effects of the second example of the second form of coefficient coding]

[0602] According to Figure 53 In the example of the first coding mode, even in a case where the syntax of coding the coefficient information flag is different depending on the presence or absence of the orthogonal transform, the syntax of coding the 16 coefficients within the sub-block after the CCB exceeds the thres can be commonized regardless of the presence or absence of the orthogonal transform. Therefore, in the case where the orthogonal transform is applied and in the case where the orthogonal transform is not applied, a part of the circuit can be commonized, and the circuit size can be reduced.

[0603] [Third form of coefficient coding]

[0604] Figure 54 is a syntax diagram showing the first coding mode of the third form. Figure 54 The syntax shown corresponds to Figure 47 is an example of the syntax of the first coding mode shown. Basically, the first coding mode is used in the case where the orthogonal transform is applied.

[0605] Here, the coefficient information flags and the parameters are the same as those shown in the first form. Further, the plurality of coefficient information flags shown here is an example, and other plurality of coefficient information flags can also be coded. For example, a part of the coefficient information flags can not be coded. Further, the coefficient information flags shown here can be replaced with coefficient information flags or parameters having other meanings.

[0606] Figure 54 The initial for loop in the example of the first coding mode corresponds to Figure 48The first loop process in the example of the first mode. In the first for loop, if CCB remains, i.e., if CCB does not exceed the threshold, the coefficient information flag such as sig_flag is encoded by CABAC encoding. If CCB does not remain, the coefficient information flag is not encoded. In addition, in the present example, the second example of the second mode can also be applied. That is, if CCB does not remain, the coefficient information flag can be encoded by bypass encoding.

[0607] The second for loop from the top and the third for loop from the top correspond to Figure 48 The second loop process in the example of the first mode. In the second for loop from the top, for the coefficient for which the coefficient information flag has been encoded, the residual value is encoded by Golomb encoding. In the third for loop from the top, for the coefficient for which the coefficient information flag has not been encoded, the coefficient is encoded by Golomb encoding. In addition, by applying the second example of the second mode to the present example, the residual value can always be encoded by Golomb encoding.

[0608] In the fourth for loop from the top, sign_flag is encoded by bypass encoding.

[0609] The syntax explained in the present mode can be applied to Figure 47 , Figure 48 , Figure 50 , Figure 51 and Figure 52 each example of the first mode.

[0610] Figure 55 is a syntax diagram indicating the basic second encoding mode of the third mode. Figure 55 The syntax shown in FIG. 8 corresponds to an example of the syntax of the second encoding mode of Figure 47 . Basically, the second encoding mode is used without applying orthogonal transformation.

[0611] Furthermore, the plurality of coefficient information flags shown here is an example, and other plurality of coefficient information flags can be encoded. For example, a part of the coefficient information flags can not be encoded. Furthermore, the coefficient information flags shown here can be replaced with coefficient information flags or parameters having other meanings.

[0612] Figure 55 The first five for loops in the example of the first mode correspond to Figure 49The first loop in the example of is processed. If CCBs remain in the first five for loops, that is, if the CCBs do not exceed the threshold, coefficient information flags such as sig_flag are encoded using CABAC coding. If no CCBs remain, the coefficient information flags are encoded using bypass coding. Furthermore, this example can also be applied to the first example of the second aspect. That is, if no CCBs remain, the coefficient information flags do not need to be encoded.

[0613] The sixth for loop from the top corresponds to Figure 49 The second loop in the example is processed. In the sixth for loop from the top, the residual value is encoded using Golombirai coding. Furthermore, in this example, the first example of the second aspect can also be applied. That is, for coefficients whose coefficient information flags are encoded, the residual value can be encoded using Golombirai coding. Then, for coefficients whose coefficient information flags are not encoded, the coefficients can be encoded using Golombirai coding.

[0614] The syntax described in this form can be applied to Figure 47 、 Figure 49 、 Figure 50 、 Figure 51 and Figure 53 Each example in .

[0615] Without applying orthogonal transformation ( Figure 55 ), the number of loop processes for encoding the coefficient information flag is greater than that in the case of applying orthogonal transform ( Figure 54 ). Therefore, when orthogonal transform is not applied, the hardware processing load may increase compared to when orthogonal transform is applied. In addition, since the syntax of coefficient coding differs depending on whether or not an orthogonal transform is performed, circuits may need to be prepared separately. As a result, the circuitry may become complex.

[0616] [First example of the third aspect of coefficient encoding]

[0617] Figure 56 This is a syntax diagram showing the second encoding method of the first example of the third aspect. Figure 56 The syntax shown corresponds to Figure 47 An example of the second encoding method. Figure 47 The first encoding method can also be used Figure 54 In addition, this example can be combined with other examples of the third form, and can also be combined with other forms.

[0618] Figure 56 The initial for loop in the example corresponds to Figure 49The first loop in the example is processed. If, in this initial for loop, 8 or more CCBs remain, that is, if the CCB after adding 8 does not exceed the threshold, up to 8 coefficient information flags are encoded using CABAC encoding for each coefficient, and the CCB is incremented up to 8 times. If the CCB does not remain at least 8, that is, if the CCB after adding 8 exceeds the threshold, 8 coefficient information flags are encoded using bypass encoding for each coefficient.

[0619] In other words, before encoding the eight coefficient information flags, it is comprehensively determined whether the eight coefficient information flags can be encoded using CABAC encoding. Then, if the eight coefficient information flags can be encoded using CABAC encoding, a maximum of eight coefficient information flags are encoded using CABAC encoding.

[0620] Furthermore, in this example, the first example of the second aspect can also be applied. That is, if there are not more than eight CCBs remaining, the eight coefficient information flags do not need to be encoded. In other words, in this case, rather than encoding the eight coefficient information flags through bypass coding, the encoding of the eight coefficient information flags can be skipped.

[0621] In addition, if Figure 56 As shown, the coding of one or more coefficient information flags among the eight coefficient information flags may be omitted depending on the value of the coefficient. For example, when sig_flag is 0, the coding of the remaining seven coefficient information flags may be omitted.

[0622] The second for loop from the top corresponds to Figure 49 The second loop in the example is processed. In the second for loop from the top, the residual value is encoded using Golombre coding. Furthermore, in this example, the first example of the second aspect can also be applied. That is, for the coefficients whose eight coefficient information flags are encoded, the residual value can be encoded using Golombre coding, and for the coefficients whose eight coefficient information flags are not encoded, the coefficient can be encoded using Golombre coding.

[0623] The multiple coefficient information flags shown here are examples, and other multiple coefficient information flags may be encoded. For example, some coefficient information flags may not be encoded. Furthermore, the coefficient information flags shown here may be replaced with coefficient information flags or parameters having other meanings.

[0624] In addition, you can also combine Figure 55 Examples and Figure 56 For example, in Figure 55In the example, before the four coefficient information flags such as sig_flag and sign_flag are encoded, it can be comprehensively determined whether the four coefficient information flags can be encoded by CABAC encoding.

[0625] [Effect of the first example of the third aspect of coefficient encoding]

[0626] exist Figure 56 In the example of , all coefficient information flags encoded by CABAC encoding are encoded in one loop process. Figure 55 Compared to the example in Figure 56 In the example above, the number of loop processes is small. Therefore, the processing volume may be reduced.

[0627] In addition, Figure 54 Examples and Figure 56 In the example of , the number of loop processes for encoding a plurality of coefficient information flags by CABAC encoding is the same. Figure 54 Examples and Figure 55 Compared to the combination of examples in Figure 54 Examples and Figure 56 In the combination of the examples, the number of circuit changes may be reduced.

[0628] Furthermore, before a plurality of coefficient information flags are encoded, it is collectively determined whether or not the plurality of coefficient information flags can be encoded by CABAC encoding. This simplifies processing and may reduce processing delay.

[0629] Furthermore, whether or not the plurality of coefficient information flags can be encoded using CABAC coding can be determined collectively before the plurality of coefficient information flags are encoded, both in the case where orthogonal transform is applied and in the case where orthogonal transform is not applied. Consequently, the difference between the encoding method used in a block to which orthogonal transform is applied and the encoding method used in a block not to which orthogonal transform is applied is further reduced, and the circuit scale can be further reduced.

[0630] In addition, Figure 56 In the example, sig_flag to abs_gt9_flag are included in one loop, but the encoding method is not limited to this. Multiple loops (for example, two loops) can be used, and it can also be determined in general whether it is possible to encode multiple coefficient information flags of each loop by CABAC encoding. Compared with one loop, the processing increases, but it is the same as Figure 55 Compared with the example, the same processing reduction effect can be obtained.

[0631] [Second example of the third aspect of coefficient encoding]

[0632] Figure 57is a syntax diagram of the second encoding method of the second example of the third modality. Figure 57 The syntax shown corresponds to Figure 47 an example of the second encoding method of Figure 47 The syntax shown can also be used in the first encoding method of Figure 54 In addition, the present example can be combined with other examples of the third modality, or can be combined with other modalities.

[0633] Figure 57 The initial for loop in the example of Figure 49 The first loop process in the example of If the CCB remains 7 or more in the initial for loop, that is, if the CCB added to 7 does not exceed the threshold value, the coefficient information flags for up to 7 coefficients are encoded by CABAC encoding, and the CCB is incremented by up to 7. If the CCB does not remain 7 or more, that is, if the CCB added to 7 exceeds the threshold value, the 7 coefficient information flags are encoded by bypass encoding.

[0634] In other words, before the 7 coefficient information flags are encoded, it is determined in general whether it is possible to encode the 7 coefficient information flags by CABAC encoding. Then, in the case where it is possible to encode the 7 coefficient information flags by CABAC encoding, the 7 coefficient information flags are encoded by CABAC encoding.

[0635] In addition, in the present example, the first example of the second modality can also be applied. That is, if the CCB remains 7 or more, the 7 coefficient information flags can also not be encoded. That is, in this case, the 7 coefficient information flags can also be skipped without being encoded by bypass encoding.

[0636] In addition, as shown in Figure 57 the encoding of one or more of the 7 coefficient information flags can also be omitted according to the values of the coefficients. For example, in the case where the sig_flag is 0, the encoding of the remaining 6 coefficient information flags can be omitted.

[0637] The second for loop from the top corresponds to the second loop process in the example of Figure 49 In the second for loop from the top, the remaining values are encoded by Golomb-Rice encoding. In addition, in the present example, the first example of the second modality can also be applied. That is, for the coefficients for which the 7 coefficient information flags are encoded, the remaining values can be encoded by Golomb-Rice encoding, and for the coefficients for which the 7 coefficient information flags are not encoded, the coefficients can be encoded by Golomb-Rice encoding.

[0638] In the 3rd for loop from above, if CCB remains, that is, CCB does not exceed the threshold, sign_flag is encoded by CABAC encoding, and CCB is incremented. If CCB does not remain, sign_flag is encoded by bypass encoding. Further, as in the example of Figure 54 , sign_flag can be always encoded by bypass encoding.

[0639] The plurality of coefficient information flags shown here is an example, and other plurality of coefficient information flags can be encoded. For example, a part of the coefficient information flags can not be encoded. Further, the coefficient information flags shown here can be replaced with coefficient information flags or parameters having other meanings.

[0640] In addition, the example of Figure 55 and the example of Figure 57 may be combined. For example, in the example of Figure 55 , before the 4 coefficient information flags of sig_flag and sign_flag and the like are encoded, it can be generally determined whether it is possible to encode the 4 coefficient information flags by CABAC encoding.

[0641] [Effects of the 2nd example of the 3rd form of coefficient encoding]

[0642] As in the example of Figure 56 , in the example of Figure 57 , the plurality of coefficient information flags (specifically, abs_gt3_flag and abs_gt5_flag and the like) for representing the size of the coefficient by comparing with the threshold are encoded in 1 loop processing. Therefore, compared with the example of Figure 55 , in the example of Figure 57 , the number of loop processing is small. Therefore, the processing amount can be reduced.

[0643] Compared with the example of Figure 56 , in the example of Figure 57 , the number of loop processing for encoding the plurality of coefficient information flags increases. However, compared with the example of Figure 56 , in the example of Figure 57 , there is a similar part to the example of Figure 54 . For example, sign_flag has been finally encoded. Therefore, compared with the combination of the example of Figure 54 and the example of Figure 56 , in the combination of the example of Figure 54 and the example of Figure 57 , the changed part of the circuit can be less.

[0644] In addition, before the plurality of coefficient information flags are encoded, it is determined in general whether it is possible to encode the plurality of coefficient information flags by CABAC encoding, so the processing is simplified, and the processing delay can be reduced.

[0645] In both cases of the case where the orthogonal transform is applied and the case where the orthogonal transform is not applied, it is also possible to determine in general whether it is possible to encode the plurality of coefficient information flags by CABAC encoding before the plurality of coefficient information flags are encoded. Thus, the difference between the encoding method used in the block where the orthogonal transform is applied and the encoding method used in the block where the orthogonal transform is not applied is further reduced, and the circuit size can be further reduced.

[0646] Further, in the example of Figure 57 From sig_flag to abs_gt9_flag is included in 1 loop, but the encoding method is not limited to this. A plurality of loops (for example, 2 loops) can be used, and it is also possible to determine in general whether it is possible to encode the plurality of coefficient information flags of each loop by CABAC encoding. Compared with 1 loop, the processing is increased, but compared with the example of Figure 55 The effect of processing reduction can also be obtained.

[0647] [Variation of coefficient encoding]

[0648] Any of the plurality of modes and the plurality of examples related to the above-described coefficient encoding can be combined. Further, any of the plurality of modes, the plurality of examples, and any of the plurality of combinations thereof related to the above-described coefficient encoding can be applied to the block of the luminance or the block of the color difference. At this time, different thres can be used in the block of the luminance and the block of the color difference.

[0649] In addition, any of the plurality of modes, the plurality of examples, and any of the plurality of combinations thereof related to the above-described coefficient encoding can be used for a block that is a block where the orthogonal transform is not applied and is a block where BDPCM (Block-based Delta Pulse Code Modulation) is applied. In the block where the BDPCM is applied, each residual signal within the block is reduced in the amount of information by subtracting a residual signal vertically or horizontally adjacent to the residual signal from the residual signal.

[0650] Further, any of the plurality of modes, the plurality of examples, and any of the plurality of combinations thereof related to the above-described coefficient encoding can also be used for a block that is a block where the BDPCM is applied and is a block of the color difference.

[0651] Furthermore, for a block to which ISP (Intra Sub-Partitions) is applied, one of the plurality of modes, the plurality of examples, and the arbitrary combination of them related to the coefficient coding described above can be used. In ISP, an intra block is divided vertically or horizontally, and the pixel values of the sub-blocks adjacent to each sub-block are used for intra prediction of the sub-block.

[0652] Furthermore, for a block to which ISP is applied and which is a color difference block, one of the plurality of modes, the plurality of examples, and the arbitrary combination of them related to the coefficient coding described above can be used.

[0653] Furthermore, in a case where Chroma Joint Coding is used as an encoding mode of a color difference block, one of the plurality of modes, the plurality of examples, and the arbitrary combination of them related to the coefficient coding described above can be used. Here, Chroma Joint Coding is an encoding method in which the value of Cr is derived from the value of Cb.

[0654] Furthermore, the value of thres in a case where the orthogonal transform is applied can be twice the value of thres in a case where the orthogonal transform is not applied. Alternatively, the value of thres in a case where the orthogonal transform is not applied can be twice the value of thres in a case where the orthogonal transform is applied.

[0655] Furthermore, the value of thres of CCB in a case where the orthogonal transform is applied can be twice the value of thres of CCB in a case where the orthogonal transform is not applied, only in a case where Chroma Joint Coding is used. Alternatively, the value of thres of CCB in a case where the orthogonal transform is not applied can be twice the value of thres of CCB in a case where the orthogonal transform is applied, only in a case where Chroma Joint Coding is used.

[0656] In addition, among the plurality of modes and the plurality of examples related to the coefficient coding described above, the scan order of the plurality of coefficients in a block in which the orthogonal transform is not applied can be the same as the scan order of the plurality of coefficients in a block in which the orthogonal transform is applied.

[0657] Furthermore, although several examples of syntax are shown in the 3rd mode and the plurality of examples of the 3rd mode, the syntax to be applied is not limited to these examples. For example, in a plurality of modes different from the 3rd mode and in a plurality of examples of them, a syntax different from any one of the plurality of syntaxes shown in the 3rd mode and the plurality of examples thereof can be used. Various syntaxes for encoding 16 coefficients can be applied.

[0658] In addition, the various forms and examples of coefficient coding illustrate the encoding process flow, but the decoding process flow is essentially the same as the encoding process flow, except for the difference in whether the bit stream is transmitted or received. For example, decoding apparatus 200 can perform inverse orthogonal transform and decoding corresponding to the orthogonal transform and encoding performed by encoding apparatus 100.

[0659] Furthermore, each flowchart related to a plurality of forms and a plurality of examples of coefficient coding is merely an example, and conditions or processes may be newly added to each flowchart, deleted, or modified.

[0660] Here, coefficients are values ​​that constitute an image, such as a block or sub-block. Specifically, the multiple coefficients that constitute an image can be obtained from multiple pixel values ​​in the image via an orthogonal transform. Alternatively, the multiple coefficients that constitute an image can be obtained from multiple pixel values ​​in the image without undergoing an orthogonal transform. In other words, the multiple coefficients that constitute an image can be the multiple pixel values ​​themselves. Furthermore, each pixel value can be a pixel value of the original image or a prediction residual value. Furthermore, the coefficients can be quantized.

[0661] [Representative examples of structure and processing]

[0662] Representative examples of the configuration and processing of the encoding apparatus 100 and decoding apparatus 200 described above are described below.

[0663] Figure 58 is a flowchart showing the operation of the encoding device 100. For example, the encoding device 100 includes a circuit and a memory connected to the circuit. The circuit and memory included in the encoding device 100 may also correspond to Figure 40 The circuit of the encoding device 100 performs Figure 58 Specifically, the circuit of the encoding device 100 encodes a block of an image during operation ( S211 ).

[0664] In one example, the circuitry of encoding apparatus 100 may encode a block of an image by limiting the number of context-adaptive encoding processes. Furthermore, the sub-block flag encoding process may be performed without including it in the number of processes subject to the limitation, both when an orthogonal transform is applied to the block and when an orthogonal transform is not applied to the block.

[0665] Here, the subblock flag encoding process is a process of encoding a subblock flag indicating whether a subblock included in a block includes a non-zero coefficient through context adaptive encoding.

[0666] Thus, the sub-block flag can be encoded by the context adaptive coding regardless of whether or not the orthogonal transform is applied and regardless of whether or not the number of times of the processing is limited. Therefore, the amount of encoding can be reduced. Further, a difference between the encoding method used in the block in which the orthogonal transform is applied and the encoding method used in the block in which the orthogonal transform is not applied can be made small, and the circuit size can be made small.

[0667] Further, the circuit of the encoding apparatus 100 can perform the position parameter encoding processing without including the position parameter encoding processing in the number of times of the processing to be limited, in a case where the orthogonal transform is applied to the block. Here, the position parameter encoding processing is processing of encoding a parameter indicating a position of a first non-zero coefficient in the block in the scanning order by the context adaptive coding.

[0668] Thus, in the case where the orthogonal transform is applied, the parameter indicating the position of the first non-zero coefficient can be encoded by the context adaptive coding regardless of whether or not the number of times of the processing is limited. Therefore, the amount of encoding can be reduced.

[0669] Further, the circuit of the encoding apparatus 100 can determine the range of the number of times of the processing to be limited in accordance with the position of the first non-zero coefficient, in the case where the orthogonal transform is applied to the block. Thus, in the case where the orthogonal transform is applied, it is possible to appropriately determine the number of times of the processing to be limited. Therefore, it is possible to appropriately adjust a balance between reduction of the amount of encoding and reduction of the processing delay.

[0670] In another example, the circuit of the encoding apparatus 100 can act as follows in both a case where the orthogonal transform is applied to the block of the encoding target image and a case where the orthogonal transform is not applied to the block.

[0671] Specifically, in both the cases, in a case where the number of times of the context adaptive coding is within the range of the number of times of the processing to be limited, the circuit of the encoding apparatus 100 can encode the coefficient information flag by the context adaptive coding. Further, in both the cases, in a case where the number of times of the context adaptive coding is not within the range of the number of times of the processing to be limited, the circuit of the encoding apparatus 100 can also skip the encoding of the coefficient information flag. Here, the coefficient information flag indicates a property of the coefficient included in the block.

[0672] Then, in a case where the coefficient information flag is encoded, the circuit of the encoding apparatus 100 can encode the residual value information by the Golomb coding. Here, the residual value information is information for reconstructing a value of the coefficient using the coefficient information flag. Further, in a case where the encoding of the coefficient information flag is skipped, the circuit of the encoding apparatus 100 can encode the value of the coefficient by the Golomb coding.

[0673] Thus, it is possible to skip the encoding of the coefficient information flag, regardless of whether or not the orthogonal transform is applied, in accordance with the limit on the number of times of processing of context adaptive encoding. Therefore, it is possible to suppress an increase in processing delay and an increase in the amount of encoding. Furthermore, it is possible to make the difference between the encoding method used in the block in which the orthogonal transform is applied and the encoding method used in the block in which the orthogonal transform is not applied smaller, and to make the circuit size smaller.

[0674] In addition, the coefficient information flag can be a flag indicating whether or not the value of the coefficient is greater than 1. Thus, it is possible to skip the encoding of the coefficient information flag indicating whether or not the value of the coefficient is greater than 1, regardless of whether or not the orthogonal transform is applied, in accordance with the limit on the number of times of processing of context adaptive encoding. Therefore, it is possible to suppress an increase in processing delay and an increase in the amount of encoding.

[0675] In yet another example, the circuit of the encoding apparatus 100 can limit the number of times of processing of context adaptive encoding, and encode the block of the image. Also, in a case where the orthogonal transform is not applied to the block, it can be determined whether or not a plurality of coefficient information flags respectively indicating a plurality of properties of coefficients included in the block satisfy a processing condition. Then, in a case where it is determined that the processing condition is satisfied, the plurality of coefficient information flags can be encoded by context adaptive encoding.

[0676] Here, the processing condition is a condition in which the number of times of processing to which the number of the plurality of coefficient information flags is added is within the limit on the number of times of processing.

[0677] Thus, in a case where the orthogonal transform is not applied, it is possible to collectively determine whether or not it is possible to use context adaptive encoding for the plurality of coefficient information flags. Therefore, it is possible to simplify the processing, and to reduce the processing delay. Furthermore, in a case where similar processing is performed on the block in which the orthogonal transform is applied, it is possible to make the difference between the encoding method used in the block in which the orthogonal transform is applied and the encoding method used in the block in which the orthogonal transform is not applied smaller, and to make the circuit size smaller.

[0678] Furthermore, the plurality of coefficient information flags can include a coefficient information flag indicating whether or not the value of the coefficient is greater than 3 and a coefficient information flag indicating whether or not the value of the coefficient is greater than 5. Thus, it is possible to make a collective determination for the plurality of coefficient information flags including the coefficient information flag indicating whether or not the value of the coefficient is greater than 3 and the coefficient information flag indicating whether or not the value of the coefficient is greater than 5. Therefore, it is possible to simplify the processing, and to reduce the processing delay.

[0679] Furthermore, the plurality of coefficient information flags may include a coefficient information flag indicating whether the coefficient value is greater than 7 and a coefficient information flag indicating whether the coefficient value is greater than 9. This makes it possible to collectively determine the plurality of coefficient information flags including four coefficient information flags: whether the coefficient value is greater than 3, whether the coefficient value is greater than 5, whether the coefficient value is greater than 7, and whether the coefficient value is greater than 9. Consequently, processing can be simplified and processing delay can be reduced.

[0680] Note that the above-described operations performed by the circuits of the encoding device 100 may also be performed by the entropy encoding unit 110 of the encoding device 100 .

[0681] Figure 59 is a flowchart showing the operation of the decoding device 200. For example, the decoding device 200 includes a circuit and a memory connected to the circuit. The circuit and memory included in the decoding device 200 may correspond to Figure 46 The circuit of the decoding device 200 performs Figure 59 Specifically, the circuit of the decoding device 200 decodes a block of an image during operation ( S221 ).

[0682] In one example, the circuitry of decoding apparatus 200 may limit the number of context-adaptive decoding processes and decode the image block. Furthermore, the circuitry may also perform sub-block flag decoding without including it in the target number of processes being limited, both in cases where an inverse orthogonal transform is applied to the block and in cases where an inverse orthogonal transform is not applied to the block.

[0683] Here, the sub-block flag decoding process is a process of decoding a sub-block flag indicating whether a sub-block included in a block includes a non-zero coefficient by context-adaptive decoding.

[0684] This makes it possible to decode sub-block flags using context-adaptive decoding, regardless of whether an inverse orthogonal transform is applied or whether the number of context-adaptive decoding processes is limited. Consequently, the amount of code can be reduced. Furthermore, the difference between the decoding method used in blocks where an inverse orthogonal transform is applied and the decoding method used in blocks where an inverse orthogonal transform is not applied can be minimized, leading to a reduction in circuit size.

[0685] Furthermore, when applying an inverse orthogonal transform to a block, the circuitry of decoding apparatus 200 may perform position parameter decoding processing without including it in the number of processing times subject to the restriction. Here, position parameter decoding processing is processing for decoding parameters indicating the position of the first non-zero coefficient in the block in scanning order using context-adaptive decoding.

[0686] Thus, in the case where the inverse orthogonal transform is applied, the parameter indicating the position of the first nonzero coefficient is likely to be decoded by the context-adaptive decoding regardless of whether the number of times of processing of the context-adaptive decoding is limited. Therefore, the amount of coding is likely to be reduced.

[0687] Further, the circuit of the decoding device 200 can determine the range of limitation of the number of times of processing according to the position of the first nonzero coefficient in the case where the inverse orthogonal transform is applied to the block. Thus, in the case where the inverse orthogonal transform is applied, the number of times of limitation of processing is likely to be appropriately determined. Therefore, the balance between the reduction of the amount of coding and the reduction of the processing delay is likely to be appropriately adjusted.

[0688] In another example, the circuit of the decoding device 200 can act as follows in both the case where the inverse orthogonal transform is applied to the block of the decoded image and the case where the inverse orthogonal transform is not applied to the block.

[0689] Specifically, in both the cases, in the case where the number of times of processing of the context-adaptive decoding is within the range of limitation of the number of times of processing, the circuit of the decoding device 200 can also decode the coefficient information flag by the context-adaptive decoding. Further, in both the cases, in the case where the number of times of processing of the context-adaptive decoding is not within the range of limitation of the number of times of processing, the circuit of the decoding device 200 can also skip the decoding of the coefficient information flag. Here, the coefficient information flag indicates the attribute of the coefficient included in the block.

[0690] Then, in the case where the coefficient information flag is decoded, the circuit of the decoding device 200 can decode the residual value information by the Golomb decoding. Here, the residual value information is information for reconstructing the value of the coefficient using the coefficient information flag. Further, in the case where the decoding of the coefficient information flag is skipped, the circuit of the decoding device 200 can decode the value of the coefficient by the Golomb decoding.

[0691] Thus, in both the case where the inverse orthogonal transform is applied and the case where the inverse orthogonal transform is not applied, the decoding of the coefficient information flag is likely to be skipped according to the limitation of the number of times of processing of the context-adaptive decoding. Therefore, the increase in the processing delay and the increase in the amount of coding are likely to be suppressed. Further, the difference between the decoding method used in the block to which the inverse orthogonal transform is applied and the decoding method used in the block to which the inverse orthogonal transform is not applied is likely to be small, and the circuit scale is likely to be small.

[0692] In addition, the coefficient information flag can also be a flag indicating whether the value of the coefficient is greater than 1. Thus, in both the case where the inverse orthogonal transform is applied and the case where the inverse orthogonal transform is not applied, the decoding of the coefficient information flag indicating whether the value of the coefficient is greater than 1 is likely to be skipped according to the limitation of the number of times of processing of the context-adaptive decoding. Therefore, the increase in the processing delay and the increase in the amount of coding are likely to be suppressed.

[0693] In yet another example, the circuit of the decoding device 200 can decode the block of the image while limiting the number of times of processing of context-adaptive decoding. Then, in a case where the inverse orthogonal transform is not applied to the block, it can be determined whether a plurality of coefficient information flags respectively indicating a plurality of properties of the coefficients included in the block satisfy a processing condition. Then, in a case where it is determined that the processing condition is satisfied, the plurality of coefficient information flags can also be decoded by the context-adaptive decoding.

[0694] Here, the processing condition is a condition in which the number of times of processing to which the number of the plurality of coefficient information flags is added is within the limit of the number of times of processing.

[0695] Thus, in a case where the inverse orthogonal transform is not applied, it is possible to generally determine whether it is possible to use the context-adaptive decoding for the plurality of coefficient information flags. Therefore, the processing can be simplified, and the processing delay can be reduced. Further, in a case where similar processing is performed on a block to which the inverse orthogonal transform is applied, the difference between the decoding method used in the block to which the inverse orthogonal transform is applied and the decoding method used in the block to which the inverse orthogonal transform is not applied can be small, and the circuit size can be small.

[0696] Further, the plurality of coefficient information flags can include a coefficient information flag indicating whether the value of the coefficient is greater than 3 and a coefficient information flag indicating whether the value of the coefficient is greater than 5. Thus, it is possible to generally determine the plurality of coefficient information flags including the coefficient information flag indicating whether the value of the coefficient is greater than 3 and the coefficient information flag indicating whether the value of the coefficient is greater than 5. Therefore, the processing can be simplified, and the processing delay can be reduced.

[0697] Further, the plurality of coefficient information flags can include a coefficient information flag indicating whether the value of the coefficient is greater than 7 and a coefficient information flag indicating whether the value of the coefficient is greater than 9. Thus, it is possible to generally determine the plurality of coefficient information flags including the four coefficient information flags of whether the value of the coefficient is greater than 3, whether the value of the coefficient is greater than 5, whether the value of the coefficient is greater than 7, and whether the value of the coefficient is greater than 9. Therefore, the processing can be simplified, and the processing delay can be reduced.

[0698] In addition, the above-described operations performed by the circuit of the decoding device 200 can be performed by the entropy decoding section 202 of the decoding device 200.

[0699] [Other Examples]

[0700] The encoding device 100 and the decoding device 200 in the above-described examples can each be used as an image encoding device and an image decoding device, respectively, or as a moving image encoding device and a moving image decoding device, respectively.

[0701] Moreover, the encoding apparatus 100 and the decoding apparatus 200 can perform only a part of the above-described operations, and other apparatuses can perform other operations. Furthermore, the encoding apparatus 100 and the decoding apparatus 200 can have only a part of the above-described constituent elements, and other apparatuses can have other constituent elements.

[0702] Furthermore, at least a part of each of the above-described examples can be utilized as an encoding method or a decoding method, or can be utilized as other methods.

[0703] Each of the constituent elements can be configured by a dedicated hardware or can be realized by executing a software program suitable for each of the constituent elements. Each of the constituent elements can be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[0704] Specifically, the encoding apparatus 100 and the decoding apparatus 200 can each have a processing circuitry and a storage device capable of being accessed from the processing circuitry, which is electrically connected to the processing circuitry. For example, the processing circuitry corresponds to the processor a1 or b1, and the storage device corresponds to the memory a2 or b2.

[0705] The processing circuitry includes at least one of a dedicated hardware and a program execution unit, and performs processing using the storage device. In addition, the storage device stores a software program executed by the program execution unit in a case where the processing circuitry includes the program execution unit.

[0706] Here, a software realizing the above-described encoding apparatus 100 or decoding apparatus 200 or the like is a program as follows.

[0707] For example, the program can cause a computer to execute an encoding method of encoding a block of an image while limiting a number of times of processing of context adaptive coding, and in encoding of the block, performing sub-block flag coding processing by context adaptive coding to encode a sub-block flag indicating whether a sub-block included in the block includes a non-zero coefficient without including the sub-block flag coding processing in the number of times of processing in both a case where an orthogonal transform is applied to the block and a case where the orthogonal transform is not applied to the block.

[0708] Further, for example, the program can cause the computer to execute a decoding method of decoding a block of an image while limiting a number of times of processing of context-adaptive decoding, in decoding of the block, in both a case where an inverse orthogonal transform is applied to the block and a case where the inverse orthogonal transform is not applied to the block, performing sub-block flag decoding processing of decoding, by context-adaptive decoding, a sub-block flag indicating whether or not a sub-block included in the block includes a non-zero coefficient without including the sub-block flag decoding processing in the number of times of processing.

[0709] Further, for example, the program can cause the computer to execute an encoding method of, in both a case where an orthogonal transform is applied to a block of an encoding target image and a case where the orthogonal transform is not applied to the block, in a case where a number of times of processing of context-adaptive encoding is within a limit range of the number of times of processing, encoding, by context-adaptive encoding, a coefficient information flag indicating a property of a coefficient included in the block, in a case where the number of times of processing is not within the limit range of the number of times of processing, skipping encoding of the coefficient information flag, in a case where the coefficient information flag is encoded, encoding, by Golomb encoding, residual value information for reconstructing a value of the coefficient using the coefficient information flag, in a case where encoding of the coefficient information flag is skipped, encoding, by Golomb encoding, the value of the coefficient.

[0710] Further, for example, the program can cause the computer to execute a decoding method of, in both a case where an inverse orthogonal transform is applied to a block of a decoding target image and a case where the inverse orthogonal transform is not applied to the block, in a case where a number of times of processing of context-adaptive decoding is within a limit range of the number of times of processing, decoding, by context-adaptive decoding, a coefficient information flag indicating a property of a coefficient included in the block, in a case where the number of times of processing is not within the limit range of the number of times of processing, skipping decoding of the coefficient information flag, in a case where the coefficient information flag is decoded, decoding, by Golomb decoding, residual value information for reconstructing a value of the coefficient using the coefficient information flag, in a case where decoding of the coefficient information flag is skipped, decoding, by Golomb decoding, the value of the coefficient.

[0711] Further, for example, the program can cause the computer to execute an encoding method of encoding a block of an image while limiting the number of times of processing of context-adaptive encoding, in the encoding of the block, in a case where an orthogonal transform is not applied to the block, determining whether a plurality of coefficient information flags respectively indicating a plurality of attributes of coefficients included in the block satisfy a processing condition, in a case where it is determined that the processing condition is satisfied, encoding the plurality of coefficient information flags by context-adaptive encoding, the processing condition being a condition in which the number of times of processing, to which the number of the plurality of coefficient information flags is added, is within a limit range of the number of times of processing.

[0712] Further, for example, the program can cause the computer to execute an encoding method of encoding a block of an image while limiting the number of times of processing of context-adaptive encoding, in the encoding of the block, in a case where an orthogonal transform is not applied to the block, determining whether a plurality of coefficient information flags respectively indicating a plurality of attributes of coefficients included in the block satisfy a processing condition, in a case where it is determined that the processing condition is satisfied, encoding the plurality of coefficient information flags by context-adaptive encoding, the processing condition being a condition in which the number of times of processing, to which the number of the plurality of coefficient information flags is added, is within a limit range of the number of times of processing.

[0713] Further, as described above, each of the constituent elements can be a circuit. These circuits can be a single circuit or different circuits. Further, each of the constituent elements can be implemented by a general-purpose processor or a dedicated processor.

[0714] Further, the processing performed by a specific constituent element can be performed by another constituent element. Further, the order of the processing can be changed, and a plurality of processes can be performed simultaneously. Further, the coding and decoding device can include the encoding device 100 and the decoding device 200.

[0715] Further, the first and second ordinal numbers and the like used in the description can be appropriately exchanged. Further, ordinal numbers can be newly assigned to the constituent elements, or the ordinal numbers can be removed.

[0716] The above describes the configurations of the encoding device 100 and the decoding device 200 based on a plurality of examples, but the configurations of the encoding device 100 and the decoding device 200 are not limited to these examples. As long as the gist of the present application is not deviated from, various modified configurations that can be thought of by those skilled in the art can be implemented for each example, and configurations in which constituent elements in different examples are combined can also be included in the range of the configurations of the encoding device 100 and the decoding device 200.

[0717] One or more of the aspects disclosed herein can be implemented in combination with at least a portion of another aspect of the present disclosure. In addition, a portion of the processing, a portion of the structure of the device, a portion of the syntax, and the like described in the flowchart of one or more of the aspects disclosed herein can be implemented in combination with another aspect.

[0718] [Implementation and Use]

[0719] In each of the above embodiments, each functional block or a block that functions can be realized by an MPU (micro processing unit) and a memory, and the like. In addition, it can be that the processing of each functional block is realized by a program execution section such as a processor that reads and executes software (program) recorded in a recording medium such as a ROM. The software can be distributed. The software can also be recorded in various recording media such as a semiconductor memory. In addition, each functional block can be realized by hardware (a dedicated circuit). Various combinations of hardware and software can be adopted.

[0720] The processing described in each of the embodiments can be realized by centralized processing using a single device (system), or can be realized by distributed processing using a plurality of devices. In addition, the processor that executes the above-described program can be a single one, or can be a plurality of processors. That is, centralized processing or distributed processing can be performed.

[0721] The aspects of the present disclosure are not limited to the above-described embodiments, and various modifications can be made thereto. Such modifications are also within the scope of the aspects of the present disclosure.

[0722] Further, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) represented in each of the above-described embodiments, and various systems that implement the application examples are described herein. Such a system can be characterized by having an image encoding device that uses the image encoding method, an image decoding device that uses the image decoding method, or an image encoding / decoding device that has both. Other structures of such a system can be appropriately changed depending on the situation.

[0723] [Use Example]

[0724] Figure 60 is a diagram that shows the overall structure of an appropriate content supply system ex100 that implements a content distribution service. The provision area of the communication service is divided into desired sizes, and a base station ex106, ex107, ex108, ex109, ex110, which is a fixed wireless station in the illustrated example, is provided in each unit.

[0725] In the content supply system ex100, each of the devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smart phone ex115 is connected via the Internet service provider ex102 or the communication network ex104, and the base stations ex106 to ex110 on the Internet ex101. The content supply system ex100 can also be connected by combining some of the above-described devices. In various implementations, the devices can also be directly or indirectly connected to each other via a telephone network or close proximity wireless, or the like, without passing through the base stations ex106 to ex110. Also, the streaming server ex103 can be connected to each of the devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smart phone ex115 via the Internet ex101 or the like. Further, the streaming server ex103 can be connected to a terminal or the like in a hotspot in the airplane ex117 via the satellite ex116.

[0726] In addition, a wireless access point or a hotspot, or the like can be used instead of the base stations ex106 to ex110. Further, the streaming server ex103 can be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, or can be directly connected to the airplane ex117 without passing through the satellite ex116.

[0727] The camera ex113 is a device such as a digital camera that can take still images and moving images. Further, the smart phone ex115 is a smart phone, a portable telephone, or a PHS (Personal Handy-phone System), or the like, which corresponds to a mode of a mobile communication system called 2G, 3G, 3.9G, 4G, and 5G to be called in future.

[0728] The home appliance ex114 is a device such as a refrigerator or a device included in a household fuel cell cogeneration system.

[0729] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 via the base station ex106 or the like, whereby live distribution or the like can be performed. In the live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smart phone ex115, and a terminal in the airplane ex117, or the like) can perform the encoding processing described in the above-described embodiments on still image or moving image content taken by a user using the terminal, can multiplex video data obtained by the encoding and audio data obtained by encoding sound corresponding to the video, and can transmit the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present application.

[0730] On the other hand, the streaming server ex103 performs streaming distribution of content data transmitted from a client having a request. The client is a computer ex111, a game machine ex112, a camera ex113, a household electrical appliance ex114, a smart phone ex115, or a terminal in an airplane ex117, which is capable of decoding the data subjected to the above-described encoding process. Each device that receives the distributed data can also perform decoding process on the received data and reproduce it. That is, each device can also function as an image decoding apparatus of one aspect of the present application.

[0731] [Decentralized processing]

[0732] Further, the streaming server ex103 can be a plurality of servers or a plurality of computers, which perform decentralized processing or recording of data and distribute it. For example, the streaming server ex103 can be realized by a CDN (Contents Delivery Network), which performs content distribution by connecting a plurality of edge servers distributed in the world and a network between the edge servers. In the CDN, a physically closer edge server can be dynamically assigned according to a client. Further, by caching and distributing content to the edge server, it is possible to reduce delay. Further, in the case where several types of errors occur or the communication state changes due to an increase in traffic or the like, it is possible to decentralize processing by a plurality of edge servers, or switch the distribution subject to another edge server, or continue distribution by bypassing a part of the network in which a failure has occurred, so that high-speed and stable distribution can be realized.

[0733] Further, not limited to decentralized processing of distribution itself, the encoding process of the captured data can be performed by each terminal, on the server side, or shared by them. As an example, two processing cycles are generally performed in the encoding process. In the first cycle, the complexity or the amount of encoding of an image of a frame or a scene unit is detected. Further, in the second cycle, processing that maintains the quality and improves the encoding efficiency is performed. For example, by performing the first encoding process by a terminal and the second encoding process by the server side that receives content, it is possible to reduce the processing load in each terminal while improving the quality and efficiency of content. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding by the terminal can also be received and reproduced by another terminal, so that more flexible real-time distribution can also be performed.

[0734] As other examples, the camera ex113 or the like extracts a feature amount (feature or amount of feature) from an image, compresses and transmits data on the feature amount as metadata to the server. The server switches quantization accuracy or the like, for example, according to the importance of the target judged from the feature amount, and performs compression corresponding to the meaning (or importance of content) of the image. The feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction at the time of re-compression in the server. In addition, it is also possible to perform simple encoding such as VLC (Variable Length Coding) by the terminal, and perform processing with a large load such as CABAC (Context Adaptive Binary Arithmetic Coding) by the server.

[0735] As other examples, in a stadium, a shopping center, or a factory, or the like, there are cases where a plurality of image data obtained by a plurality of terminals photographing substantially the same scene. In this case, using a plurality of terminals that have performed photographing, and using other terminals that have not performed photographing and a server as necessary, for example, respectively distribute and perform distributed processing of encoding processing in units of GOP (Group of Picture), units of a picture, or units of a tile obtained by dividing a picture, or the like. Thereby, it is possible to reduce delay and more preferably realize real-time performance.

[0736] Since a plurality of image data is substantially the same scene, it is also possible to manage and / or instruct by the server to refer to image data photographed by each terminal to each other. In addition, it is also possible for the server to receive encoded data from each terminal and change the reference relationship between a plurality of data, or to correct or replace a picture itself and re-encode. Thereby, it is possible to generate a stream that improves the quality and efficiency of one data.

[0737] Furthermore, the server can also perform transcoding to change the encoding method of the image data and distribute the image data. For example, the server can also change the encoding method of the MPEG type to the VP type (for example, VP9), and can also change H.264 to H.265 or the like.

[0738] In this way, the encoding processing can be performed by the terminal or one or more servers. Therefore, the following description uses "server" or "terminal" or the like as the subject of the processing, but it is also possible to perform part or all of the processing by the server with the terminal, and it is also possible to perform part or all of the processing by the terminal with the server. In addition, as for these, the same is true for the decoding processing.

[0739] [3D, multi-angle]

[0740] The number of cases in which images or videos taken by a plurality of terminals such as cameras ex113 and / or smartphones ex115 and the like that are roughly synchronized with each other are merged and used is increasing. The videos taken by the respective terminals can be merged based on the relative positional relationship between the terminals obtained separately, or a region in which feature points coincide in the videos, and the like.

[0741] The server not only encodes two-dimensional moving images, but also automatically or at a time specified by the user encodes still images based on scene analysis of the moving images and the like and transmits them to the receiving terminal. The server also generates a three-dimensional shape of a scene based on videos of the same scene taken from different angles, not only two-dimensional moving images, when the relative positional relationship between the terminals that take the videos can be obtained. The server can also separately encode three-dimensional data generated by a point cloud and the like, and can select or reconstruct videos transmitted to the receiving terminal from the videos taken by the plurality of terminals based on the results of recognizing or tracking a person or a target using the three-dimensional data.

[0742] In this way, the user can not only arbitrarily select each video corresponding to each terminal and enjoy the scene, but also enjoy the content of a video in which a viewpoint is selected from three-dimensional data reconstructed using a plurality of images or videos. Furthermore, sound can also be collected from a plurality of different angles together with the video, and the server multiplexes sound from a specific angle or space with the corresponding video and transmits the multiplexed video and sound.

[0743] In addition, in recent years, contents that correspond the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) are becoming widespread. In the case of a VR image, the server creates viewpoint images for the right eye and the left eye, respectively, and can encode them so as to allow reference between the respective viewpoint videos by Multi-View Coding (MVC) and the like, or can encode them as different streams without reference to each other. At the time of decoding of the different streams, they can be reproduced in synchronization with each other according to the viewpoint of the user, so as to reproduce a virtual three-dimensional space.

[0744] In the case of the AR image, the server can also superimpose the virtual object information on the virtual space on the camera information of the real space based on the three-dimensional position or movement of the viewpoint of the user. The decoding device acquires or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the viewpoint of the user, and creates superimposition data by smoothly connecting them. Alternatively, the decoding device can transmit the movement of the viewpoint of the user to the server in addition to the commission of the virtual object information. The server can create superimposition data according to the three-dimensional data held in the server, match the received movement of the viewpoint, encode the superimposition data, and distribute it to the decoding device. Typically, the superimposition data has an alpha value indicating the degree of transparency in addition to RGB, and the server sets the alpha value of a portion other than the target created according to the three-dimensional data to 0 or the like and encodes it in a state of being transparent at that portion. Alternatively, the server can set the RGB value of a predetermined value as the background, as in chroma keying, and generate data in which the portion other than the target is set to the background color. The RGB value of the predetermined value can be determined in advance.

[0745] Likewise, the decoding process of the distributed data can be performed by the client (for example, the terminal), on the server side, or shared with each other. As an example, a certain terminal can first transmit a reception request to the server, another terminal can receive and decode the content corresponding to the request, and the decoded signal can be transmitted to the device having a display. By dispersing the process regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. Further, as another example, a large-size image data can be received by a TV or the like, and a part of the region such as a tile after the picture is divided can be decoded and displayed by the personal terminal of the viewer. Thus, it is possible to confirm one's own responsible area or an area that one wants to confirm in more detail while sharing the entire image.

[0746] In a situation where a plurality of wireless communications of short, medium, or long distances indoors and outdoors can be used, it is possible to seamlessly receive the content using the distribution system standards such as MPEG-DASH. The user can also switch in real time while freely selecting the user's terminal, the decoding device or the display device such as a display arranged indoors and outdoors. Further, it is possible to switch and decode the terminal of the decoding and the terminal of the display using one's own position information or the like. Thus, it is also possible to map and display information on a part of the wall surface or the floor of a building next to a device that can display during the user's movement to the destination. Further, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, which is cached in a server that can be accessed from the receiving terminal in a short time, or duplicated in an edge server of a content distribution service, or the like.

[0747] [Scalable Coding]

[0748] To switch content, use Figure 61 The example illustrates a scalable stream compressed and encoded using the moving picture coding method described in the above embodiments. For the server, multiple streams with the same content but different qualities can be provided as a single stream. Alternatively, a structure can be employed to switch content by leveraging the temporally and spatially scalable nature of streams achieved through layered coding, as shown in the figure. Specifically, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. This allows the decoding side to freely switch between low-resolution and high-resolution content. For example, if a user wants to watch a video viewed on a smartphone ex115 while on the go, then watch it later on a device such as an internet TV at home, the device can simply decode the same stream into different layers, reducing the burden on the server.

[0749] Furthermore, in addition to the hierarchical structure of encoding pictures per layer and implementing enhancement layers above the base layer as described above, the enhancement layers may also include metadata such as statistical information based on the image. Alternatively, the decoding side may generate high-definition content by super-resolutioning the base layer pictures based on this metadata. Super-resolution can improve the signal-to-noise ratio while maintaining and / or increasing resolution. Meta-information includes information for determining linear or nonlinear filter coefficients used in super-resolution processing, as well as information for determining parameter values ​​in filtering, machine learning, or least-squares operations used in super-resolution processing.

[0750] Alternatively, a structure can be provided that divides a picture into tiles according to the meaning of the object in the image. The decoding side decodes only a part of the area by selecting the tile to be decoded. Moreover, by storing the attributes of the object (people, cars, balls, etc.) and the position in the image (coordinate position in the same image, etc.) as meta information, the decoding side can determine the position of the desired object based on the meta information and decide the tile that includes the object. For example, Figure 62 As shown, SEI (supplemental enhancement information) messages in HEVC, which are different from pixel data, can also be used to store meta-information. This meta-information indicates, for example, the position, size, or color of the main object.

[0751] The meta information can also be saved in units of a plurality of pictures such as a stream, a sequence, or a random access unit. The decoding side can acquire the time when a specific person appears in the video, and by matching the information of the picture unit and the time information, the picture in which the target exists can be determined, and the position of the target in the picture can be decided.

[0752] [Optimization of Web Page]

[0753] Figure 63 Fig. 1 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. Figure 64 Fig. 2 is a diagram showing an example of a display screen of a web page in a smart phone ex115 or the like. As shown in Figs. 1 and 2, a web page can include a plurality of link images as links to image contents. Figure 63 Figure 64 As shown in Figs. 1 and 2, a web page can include a plurality of link images as links to image contents. The visibility of the link images can differ depending on the device on which the web page is viewed. In a case where a plurality of link images are visible on the screen, before the user explicitly selects a link image, or before the link image approaches the vicinity of the center of the screen or the entirety of the link image enters the screen, the display device (decoding device) can display a still image or an I picture possessed by each content as a link image, can display a moving image such as a gif animation using a plurality of still images or I pictures, or can decode and display the moving image using only the base layer.

[0754] In a case where a link image is selected by the user, the display device decodes the base layer with the highest priority, for example. If there is information indicating that the content is scalable in the HTML constituting the web page, the display device can decode up to the enhancement layer. Further, in a case where the selection is made in order to ensure real-time performance or the communication band is very tight, the display device can reduce the delay between the decoding time and the display time of the initial picture (the delay from the start of decoding to the start of display) by decoding and displaying only the pictures referred to in the front (I pictures, P pictures, and B pictures that refer only to the front). Furthermore, the display device can forcibly ignore the reference relationship of the pictures and decode all the B pictures and P pictures as referring to the front, and as the number of pictures received increases with the passage of time, perform normal decoding.

[0755] [Automatic Travel]

[0756] Further, in a case where still images or moving image data such as two-dimensional or three-dimensional map information are transmitted and received for the automatic travel or travel assistance of a vehicle, the receiving terminal can receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them in correspondence. The meta information can belong to a layer or can be multiplexed only with the image data.

[0757] ​In this case, since the vehicle, the drone, or the airplane, etc. including the reception terminal is moving, the reception terminal can perform seamless reception and decoding at the time of switching the base stations ex106 to ex110 by transmitting the position information of the reception terminal. Further, the reception terminal can dynamically switch the degree of reception of the meta information or the degree of update of the map information according to the selection of the user, the situation of the user, and / or the state of the communication band.

[0758] In the content supply system ex100, the client can receive and decode the encoded information transmitted by the user in real time, and reproduce it.

[0759] [Distribution of Personal Contents]

[0760] Further, in the content supply system ex100, not only the high-quality, long-time contents provided by the video distribution business operator, but also the low-quality, short-time contents provided by the individual can be distributed by unicast or multicast. It is conceivable that such personal contents will increase in the future. In order to make the personal contents better contents, the server can perform the encoding process after performing the editing process. This can be realized, for example, by the following structure.

[0761] After the shooting, the server performs the recognition process of the shooting error, the scene search, the meaning analysis, and the target detection, etc. based on the original image data or the encoded data in real time or cumulatively. Further, the server performs the editing of the correction of the focus deviation or the hand shake, etc. manually or automatically, or the deletion of the scene which is less important than the other pictures or the scene in which the focus is not aligned, or the edge emphasis of the target, or the change of the color tone, etc. based on the recognition result. The server encodes the edited data based on the editing result. Further, it is known that the viewing rate decreases if the shooting time is too long, so the server can automatically limit the scene which is less important than the above, or the scene in which the motion is less, etc. based on the image processing result to be the contents within a specific time range according to the shooting time. Alternatively, the server can generate and encode the summary based on the result of the meaning analysis of the scene.

[0762] There are cases where personal content in its original state may be written with content that infringes copyright, author's personality rights or portrait rights, etc., or there are cases where the scope of sharing exceeds the desired scope, which is inconvenient for individuals. Therefore, for example, the server can also forcibly change the faces of people in the peripheral part of the screen, or the home, etc. to an out-of-focus image for encoding. In addition, the server can also identify whether the face of a person different from the pre-registered person is captured in the image to be encoded, and if so, perform processing such as applying mosaics to the face part. Alternatively, as pre-processing or post-processing for encoding, the user can also specify the person or background area that he wants to process the image from the perspective of copyright, etc. The server can also replace the specified area with another image, or blur the focus, etc. If it is a person, it can track the person in the moving image and replace the image of the person's face.

[0763] The viewing of personal content with a small amount of data has a strong demand for real-time performance, so although it also depends on the bandwidth, the decoding device can also receive, decode, and reproduce the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the playback is repeated more than twice, such as in the case of looped playback, the enhancement layer is also included in the playback of high-definition images. In this way, if the stream is scalable, it can provide an experience in which the moving image is relatively rough when it is not selected or at the beginning of viewing, but the stream gradually becomes smoother and the image becomes better. In addition to scalable coding, the same experience can be provided when the relatively rough stream played back the first time and the second stream encoded with reference to the moving image of the first time are composed of a single stream.

[0764] [Other application examples]

[0765] In addition, these encoding and decoding processes are usually processed in the LSI ex500 of each terminal. Figure 60 ) can be a single chip or a multi-chip configuration. Alternatively, video encoding or decoding software can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) readable by the computer ex111 or the like, and encoding and decoding can be performed using this software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera can be transmitted. In this case, the video data can be encoded by the LSI ex500 in the smartphone ex115.

[0766] Alternatively, the LSIex500 can be configured to download and activate application software. In this case, the terminal first determines whether it is compatible with the content's encoding scheme or has the capability to perform a specific service. If the terminal is not compatible with the content's encoding scheme or does not have the capability to perform a specific service, it can download the codec or application software to retrieve and play the content.

[0767] Furthermore, the content delivery system ex100 is not limited to the content delivery system ex100 via the Internet ex101; at least one of the video encoding devices (image encoding devices) or video decoding devices (image decoding devices) described in the above-mentioned embodiments can also be incorporated into a digital broadcasting system. Since multiplexed data containing multiplexed video and audio is transmitted and received over broadcast radio waves using satellites, the content delivery system ex100 differs from the unicast-friendly structure of the content delivery system ex100 in that it is suitable for multicast. However, the encoding and decoding processes can be applied in the same manner.

[0768] [Hardware structure]

[0769] Figure 65 Is a further detailed representation Figure 60 FIG. 1 shows a diagram of the smartphone ex115. Figure 66 This diagram shows an example of the configuration of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of capturing both video and still images, and a display unit ex458 that displays images captured by the camera unit ex465 and decoded data such as images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio and sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded or decoded data such as captured videos or still images, recorded audio, received videos or still images, and emails, and a slot unit ex464 that serves as an interface with a SIM card ex468 for identifying users and authenticating access to various data, including the network. Alternatively, an external memory card can be used in place of the memory unit ex467.

[0770] The main control section ex460 capable of comprehensively controlling the display section ex458, the operation section ex466, and the like, and the power supply circuit section ex461, the operation input control section ex462, the image signal processing section ex455, the camera interface section ex463, the display control section ex459, the modulation / demodulation section ex452, the multiplexing / demultiplexing section ex453, the sound signal processing section ex454, the slot section ex464, and the memory section ex467 are connected to each other via a bus ex470 in synchronization.

[0771] The power supply circuit section ex461 activates the smartphone ex115 to an operable state if the power key is brought to an on state by the user's operation, and supplies electric power from a battery pack to each section.

[0772] The smartphone ex115 performs processing of a call and data communication and the like based on the control of the main control section ex460 having a CPU, a ROM, a RAM, and the like. At the time of a call, a sound signal of a sound collected by the sound input section ex456 is converted into a digital sound signal by the sound signal processing section ex454, a spectrum spread process is performed by the modulation / demodulation section ex452, a digital-analog conversion process and a frequency conversion process are performed by the transmission / reception section ex451, and the resultant signal is transmitted via the antenna ex450. Further, received data is amplified and subjected to a frequency conversion process and an analog-digital conversion process, a spectrum inverse spread process is performed by the modulation / demodulation section ex452, and an analog sound signal is converted by the sound signal processing section ex454, after which it is output from the sound output section ex457. At the time of data communication, a text, a still image, or image data can be sent out under the control of the main control section ex460 via the operation input control section ex462 based on the operation of the operation section ex466 of the main body section and the like. The same transmission and reception processing is performed. At the time of data communication, in the case of transmitting an image, a still image, or image and sound, the image signal processing section ex455 compressively encodes an image signal saved in the memory section ex467 or an image signal input from the camera section ex465 by the moving image encoding method indicated in each of the embodiments described above, and sends the encoded image data to the multiplexing / demultiplexing section ex453. The sound signal processing section ex454 encodes a sound signal collected by the sound input section ex456 in the process of photographing an image, a still image, by the camera section ex465, and sends the encoded sound data to the multiplexing / demultiplexing section ex453. The multiplexing / demultiplexing section ex453 multiplexes the encoded image data and the encoded sound data in a prescribed manner, performs a modulation process and a conversion process by the modulation / demodulation section (modulation / demodulation circuit section) ex452 and the transmission / reception section ex451, and transmits them via the antenna ex450. The prescribed manner can be determined in advance.

[0773] In a case where an image attached to an e-mail or a chat tool or an image linked on a Web page is received, in order to decode multiplexed data received via the antenna ex450, the multiplexing / demultiplexing section ex453 separates the multiplexed data into a bit stream of image data and a bit stream of sound data by demultiplexing the multiplexed data, supplies the encoded image data to the image signal processing section ex455 via the synchronous bus ex470, and supplies the encoded sound data to the sound signal processing section ex454. The image signal processing section ex455 decodes the image signal by a moving image decoding method corresponding to the moving image encoding method indicated in each of the above-described embodiments, and displays an image or a still image included in a linked moving image file from the display section ex458 via the display control section ex459. The sound signal processing section ex454 decodes the sound signal, and outputs the sound from the sound output section ex457. Since real-time streaming media is becoming more popular, depending on the situation of the user, there can be a case where reproduction of sound is not socially appropriate. Therefore, it is also possible that, as an initial value, a structure in which the sound signal is not reproduced and only the image data is reproduced is preferable, and the sound is reproduced in synchronization only when the user performs a click or the like on the image data.

[0774] Furthermore, although the smart phone ex115 is described here as an example, as a terminal, in addition to a transceiving type terminal having both an encoder and a decoder, a transmission terminal having only an encoder, a reception terminal having only a decoder, and the like can be considered. In a system for digital broadcasting, a case where multiplexed data in which sound data is multiplexed with image data is received and transmitted is described. However, in addition to sound data, character data associated with an image or the like can be multiplexed in the multiplexed data. In addition, instead of multiplexed data, the image data itself can be received or transmitted.

[0775] In addition, although a case where the main control section ex460 including a CPU controls the encoding or decoding process is described, a case where various terminals are provided with a GPU is also common. Therefore, a structure in which a larger region is processed together using the performance of a GPU by a memory shared by the CPU and the GPU or a memory in which addresses are managed in a manner that can be commonly used can be made. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, if the processes of motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization are not performed by the CPU but are performed by the GPU together in units of pictures or the like, the processes are more efficient.

[0776] Industrial Applicability

[0777] The present application can be used in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conference system, or an electronic mirror.

[0778] Reference Signs List

[0779] 100 encoding device

[0780] 102 division section

[0781] 104 subtraction section

[0782] 106 conversion section

[0783] 108 quantization section

[0784] 110 entropy encoding section

[0785] 112, 204 inverse quantization section

[0786] 114, 206 inverse conversion section

[0787] 116, 208 addition section

[0788] 118, 210 block memory

[0789] 120, 212 loop filter section

[0790] 122, 214 frame memory

[0791] 124, 216 intra prediction section

[0792] 126, 218 inter prediction section

[0793] 128, 220 prediction control section

[0794] 200 decoding device

[0795] 202 entropy decoding section

[0796] 1201 boundary determination section

[0797] 1202, 1204, 1206 switch

[0798] 1203 filter determination section

[0799] 1205 filter processing section

[0800] 1207 filter characteristic decision section

[0801] 1208 processing determination section

[0802] a1, b1 processor

[0803] a2, b2 memory

Claims

1. An encoding device, wherein: have: circuits; and a memory connected to the circuit, The circuit performs residual coding of the current block.

2. A decoding device, wherein: have: circuits; and a memory connected to the circuit, The circuit performs decoding of the current block.

3. A computer-readable non-transitory storage medium storing a bitstream, wherein: The bit stream includes information for causing a decoding device to perform a decoding process.