Coding device, decoding device, coding method, decoding method and storage medium
By independently encoding and decoding HRD parameters in a video encoding device and a decoding device, the problem of low HRD parameter processing efficiency in the prior art is solved, the encoding efficiency and image quality are improved, the processing amount and circuit scale are reduced, and the processing speed is increased.
Patent Information
- Application Number
- CN202080059896.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2020-09-11
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2040-09-11
AI Technical Summary
Existing video coding technologies require improvement in coding efficiency, image quality, processing capacity, circuit size, and processing speed. In particular, there is a problem of low efficiency in the processing of HRD parameters.
By encoding HRD parameters independently of sequence parameter sets into SEI messages of buffering periods, picture timings, and decoding units in encoding and decoding devices, independent parsing of HRD parameters is achieved, thereby improving encoding efficiency and decoding speed.
It improves coding efficiency, image quality, reduces processing workload and circuit size, accelerates processing speed, and realizes timely analysis and independent processing of HRD parameters.
Smart Images

Figure CN114731441B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding, and in particular to systems, components, and methods for encoding and decoding moving images. Background Art
[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). This advancement has led to a growing need for improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data used in a variety of applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding.
[0003] Furthermore, Non-Patent Document 1 relates to an example of existing standards related to the above-mentioned video encoding technology.
[0004] Prior art literature
[0005] Non-patent literature
[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention
[0007] Problems to be solved by the invention
[0008] Regarding the coding methods described above, it is expected that new methods will be proposed to improve coding efficiency, improve image quality, reduce processing volume, reduce circuit scale, or appropriately select elements or actions such as filters, blocks, sizes, motion vectors, reference pictures or reference blocks.
[0009] The present disclosure provides a structure or method that can contribute to one or more of the following: improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, the present disclosure may include structures or methods that can contribute to benefits other than those listed above.
[0010] Means for solving problems
[0011] For example, the encoding device of an embodiment of the present disclosure includes a circuit and a memory connected to the circuit, and the circuit, when in operation, encodes HRD parameters related to the decoding unit and the HRD independently of the sequence parameter set into the SEI message during buffering, wherein the HRD (Hypothetical Reference Decoder) is a hypothetical reference decoder and the SEI (Supplemental Enhancement Information) is supplementary enhancement information.
[0012] In video coding technology, new methods are expected to be proposed in order to improve coding efficiency, improve image quality, reduce circuit size, and the like.
[0013] The structures or methods of each embodiment or part thereof in the present disclosure can achieve, for example, at least one of improved coding efficiency, improved image quality, reduced encoding / decoding processing load, reduced circuit scale, or improved encoding / decoding processing speed. Alternatively, the structures or methods of each embodiment or part thereof in the present disclosure can perform appropriate selection of components / actions such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks during encoding and decoding. In addition, the present disclosure also includes the disclosure of structures or methods that can provide benefits other than those described above. For example, structures or methods that improve coding efficiency while suppressing an increase in processing load can be used.
[0014] The description and drawings clarify further advantages and effects of one embodiment of the present disclosure. These advantages and / or effects are achieved through several embodiments and the features described in the description and drawings, but all of them do not necessarily need to be provided in order to achieve one or more advantages and / or effects.
[0015] In addition, these general or specific forms can be implemented by systems, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0016] Effects of the Invention
[0017] One aspect of the disclosed structure or method can contribute to one or more of the following: improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, one aspect of the disclosed structure or method can also contribute to benefits other than those listed above. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a schematic diagram showing an example of the configuration of a transmission system according to an embodiment.
[0019] Figure 2 This is a diagram showing an example of the hierarchical structure of data in a stream.
[0020] Figure 3 This is a diagram showing an example of the structure of a slice.
[0021] Figure 4 This is a diagram showing an example of a tile structure.
[0022] Figure 5 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0023] Figure 6 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0024] Figure 7 This is a block diagram showing an example of the functional structure of the encoding device according to the embodiment.
[0025] Figure 8 This is a block diagram showing an example of an implementation of an encoding device.
[0026] Figure 9 This is a flowchart showing an example of the overall encoding process performed by the encoding device.
[0027] Figure 10 This is a diagram showing an example of block division.
[0028] Figure 11 This is a diagram showing an example of the functional configuration of a dividing unit.
[0029] Figure 12 A diagram showing an example of a segmentation pattern.
[0030] Figure 13A This is a diagram showing an example of a syntax tree of a segmentation pattern.
[0031] Figure 13B This is a diagram showing another example of a syntax tree of a segmentation pattern.
[0032] Figure 14 This is a table showing the transformation basis functions corresponding to each transformation type.
[0033] Figure 15 This is a diagram showing an example of SVT.
[0034] Figure 16 This is a flowchart showing an example of processing performed by the conversion unit.
[0035] Figure 17 This is a flowchart showing another example of the processing performed by the conversion unit.
[0036] Figure 18This is a block diagram showing an example of the functional configuration of a quantization unit.
[0037] Figure 19 This is a flowchart showing an example of quantization performed by the quantization unit.
[0038] Figure 20 This is a block diagram showing an example of the functional structure of the entropy coding unit.
[0039] Figure 21 This is a diagram showing the CABAC process in the entropy coding unit.
[0040] Figure 22 This is a block diagram showing an example of the functional configuration of a loop filter unit.
[0041] Figure 23A This is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).
[0042] Figure 23B FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0043] Figure 23C FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0044] Figure 23D This is a diagram showing an example in which a Y sample (first component) is used for CCALF of Cb and CCALF of Cr (a plurality of components different from the first component).
[0045] Figure 23E A diagram showing a diamond filter.
[0046] Figure 23F It is a diagram showing an example of JC-CCALF.
[0047] Figure 23G This is a diagram showing examples of weight_index candidates of JC-CCALF.
[0048] Figure 24 This is a block diagram showing an example of a detailed configuration of a loop filter unit functioning as a DBF.
[0049] Figure 25 This is a diagram showing an example of deblocking filtering having filter characteristics that are symmetric with respect to block boundaries.
[0050] Figure 26 This is a diagram for explaining an example of a block boundary on which a deblocking filtering process is performed.
[0051] Figure 27 This is a diagram showing an example of the Bs value.
[0052] Figure 28 This is a flowchart showing an example of processing performed by the prediction unit of the encoding device.
[0053] Figure 29 This is a flowchart showing another example of processing performed by the prediction unit of the encoding device.
[0054] Figure 30 This is a flowchart showing another example of processing performed by the prediction unit of the encoding device.
[0055] Figure 31 This is a diagram showing an example of 67 intra prediction modes in intra prediction.
[0056] Figure 32 This is a flowchart showing an example of processing performed by the intra prediction unit.
[0057] Figure 33 A diagram showing an example of each reference picture.
[0058] Figure 34 This is a conceptual diagram showing an example of a reference picture list.
[0059] Figure 35 This is a flowchart showing the flow of basic processing of inter-frame prediction.
[0060] Figure 36 This is a flowchart showing an example of MV derivation.
[0061] Figure 37 This is a flowchart showing another example of MV derivation.
[0062] Figure 38A This is a diagram showing an example of classification of each mode in MV derivation.
[0063] Figure 38B This is a diagram showing an example of classification of each mode in MV derivation.
[0064] Figure 39 This is a flowchart showing an example of inter prediction based on the normal inter mode.
[0065] Figure 40 This is a flowchart showing an example of inter-frame prediction based on the normal merge mode.
[0066] Figure 41 This is a diagram for explaining an example of MV derivation processing based on the normal merge mode.
[0067] Figure 42 This is a diagram for explaining an example of MV derivation processing based on the HMVP mode.
[0068] Figure 43This is a flowchart showing an example of FRUC (frame rate up conversion).
[0069] Figure 44 This is a diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0070] Figure 45 This is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0071] Figure 46A This is a diagram for explaining an example of derivation of a sub-block-based MV in an affine mode using two control points.
[0072] Figure 46B This is a diagram for explaining an example of derivation of MVs in sub-block units in the affine mode using three control points.
[0073] Figure 47A This is a conceptual diagram for explaining an example of MV derivation of control points in the affine mode.
[0074] Figure 47B This is a conceptual diagram for explaining an example of MV derivation of control points in the affine mode.
[0075] Figure 47C This is a conceptual diagram for explaining an example of MV derivation of control points in the affine mode.
[0076] Figure 48A A diagram for explaining an affine pattern having two control points.
[0077] Figure 48B A diagram for explaining an affine pattern having three control points.
[0078] Figure 49A This is a conceptual diagram for explaining an example of a method for deriving MVs of control points when the number of control points in an already coded block and a current block is different.
[0079] Figure 49B This is a conceptual diagram for explaining another example of a method for deriving an MV of a control point when the number of control points in an already coded block and a current block is different.
[0080] Figure 50 This is a flowchart showing an example of processing in the affine merge mode.
[0081] Figure 51 This is a flowchart showing an example of processing in the affine inter mode.
[0082] Figure 52AThis is a diagram for explaining the generation of predicted images of two triangles.
[0083] Figure 52B This is a conceptual diagram showing an example of the first portion of the first partition and the first and second sample sets.
[0084] Figure 52C This is a conceptual diagram showing the first part of the first partition.
[0085] Figure 53 This is a flowchart showing an example of the triangle mode.
[0086] Figure 54 This is a diagram showing an example of the ATMVP mode for deriving MVs in sub-block units.
[0087] Figure 55 This is a diagram showing the relationship between merge mode and DMVR (dynamic motion vector refreshing).
[0088] Figure 56 This is a conceptual diagram for explaining an example of DMVR.
[0089] Figure 57 This is a conceptual diagram for explaining another example of DMVR for determining MV.
[0090] Figure 58A This is a diagram showing an example of motion search in DMVR.
[0091] Figure 58B This is a flowchart showing an example of motion search in DMVR.
[0092] Figure 59 This is a flowchart showing an example of generating a predicted image.
[0093] Figure 60 This is a flowchart showing another example of generating a predicted image.
[0094] Figure 61 This is a flowchart for explaining an example of predicted image correction processing based on OBMC (overlapped block motion compensation).
[0095] Figure 62 This is a conceptual diagram for explaining an example of predicted image correction processing based on OBMC.
[0096] Figure 63 This is a diagram for explaining a model assuming uniform linear motion.
[0097] Figure 64This is a flowchart showing an example of inter-frame prediction according to BIO.
[0098] Figure 65 This is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.
[0099] Figure 66A This is a diagram for explaining an example of a method for generating a predicted image using a brightness correction process based on LIC (local illumination compensation).
[0100] Figure 66B This is a flowchart showing an example of a method for generating a predicted image using a brightness correction process based on LIC.
[0101] Figure 67 This is a block diagram showing the functional structure of a decoding device according to an embodiment.
[0102] Figure 68 This is a block diagram showing an implementation example of a decoding device.
[0103] Figure 69 This is a flowchart showing an example of the overall decoding process performed by the decoding device.
[0104] Figure 70 This is a diagram showing the relationship between the division determination unit and other components.
[0105] Figure 71 This is a block diagram showing an example of the functional structure of the entropy decoding unit.
[0106] Figure 72 This is a diagram showing the CABAC process in the entropy decoding unit.
[0107] Figure 73 This is a block diagram showing an example of the functional configuration of an inverse quantization unit.
[0108] Figure 74 This is a flowchart showing an example of inverse quantization performed by the inverse quantization unit.
[0109] Figure 75 This is a flowchart showing an example of processing performed by the inverse transformation unit.
[0110] Figure 76 This is a flowchart showing another example of the processing performed by the inverse transformation unit.
[0111] Figure 77 This is a block diagram showing an example of the functional configuration of a loop filter unit.
[0112] Figure 78 This is a flowchart showing an example of processing performed by the prediction unit of the decoding device.
[0113] Figure 79 This is a flowchart showing another example of processing performed by the prediction unit of the decoding device.
[0114] Figure 80A This is a flowchart showing a part of another example of processing performed by the prediction unit of the decoding device.
[0115] Figure 80B This is a flowchart showing the remaining portion of another example of processing performed by the prediction unit of the decoding device.
[0116] Figure 81 This is a diagram showing an example of processing performed by the intra prediction unit of the decoding device.
[0117] Figure 82 This is a flowchart showing an example of MV derivation in the decoding device.
[0118] Figure 83 This is a flowchart showing another example of MV derivation in the decoding device.
[0119] Figure 84 This is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in a decoding device.
[0120] Figure 85 This is a flowchart illustrating an example of inter-frame prediction based on the normal merge mode in a decoding device.
[0121] Figure 86 This is a flowchart showing an example of inter-frame prediction based on the FRUC mode in the decoding device.
[0122] Figure 87 This is a flowchart illustrating an example of inter-frame prediction based on the affine merge mode in a decoding device.
[0123] Figure 88 This is a flowchart illustrating an example of inter prediction based on the affine inter mode in a decoding device.
[0124] Figure 89 This is a flowchart showing an example of inter-frame prediction based on the triangular mode in a decoding device.
[0125] Figure 90 This is a flowchart showing an example of a DMVR-based motion search in a decoding device.
[0126] Figure 91 This is a flowchart showing a detailed example of a motion search based on DMVR in a decoding device.
[0127] Figure 92 This is a flowchart showing an example of generation of a predicted image in a decoding device.
[0128] Figure 93 This is a flowchart showing another example of generation of a predicted image in a decoding device.
[0129] Figure 94 This is a flowchart showing an example of correction of a predicted image based on OBMC in a decoding device.
[0130] Figure 95 This is a flowchart showing an example of correction of a predicted image based on BIO in a decoding device.
[0131] Figure 96 This is a flowchart showing an example of modification of a predicted image based on LIC in a decoding device.
[0132] Figure 97 This diagram shows the syntax structure of information related to the usage of decoding units included in HRD (Hypothetical Reference Decoder) encoded in the buffering period SEI (Supplemental Enhancement Information).
[0133] Figure 98 This is a flowchart showing an example of a process of checking bitstream consistency using HRD in the first form of the embodiment.
[0134] Figure 99 This diagram shows the syntax structure of information related to the usage of decoding units included in the HRD encoded in the picture timing SEI.
[0135] Figure 100 This is a flowchart showing the operation of the encoding device according to the embodiment.
[0136] Figure 101 This is a flowchart showing the operation of the decoding device according to the embodiment.
[0137] Figure 102 This is a diagram showing the overall structure of a content provision system that implements content distribution services.
[0138] Figure 103 This is a diagram showing an example of a display screen of a web page.
[0139] Figure 104 This is a diagram showing an example of a display screen of a web page.
[0140] Figure 105 This is a diagram showing an example of a smart phone.
[0141] Figure 106 This is a block diagram showing a configuration example of a smartphone. DETAILED DESCRIPTION
[0142] [Introduction]
[0143] In video coding, the HRD defines specifications to verify compatibility between the bitstream and the decoding device. To perform this compatibility test, information related to the HRD buffer and picture timing is encoded in the bitstream, including the SPS (Sequence Parameter Set), the buffering period SEI message, the picture timing SEI message, and the decoding unit information SEI message when the HRD processes each decoding unit.
[0144] Since SEI messages tend to be sent before the slice header of a slice of a picture, and information about the SPS applied to the picture is contained in the slice header, it is expected that the SEI associated with the above-mentioned HRD, that is, all HRD-associated SEI messages can be parsed independently of the SPS.
[0145] Therefore, the encoding device in the embodiment of the present invention includes a circuit and a memory connected to the circuit, and the circuit, during operation, encodes one or more HRD parameters independently of the sequence parameter set into one or more HRD-associated SEI messages, wherein the one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI (Supplemental Enhancement Information) messages associated with HRD.
[0146] Therefore, the encoding device in the embodiment of the present disclosure can encode the HRD parameters so that they are parsed independently from the SPS. Therefore, the encoding device can encode the HRD parameters so that they are parsed immediately after receiving the SEI message.
[0147] In addition, in the encoding device in the embodiment of the present disclosure, the one or more HRD-associated SEI messages include a buffering period SEI message, and the circuit encodes at least one of the one or more HRD parameters into the buffering period SEI message independently of the sequence parameter set and other HRD-associated SEI messages.
[0148] Thus, the encoding device in the embodiment of the present disclosure can encode the HRD parameters and the SPS independently into the buffering period SEI message. Thus, the encoding device can encode the HRD parameters so that they can be parsed immediately after receiving the SEI message.
[0149] In addition, in the encoding device in the embodiment of the present disclosure, the one or more HRD-associated SEI messages include a picture timing SEI message, and the circuit encodes at least one of the one or more HRD parameters into the picture timing SEI message only depending on the buffer period SEI message in the one or more HRD-associated SEI messages.
[0150] Thus, the encoding device in the embodiment of the present disclosure can encode the HRD parameters into the picture timing SEI message independently of the SPS and relying only on the buffering period SEI message. Thus, the encoding device can encode the HRD parameters so that they can be parsed immediately after receiving the SEI message.
[0151] In addition, in the encoding device in the embodiment of the present disclosure, the one or more HRD-associated SEI messages include a decoding unit information SEI message, and the circuit encodes at least one of the one or more HRD parameters into the decoding unit information SEI message only depending on the buffering period SEI message in the one or more HRD-associated SEI messages.
[0152] Thus, the encoding device in the embodiment of the present disclosure can encode the HRD parameters into the decoding unit information SEI message independently of the SPS and relying only on the buffering period SEI message. Thus, the encoding device can encode the HRD parameters so that they can be parsed immediately after receiving the SEI message.
[0153] In addition, in the encoding device in an embodiment of the present disclosure, the circuit encodes at least one of the one or more HRD parameters into the sequence parameter, and encodes the at least one of the one or more HRD parameters into at least one of the one or more HRD-associated SEI messages independently of the sequence parameter set.
[0154] Thus, the encoding device in the embodiment of the present disclosure can encode the HRD parameters encoded in the sequence parameter set into the HRD-associated SEI message. Thus, the encoding device can encode the HRD parameters so that they can be parsed immediately after receiving the SEI message.
[0155] In addition, in the encoding device in an embodiment of the present disclosure, the circuit encodes information about whether there is a CPB (Coded Picture Buffer) parameter for the decoding unit as one of the one or more HRD parameters in the picture timing SEI message or the decoding unit information SEI message.
[0156] Thus, the encoding device in the embodiment of the present disclosure can encode the CPB parameter indicating whether the HRD parameter for the decoding unit exists in the buffering period SEI message. Thus, the encoding device can encode the HRD parameter so that it can be parsed immediately after receiving the SEI message.
[0157] In addition, in the encoding device in the embodiment of the present disclosure, when there are multiple HRD parameters, the circuit encodes a parameter indicating whether the multiple HRD parameters exist in the picture timing SEI message or the decoding unit information SEI message into the buffering period SEI message.
[0158] Thus, the encoding device in the embodiment of the present disclosure can encode the parameter indicating whether the HRD parameter exists in the picture timing SEI message or the decoding unit information SEI message. Thus, the encoding device can encode the HRD parameter so that it can be parsed immediately after receiving the SEI message.
[0159] In addition, the decoding device in the embodiment of the present disclosure includes a circuit and a memory connected to the circuit, and the circuit, during operation, decodes one or more HRD parameters from one or more HRD-associated SEI messages independently of the sequence parameter set. The one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with HRD.
[0160] Therefore, the decoding apparatus in the embodiment of the present disclosure can decode the HRD parameters independently of the SPS. Therefore, the decoding apparatus can decode the HRD parameters immediately after receiving the SEI message.
[0161] In addition, in the decoding device in the embodiment of the present disclosure, the one or more HRD-related SEI messages include a buffering period SEI message, and the circuit decodes at least one of the one or more HRD parameters from the buffering period SEI message independently of the sequence parameter set and other HRD-related SEI messages.
[0162] Thus, the decoding device in the embodiment of the present disclosure can decode the HRD parameters independently from the SPS from the SEI message during buffering. Thus, the decoding device can decode the HRD parameters so that they can be parsed immediately after receiving the SEI message.
[0163] In addition, in the decoding device in the embodiment of the present disclosure, the one or more HRD-associated SEI messages include a picture timing SEI message, and the circuit decodes at least one of the one or more HRD parameters from the picture timing SEI message only depending on the buffer period SEI message in the one or more HRD-associated SEI messages.
[0164] Thus, the decoding device in the embodiment of the present disclosure can decode the HRD parameters from the picture timing SEI message independently of the SPS and relying only on the buffering period SEI message. Thus, the decoding device can decode the HRD parameters immediately after receiving the SEI message.
[0165] In addition, in the decoding device in the embodiment of the present disclosure, the one or more HRD-associated SEI messages include a decoding unit information SEI message, and the circuit decodes at least one of the one or more HRD parameters from the decoding unit information SEI message, relying only on the buffer period SEI message in the one or more HRD-associated SEI messages.
[0166] Thus, the decoding device in the embodiment of the present disclosure can decode the HRD parameters from the decoding unit information SEI message independently of the SPS and relying only on the buffering period SEI message. Thus, the decoding device can decode the HRD parameters immediately after receiving the SEI message.
[0167] In addition, in the decoding device in the embodiment of the present disclosure, the circuit decodes at least one of the one or more HRD parameters from the sequence parameter set, and decodes the at least one of the one or more HRD parameters from at least one of the one or more HRD-associated SEI messages independently of the sequence parameter set.
[0168] Thus, the decoding apparatus in the embodiment of the present disclosure can decode the HRD parameters decoded from the sequence parameter set from the HRD-associated SEI message. Thus, the decoding apparatus can decode the HRD parameters immediately after receiving the SEI message.
[0169] Furthermore, in the decoding device in the embodiment of the present disclosure, the circuit decodes information on whether CPB parameters for the decoding unit exist in the picture timing SEI message or the decoding unit information SEI message from within the buffering period SEI message or the picture timing SEI message.
[0170] Thus, the decoding device in the embodiment of the present disclosure can decode the CPB parameter indicating whether the HRD parameter for the decoding unit exists from the buffering period SEI message. Thus, the decoding device can decode the HRD parameter immediately after receiving the SEI message.
[0171] Furthermore, in the decoding device according to the embodiment of the present disclosure, the circuit decodes a parameter indicating whether a plurality of HRD parameters related to the decoding unit exist from within the buffering period SEI message as one of the one or more HRD parameters.
[0172] Thus, the decoding device in the embodiment of the present disclosure can decode the parameter indicating whether the HRD parameter exists in the picture timing SEI message or the decoding unit information SEI message. Thus, the decoding device can decode the HRD parameter immediately after receiving the SEI message.
[0173] In addition, in the decoding device in the embodiment of the present disclosure, when the multiple HRD parameters exist, the circuit decodes the parameter indicating whether the multiple HRD parameters exist in the picture timing SEI message or the decoding unit information SEI message from the buffering period SEI message.
[0174] Thus, the decoding device in the embodiment of the present disclosure can decode the parameter indicating whether the HRD parameter exists in the picture timing SEI message or the decoding unit information SEI message. Thus, the decoding device can decode the HRD parameter so that it is parsed immediately after receiving the SEI message.
[0175] In addition, the encoding method in the embodiment of the present disclosure encodes one or more HRD parameters independently of the sequence parameter set into one or more HRD-associated SEI messages, wherein the one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with HRD.
[0176] Therefore, the encoding method in the embodiment of the present disclosure can achieve the same effect as the above-mentioned encoding device.
[0177] In addition, the decoding method in the embodiment of the present invention decodes one or more HRD parameters from one or more HRD-associated SEI messages independently of the sequence parameter set, wherein the one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with HRD.
[0178] Therefore, the decoding method in the embodiment of the present disclosure can achieve the same effect as the above-mentioned decoding device.
[0179] Furthermore, for example, an encoding device according to one aspect of the present disclosure may include a partitioning unit, an intra-frame prediction unit, an inter-frame prediction unit, a loop filter unit, a transform unit, a quantization unit, and an entropy encoding unit.
[0180] The segmentation unit may segment the picture into a plurality of blocks. The intra-frame prediction unit may perform intra-frame prediction on a block included in the plurality of blocks. The inter-frame prediction unit may perform inter-frame prediction on the block. The transformation unit may transform a prediction error between a predicted image obtained by the intra-frame prediction or the inter-frame prediction and an original image to generate transformation coefficients. The quantization unit may quantize the transformation coefficients to generate quantized coefficients. The entropy coding unit may encode the quantized coefficients to generate a coded bitstream. The loop filtering unit may apply a filter to the reconstructed image of the block.
[0181] Furthermore, for example, the encoding device may be an encoding device that encodes a moving image including a plurality of pictures.
[0182] Moreover, it may also be that the entropy coding unit encodes one or more HRD parameters independently of the sequence parameter set into one or more HRD-associated SEI messages, wherein the one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with HRD.
[0183] Furthermore, for example, a decoding device according to one aspect of the present disclosure may include an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, and a loop filter unit.
[0184] The entropy decoding unit may decode quantized coefficients of a block within a picture from a coded bitstream. The inverse quantization unit may inversely quantize the quantized coefficients to obtain transform coefficients. The inverse transform unit may inversely transform the transform coefficients to obtain a prediction error. The intra-frame prediction unit may perform intra-frame prediction on the block. The inter-frame prediction unit may perform inter-frame prediction on the block. The filtering unit may apply a filter to a reconstructed image generated using a predicted image obtained through the intra-frame prediction or the inter-frame prediction and the prediction error.
[0185] Furthermore, for example, the decoding device may be a decoding device that decodes a moving image including a plurality of pictures.
[0186] Moreover, it may also be that the entropy decoding unit decodes one or more HRD parameters from one or more HRD-associated SEI messages independently of the sequence parameter set, wherein the one or more HRD parameters are one or more parameters of HRD (Hypothetical Reference Decoder) related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with HRD.
[0187] Moreover, these inclusive or specific forms can also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0188] [Definition of terms]
[0189] As an example, each term can be defined as follows.
[0190] (1) Image
[0191] A unit of data consisting of a collection of pixels, consisting of a picture or a block smaller than a picture, and includes not only moving images but also still images.
[0192] (2) Pictures
[0193] A unit of image processing consisting of a collection of pixels, sometimes called a frame or field.
[0194] (3) Block
[0195] A processing unit is a collection of a specific number of pixels, as listed in the following examples. The name is not limited. Furthermore, there are no restrictions on the shape. For example, it includes a rectangle composed of M×N pixels and a square composed of M×M pixels, as well as triangles, circles, and other shapes.
[0196] (Block example)
[0197] Slices / Tiles / Bricks
[0198] CTU / super block / basic partition unit
[0199] VPDU / Hardware processing division unit
[0200] CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transform Block Unit (TU) / Unit
[0201] ·Sub-block
[0202] (4) Pixels / sample
[0203] A dot is the smallest unit constituting an image, and includes not only pixels at integer positions but also pixels at decimal positions generated based on pixels at integer positions.
[0204] (5) Pixel value / sample value
[0205] It is the inherent value of a pixel, including brightness value, color difference value, RGB grayscale, depth value, or binary value of 0 or 1.
[0206] (6) Logo
[0207] In addition to 1 bit, multiple bits are also possible, for example, parameters or indexes of 2 or more bits may be used. In addition, not only binary values but also multiple values using numbers in other bases may be used.
[0208] (7) Signal
[0209] In order to transmit information, it is coded and encoded, and includes not only discrete digital signals but also analog signals that take continuous values.
[0210] (8) Stream / Bitstream
[0211] A stream or bitstream is a sequence of digital data or a stream of digital data. A stream or bitstream can be a single stream or can be composed of multiple streams divided into multiple layers. Furthermore, this includes transmission over a single transmission path using serial communication as well as transmission over multiple transmission paths using packet communication.
[0212] (9) Difference / Differential
[0213] In the case of scalars, in addition to simple difference (xy), any difference operation can be performed, including the absolute value of the difference (|xy|), square difference (x^2-y^2), square root of the difference (√(xy)), weighted difference (ax-by: a and b are constants), and offset difference (x-y+a: a is the offset).
[0214] (10) and
[0215] In the case of scalars, in addition to the simple sum (x+y), any sum operation can be performed, including the absolute value of the sum (|x+y|), the square sum (x^2+y^2), the square root of the sum (√(x+y)), the weighted sum (ax+by: a and b are constants), and the offset sum (x+y+a: a is the offset).
[0216] (11) based on
[0217] This also includes cases where elements other than the one being based are added. In addition, in addition to cases where a direct result is obtained, this also includes cases where a result is obtained via an intermediate result.
[0218] (12) Use (used, using)
[0219] This also includes the case where elements other than the ones to be used are added. In addition, in addition to the case where a direct result is obtained, the case where a result is obtained via an intermediate result is also included.
[0220] (13) prohibit (prohibit, forbid)
[0221] It can also be called non-permission. In addition, non-prohibition or permission does not necessarily mean obligation.
[0222] (14) Restriction (limit, restriction / restrict / restricted)
[0223] It can also be called "not allowed." Furthermore, not prohibiting or allowing something does not necessarily mean an obligation. Furthermore, a partial prohibition in terms of quantity or quality is sufficient, and this includes a complete prohibition.
[0224] (15) Chromatic aberration (chroma)
[0225] An adjective denoted by the symbols Cb and Cr that specifies that a sample array or a single sample represents one of two color difference signals associated with a primary color. Furthermore, the term "chrominance" can be used instead of the term "chroma."
[0226] (16) Luma
[0227] An adjective that specifies that a sample arrangement or a single sample represents a monochromatic signal associated with a primary color, and is represented by the symbol or subscript Y or L. The term luminance can also be used instead of luma.
[0228] [Explanation related to the record]
[0229] In the drawings, the same reference numerals represent the same or similar components. In addition, the sizes and relative positions of the components in the drawings are not necessarily drawn to scale.
[0230] The following embodiments are described in detail with reference to the accompanying drawings. The embodiments described below are intended to be inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connections of components, steps, and the relationships and order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.
[0231] The following describes embodiments of encoding and decoding devices. The embodiments are examples of encoding and decoding devices to which the processing and / or structures described in each aspect of this disclosure can be applied. The processing and / or structures can also be implemented in encoding and decoding devices that differ from the embodiments. For example, the processing and / or structures applied to the embodiments may include any of the following.
[0232] (1) Any of the multiple components of the encoding device or decoding device described in the embodiments of the present disclosure may be replaced by another component described in any of the embodiments of the present disclosure, or these components may be combined.
[0233] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be modified by adding, replacing, deleting, or other arbitrary changes. For example, any function or process may be replaced by another function or process described in any of the various aspects of the present disclosure, or these functions or processes may be combined.
[0234] (3) In the method implemented by the encoding device or decoding device of the embodiment, any changes such as addition, replacement, or deletion may be made to a portion of the multiple processes included in the method. For example, any process in the method may be replaced with another process described in any of the various aspects of the present disclosure, or they may be combined.
[0235] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with a component described in any of the aspects of the present disclosure, may be combined with a component having a portion of the functions described in any of the aspects of the present disclosure, or may be combined with a component that implements a portion of the processing implemented by a component described in any of the aspects of the present disclosure.
[0236] (5) A component having a portion of the functions of the encoding device or decoding device of the embodiment, or a component implementing a portion of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced with a component described in any of the aspects of the present disclosure, a component having a portion of the functions described in any of the aspects of the present disclosure, or a component implementing a portion of the processing described in any of the aspects of the present disclosure;
[0237] (6) In a method implemented by an encoding device or decoding device of an embodiment, one of the multiple processes included in the method is replaced with a process described in each aspect of the present disclosure or a similar process, or a combination of these processes;
[0238] (7) Some of the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the process described in any of the aspects of the present disclosure.
[0239] (8) The implementation of the processes and / or structures described in the various aspects of this disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or structures may be implemented in a device used for a purpose different from that of the video encoding or video decoding disclosed in the embodiments.
[0240] [System Structure]
[0241] Figure 1 This is a schematic diagram showing an example of the configuration of the transmission system according to this embodiment.
[0242] The transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs is, for example, Figure 1 As shown, it includes an encoding device 100, a network Nw and a decoding device 200.
[0243] An image is input to the encoding device 100. The encoding device 100 encodes the input image to generate a stream and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. This encoding compresses the image.
[0244] In addition, the original image input to the encoding device 100 before being encoded is also called the original image, original signal, or original sample. In addition, the image can be a moving image or a still image. In addition, the image is a general concept of sequence, picture, block, etc., and is not restricted by spatial and temporal regions unless otherwise specified. In addition, an image is composed of an arrangement of pixels or pixel values, and the signal or pixel value representing the image is also called a sample. In addition, a stream can be called a bit stream, a coded bit stream, a compressed bit stream, or a coded signal. Furthermore, the encoding device can also be called an image encoding device or a moving image encoding device, and the encoding method of the encoding device 100 can also be called an encoding method, an image encoding method, or a moving image encoding method.
[0245] Network Nw transmits the stream generated by encoding device 100 to decoding device 200. Network Nw can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network Nw is not necessarily limited to a bidirectional communication network and can also be a unidirectional communication network that transmits broadcast waves, such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, network Nw can be replaced by a storage medium that records streams, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0246] The decoding device 200 decodes a stream transmitted via the network Nw to generate a decoded image, for example, an uncompressed image. For example, the decoding device decodes the stream using a decoding method corresponding to the encoding method of the encoding device 100 .
[0247] In addition, the decoding device may also be referred to as an image decoding device or a moving image decoding device, and the decoding method of the decoding device 200 may also be referred to as a decoding method, an image decoding method, or a moving image decoding method.
[0248] [Data structure]
[0249] Figure 2 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. The video sequence is, for example, Figure 2 As shown in (a), it includes VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information) and multiple pictures.
[0250] The VPS includes coding parameters common to the multiple layers in a moving picture composed of multiple layers and coding parameters associated with the multiple layers included in the moving picture or with each layer.
[0251] The SPS includes parameters used for the sequence, that is, encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may also represent the width or height of the picture. In addition, there may be multiple SPSs.
[0252] The PPS contains parameters used for pictures, namely, coding parameters that decoding apparatus 200 references to decode each picture in a sequence. For example, these coding parameters may include a reference value for the quantization width used in picture decoding and a flag indicating the use of weighted prediction. Furthermore, multiple PPSs may exist. SPSs and PPSs are sometimes referred to simply as parameter sets.
[0253] like Figure 2 As shown in (b) of FIG. 1 , a picture may include a picture header and one or more slices. The picture header includes coding parameters that the decoding apparatus 200 refers to in order to decode the one or more slices.
[0254] like Figure 2 As shown in (c), a slice includes a slice header and one or more tiles. The slice header includes coding parameters that the decoding device 200 refers to in order to decode the one or more tiles.
[0255] like Figure 2 As shown in (d), a brick includes one or more CTUs (Coding Tree Units).
[0256] Furthermore, a picture may not contain slices but may contain tile groups instead of slices. In this case, a tile group contains one or more tiles. Alternatively, a tile may contain a slice.
[0257] CTU is also called super block or basic partition unit. Figure 2 As shown in (e) of FIG. 1 , such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes coding parameters that the decoding apparatus 200 refers to in order to decode one or more CUs.
[0258] A CU can also be split into multiple small CUs. Figure 2As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information used to predict the CU, and the residual coefficient information is information representing the prediction residual described later. In addition, the CU is basically the same as the PU (Prediction Unit) and the TU (Transform Unit), but for example, in the SBT described later, it may also include multiple TUs smaller than the CU. In addition, the CU may also process each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. The VPDU is a fixed unit that can be processed in one stage when pipeline processing is performed in hardware, for example.
[0259] In addition, the stream may not have Figure 2 The hierarchy of any part of the hierarchies shown. In addition, the order of these hierarchies can be swapped, and any hierarchy can be replaced by another hierarchy. In addition, the picture of the object being processed by the apparatus such as the encoding apparatus 100 or the decoding apparatus 200 at the current time point is referred to as the current picture. If the processing is encoding, the current picture is synonymous with the encoding object picture, and if the processing is decoding, the current picture is synonymous with the decoding object picture. In addition, the block of the object such as a CU or a CU etc. being processed by the apparatus such as the encoding apparatus 100 or the decoding apparatus 200 at the current time point is referred to as the current block. If the processing is encoding, the current block is synonymous with the encoding object block, and if the processing is decoding, the current block is synonymous with the decoding object block.
[0260] [Structural slices / tiles of the image]
[0261] In order to decode pictures in parallel, pictures are sometimes composed of slice units or tile units.
[0262] A slice is the basic coding unit that makes up a picture. For example, a picture is composed of one or more slices. In addition, a slice is composed of one or more consecutive CTUs.
[0263] Figure 3This is a diagram showing an example of the structure of a slice. For example, a picture contains 11×8 CTUs and is divided into 4 slices (slices 1-4). Slice 1 is composed of 16 CTUs, slice 2 is composed of 21 CTUs, slice 3 is composed of 29 CTUs, and slice 4 is composed of 22 CTUs. Here, each CTU in the picture belongs to a certain slice. The shape of the slice is a shape that divides the picture in the horizontal direction. The boundary of the slice does not need to be the end of the picture, but can also be somewhere in the boundary of the CTU in the picture. The processing order (encoding order or decoding order) of the CTU in the slice is, for example, a raster scan order. In addition, a slice includes a slice header and encoded data. The slice header can also record the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.
[0264] A tile is a unit of rectangular area that constitutes a picture. Each tile may be assigned a number called a TileId in raster scan order.
[0265] Figure 4 This is a diagram showing an example of a tile structure. For example, a picture contains 11×8 CTUs and is divided into four rectangular area tiles (tiles 1-4). When tiles are used, the processing order of CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed, for example, in raster scan order. When tiles are used, at least one CTU is processed in each of the multiple tiles, for example, in raster scan order. For example, Figure 4 As shown, the processing order of multiple CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.
[0266] In addition, one tile may contain more than one slice, and one slice may contain more than one tile.
[0267] In addition, a picture may also be composed of tile set units. A tile set may contain one or more tile groups, or may contain one or more tiles. A picture may be composed of only one of a tile set, a tile group, and a tile. For example, the order in which a plurality of tiles are scanned in raster order for each tile set is set as the basic encoding order of the tiles. A collection of one or more tiles having a continuous basic encoding order in each tile set is set as a tile group. Such a picture may also be generated by the segmentation unit 102 described later (see Figure 7 )constitute.
[0268] [Scalable Coding]
[0269] Figure 5 and Figure 6This is a diagram showing an example of the structure of a scalable stream.
[0270] like Figure 5 As shown, encoding device 100 can generate a temporally and spatially scalable stream by encoding multiple pictures by dividing them into one of multiple layers. For example, encoding device 100 achieves scalability by encoding each picture layer, with enhancement layers positioned above a base layer. This type of encoding of each picture is called scalable encoding. This allows decoding device 200 to switch the image quality of the image displayed by decoding the stream. Specifically, decoding device 200 determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication band. As a result, decoding device 200 can freely switch between decoding the same content at low and high resolutions. For example, a user of this stream might be viewing a moving image on a smartphone while on the move, then viewing the remainder of the moving image on a device such as an internet TV after returning home. Furthermore, the smartphone and the device may each be equipped with a decoding device 200 with either the same or different performance. In this case, if the device decodes the higher-level layer of the stream, the user can view high-quality moving images upon returning home. This eliminates the need for the encoding device 100 to generate multiple streams having the same content but different image quality, thereby reducing the processing load.
[0271] Furthermore, the enhancement layer may also include metadata such as statistical information based on the image. Alternatively, the decoding device 200 may generate a high-definition moving image by super-resolutioning the base layer picture based on the metadata. Super-resolution may be achieved by either increasing the signal-to-noise ratio at the same resolution or increasing the resolution. Meta-information includes information for determining linear or nonlinear filter coefficients used in the super-resolution process, as well as information for determining parameter values in filtering, machine learning, or least-squares operations used in the super-resolution process.
[0272] Alternatively, the image may be divided into tiles according to the meaning of each object in the image. In this case, the decoding device 200 can decode only a portion of the area in the image by selecting the tile to be decoded. Furthermore, the attributes of the object (person, car, ball, etc.) and the position in the image (coordinate position in the same image, etc.) can be saved as meta information. In this case, the decoding device 200 can determine the position of the desired object based on the meta information and decide the tile containing the object. For example, Figure 6 As shown, a data storage structure different from pixel data such as SEI in HEVC can also be used to store meta-information. This meta-information, for example, indicates the position, size, or color of the main object.
[0273] Alternatively, the meta-information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. This allows the decoding device 200 to determine the time at which a specific person appears in a moving image, and, using this time and picture-based information, to identify the picture in which the object appears and the location of the object within that picture.
[0274] [Encoding device]
[0275] Next, the encoding device 100 according to the embodiment will be described. Figure 7 1 is a block diagram showing an example of the functional configuration of the encoding device 100 according to the embodiment. The encoding device 100 encodes an image in units of blocks.
[0276] like Figure 7 As shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. The intra-frame prediction unit 124 and the inter-frame prediction unit 126 each constitute a part of a prediction processing unit.
[0277] [Encoding device installation example]
[0278] Figure 8 1 is a block diagram showing an implementation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. For example, Figure 7 The multiple components of the encoding device 100 shown are Figure 8 The processor a1 and the memory a2 shown are implemented.
[0279] Processor a1 is a circuit that processes information and can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes images. Processor a1 can also be a processor such as a CPU. In addition, processor a1 can also be a collection of multiple electronic circuits. In addition, for example, processor a1 can also play a role in Figure 7 The functions of the multiple components of the encoding device 100 excluding the components for storing information are shown.
[0280] Memory a2 is a dedicated or general-purpose memory that stores information used by processor a1 to encode images. Memory a2 can be an electronic circuit or connected to processor a1. Alternatively, memory a2 can be included in processor a1. Alternatively, memory a2 can be a collection of multiple electronic circuits. Memory a2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory a2 can be either non-volatile or volatile memory.
[0281] For example, the memory a2 may store coded images or a stream corresponding to the coded images. In addition, the memory a2 may store a program for the processor a1 to encode images.
[0282] In addition, for example, memory a2 can also serve as Figure 7 The memory a2 can be used to store information among the multiple components of the encoding device 100 shown in FIG. Figure 7 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed images (specifically, reconstructed blocks and reconstructed pictures, etc.).
[0283] In addition, in the encoding device 100, it is not necessary to install Figure 7 Not all of the multiple components shown may perform all of the multiple processes described above. Figure 7 Some of the multiple components shown may be included in other devices, and some of the multiple processes described above may be executed by other devices.
[0284] Below, after describing the overall processing flow of the encoding device 100 , each component included in the encoding device 100 will be described.
[0285] [Overall flow of encoding processing]
[0286] Figure 9 3 is a flowchart showing an example of the overall encoding process performed by the encoding device 100 .
[0287] First, the segmentation unit 102 of the encoding device 100 segments the picture contained in the original image into a plurality of fixed-size blocks (128×128 pixels) (step Sa_1). The segmentation unit 102 then selects a segmentation pattern for the fixed-size blocks (step Sa_2). Specifically, the segmentation unit 102 further segments the fixed-size blocks into a plurality of blocks that conform to the selected segmentation pattern. The encoding device 100 then performs steps Sa_3 through Sa_9 on each of the plurality of blocks.
[0288] The prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126 and the prediction control unit 128 generates a predicted image of the current block (step Sa_3). The predicted image is also called a prediction signal, a prediction block, or a prediction sample.
[0289] Next, the subtraction unit 104 generates the difference between the current block and the predicted image as a prediction residual (step Sa_4). The prediction residual is also called a prediction error.
[0290] Next, the transform unit 106 and the quantization unit 108 transform and quantize the predicted image to generate a plurality of quantization coefficients (step Sa_5).
[0291] Next, the entropy coding unit 110 generates a stream by encoding (specifically, entropy coding) the plurality of quantized coefficients and prediction parameters related to generation of a predicted image (step Sa_6).
[0292] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore the prediction residual by performing inverse quantization and inverse transformation on the plurality of quantized coefficients (step Sa_7).
[0293] Next, the adder 116 reconstructs the current block by adding the predicted image to the restored prediction residual (step Sa_8). This generates a reconstructed image. Furthermore, a reconstructed image is also referred to as a reconstructed block. In particular, a reconstructed image generated by the encoding device 100 is also referred to as a locally decoded block or locally decoded image.
[0294] When the reconstructed image is generated, the loop filter unit 120 filters the reconstructed image as necessary (step Sa_9).
[0295] Then, the encoding device 100 determines whether encoding of the entire picture is completed (step Sa_10 ), and if it is determined that encoding is not completed (No in step Sa_10 ), it repeats the process from step Sa_2 .
[0296] Furthermore, in the above example, encoding device 100 selects a single partitioning pattern for fixed-size blocks and encodes each block according to that partitioning pattern. However, encoding device 100 may also encode each block according to each of multiple partitioning patterns. In this case, encoding device 100 may evaluate the cost for each of the multiple partitioning patterns and, for example, select the stream obtained by encoding according to the partitioning pattern with the lowest cost as the final output stream.
[0297] Furthermore, the processing of steps Sa_1 to Sa_10 may be performed sequentially by the encoding device 100 , or a portion or a plurality of these processes may be performed in parallel or in a different order.
[0298] The encoding process performed by encoding device 100 is a hybrid encoding process using predictive coding and transform coding. Predictive coding is performed using a coding loop comprised of a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128. Specifically, the prediction processing unit comprised of the intra-frame prediction unit 124 and the inter-frame prediction unit 126 constitutes part of the coding loop.
[0299] [Division]
[0300] The segmentation unit 102 divides each picture included in the original image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first divides the picture into blocks of a fixed size (for example, 128×128 pixels). The fixed-size blocks are sometimes called coding tree units (CTUs). Furthermore, the segmentation unit 102 divides each fixed-size block into blocks of a variable size (for example, less than 64×64 pixels) based on, for example, recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. The variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be used as processing units of CUs, PUs, or TUs.
[0301] Figure 10 : is a diagram showing an example of block division in an embodiment. Figure 10 In the figure, the solid line represents the block boundary based on quadtree block partitioning, and the dotted line represents the block boundary based on binary tree block partitioning.
[0302] Here, the block 10 is a square block of 128 x 128 pixels. The block 10 is first partitioned into four square blocks of 64 x 64 pixels (quadtree block partitioning).
[0303] The upper left 64×64 pixel square block is further vertically divided into two rectangular blocks of 32×64 pixels each, and the left 32×64 pixel square block is further vertically divided into two rectangular blocks of 16×64 pixels each (binary tree block division). As a result, the upper left 64×64 pixel square block is divided into two 16×64 pixel rectangular blocks 11 and 12, and a 32×64 pixel rectangular block 13.
[0304] The upper right square block of 64×64 pixels is horizontally divided into two rectangular blocks 14 and 15 each consisting of 64×32 pixels (binary tree block division).
[0305] The 64×64 pixel square block in the lower left is divided into four square blocks of 32×32 pixels each (quadtree block division). The upper left block and the lower right block of the four square blocks of 32×32 pixels each are further divided. The 32×32 pixel square block in the upper left is vertically divided into two rectangular blocks of 16×32 pixels each, and the rectangular block of 16×32 pixels on the right is further horizontally divided into two square blocks of 16×16 pixels each (binary tree block division). The 32×32 pixel square block in the lower right is horizontally divided into two rectangular blocks of 32×16 pixels each (binary tree block division). As a result, the 64×64 pixel square block in the lower left corner is divided into a rectangular block 16 of 16×32 pixels, two square blocks 17 and 18 of 16×16 pixels each, two square blocks 19 and 20 of 32×32 pixels each, and two rectangular blocks 21 and 22 of 32×16 pixels each.
[0306] The lower right block 23 consisting of 64×64 pixels is not divided.
[0307] As above, in Figure 10 In FIG, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. This type of partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.
[0308] In addition, Figure 10 In the example above, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). Division including this ternary tree division is called MBT (multi type tree) division.
[0309] Figure 11 1 is a diagram showing an example of the functional structure of the segmentation unit 102. Figure 11 As shown, the division unit 102 may include a block division determination unit 102a. As an example, the block division determination unit 102a may perform the following processing.
[0310] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122 and determines the above-mentioned division pattern based on the block information. The division unit 102 divides the original image according to the division pattern and outputs one or more blocks obtained by the division to the subtraction unit 104.
[0311] Furthermore, the block division determination unit 102a outputs, for example, parameters indicating the above-described division pattern to the transformation unit 106, the inverse transformation unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy coding unit 110. The transformation unit 106 can transform the prediction residual based on the parameters, and the intra prediction unit 124 and the inter prediction unit 126 can generate a predicted image based on the parameters. Furthermore, the entropy coding unit 110 can also perform entropy coding on the parameters.
[0312] As an example, parameters related to the segmentation pattern may be written into the stream as follows.
[0313] Figure 12 This figure shows examples of partitioning patterns. Examples include: four-partitioning (QT), which divides a block into two in both the horizontal and vertical directions; three-partitioning (HT or VT), which divides a block in the same direction at a ratio of 1:2:1; two-partitioning (HB or VB), which divides a block in the same direction at a ratio of 1:1; and no partitioning (NS).
[0314] In the case of four-way division and no division, the division pattern does not have the block division direction, while in the case of two-way division and three-way division, the division pattern has division direction information.
[0315] Figure 13A and Figure 13B This is a diagram showing an example of a syntax tree of a segmentation pattern. Figure 13A In the example, first, there is information indicating whether to split (S: Split flag), then there is information indicating whether to split into four (QT: QT flag). Next, there is information indicating whether to split into three or two (TT: TT flag or BT: BT flag), and finally there is information indicating the direction of splitting (Ver: Vertical flag or Hor: Horizontal flag). In addition, for each of more than one blocks obtained by such splitting based on the splitting style, the same process can be repeatedly applied to splitting. That is, as an example, it is also possible to recursively implement the judgment of whether to split, whether to split into four, whether the splitting method is horizontal or vertical, and whether to split into three or two, and the judgment result of the implementation is calculated according to the result of the splitting. Figure 13A The encoding order disclosed by the syntax tree shown is encoded into the stream.
[0316] In addition, Figure 13A In the syntax tree shown, these information are arranged in the order of S, QT, TT, Ver, but they can also be arranged in the order of S, QT, Ver, BT. Figure 13BIn the example, first, there is information indicating whether to split (S: Split flag), then information indicating whether to split into four (QT: QT flag). Next, there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag), and finally, there is information indicating whether to split into two or three (BT: BT flag or TT: TT flag).
[0317] The division patterns described here are merely examples, and a division pattern other than the described division pattern may be used, or only a part of the described division pattern may be used.
[0318] [Subtraction Department]
[0319] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in units of blocks input from and divided by the division unit 102. Specifically, the subtraction unit 104 calculates a prediction residual for the current block. The subtraction unit 104 then outputs the calculated prediction residual to the transformation unit 106.
[0320] The original image is an input signal of the encoding device 100 , and is, for example, a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two chroma signals).
[0321] [Conversion Unit]
[0322] The transform unit 106 transforms the spatial domain prediction residual into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction residual.
[0323] Alternatively, the transform unit 106 may adaptively select a transform type from a plurality of transform types and transform the prediction residual into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes referred to as an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT). Furthermore, the transform basis function is sometimes referred to as simply a basis.
[0324] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Furthermore, these transform types may be described as DCT2, DCT5, DCT8, DST1, and DST7, respectively. Figure 14is a table showing the transformation basis functions corresponding to each transformation type. Figure 14 Where N represents the number of input pixels. The selection of a transform type from among these multiple transform types may depend on, for example, the type of prediction (intra-frame prediction, inter-frame prediction, etc.) or the intra-frame prediction mode.
[0325] Information indicating whether such EMT or AMT is applied (e.g., an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. However, signaling of this information is not limited to the CU level and may also be performed at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0326] In addition, the transform unit 106 may also re-transform the transform coefficients (i.e., the transform results). Such re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 re-transforms each sub-block (e.g., a sub-block of 4×4 pixels) contained in the block of transform coefficients corresponding to the intra-frame prediction residual. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but may also be other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).
[0327] The transform unit 106 can also apply separable transforms and non-separable transforms. Separable transforms are a method of performing multiple transforms in each direction corresponding to the number of input dimensions. Non-separable transforms are a method of treating two or more dimensions as a single dimension and performing the transforms together when the input is multidimensional.
[0328] For example, as an example of non-separable transformation, when a 4×4 pixel block is input, it is considered as an array of 16 elements, and the array is transformed using a 16×16 transformation matrix.
[0329] In a further example of the non-separable transformation, a 4×4 pixel input block may be regarded as an array of 16 elements, and then a transformation (Hypercube Givens Transform) may be performed by performing multiple Givens rotations on the array.
[0330] In the transformation performed by the transformation unit 106 , the type of transformation of the transformation basis function to be transformed into the frequency domain can be switched according to the region within the CU. An example of this type of transformation is SVT (Spatially Varying Transform).
[0331] Figure 15 This is a diagram showing an example of SVT.
[0332] In SVT, Figure 15 As shown, the CU is divided into two equal parts in the horizontal direction or the vertical direction, and only the area on either side is transformed into the frequency domain. The transformation type can be set for each area, for example, using DST7 and DCT8. For example, for the area at position 0 in the two areas obtained by dividing the CU into two equal parts in the vertical direction, DST7 and DCT8 can be used. Or, for the area at position 1 in the two areas, DST7 is used. Similarly, for the area at position 0 in the two areas obtained by dividing the CU into two equal parts in the horizontal direction, DST7 and DCT8 are used. Or, for the area at position 1 in the two areas, DST7 is used. In such a Figure 15 In the example shown, only one of the two regions within the CU is transformed, while the other is not. However, it is also possible to transform both regions separately. In addition, the partitioning method is not limited to bisection; it can also be quartered. In addition, it is possible to be more flexible and encode the information indicating the partitioning method and perform signaling in the same way as CU partitioning. In addition, SVT is sometimes also called SBT (Sub-block Transform).
[0333] The aforementioned AMT and EMT can also be referred to as MTS (Multiple Transform Selection). When MTS is applied, a transform type such as DST7 or DCT8 can be selected, and information indicating the selected transform type can be encoded as index information for each CU. On the other hand, as a process of selecting the transform type used in the orthogonal transform based on the shape of the CU without encoding the index information, there is a process called IMTS (Implicit MTS). When IMTS is applied, for example, if the shape of the CU is rectangular, DST7 is used on the short side of the rectangle and DCT2 is used on the long side, and orthogonal transforms are performed separately. In addition, for example, when the shape of the CU is square, if MTS is valid within the sequence, DCT2 is used for orthogonal transform, and if MTS is invalid, DST7 is used for orthogonal transform. DCT2 and DST7 are examples, and other transform types can be used, or the combination of transform types used can be set to different combinations. IMTS can be used only in intra-frame predicted blocks, or it can be used together in intra-frame predicted blocks and inter-frame predicted blocks.
[0334] In the above, the three processes of MTS, SBT, and IMTS have been described as the selection process for selectively switching the transform type used in the orthogonal transform. However, all three selection processes may be valid, or only some of the selection processes may be selectively valid. Whether each selection process is valid can be identified by flag information in the header such as the SPS. For example, if all three selection processes are valid, one of the three selection processes is selected for orthogonal transform in units of CUs. In addition, as long as the selection process for selectively switching the transform type can achieve at least one of the following four functions [1] to [4], a selection process different from the above three selection processes may be used, or each of the above three selection processes may be replaced by another process. Function [1] is a function of performing an orthogonal transform on the entire range within the CU and encoding information indicating the transform type used in the transform. Function [2] is a function of performing an orthogonal transform on the entire range of the CU, not encoding information indicating the transform type, and determining the transform type based on a prescribed rule. Function [3] is a function of performing an orthogonal transform on a part of the CU and encoding information indicating the transform type used in the transform. Function [4] is a function of performing an orthogonal transform on a part of the CU area, and determining the transform type based on a predetermined rule without encoding information indicating the transform type used in the transform.
[0335] Furthermore, the presence or absence of each of MTS, IMTS, and SBT can be determined for each processing unit, for example, in sequence units, picture units, tile units, slice units, CTU units, or CU units.
[0336] In addition, the tool for selectively switching the transformation type in the present disclosure can also be renamed as a method for adaptively selecting a basis used in the transformation process, a selection process, or a process for selecting a basis. In addition, the tool for selectively switching the transformation type can also be renamed as a mode for adaptively selecting the transformation type.
[0337] Figure 16 This is a flowchart showing an example of processing performed by the conversion unit 106.
[0338] For example, the transform unit 106 determines whether to perform an orthogonal transform (step St_1). If the decision is to perform an orthogonal transform (yes in step St_1), the transform unit 106 selects a transform type from a plurality of transform types (step St_2). Next, the transform unit 106 performs an orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transform unit 106 then outputs information indicating the selected transform type to the entropy coding unit 110, which encodes the information (step St_4). On the other hand, if the decision is not to perform an orthogonal transform (no in step St_1), the transform unit 106 outputs information indicating that an orthogonal transform is not performed to the entropy coding unit 110, which encodes the information (step St_5). The decision to perform an orthogonal transform in step St_1 can be based on, for example, the size of the transform block, the prediction mode applied to the CU, and so on. Alternatively, information indicating the transform type used for the orthogonal transform may not be encoded, and an orthogonal transform may be performed using a predetermined transform type.
[0339] Figure 17 : is a flowchart showing another example of the processing performed by the conversion unit 106. Figure 17 The example shown is similar to Figure 16 The example shown is also an example of orthogonal transform in the case of applying a method of selectively switching the transform type used in the orthogonal transform.
[0340] As an example, the first transform type group may include DCT2, DCT7, and DCT8. Furthermore, as an example, the second transform type group may include DCT2. Furthermore, the transform types included in the first and second transform type groups may overlap in part or may be completely different.
[0341] Specifically, the transform unit 106 determines whether the transform size is less than or equal to a predetermined value (step Su_1). If it is determined to be less than or equal to the predetermined value (yes in step Su_1), the transform unit 106 performs an orthogonal transform on the prediction residual of the current block using the transform type included in the first transform type group (step Su_2). Next, the transform unit 106 outputs information indicating which transform type from the one or more transform types included in the first transform type group is to be used to the entropy coding unit 110, thereby encoding the information (step Su_3). On the other hand, if it is determined that the transform size is not less than or equal to the predetermined value (no in step Su_1), the transform unit 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4).
[0342] In step Su_3, the information indicating the transform type used for the orthogonal transform may be information indicating a combination of a transform type applied to the vertical direction and a transform type applied to the horizontal direction of the current block. Furthermore, the first transform type group may include only one transform type, and the information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and information indicating the transform type used for the orthogonal transform among one or more transform types included in the second transform type group may be encoded.
[0343] Alternatively, the transform type may be determined based solely on the transform size. Furthermore, if the transform type used for orthogonal transform is determined based on the transform size, the determination is not limited to determining whether the transform size is equal to or smaller than a predetermined value.
[0344] [Quantitative Department]
[0345] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the multiple transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the multiple quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.
[0346] The predetermined scanning order is the order used for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order is defined by ascending order (from low frequency to high frequency) or descending order (from high frequency to low frequency) of frequency.
[0347] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. In other words, if the value of the quantization parameter increases, the error in the quantization coefficient (quantization error) increases.
[0348] Quantization also sometimes uses a quantization matrix. For example, multiple quantization matrices may be used to correspond to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luma and chroma. Furthermore, quantization involves digitizing values sampled at predetermined intervals and assigning them to predetermined levels. In this technical field, rounding, rounding, or scaling are sometimes used.
[0349] Methods for using a quantization matrix include using a quantization matrix directly set on the encoding device 100 side and using a default quantization matrix (default matrix). By directly setting the quantization matrix on the encoding device 100 side, a quantization matrix suitable for the characteristics of the image can be set. However, this method has the disadvantage of increasing the amount of code due to encoding the quantization matrix. Alternatively, rather than directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix used for quantization of the current block can be generated based on the default quantization matrix or the encoded quantization matrix.
[0350] On the other hand, there is also a method of performing quantization so that the coefficients of high-frequency components and low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to using a quantization matrix (flat matrix) in which all coefficients have the same value.
[0351] The quantization matrix can be encoded at, for example, sequence level, picture level, slice level, brick level, or CTU level.
[0352] When using a quantization matrix, the quantization unit 108, for example, scales the quantization width calculated based on the quantization parameter for each transform coefficient using the value of the quantization matrix. Quantization processing performed without using a quantization matrix may also be processing that quantizes the transform coefficients based on the quantization width calculated based on the quantization parameter. Furthermore, in quantization processing without using a quantization matrix, the quantization width may be multiplied by a predetermined value common to all transform coefficients in a block.
[0353] Figure 18 This is a block diagram showing an example of the functional configuration of the quantization unit 108 .
[0354] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108 a , a predicted quantization parameter generation unit 108 b , a quantization parameter generation unit 108 c , a quantization parameter storage unit 108 d , and a quantization processing unit 108 e .
[0355] Figure 19 This is a flowchart showing an example of quantization performed by the quantization unit 108 .
[0356] As an example, the quantization unit 108 may be based on Figure 19 The flowchart shown in the figure performs quantization on a per-CU basis. Specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). If the determination is to perform quantization (yes in step Sv_1), the quantization parameter generator 108c generates a quantization parameter for the current block (step Sv_2) and stores the quantization parameter in the quantization parameter storage unit 108d (step Sv_3).
[0357] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameter generated in step Sv_2 (step Sv_4). The predicted quantization parameter generation unit 108b then obtains the quantization parameter for a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). Based on the obtained quantization parameter, the predicted quantization parameter generation unit 108b generates a predicted quantization parameter for the current block (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameter for the current block generated by the quantization parameter generation unit 108c and the predicted quantization parameter for the current block generated by the predicted quantization parameter generation unit 108b (step Sv_7). This difference calculation generates a differential quantization parameter. The differential quantization parameter generation unit 108a outputs the differential quantization parameter to the entropy coding unit 110, which encodes the differential quantization parameter (step Sv_8).
[0358] In addition, the differential quantization parameter can also be encoded at the sequence level, picture level, slice level, brick level, or CTU level. In addition, the initial value of the quantization parameter can be encoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0359] Furthermore, the quantization unit 108 may include a plurality of quantizers, and may apply dependent quantization for quantizing transform coefficients using a quantization method selected from a plurality of quantization methods.
[0360] [Entropy coding unit]
[0361] Figure 20 110 is a block diagram showing an example of the functional configuration of the entropy coding unit 110 .
[0362] The entropy coding unit 110 performs entropy coding on the quantization coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130, thereby generating a stream. This entropy coding employs, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding). Specifically, the entropy coding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization, converting multi-value signals such as quantization coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, and Fixed Length Binarization. The context control unit 110b derives a context value, that is, the probability of occurrence of a binary signal, based on the characteristics of a syntactic element or the surrounding conditions. Examples of methods for deriving this context value include bypass, syntactic element reference, upper / left neighboring block reference, and hierarchy information reference. The binary arithmetic coding unit 110 c performs arithmetic coding on the binarized signal using the derived context value.
[0363] Figure 21 3 is a diagram showing the flow of CABAC in the entropy coding unit 110 .
[0364] First, the CABAC in the entropy coding unit 110 is initialized. During this initialization, the binary arithmetic coding unit 110c is initialized and an initial context value is set. The binarization unit 110a and the binary arithmetic coding unit 110c then sequentially perform binarization and arithmetic coding on, for example, multiple quantized coefficients of a CTU. At this time, the context control unit 110b updates the context value each time arithmetic coding is performed. The context control unit 110b then backs off the context value as post-processing. This backed-off context value is used, for example, as the initial context value for the next CTU.
[0365] [Inverse quantization unit]
[0366] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. Furthermore, the inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114.
[0367] [Inverse transformation unit]
[0368] The inverse transform unit 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual for the current block by performing an inverse transform corresponding to the transform performed by the transform unit 106 on the transform coefficients. The inverse transform unit 114 then outputs the restored prediction residual to the adder unit 116.
[0369] Furthermore, the restored prediction residual generally loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual generally includes a quantization error.
[0370] [Addition Department]
[0371] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 to the predicted image input from the prediction control unit 128. As a result, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter unit 120.
[0372] [Block Memory]
[0373] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are blocks in the current picture. Specifically, the block memory 118 stores the reconstructed image output from the adder 116 .
[0374] [Frame Memory]
[0375] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120 .
[0376] [Loop filter unit]
[0377] The loop filter unit 120 performs loop filtering on the reconstructed image output by the adder unit 116, and outputs the filtered reconstructed image to the frame memory 122. Loop filtering refers to filtering used within the encoding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).
[0378] Figure 22 3 is a block diagram showing an example of the functional configuration of the loop filter unit 120 .
[0379] For example Figure 22As shown, the loop filter unit 120 includes a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a applies the above-mentioned deblocking filter processing to the reconstructed image. The SAO processing unit 120b applies the above-mentioned SAO processing to the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 120c applies the above-mentioned ALF processing to the reconstructed image after the SAO processing. Details about ALF and deblocking filtering will be described later. SAO processing is a process that improves image quality by reducing ringing (a phenomenon in which pixel values around edges are deformed in a fluctuating manner) and correcting deviations in pixel values. In the SAO processing, for example, there are edge offset processing and band offset processing. In addition, the loop filter unit 120 may not have Figure 22 The disclosed processing units may also include only a part of the processing units. In addition, the loop filter unit 120 may also be configured in accordance with Figure 22 A structure in which the above-mentioned respective processes are performed in an order different from the processing order disclosed in .
[0380] [Loop Filter Section > Adaptive Loop Filter]
[0381] In ALF, a least squares error filter is used to remove coding distortion. For example, for each 2×2 pixel sub-block in the current block, one filter is selected from multiple filters based on the direction and activity of the local gradient.
[0382] Specifically, a sub-block (e.g., a 2×2 pixel sub-block) is first classified into multiple classes (e.g., 15 or 25 classes). The sub-blocks are classified, for example, based on the direction and activity of the gradient. In a specific example, a classification value C (e.g., C = 5D + A) is calculated using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4). Based on the classification value C, the sub-blocks are then classified into multiple classes.
[0383] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.
[0384] Based on the result of such classification, a filter to be used for the sub-block is determined from among a plurality of filters.
[0385] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 23A to 23C FIG. 1 is a diagram showing a plurality of examples of filter shapes used in ALF. Figure 23A represents a 5×5 diamond-shaped filter, Figure 23B represents a 7×7 diamond-shaped filter, Figure 23C Represents a 9×9 diamond-shaped filter. Information representing the filter shape is typically signaled at the picture level. However, signaling of the filter shape information need not be limited to the picture level and can also be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0386] The on / off of ALF can also be determined at the picture level or the CU level. For example, regarding luminance, whether to use ALF can be determined at the CU level, and regarding chrominance, whether to use ALF can be determined at the picture level. The information indicating whether ALF is on / off is usually signaled at the picture level or the CU level. In addition, the signaling of the information indicating whether ALF is on / off does not need to be limited to the picture level or the CU level, and can also be at other levels (for example, the sequence level, the slice level, the brick level, or the CTU level).
[0387] As described above, one filter is selected from a plurality of filters to apply ALF processing to a sub-block. For each of these multiple filters (e.g., up to 15 or 25 filters), a coefficient set consisting of the multiple coefficients used in the filter is typically signaled at the picture level. However, the signaling of the coefficient set is not limited to the picture level and may be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0388] [Loop Filter > Cross Component Adaptive Loop Filter]
[0389] Figure 23D This is a diagram showing an example in which a Y sample (first component) is used for CCALF of Cb and CCALF of Cr (a plurality of components different from the first component). Figure 23E A diagram showing a diamond filter.
[0390] An example of CC-ALF is to use a linear diamond filter ( Figure 23D 、 Figure 23E) is applied to the luminance channel of each chroma component. For example, the filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed decimal point representation. The application of the filter is controlled to a variable block size and is notified by a context-coded flag received by each block of samples. The block size and CC-ALF validation flag are received at the slice level for each chroma component. The syntax and semantics of CC-ALF are provided in the Appendix. In this document, block sizes of 16x16, 32x32, 64x64, and 128x128 are supported (in chroma samples).
[0391] [Loop Filtering > Joint Chroma Cross Component Adaptive Loop Filter]
[0392] Figure 23F It is a diagram showing an example of JC-CCALF. Figure 23G This is a diagram showing examples of weight_index candidates of JC-CCALF.
[0393] One example of JC-CCALF uses only one CCALF filter, generating a single CCALF filter output as the color difference adjustment signal for just one color component, and then applying appropriately weighted versions of the same color difference adjustment signal to the other color components. This reduces the complexity of conventional CCALF to approximately half.
[0394] The weight value is encoded as a code (sign) and a weight index. The weight index (denoted as weight_index) is encoded as 3 bits and specifies the size of the JC-CCALF weight JcCcWeight. It cannot be equal to 0. The size of JcCcWeight is determined as follows.
[0395] When weight_index is 4 or less, JcCcWeight is equal to weight_index>>2.
[0396] In other cases, JcCcWeight is equal to 4 / (weight_index-4).
[0397] Block-level on / off control for the Cb and Cr ALF filters is separate. Similar to CCALF, two separate sets of block-level on / off control flags are encoded. Unlike CCALF, the on / off control blocks for Cb and Cr are the same size, so only a single block size variable is encoded.
[0398] [Loop Filter Section > Deblocking Filter]
[0399] In the deblocking filtering process, the loop filter unit 120 performs filtering on block boundaries of the reconstructed image to reduce distortion generated at the block boundaries.
[0400] Figure 24 This is a block diagram showing an example of a detailed configuration of the deblocking filter processing unit 120 a .
[0401] The deblocking filter processing unit 120 a includes, for example, a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 , and 1206 .
[0402] The boundary determination unit 1201 determines whether there is a pixel (ie, target pixel) to be subjected to deblocking filtering near a block boundary and outputs its determination result to the switch 1202 and the processing determination unit 1208 .
[0403] When the boundary determination unit 1201 determines that the target pixel exists near a block boundary, the switch 1202 outputs the image before filtering to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel does not exist near a block boundary, the switch 1202 outputs the image before filtering to the switch 1206. The image before filtering is composed of the target pixel and at least one surrounding pixel located around the target pixel.
[0404] The filter determination unit 1203 determines whether to perform deblocking filtering on the target pixel based on the pixel value of at least one surrounding pixel located around the target pixel, and then outputs the determination result to the switch 1204 and the processing determination unit 1208 .
[0405] When the filter determination unit 1203 determines that the deblocking filtering process is to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that the deblocking filtering process is not to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the switch 1206.
[0406] When the pre-filtered image is obtained via switches 1202 and 1204 , the filter processing unit 1205 performs deblocking filtering on the target pixel using the filter characteristics determined by the filter characteristic determination unit 1207 . The filter processing unit 1205 then outputs the filtered pixel to the switch 1206 .
[0407] According to the control of the processing determination unit 1208 , the switch 1206 selectively outputs pixels that have not been processed by the deblocking filter and pixels that have been processed by the deblocking filter by the filter processing unit 1205 .
[0408] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the object pixel exists near the block boundary and the filtering determination unit 1203 determines that the deblocking filtering process is to be performed on the object pixel, the processing determination unit 1208 outputs the pixel processed by the deblocking filtering from the switch 1206. In addition, in cases other than the above-mentioned cases, the processing determination unit 1208 outputs the pixel that has not been processed by the deblocking / filtering process from the switch 1206. By repeatedly outputting such pixels, the image processed by the filtering process is output from the switch 1206. In addition, Figure 24 The illustrated configuration is an example of the configuration of the deblocking filter processing unit 120 a , and the deblocking filter processing unit 120 a may have another configuration.
[0409] Figure 25 This is a diagram showing an example of a deblocking filter having a filter characteristic that is symmetric with respect to a block boundary.
[0410] In the deblocking filter process, for example, using pixel values and quantization parameters, one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter, is selected. Figure 25 As shown, when pixels p0 to p2 and pixels q0 to q2 exist across a block boundary, the pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the calculation shown in the following equation.
[0411] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0412] q'1=(p0+q0+q1+q2+2) / 4
[0413] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0414] In the above equations, p0 through p2 and q0 through q2 are the pixel values of pixels p0 through p2 and q0 through q2, respectively. Furthermore, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the side opposite the block boundary. On the right side of each of the above equations, the coefficients multiplied by the pixel values of each pixel used in the deblocking filter process are the filter coefficients.
[0415] Furthermore, during the deblocking filter process, clipping can be performed so that the calculated pixel value remains unchanged even if it exceeds a threshold. In this clipping process, the calculated pixel value based on the above equation is clipped to "pre-calculated pixel value ± 2 × threshold value" using a threshold determined by the quantization parameter. This prevents excessive smoothing.
[0416] Figure 26 This is a diagram for explaining an example of a block boundary on which a deblocking filtering process is performed. Figure 27 This is a diagram showing an example of BS value.
[0417] The block boundary for deblocking filtering is, for example, Figure 26 The deblocking filter is performed in units of, for example, 4 rows or 4 columns. First, for Figure 26 The block P and block Q shown are as follows: Figure 27 That determines the Bs (Boundary Strength) value.
[0418] according to Figure 27 The Bs value can determine whether to perform deblocking filtering with different strengths even at block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering is performed on the color difference signal. When the Bs value is 1 or more and the specified conditions are met, deblocking filtering is performed on the luminance signal. In addition, the determination conditions of the Bs value are not limited to Figure 27 The conditions shown can also be determined based on other parameters.
[0419] [Prediction Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]
[0420] Figure 28 This is a flowchart showing an example of processing performed by the prediction unit of the encoding device 100. Furthermore, as an example, the prediction unit is composed of all or part of the components of the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra-frame prediction unit 124 and the inter-frame prediction unit 126.
[0421] The prediction unit generates a predicted image for the current block (step Sb_1). Predicted images include, for example, intra-frame predicted images (intra-frame prediction signals) or inter-frame predicted images (inter-frame prediction signals). Specifically, the prediction unit generates a predicted image for the current block using a reconstructed image obtained by generating predicted images for other blocks, generating prediction residuals, generating quantized coefficients, restoring the prediction residuals, and adding the predicted images.
[0422] The reconstructed image may be, for example, an image of a reference picture or an image including the current block, that is, an image of an encoded block in the current picture (ie, the aforementioned other block). The encoded block in the current picture may be, for example, an adjacent block of the current block.
[0423] Figure 29 This is a flowchart showing another example of processing performed by the prediction unit of the encoding device 100.
[0424] The prediction unit generates a predicted image using the first method (step Sc_1a), generates a predicted image using the second method (step Sc_1b), and generates a predicted image using the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating predicted images, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may also be used in these prediction methods.
[0425] Next, the prediction unit evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, respectively (step Sc_2). For example, the prediction unit evaluates the predicted images by calculating the cost C for each of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, and comparing the costs C of these predicted images. In addition, the cost C is calculated by an equation of the RD optimization model, such as C=D+λ×R. In this equation, D is the coding distortion of the predicted image, and is represented by, for example, the sum of the absolute values of the differences between the pixel values of the current block and the pixel values of the predicted image. In addition, R is the bit rate of the stream. In addition, λ is an undetermined multiplier, such as Lagrangian.
[0426] Next, the prediction unit selects one of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c, respectively (step Sc_3). That is, the prediction unit selects a method or mode for obtaining the final prediction image. For example, the prediction unit selects a prediction image with the smallest cost C based on the cost C calculated for these prediction images. Alternatively, the evaluation of step Sc_2 and the selection of the prediction image in step Sc_3 may also be performed based on the parameters used in the encoding process. The encoding device 100 may signal information for determining the selected prediction image, method, or mode into a stream. The information may be, for example, a flag, etc. Thus, the decoding device 200 can generate a prediction image in accordance with the method or mode selected in the encoding device 100 based on the information. In addition, in Figure 29 In the example shown, the prediction unit selects one of the predicted images after generating them using various methods. However, before generating these predicted images, the prediction unit may select a method or mode based on the parameters used in the encoding process described above and generate the predicted images based on that method or mode.
[0427] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction unit may select a final prediction image for the current block from prediction images generated according to these prediction methods.
[0428] Figure 30 This is a flowchart showing another example of processing performed by the prediction unit of the encoding device 100.
[0429] First, the prediction unit generates a predicted image by intra-frame prediction (step Sd_1a), and generates a predicted image by inter-frame prediction (step Sd_1b). In addition, the predicted image generated by intra-frame prediction is also called an intra-frame predicted image, and the predicted image generated by inter-frame prediction is also called an inter-frame predicted image.
[0430] Next, the prediction unit evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). The aforementioned cost C can also be used in this evaluation. The prediction unit then selects the prediction image that yields the minimum cost C from among the intra-frame prediction image and the inter-frame prediction image as the final prediction image for the current block (step Sd_3). In other words, the prediction method or mode is selected for generating the prediction image for the current block.
[0431] [Intra-frame prediction unit]
[0432] The intra prediction unit 124 performs intra prediction (also called intra-screen prediction) on the current block by referring to blocks in the current picture stored in the block memory 118, thereby generating a predicted image (i.e., an intra-predicted image) for the current block. Specifically, the intra prediction unit 124 performs intra prediction by referring to pixel values (e.g., luminance values and chrominance values) of blocks adjacent to the current block to generate the intra-predicted image, and outputs the intra-predicted image to the prediction control unit 128.
[0433] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0434] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.
[0435] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. Furthermore, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 31This figure shows all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions specified by the H.265 / HEVC specification, and the dotted arrows represent the additional 32 directions (2 non-directional prediction modes in Figure 31 (not shown in the figure).
[0436] In various implementations, intra-frame prediction of chrominance blocks can also refer to luma blocks. That is, the chrominance component of the current block can be predicted based on the luma component of the current block. This type of intra-frame prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra-frame prediction mode for chrominance blocks that refers to luma blocks (e.g., CCLM mode) can also be added as one of the intra-frame prediction modes for chrominance blocks.
[0437] The intra-frame prediction unit 124 may also modify the intra-predicted pixel values based on the gradients of reference pixels in the horizontal and vertical directions. Intra-frame prediction with such modification is sometimes called PDPC (position-dependent intraprediction combination). Information indicating whether PDPC is used (e.g., a PDPC flag) is typically signaled at the CU level. However, this information is not necessarily signaled at the CU level and may be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0438] Figure 32 This is a flowchart showing an example of processing performed by the intra prediction unit 124 .
[0439] The intra-frame prediction unit 124 selects one intra-frame prediction mode from a plurality of intra-frame prediction modes (step Sw_1). Then, the intra-frame prediction unit 124 generates a predicted image according to the selected intra-frame prediction mode (step Sw_2). Next, the intra-frame prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM is composed of, for example, 6 intra-frame prediction modes. Two of the 6 intra-frame prediction modes can be a Planar prediction mode and a DC prediction mode, and the remaining 4 modes can be directional prediction modes. Then, the intra-frame prediction unit 124 determines whether the intra-frame prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0440] Here, if it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). In addition, the MPM flag set to 1 and the information indicating the intra prediction mode are encoded as prediction parameters by the entropy coding unit 110.
[0441] On the other hand, if it is determined that the selected intra-frame prediction mode is not included in the MPM (No in step Sw_4), the intra-frame prediction unit 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra-frame prediction unit 124 does not set the MPM flag. Then, the intra-frame prediction unit 124 generates information indicating the selected intra-frame prediction mode from one or more intra-frame prediction modes not included in the MPM (step Sw_8). In addition, the MPM flag set to 0 and the information indicating the intra-frame prediction mode are each encoded as prediction parameters by the entropy coding unit 110. The information indicating the intra-frame prediction mode represents, for example, any value from 0 to 60.
[0442] [Inter-frame prediction unit]
[0443] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-screen prediction) on the current block, referring to a reference picture different from the current picture stored in the frame memory 122, to generate a predicted image (inter-frame prediction image). Inter-frame prediction is performed on the current block or the current sub-block within the current block. A sub-block is contained within a block and is a smaller unit than a block. The size of a sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of a sub-block can also be switched on a slice, tile, or picture basis.
[0444] For example, the inter-frame prediction unit 126 performs motion estimation on the current block or sub-block within a reference picture to find the reference block or sub-block that best matches the current block or sub-block. Furthermore, the inter-frame prediction unit 126 obtains motion information (e.g., a motion vector) to compensate for the motion or change from the reference block or sub-block to the current block or sub-block. Based on this motion information, the inter-frame prediction unit 126 performs motion compensation (or motion prediction) to generate an inter-frame predicted image for the current block or sub-block. The inter-frame prediction unit 126 then outputs the generated inter-frame predicted image to the prediction control unit 128.
[0445] The motion information used in motion compensation can be signaled as an inter-frame prediction image in various forms. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.
[0446] [See image list]
[0447] Figure 33 is a diagram showing an example of each reference picture, Figure 34 1 is a conceptual diagram showing an example of a reference picture list. The reference picture list is a list showing one or more reference pictures stored in the frame memory 122. Figure 33 In the figure, rectangles represent pictures, arrows represent reference relationships between pictures, the horizontal axis represents time, I, P and B in the rectangles represent intra-frame prediction pictures, single-prediction pictures and double-prediction pictures respectively, and the numbers in the rectangles represent the decoding order. Figure 33 As shown in , the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. Figure 34 As shown in the figure, the reference picture list is a list of candidates for reference pictures. For example, one picture (or slice) can have more than one reference picture list. For example, if the current picture is a single-predicted picture, one reference picture list is used. If the current picture is a double-predicted picture, two reference picture lists are used. Figure 33 and Figure 34 In the example, picture B3, which is the current picture currPic, has two reference picture lists, namely the L0 list and the L1 list. When the current picture currPic is picture B3, the candidates for the reference pictures of the current picture currPic are I0, P1 and B2, and each reference picture list (i.e., the L0 list and the L1 list) represents these pictures. The inter-frame prediction unit 126 or the prediction control unit 128 specifies which picture in each reference picture list to actually refer to by referring to the picture index refidxLx. Figure 34 , reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.
[0448] Such a reference picture list can be generated in sequence units, picture units, slice units, tile units, CTU units, or CU units. Furthermore, a reference picture index indicating a reference picture used for inter prediction among the reference pictures shown in the reference picture list can be encoded at the sequence level, picture level, slice level, tile level, CTU level, or CU level. Furthermore, a common reference picture list can be used in multiple inter prediction modes.
[0449] [Basic process of inter-frame prediction]
[0450] Figure 35 This is a flowchart showing the basic process of inter-frame prediction.
[0451] The inter prediction unit 126 first generates a predicted image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4).
[0452] Here, when generating a predicted image, the inter-frame prediction unit 126 generates the predicted image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Furthermore, when determining the MV, the inter-frame prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The candidate MV is selected by, for example, the inter-frame prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Previously derived MVs may be added to the candidate MV list as candidate MVs. Furthermore, when deriving the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from the at least one candidate MV and determine the selected at least one candidate MV as the MV for the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV for the current block by searching the region of the reference picture indicated by each of the at least one selected candidate MVs. In addition, the operation of searching for the region of the reference picture may also be referred to as motion estimation.
[0453] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126 . However, for example, the processing of step Se_1 or step Se_2 may be performed by other components included in the encoding device 100 .
[0454] In addition, a candidate MV list may be prepared for each process in each inter-frame prediction mode, or a common candidate MV list may be used in multiple inter-frame prediction modes. Figure 9 The processing of steps Sa_3 and Sa_4 shown in FIG. In addition, the processing of step Se_3 is equivalent to Figure 30 The processing of step Sd_1b.
[0455] [MV export process]
[0456] Figure 36 This is a flowchart showing an example of MV derivation.
[0457] The inter-frame prediction unit 126 can derive the MV of the current block in a mode of encoding motion information (e.g., MV). In this case, for example, the motion information can be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.
[0458] Alternatively, the inter prediction unit 126 may derive the MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the stream.
[0459] Here, the MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the modes that encode motion information include the normal inter mode, normal merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Furthermore, motion information may include not only the MV but also the predicted MV selection information described later. Furthermore, modes that do not encode motion information include the FRUC mode. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses this selected mode to derive the MV of the current block.
[0460] Figure 37 This is a flowchart showing another example of MV derivation.
[0461] The inter-frame prediction unit 126 can derive the MV of the current block by encoding a differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. In other words, the encoded differential MV is included in the stream. This differential MV is the difference between the MV of the current block and its predicted MV. Furthermore, the predicted MV is a predicted motion vector.
[0462] Alternatively, the inter prediction unit 126 may derive the MV in a mode in which the difference MV is not encoded. In this case, the encoded difference MV is not included in the stream.
[0463] As described above, the MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the normal inter mode and affine mode (specifically, affine inter mode) encode the differential MV. Furthermore, the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode) do not encode the differential MV. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0464] [MV export mode]
[0465] Figure 38A and Figure 38B is a diagram showing an example of the classification of each mode derived from MV. Figure 38AAs shown in FIG, the MV derivation mode is classified into three major modes according to whether the motion information is encoded and whether the differential MV is encoded. The three modes are inter-frame mode, merge mode, and FRUC (frame rate up-conversion) mode. The inter-frame mode is a mode for performing motion search and encoding the motion information and differential MV. For example, Figure 38B As shown, inter-frame mode includes affine inter-frame mode and normal inter-frame mode. Merge mode is a mode that does not perform motion search, but selects MV from the surrounding coded blocks and uses the MV to derive the MV of the current block. This merge mode is basically a mode that encodes motion information without encoding the differential MV. For example, Figure 38B As shown, the merge modes include normal merge mode (sometimes also called normal merge mode or regular merge mode), MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined inter merge / intraprediction) mode, triangle mode, ATMVP mode, and affine merge mode. Among the modes included in the merge mode, MMVD mode exceptionally encodes the differential MV. Furthermore, the aforementioned affine merge mode and affine inter mode are included in the affine mode. The affine mode assumes an affine transformation and derives the MV of each of the multiple sub-blocks constituting the current block as the MV of the current block. The FRUC mode derives the MV of the current block by searching between already coded regions, and does not encode either motion information or the differential MV. Details of these modes will be described later.
[0466] in addition, Figure 38A and Figure 38B The classification of each mode shown is an example and is not limited to this. For example, when encoding the differential MV in the CIIP mode, the CIIP mode is classified as the inter mode.
[0467] [MV Export > Normal Interframe Mode]
[0468] Normal inter mode is an inter prediction mode that derives the MV of the current block by finding a block similar to the image of the current block from an area of a reference picture indicated by a candidate MV. In addition, in this normal inter mode, a differential MV is encoded.
[0469] Figure 39 This is a flowchart showing an example of inter prediction based on the normal inter mode.
[0470] First, the inter prediction unit 126 obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter prediction unit 126 creates a candidate MV list.
[0471] Next, the inter-frame prediction unit 126 extracts N (N is an integer greater than or equal to 2) candidate MVs from the multiple candidate MVs obtained in step Sg_1 as prediction MV candidates, according to a predetermined priority order (step Sg_2). Note that this priority order is predetermined for each of the N candidate MVs.
[0472] Next, the inter-frame prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_3). At this point, the inter-frame prediction unit 126 encodes prediction MV selection information identifying the selected prediction MV into the stream. Specifically, the inter-frame prediction unit 126 outputs the prediction MV selection information as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0473] Next, the inter-frame prediction unit 126 refers to the coded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference between the derived MV and the predicted MV as a differential MV into the stream. In other words, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130. The coded reference picture is a picture composed of multiple blocks reconstructed after encoding.
[0474] Finally, the inter prediction unit 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the coded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are performed for each block. For example, when steps Sg_1 to Sg_5 are performed for all blocks included in a slice, inter prediction using the normal inter mode for that slice is completed. Alternatively, when steps Sg_1 to Sg_5 are performed for all blocks included in a picture, inter prediction using the normal inter mode for that picture is completed. Alternatively, when steps Sg_1 to Sg_5 are performed for a portion of the blocks rather than all blocks included in a slice, inter prediction using the normal inter mode for that slice is completed. Similarly, when steps Sg_1 to Sg_5 are performed for a portion of the blocks included in a picture, inter prediction using the normal inter mode for that picture is completed.
[0475] The predicted image is the inter prediction signal described above. Information indicating the inter prediction mode (in the above example, the normal inter mode) used to generate the predicted image, contained in the coded signal, is coded as, for example, prediction parameters.
[0476] The candidate MV list can also be used in conjunction with lists used in other modes. Furthermore, processing related to the candidate MV list can be applied to processing related to lists used in other modes. Processing related to the candidate MV list includes, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.
[0477] [MV Export > Normal Merge Mode]
[0478] Normal merge mode is an inter-frame prediction mode that derives the MV of the current block by selecting a candidate MV from a candidate MV list as the MV of the current block. Normal merge mode is a narrowly defined merge mode, sometimes also referred to simply as merge mode. In this embodiment, a distinction is made between normal merge mode and merge mode, with merge mode used in a broad sense.
[0479] Figure 40 This is a flowchart showing an example of inter-frame prediction based on the normal merge mode.
[0480] First, the inter prediction unit 126 obtains a plurality of candidate MVs for the current block based on information on a plurality of coded block MVs located temporally or spatially around the current block (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.
[0481] Next, the inter-frame prediction unit 126 selects one candidate MV from the multiple candidate MVs obtained in step Sh_1 to derive the MV for the current block (step Sh_2). At this point, the inter-frame prediction unit 126 encodes MV selection information identifying the selected candidate MV into the stream. Specifically, the inter-frame prediction unit 126 outputs the MV selection information as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0482] Finally, the inter-frame prediction unit 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processing of steps Sh_1 to Sh_3 is performed on each block, for example. For example, when the processing of steps Sh_1 to Sh_3 is performed on all blocks included in a slice, the inter-frame prediction using the normal merge mode for the slice is completed. In addition, when the processing of steps Sh_1 to Sh_3 is performed on all blocks included in a picture, the inter-frame prediction using the normal merge mode for the picture is completed. In addition, the processing of steps Sh_1 to Sh_3 may be performed on a portion of the blocks instead of all blocks included in the slice, and the inter-frame prediction using the normal merge mode for the slice is completed. Similarly, when the processing of steps Sh_1 to Sh_3 is performed on a portion of the blocks included in the picture, the inter-frame prediction using the normal merge mode for the picture is completed.
[0483] Furthermore, information indicating the inter prediction mode (in the above example, the normal merge mode) used in generating the predicted image contained in the stream is encoded as, for example, prediction parameters.
[0484] Figure 41 This is a diagram for explaining an example of MV derivation processing of the current picture based on the normal merge mode.
[0485] First, the inter-frame prediction unit 126 generates a candidate MV list containing candidate MVs. These candidate MVs include: spatially neighboring candidate MVs, which are MVs belonging to multiple coded blocks spatially surrounding the current block; temporally neighboring candidate MVs, which are MVs belonging to blocks projected onto the position of the current block in the coded reference picture; combined candidate MVs, which are MVs generated by combining the MV values of spatially neighboring candidate MVs and temporally neighboring candidate MVs; and zero candidate MVs, which are MVs with a value of zero.
[0486] Next, the inter prediction unit 126 selects one candidate MV from the plurality of candidate MVs registered in the candidate MV list, and determines the one candidate MV as the MV of the current block.
[0487] Then, the entropy coding unit 110 encodes merge_idx, which is a signal indicating which candidate MV is selected, in the stream.
[0488] In addition, Figure 41 The candidate MVs registered in the candidate MV list described in the figure are just examples, and the number may be different from that in the figure, or the structure may not include some of the types of candidate MVs in the figure, or the structure may add candidate MVs other than the types of candidate MVs in the figure.
[0489] The MV of the current block derived by the normal merge mode can also be used to perform DMVR (dynamic motion vector refreshing) to determine the final MV. In addition, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same way as the normal merge mode, but encodes the differential MV. Figure 38B As shown, such MMVD can also be classified as a merge mode along with the normal merge mode. In addition, the differential MV in MMVD mode may not be the same as the differential MV used in inter-frame mode. For example, the derivation of the differential MV in MMVD mode may require less processing than the derivation of the differential MV in inter-frame mode.
[0490] Alternatively, a CIIP (Combined inter merge / intra prediction) mode may be performed in which a predicted image generated by inter prediction and a predicted image generated by intra prediction are superimposed to generate a predicted image for the current block.
[0491] In addition, the candidate MV list may also be referred to as a candidate list. In addition, merge_idx is MV selection information.
[0492] [MV Export > HMVP Mode]
[0493] Figure 42 This is a diagram for explaining an example of MV derivation processing of the current picture based on the HMVP mode.
[0494] In normal merge mode, one candidate MV is selected from a candidate MV list generated with reference to an already coded block (e.g., a CU), thereby determining the MV of the current block (e.g., a CU). Other candidate MVs can also be registered in the candidate MV list. This mode of registering other candidate MVs is called HMVP mode.
[0495] In the HMVP mode, candidate MVs are managed using a FIFO (First-In First-Out) buffer for HMVP, separate from the candidate MV list of the normal merge mode.
[0496] In the FIFO buffer, motion information such as the MV of the previously processed blocks is sequentially saved from the new FIFO buffer. In the management of this FIFO buffer, each time a block is processed, the MV of the latest block (i.e., the CU processed immediately before) is saved in the FIFO buffer, and the MV of the oldest CU in the FIFO buffer (i.e., the CU processed first) is deleted from the FIFO buffer. Figure 42In the example shown, HMVP1 is the MV of the newest block, and HMVP5 is the MV of the oldest block.
[0497] Then, for example, the inter-frame prediction unit 126 sequentially checks, starting with HMVP1, for each MV managed in the FIFO buffer to see if the MV is different from all candidate MVs registered in the candidate MV list for normal merge mode. Furthermore, if the inter-frame prediction unit 126 determines that the MV managed in the FIFO buffer is different from all candidate MVs, it may add the MV as a candidate MV to the candidate MV list for normal merge mode. In this case, the number of candidate MVs registered in the FIFO buffer may be one or more.
[0498] In this way, by using the HMVP mode, not only can the MVs of blocks that are spatially or temporally adjacent to the current block be added to the candidate, but also the MVs of blocks processed in the past can be added to the candidate. As a result, by expanding the variation of candidate MVs in the normal merge mode, the possibility of improving coding efficiency becomes higher.
[0499] Furthermore, the MV can also be motion information. That is, the candidate MV list and the information stored in the FIFO buffer can include not only the MV value but also information indicating the reference picture, the reference direction, the number of pictures, etc. Furthermore, the block can be, for example, a CU.
[0500] in addition, Figure 42 The candidate MV list and FIFO buffer are one example, and the candidate MV list and FIFO buffer can also be Figure 42 Lists or buffers of different sizes, or in the same Figure 42 A structure for registering candidate MVs in a different order. The processing described here is common to both the encoding device 100 and the decoding device 200 .
[0501] Furthermore, HMVP mode can be applied to modes other than normal merge mode. For example, motion information such as MVs of blocks previously processed in affine mode can be sequentially stored in a new FIFO buffer and used as candidate MVs. The application of HMVP mode to affine mode is also referred to as historical affine mode.
[0502] [MV Export > FRUC Mode]
[0503] Motion information may not be signaled on the encoding device 100 side, but may be derived on the decoding device 200 side. For example, motion information may be derived by performing a motion search on the decoding device 200 side. In this case, the decoding device 200 side performs a motion search without using pixel values of the current block. Examples of such motion search modes performed on the decoding device 200 side include the FRUC (frame rate up-conversion) mode and the PMMVD (pattern matched motion vector derivation) mode.
[0504] Figure 43 An example of FRUC processing is shown in . First, referring to the MVs of each coded block spatially or temporally adjacent to the current block, a list is generated that represents these MVs as candidate MVs (i.e., a candidate MV list, which can also be shared with the candidate MV list for normal merge mode) (step Si_1). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and one candidate is selected as the best candidate MV based on the evaluation value. Then, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is derived as is as the MV for the current block. Alternatively, for example, the MV for the current block can be derived by performing pattern matching in the surrounding area of the position in the reference picture corresponding to the selected best candidate MV. Specifically, the surrounding area of the best candidate MV can be searched using pattern matching and evaluation values in the reference picture. If an MV with a better evaluation value is found, the best candidate MV is updated to that MV and used as the final MV for the current block. It is not necessary to perform the update to the MV having a better evaluation value.
[0505] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processing of steps Si_1 to Si_5 is performed on each block, for example. For example, when the processing of steps Si_1 to Si_5 is performed on all blocks included in the slice, the inter-frame prediction using the FRUC mode for the slice is completed. In addition, when the processing of steps Si_1 to Si_5 is performed on all blocks included in the picture, the inter-frame prediction using the FRUC mode for the picture is completed. In addition, the processing of steps Si_1 to Si_5 may be performed on a part of the blocks instead of all blocks included in the slice, and the inter-frame prediction using the FRUC mode for the slice is completed. Similarly, when the processing of steps Si_1 to Si_5 is performed on a part of the blocks included in the picture, the inter-frame prediction using the FRUC mode for the picture is completed.
[0506] Even when processing is performed in sub-block units, the same processing as that in the block unit may be performed.
[0507] The evaluation value can be calculated using various methods. For example, a reconstructed image of a region within a reference picture corresponding to the MV can be compared with a reconstructed image of a predetermined region (e.g., as shown below, this region can be a region of another reference picture or a region of an adjacent block in the current picture). The difference in pixel values between the two reconstructed images can then be calculated and used to evaluate the MV. Furthermore, the evaluation value can be calculated using other information besides the difference.
[0508] Next, we'll explain pattern matching in detail. First, a candidate MV from the candidate MV list (also called a merge list) is selected as the starting point for a pattern matching search. Pattern matching can be done using either primary or secondary pattern matching. Primary and secondary pattern matching are also known as bilateral matching and template matching, respectively.
[0509] [MV Export > FRUC > Bidirectional Matching]
[0510] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, the area in the other reference pictures along the motion trajectory of the current block is used as the predetermined area for calculating the candidate MV evaluation value.
[0511] Figure 44: is a diagram for explaining an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along the motion trajectory. Figure 44 As shown, in the first pattern matching, two MVs (MV0, MV1) are derived by searching for the best-matching pair among pairs of blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by a symmetric MV scaled by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.
[0512] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, mirror-symmetric MVs in both directions are derived in the first pattern matching.
[0513] [MV Export > FRUC > Template Matching]
[0514] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (blocks adjacent to the current block in the current picture (e.g., blocks above and / or to the left)) and a block in the reference picture. Therefore, in the second pattern matching, blocks adjacent to the current block in the current picture are used as the specified area for calculating the candidate MV evaluation value.
[0515] Figure 45 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 45 As shown, in the second pattern matching, the MV of the current block is derived by searching the reference picture (Ref0) for a block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded area adjacent to the left and above, or one of the two, is derived, and the reconstructed image at the same position in the coded reference picture (Ref0) specified by the candidate MV is used to calculate the evaluation value. The candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.
[0516] Such information indicating whether FRUC mode is adopted (e.g., called a FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is adopted (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (first pattern matching or second pattern matching) is signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and may be performed at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0517] [MV Export > Affine Mode]
[0518] Affine mode is a mode that generates MVs using affine transform. For example, MVs can be derived in sub-block units based on the MVs of multiple adjacent blocks. This mode is sometimes called affine motion compensation prediction mode.
[0519] Figure 46A This is a diagram for explaining an example of deriving the MV of a sub-block unit based on the MVs of a plurality of adjacent blocks. Figure 46A In the example, the current block includes 16 sub-blocks consisting of 4×4 pixels. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the MV of the adjacent sub-block. Then, according to the following formula (1A), the two motion vectors v0 and v1 are projected to derive the motion vector (v x , v y ).
[0520]
Formula 1
[0521]
[0522] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predetermined weight coefficient.
[0523] Information indicating this affine mode (e.g., called an affine flag) can be signaled at the CU level. Furthermore, the signaling of information indicating this affine mode need not be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).
[0524] Furthermore, such affine modes may include several modes with different methods for deriving the MVs of the upper left and upper right control points. For example, there are two affine modes: affine inter (also called affine normal inter) mode and affine merge mode.
[0525] Figure 46B This is a diagram for explaining an example of deriving MVs in sub-block units in an affine mode using three control points. Figure 46B In the example, the current block consists of 16 4×4 pixel sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the MV of the adjacent block, and the motion vector v2 of the lower left corner control point of the current block is derived based on the MV of the adjacent block. Then, according to the following formula (1B), the three motion vectors v0, v1 and v2 are projected to derive the motion vectors (v x , v y ).
[0526]
Formula 2
[0527]
[0528] Here, x and y represent the horizontal and vertical positions of the sub-block center, respectively, and w and h represent predetermined weight coefficients. Alternatively, w represents the width of the current block, and h represents the height of the current block.
[0529] Affine modes using different numbers of control points (e.g., 2 and 3) can also be switched and signaled at the CU level. In addition, information indicating the number of control points of the affine mode used at the CU level can also be signaled at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).
[0530] Furthermore, in such an affine mode with three control points, several modes may be included, each with different methods for deriving the MVs of the upper left, upper right, and lower left control points. For example, in an affine mode with three control points, similar to the affine mode with two control points, there are two modes: affine inter mode and affine merge mode.
[0531] Furthermore, in the affine mode, the size of each sub-block included in the current block is not limited to 4×4 pixels, but may be other sizes. For example, the size of each sub-block may be 8×8 pixels.
[0532] [MV Export > Affine Mode > Control Points]
[0533] Figure 47A 、 Figure 47B and Figure 47C This is a conceptual diagram for explaining an example of MV derivation of control points in the affine mode.
[0534] In affine mode, such as Figure 47AAs shown, for example, based on multiple MVs corresponding to blocks coded in affine mode among the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, a predicted MV for each of the control points of the current block is calculated. Specifically, the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in this order to determine the first valid block coded in affine mode. The MVs of the control points of the current block are calculated based on the multiple MVs corresponding to the determined blocks.
[0535] For example, Figure 47B As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 are derived that are projected onto the upper left corner and upper right corner of the encoded block including block A. Then, based on the derived motion vectors v3 and v4, motion vector v0 of the upper left control point and motion vector v1 of the upper right control point of the current block are calculated.
[0536] For example, Figure 47C As shown in FIG5 , when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, motion vector v0 of the upper left control point, motion vector v1 of the upper right control point, and motion vector v2 of the lower left control point of the current block are calculated.
[0537] in addition, Figures 47A to 47C The MV derivation method shown can be used in the following Figure 50 The derivation of the MV of each control point of the current block in step Sk_1 shown in FIG. 1 can also be used in the following Figure 51 The derivation of the predicted MV of each control point of the current block in step Sj_1 is shown.
[0538] Figure 48A and Figure 48B This is a conceptual diagram for explaining another example of derivation of control point MV in the affine mode.
[0539] Figure 48A A diagram for explaining an affine pattern having two control points.
[0540] In this affine model, if Figure 48AAs shown in FIG, an MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left corner control point of the current block. Similarly, an MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right corner control point of the current block.
[0541] Figure 48B A diagram for explaining an affine pattern having three control points.
[0542] In this affine model, if Figure 48B As shown in FIG, an MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left corner control point of the current block. Similarly, an MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right corner control point of the current block. Furthermore, an MV selected from the MVs of the coded blocks F and G adjacent to the current block is used as the motion vector v2 of the lower left corner control point of the current block.
[0543] in addition, Figure 48A and Figure 48B The MV derivation method shown can be used in the following Figure 50 The derivation of the MV of each control point of the current block in step Sk_1 shown in FIG. 1 can also be used for the following Figure 51 The predicted MV of each control point of the current block is derived in step Sj_1.
[0544] Here, for example, when affine modes with different numbers of control points (eg, 2 and 3) are switched at the CU level and signaled, the number of control points may differ between the coded block and the current block.
[0545] Figure 49A and Figure 49B This is a conceptual diagram for explaining an example of a method for deriving MVs of control points when the number of control points in an already coded block and a current block is different.
[0546] For example, Figure 49A As shown, the current block has three control points: the top-left corner, the top-right corner, and the bottom-left corner. Block A, adjacent to the left of the current block, is coded in an affine mode with two control points. In this case, motion vectors v3 and v4 are derived, projected onto the top-left and top-right corners of the coded block containing block A. Then, based on the derived motion vectors v3 and v4, motion vector v0 for the top-left control point and motion vector v1 for the top-right control point of the current block are calculated. Furthermore, based on the derived motion vectors v0 and v1, motion vector v2 for the bottom-left control point is calculated.
[0547] For example, Figure 49B As shown, the current block has two control points, the upper left and upper right corners, and block A, adjacent to the left of the current block, is coded in affine mode with three control points. In this case, motion vectors v3, v4, and v5 are derived, projected onto the positions of the upper left, upper right, and lower left corners of the coded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, motion vectors v0 for the upper left control point and v1 for the upper right control point of the current block are calculated.
[0548] in addition, Figure 49A and Figure 49B The MV derivation method shown can be used in the following Figure 50 The derivation of the MV of each control point of the current block in step Sk_1 shown in FIG. 1 can also be used for the following Figure 51 The predicted MV of each control point of the current block is derived in step Sj_1.
[0549] [MV Export > Affine Mode > Affine Merge Mode]
[0550] Figure 50 This is a flowchart showing an example of the affine merge mode.
[0551] In the affine merge mode, first, the inter prediction unit 126 derives the MV of each control point of the current block (step Sk_1). Figure 46A As shown, it is the upper left and upper right corner points of the current block, or as Figure 46B As shown in FIG, these are the points at the upper left corner, upper right corner, and lower left corner of the current block. In this case, the inter prediction unit 126 may also encode MV selection information for identifying the derived two or three MVs into the stream.
[0552] For example, when using Figures 47A to 47C In the case of the MV derivation method shown, as Figure 47A As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.
[0553] The inter-frame prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control point. For example, when block A is determined and has two control points, Figure 47BAs shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block including the block A. For example, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block onto the current block.
[0554] Alternatively, in the case where block A is determined and block A has 3 control points, as Figure 47C As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block.
[0555] In addition, as mentioned above, Figure 49A As shown, determine block A, and when block A has 2 control points, calculate the MV of 3 control points, or as mentioned above Figure 49B As shown, block A is determined, and when block A has three control points, the MVs of two control points are calculated.
[0556] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two motion vectors v0 and v1 and the above-mentioned equation (1A), or uses three motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the MV of the sub-block as an affine MV (step Sk_2). Then, the inter-frame prediction unit 126 performs motion compensation on the sub-block using these affine MVs and the encoded reference picture (step Sk_3). When the processes of steps Sk_2 and Sk_3 are performed on all sub-blocks included in the current block, the process of generating a predicted image using the affine merge mode for the current block is completed. In other words, motion compensation is performed on the current block, and a predicted image for the current block is generated.
[0557] In addition, in step Sk_1, the candidate MV list mentioned above may also be generated. The candidate MV list may also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. Multiple MV derivation methods may be Figures 47A to 47C The MV derivation method shown, Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B The MV derivation method shown and any combination of other MV derivation methods.
[0558] Furthermore, the candidate MV list may include candidate MVs of modes other than the affine mode that perform prediction in sub-block units.
[0559] In addition, as a candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and a candidate MV list including candidate MVs of an affine merge mode with 3 control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and a candidate MV list including candidate MVs of an affine merge mode with 3 control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes with 2 control points and the affine merge mode with 3 control points may be generated. The candidate MVs may be, for example, the MVs of the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or the MVs of valid blocks among these blocks.
[0560] Alternatively, an index indicating which candidate MV in the candidate MV list is to be selected may be transmitted as the MV selection information.
[0561] [MV Export > Affine Mode > Affine Inter Mode]
[0562] Figure 51 This is a flowchart showing an example of the affine inter mode.
[0563] In the affine inter mode, first, the inter prediction unit 126 derives the predicted MV (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). Figure 46A or Figure 46B As shown, the control point is the point at the upper left corner, upper right corner or lower left corner of the current block.
[0564] For example, when using Figure 48A and Figure 48B In the case of the MV derivation method shown in FIG. 1 , the inter-frame prediction unit 126 selects Figure 48A or Figure 48B The inter-frame prediction unit 126 derives the predicted MV (v0, v1) or (v0, v1, v2) of the control point of the current block from the MV of a block in the coded blocks near each control point of the current block. At this time, the inter-frame prediction unit 126 encodes predicted MV selection information for identifying the selected two or three predicted MVs into the stream.
[0565] For example, the inter-frame prediction unit 126 may determine which block's MV from among the already coded blocks adjacent to the current block to select as the predicted MV for the control point by using a cost evaluation or the like, and may include a flag indicating which predicted MV was selected in the bitstream. Specifically, the inter-frame prediction unit 126 outputs the predicted MV selection information, such as the flag, as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0566] Next, the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4) while simultaneously updating the predicted MV selected or derived in step Sj_1 (step Sj_2). Specifically, the inter-frame prediction unit 126 uses the MV of each sub-block corresponding to the predicted MV to be updated as an affine MV and calculates it using equation (1A) or equation (1B) (step Sj_3). The inter-frame prediction unit 126 then performs motion compensation on each sub-block using these affine MVs and the coded reference picture (step Sj_4). Each time the predicted MV is updated in step Sj_2, the processes of steps Sj_3 and Sj_4 are repeated for all blocks within the current block. As a result, during the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted MV that yields the minimum cost as the MV of the control point (step Sj_5). At this point, the inter-frame prediction unit 126 also encodes the difference between the determined MV and the predicted MV as a differential MV into the stream. That is, the inter prediction unit 126 outputs the difference MV as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130 .
[0567] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0568] In addition, in step Sj_1, the candidate MV list mentioned above may also be generated. The candidate MV list may also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. Multiple MV derivation methods may be Figures 47A to 47C The MV derivation method shown, Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B The MV derivation method shown and any combination of other MV derivation methods.
[0569] Furthermore, the candidate MV list may include candidate MVs of modes other than the affine mode that perform prediction in sub-block units.
[0570] Furthermore, as a candidate MV list, a candidate MV list including candidate MVs of an affine inter mode with two control points and candidate MVs of an affine inter mode with three control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine inter mode with two control points and a candidate MV list including candidate MVs of an affine inter mode with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of either an affine inter mode with two control points or an affine inter mode with three control points may be generated. The candidate MVs may be, for example, the MVs of the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or the MVs of valid blocks among these blocks.
[0571] Alternatively, an index indicating which candidate MV in the candidate MV list is selected may be sent as the predicted MV selection information.
[0572] [MV Export > Triangle Mode]
[0573] In the above example, the inter-frame prediction unit 126 generates a single rectangular predicted image for a rectangular current block. However, the inter-frame prediction unit 126 may generate multiple predicted images of shapes other than a rectangle for the current rectangular block and combine these multiple predicted images to generate a final rectangular predicted image. A shape other than a rectangle may be, for example, a triangle.
[0574] Figure 52A This is a diagram for explaining the generation of predicted images of two triangles.
[0575] The inter-frame prediction unit 126 performs motion compensation on the first triangular partition within the current block using the first MV of the first partition, thereby generating a predicted image for the triangle. Similarly, the inter-frame prediction unit 126 performs motion compensation on the second triangular partition within the current block using the second MV of the second partition, thereby generating a predicted image for the triangle. Furthermore, the inter-frame prediction unit 126 combines these predicted images to generate a predicted image of the same rectangle as the current block.
[0576] Alternatively, as the predicted image for the first partition, a rectangular first predicted image corresponding to the current block may be generated using the first MV. Alternatively, as the predicted image for the second partition, a rectangular second predicted image corresponding to the current block may be generated using the second MV. Alternatively, the predicted image for the current block may be generated by performing a weighted addition of the first and second predicted images. Furthermore, the weighted addition may be performed only on a portion of the region sandwiching the boundary between the first and second partitions.
[0577] Figure 52BThis is a conceptual diagram illustrating an example of a first portion of a first partition overlapping a second partition, and a first sample set and a second sample set that can be weighted as part of a correction process. The first portion can be, for example, one-quarter the width or height of the first partition. In another example, the first portion can have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero, such as the integer 2. Figure 52B The rectangular partition represents a rectangular portion having a width that is one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion. Figure 52B The central example shows a rectangular partition having a height that is one-quarter the height of partition 1. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion. Figure 52B The example on the right shows a triangular partition of a polygonal part with a height corresponding to two samples. Here, the first sample set contains samples outside the first part and samples inside the first part, and the second sample set contains samples inside the first part.
[0578] The first portion may be a portion of the first partition that overlaps with an adjacent partition. Figure 52C This is a conceptual diagram showing the first portion of the first partition, which is a portion of the first partition that overlaps with a portion of an adjacent partition. For simplicity of explanation, a rectangular partition is shown with a portion overlapping a spatially adjacent rectangular partition. Partitions of other shapes, such as triangular partitions, can be used, and the overlapping portion may also overlap with spatially or temporally adjacent partitions.
[0579] Furthermore, an example in which predicted images are generated for each of two partitions using inter-frame prediction has been shown, but a predicted image may be generated for at least one partition using intra-frame prediction.
[0580] Figure 53 This is a flowchart showing an example of the triangle mode.
[0581] In triangular mode, the inter-frame prediction unit 126 first divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame prediction unit 126 may encode partition information, which is information related to the division into each partition, as prediction parameters into the stream. Specifically, the inter-frame prediction unit 126 may output the partition information as prediction parameters to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0582] Next, the inter-frame prediction unit 126 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks temporally or spatially surrounding the current block (step Sx_2). In other words, the inter-frame prediction unit 126 creates a candidate MV list.
[0583] Next, the inter-frame prediction unit 126 selects the candidate MV for the first partition and the candidate MV for the second partition as the first MV and the second MV, respectively, from the multiple candidate MVs obtained in step Sx_1 (step Sx_3). At this time, the inter-frame prediction unit 126 may also encode MV selection information identifying the selected candidate MVs as prediction parameters into the stream. In other words, the inter-frame prediction unit 126 may output the MV selection information as prediction parameters to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0584] Next, the inter-frame prediction unit 126 performs motion compensation using the selected first MV and the coded reference picture to generate a first predicted image (step Sx_4). Similarly, the inter-frame prediction unit 126 performs motion compensation using the selected second MV and the coded reference picture to generate a second predicted image (step Sx_5).
[0585] Finally, the inter prediction unit 126 performs weighted addition on the first predicted image and the second predicted image to generate a predicted image of the current block (step Sx_6).
[0586] In addition, Figure 52A In the example shown, the first partition and the second partition are each a triangle, but they may also be a trapezoid or have different shapes. Figure 52A In the example shown, the current block is composed of 2 partitions, but it can also be composed of 3 or more partitions.
[0587] Furthermore, the first and second partitions may overlap. That is, the first and second partitions may contain the same pixel region. In this case, the predicted image in the first partition and the predicted image in the second partition may be used to generate the predicted image for the current block.
[0588] In addition, although this example shows an example in which predicted images are generated by inter-frame prediction in both of the two partitions, a predicted image may be generated by intra-frame prediction for at least one partition.
[0589] In addition, the candidate MV list used to select the first MV and the candidate MV list used to select the second MV may be different or the same candidate MV list.
[0590] Furthermore, the partition information may include an index indicating the direction in which at least the current block is divided into multiple partitions. The MV selection information may also include an index indicating the first selected MV and an index indicating the second selected MV. A single index may also indicate multiple pieces of information. For example, a single index may be encoded that collectively indicates a portion or all of the partition information and a portion or all of the MV selection information.
[0591] [MV Export > ATMVP Mode]
[0592] Figure 54 This is a diagram showing an example of the ATMVP mode for deriving MVs in sub-block units.
[0593] The ATMVP mode is classified as a merge mode. For example, in the ATMVP mode, candidate MVs in sub-block units are registered in a candidate MV list used for the normal merge mode.
[0594] Specifically, in the ATMVP mode, first, if Figure 54 As shown, a temporal MV reference block corresponding to the current block is identified in the coded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block. Next, for each subblock within the current block, an MV to be used when encoding the region corresponding to the subblock within the temporal MV reference block is identified. The MVs thus identified are included in a candidate MV list as candidate MVs for the subblocks of the current block. When such a candidate MV is selected for each subblock from the candidate MV list, motion compensation is performed on the subblock using the candidate MV as the subblock's MV. This generates a predicted image for each subblock.
[0595] In addition, Figure 54 In the example shown, the block adjacent to the lower left of the current block is used as the neighboring MV reference block, but other blocks can also be used. Furthermore, the sub-block size can be 4x4 pixels, 8x8 pixels, or other sizes. The sub-block size can also be switched in units such as slices, tiles, or pictures.
[0596] [Sports Search>DMVR]
[0597] Figure 55 This is a diagram showing the relationship between the merge mode and DMVR.
[0598] The inter-frame prediction unit 126 derives the MV of the current block in merge mode (step S1_1). Next, the inter-frame prediction unit 126 determines whether to perform an MV search, i.e., a motion search (step S1_2). If it is determined not to perform a motion search (No in step S1_2), the inter-frame prediction unit 126 determines the MV derived in step S1_1 as the final MV for the current block (step S1_4). In other words, in this case, the MV of the current block is determined in merge mode.
[0599] On the other hand, if it is determined in step S1_1 that a motion search is to be performed (Yes in step S1_2), the inter prediction unit 126 searches the surrounding area of the reference picture indicated by the MV derived in step S1_1 to derive a final MV for the current block (step S1_3). In other words, in this case, the MV of the current block is determined by DMVR.
[0600] Figure 56 This is a conceptual diagram for explaining an example of DMVR for determining MV.
[0601] First, for example, in merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, based on candidate MV (L0), reference pixels are determined from the first reference picture (L0), a coded picture in the L0 list. Similarly, based on candidate MV (L1), reference pixels are determined from the second reference picture (L1), a coded picture in the L1 list. A template is generated by averaging these reference pixels.
[0602] Next, using this template, the surrounding areas of the candidate MVs in the first reference picture (L0) and the second reference picture (L1) are searched, and the MV with the lowest cost is determined as the final MV for the current block. Alternatively, the cost can be calculated using, for example, the difference between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values.
[0603] Even if it is not the process described here, any process can be used as long as it can search the vicinity of the candidate MV and derive the final MV.
[0604] Figure 57 This is a conceptual diagram for explaining another example of DMVR for determining MV. Figure 57 The example shown is the same as Figure 56 The example of DMVR shown is different in that the cost is calculated without generating a template.
[0605] First, the inter prediction unit 126 searches for the reference blocks included in the reference pictures of the L0 list and the L1 list based on the candidate MV, that is, the initial MV, obtained from the candidate MV list. Figure 57As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. During motion search, the inter-frame prediction unit 126 first sets a search position for the reference picture in the L0 list. The difference vector representing this set search position, specifically the difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to this search position, is MVd_L0. The inter-frame prediction unit 126 then determines a search position in the reference picture in the L1 list. This search position is represented by the difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to this search position. Specifically, the inter-frame prediction unit 126 determines MVd_L1 by mirroring the difference vector MVd_L0. That is, the inter-frame prediction unit 126 sets the search position at a position symmetrical to the position indicated by the initial MV in each of the L0 and L1 lists. The inter prediction section 126 calculates the sum of absolute differences (SAD) of pixel values within the block of the search position as a cost for each search position, and finds a search position that minimizes the cost.
[0606] Figure 58A is a diagram showing an example of motion search in DMVR. Figure 58B : is a flowchart showing an example of the motion search.
[0607] First, in Step 1, the inter-frame prediction unit 126 calculates the cost for the search position (also called the starting point) represented by the initial MV and the eight surrounding search positions. The inter-frame prediction unit 126 then determines whether the cost of a search position other than the starting point is the lowest. If the cost of a search position other than the starting point is the lowest, the inter-frame prediction unit 126 moves to the search position with the lowest cost and proceeds to Step 2. On the other hand, if the cost of the starting point is the lowest, the inter-frame prediction unit 126 skips Step 2 and proceeds to Step 3.
[0608] In Step 2, the inter-frame prediction unit 126 uses the search position moved based on the processing results of Step 1 as a new starting point and performs the same search as in Step 1. Furthermore, the inter-frame prediction unit 126 determines whether the cost of a search position other than the starting point is the lowest. If the cost of a search position other than the starting point is the lowest, the inter-frame prediction unit 126 proceeds to Step 4. On the other hand, if the cost of the starting point is the lowest, the inter-frame prediction unit 126 proceeds to Step 3.
[0609] In Step 4 , the inter prediction unit 126 treats the search position of the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a difference vector.
[0610] In Step 3, the inter-frame prediction unit 126 determines the fractional-precision pixel position with the lowest cost based on the costs at four points located above, below, and to the left and right of the starting point in Step 1 or Step 2, and sets this pixel position as the final search position. This fractional-precision pixel position is determined by weighted summing the vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) located above, below, and to the left, using the costs of the search positions of each of the four points as weights. The inter-frame prediction unit 126 then determines the difference between the position represented by the initial MV and this final search position as a difference vector.
[0611] [Motion Compensation > BIO / OBMC / LIC]
[0612] In motion compensation, there are modes for generating a predicted image and then correcting the predicted image. Examples of these modes include BIO, OBMC, and LIC, which will be described later.
[0613] Figure 59 This is a flowchart showing an example of generating a predicted image.
[0614] The inter prediction unit 126 generates a predicted image (step Sm_1 ), and corrects the predicted image using any of the above-mentioned modes (step Sm_2 ).
[0615] Figure 60 This is a flowchart showing another example of generating a predicted image.
[0616] The inter-frame prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image using this MV (step Sn_2) and determines whether to perform correction processing (step Sn_3). If it is determined that correction processing is to be performed (yes in step Sn_3), the inter-frame prediction unit 126 corrects the predicted image to generate a final predicted image (step Sn_4). Furthermore, in the LIC described later, luminance and chrominance may also be corrected in step Sn_4. On the other hand, if it is determined that correction processing is not to be performed (no in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correction (step Sn_5).
[0617] [Motion Compensation > OBMC]
[0618] Inter-frame prediction images can be generated using not only the motion information of the current block obtained through motion search, but also the motion information of neighboring blocks. Specifically, an inter-frame prediction image can be generated in sub-block units within the current block by weighted addition of a prediction image based on motion information obtained through motion search (within the reference picture) and a prediction image based on motion information of neighboring blocks (within the current picture). This type of inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or OBMC mode.
[0619] In OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) can also be signaled at the sequence level. Furthermore, information indicating whether OBMC mode is applied (e.g., OBMC flag) can also be signaled at the CU level. Furthermore, the signaling level of this information need not be limited to the sequence and CU levels and can also be at other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0620] The OBMC mode will be described in more detail. Figure 61 and Figure 62 It is a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC.
[0621] First, if Figure 62 As shown in FIG, the MV assigned to the current block is used to obtain a predicted image (Pred) based on conventional motion compensation. Figure 62 In FIG, the arrow “MV” points to the reference picture and indicates which block the current block of the current picture refers to to obtain the predicted image.
[0622] Next, the MV (MV_L) derived for the already coded left-adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is represented by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by overlaying the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0623] Similarly, the MV (MV_U) derived for the previously coded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The MV (MV_U) is represented by an arrow "MV_U" pointing from the current block to the reference picture. The predicted image Pred_U is then overlaid with the predicted image (e.g., Pred and Pred_L) that has undergone the first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image for the current block, with the boundaries with adjacent blocks blended (smoothed).
[0624] In addition, the above example is a two-path correction method using left-adjacent and upper-adjacent blocks, but the correction method can also be a three-path correction method or more than three-path correction method using right-adjacent and / or lower-adjacent blocks.
[0625] Furthermore, the area to be overlapped may not be the entire pixel area of the block, but may be only a partial area near the block boundary.
[0626] In this description, the predicted image correction process using OBMC is described as obtaining a single predicted image Pred by superimposing a single reference picture with the additional predicted images Pred_L and Pred_U. However, when correcting the predicted image based on multiple reference pictures, the same process can be applied to each of the multiple reference pictures. In this case, after performing OBMC image correction based on multiple reference pictures and obtaining a corrected predicted image from each reference picture, the final predicted image is obtained by further superimposing the obtained multiple corrected predicted images.
[0627] In addition, in OBMC, the unit of the current block can be a PU unit or a sub-block unit obtained by further dividing the PU.
[0628] One method for determining whether to apply OBMC is to use, for example, obmc_flag, a signal indicating whether OBMC is to be applied. As a specific example, encoding device 100 may also determine whether the current block belongs to a region with complex motion. If the current block belongs to a region with complex motion, encoding device 100 sets obmc_flag to a value of 1 and applies OBMC for encoding. If the current block does not belong to a region with complex motion, encoding device 100 sets obmc_flag to a value of 0 and does not apply OBMC for encoding. Meanwhile, decoding device 200 decodes obmc_flag described in the stream and switches whether to apply OBMC for decoding based on the value.
[0629] [Motion Compensation > BIO]
[0630] Next, we'll explain how to derive MVs. First, we'll explain a method for deriving MVs based on a model that assumes constant-velocity linear motion. This mode is sometimes called a BIO (bi-directional optical flow) mode. Alternatively, BDOF (bidirectional optical flow) can be used instead of BIO.
[0631] Figure 63 This is a diagram for explaining a model assuming uniform linear motion. Figure 63 In (v x , v y ) represents the velocity vector, τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.
[0632] At this time, the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (2) holds.
[0633]
Formula 3
[0634]
[0635] Here, I(k) represents the luminance value of reference image k (k = 0, 1) after motion compensation. This optical flow equation states that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Block-level motion vectors obtained from the candidate MV list, etc., can also be corrected on a pixel-by-pixel basis based on a combination of this optical flow equation and Hermite interpolation.
[0636] Alternatively, the decoding device 200 may derive an MV using a method different from the derivation of a motion vector based on a model assuming constant velocity linear motion. For example, a motion vector may be derived in sub-block units based on the MVs of a plurality of adjacent blocks.
[0637] Figure 64 This is a flowchart showing an example of inter-frame prediction according to BIO. Figure 653 is a diagram showing an example of the functional configuration of the inter-frame prediction unit 126 that performs the inter-frame prediction according to the BIO.
[0638] like Figure 65 As shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolated image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a predicted image correction unit 126f. The memory 126a may also be the frame memory 122.
[0639] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture containing the current block (Cur Pic) to derive two motion vectors (M0, M1). The inter-frame prediction unit 126 then uses these two motion vectors (M0, M1) to derive a predicted image for the current block (step Sy_1). Motion vector M0 is the motion vector (MVx0, MVy0) corresponding to reference picture Ref0, and motion vector M1 is the motion vector (MVx1, MVy1) corresponding to reference picture Ref1.
[0640] Next, the interpolation image deriving unit 126b refers to the memory 126a and derives the interpolation image I of the current block using the motion vector M0 and the reference picture L0. 0 In addition, the interpolation image deriving unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block. 1 (Step Sy_2). Here, the interpolated image I 0 The interpolated image I is the image contained in the reference image Ref0 derived for the current block. 1 It is the image included in the reference picture Ref1 derived for the current block. 0 and interpolated image I 1 Alternatively, in order to properly derive the gradient image described later, the interpolated image I 0 and the interpolated image I 1 They can be images larger than the current block. In addition, the interpolated image I 0 and I 1 It may include a predicted image derived by applying a motion vector (M0, M1), a reference picture (L0, L1), and a motion compensation filter.
[0641] In addition, the gradient image deriving unit 126c uses the interpolation image I 0 and the interpolated image I 1 , export the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1)(Step Sy_3). In addition, the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), the vertical gradient image is (Iy 0 , Iy 1 The gradient image deriving unit 126c may derive the gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image only needs to represent the spatial variation of pixel values along the horizontal or vertical direction.
[0642] Next, the optical flow deriving unit 126d uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) derives the optical flow (vx, vy) as the velocity vector (step Sy_4). The optical flow is a coefficient that corrects the spatial movement of pixels and may also be called a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, a sub-block may be a 4x4 pixel sub-CU. Furthermore, the optical flow may be derived in units other than sub-blocks, such as pixels.
[0643] Next, the inter-frame prediction unit 126 uses the optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value derivation unit 126e uses the optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_5). Furthermore, the predicted image correction unit 126f may also use the correction values to correct the predicted image of the current block (step Sy_6). The correction values may be derived for each pixel, for multiple pixels, or for sub-blocks.
[0644] In addition, the BIO process is not limited to Figure 64 The disclosed processing can be implemented only Figure 64 A part of the disclosed processing may be added to or replaced with a different processing, or may be executed in a different processing order.
[0645] [Motion Compensation > LIC]
[0646] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) will be described.
[0647] Figure 66A : is a diagram for explaining an example of a method for generating a predicted image using brightness correction processing based on LIC. Figure 66BThis is a flowchart showing an example of a method for generating a predicted image using the LIC.
[0648] First, the inter prediction unit 126 derives an MV from an already encoded reference picture and obtains a reference picture corresponding to the current block (step Sz_1 ).
[0649] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values of the current block vary between the reference picture and the current picture (step Sz_2). This extraction is performed based on the luminance pixel values of the coded left-adjacent reference region (peripheral reference region) and the coded upper-adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. The inter-frame prediction unit 126 then calculates a luminance correction parameter using this information indicating how the luminance values vary (step Sz_3).
[0650] The inter-frame prediction unit 126 applies the brightness correction parameter to the reference image within the reference picture specified by the MV, generating a predicted image for the current block (step Sz_4). Specifically, the predicted image, which is the reference image within the reference picture specified by the MV, is corrected based on the brightness correction parameter. This correction can be performed to correct brightness or color difference. Specifically, the color difference correction parameter can be calculated using information indicating how the color difference changes, and the color difference correction process can be performed.
[0651] in addition, Figure 66A The shape of the peripheral reference area in is just an example, and shapes other than this may be used.
[0652] In addition, the process of generating a predicted image based on one reference image is described here, but the same is true when generating a predicted image based on multiple reference images. The predicted image can also be generated after brightness correction processing is performed on the reference images obtained from each reference image in the same way as described above.
[0653] One method for determining whether to apply LIC is to use, for example, a method using a lic_flag, which serves as a signal indicating whether LIC is applied. As a specific example, the encoding device 100 determines whether the current block belongs to an area where luminance changes. If the current block belongs to an area where luminance changes, the lic_flag is set to a value of 1, and encoding is performed using LIC. If the current block does not belong to an area where luminance changes, the lic_flag is set to a value of 0, and encoding is performed without applying LIC. Alternatively, the decoding device 200 may decode the lic_flag described in the stream and switch whether to apply LIC based on its value for decoding.
[0654] Another method for determining whether to apply LIC is to determine whether LIC is applied to neighboring blocks. As a specific example, when the current block is processed in merge mode, the inter-frame prediction unit 126 determines whether the neighboring, already-encoded blocks selected during MV derivation in merge mode were coded using LIC. Based on the result, the inter-frame prediction unit 126 switches whether to apply LIC and performs coding. In this example, the same process also applies to the decoding device 200.
[0655] use Figure 66A and Figure 66B LIC (luminance correction processing) has been described above, and its details will be described below.
[0656] First, the inter prediction unit 126 derives an MV for acquiring a reference image corresponding to the current block from a reference picture that is an already coded picture.
[0657] Next, the inter-frame prediction unit 126 uses the luminance pixel values of the coded neighboring reference areas to the left and above the current block, as well as the luminance pixel values at the same position in the reference picture specified by the MV, to extract information indicating how the luminance values in the reference picture and the current picture change, and calculates a luminance correction parameter. For example, the luminance pixel value of a pixel in the neighboring reference area in the current picture is set to p0, and the luminance pixel value of the pixel in the neighboring reference area at the same position in the reference picture is set to p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing the solution A×p1+B=p0 for multiple pixels in the neighboring reference area as luminance correction parameters.
[0658] Next, the inter-frame prediction unit 126 uses the brightness correction parameters to perform brightness correction on the reference image within the reference picture specified by the MV to generate a predicted image for the current block. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after brightness correction is set to p3. The inter-frame prediction unit 126 calculates A×p2+B=p3 for each pixel in the reference image to generate the predicted image after brightness correction.
[0659] Alternatively, you can use Figure 66A For example, a region including a predetermined number of pixels thinned out from the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference region. In addition, the peripheral reference region is not limited to the region adjacent to the current block, and may also be a region not adjacent to the current block. Figure 66AIn the example shown, the peripheral reference region in the reference picture is a region specified by the MV of the current picture from the peripheral reference region in the current picture, but it may also be a region specified by another MV. For example, the other MV may also be the MV of the peripheral reference region in the current picture.
[0660] Note that, although the operation in the encoding device 100 is described here, the operation in the decoding device 200 is also the same.
[0661] Furthermore, LIC can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for all of them.
[0662] Furthermore, LIC can also be applied in sub-block units. For example, modification parameters can be derived using a surrounding reference area of the current sub-block and a surrounding reference area of a reference sub-block in a reference picture specified by the MV of the current sub-block.
[0663] [Prediction Control Department]
[0664] The prediction control unit 128 selects one of the intra-frame prediction image (pixels or signals output from the intra-frame prediction unit 124) and the inter-frame prediction image (pixels or signals output from the inter-frame prediction unit 126), and outputs the selected prediction image to the subtraction unit 104 and the addition unit 116.
[0665] [Prediction parameter generation unit]
[0666] The prediction parameter generation unit 130 can output information related to intra-frame prediction, inter-frame prediction, and the selection of a prediction image in the prediction control unit 128 as prediction parameters to the entropy coding unit 110. The entropy coding unit 110 can generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream and perform the same prediction processing as that performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction parameters can include a prediction selection signal (e.g., MV, prediction type, or prediction mode used by the intra-frame prediction unit 124 or the inter-frame prediction unit 126), or an index, flag, or value based on the prediction processing performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128, or indicating the prediction processing.
[0667] [Decoding device]
[0668] Next, a decoding device 200 capable of decoding the stream output from the above-described encoding device 100 will be described. Figure 672 is a block diagram showing an example of the functional configuration of the decoding device 200 according to the embodiment. The decoding device 200 is a device that decodes a coded image, ie, a stream, in units of blocks.
[0669] like Figure 67 As shown, the decoding apparatus 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a partition determination unit 224. The intra-frame prediction unit 216 and the inter-frame prediction unit 218 each constitute a part of a prediction processing unit.
[0670] [Decoding device installation example]
[0671] Figure 68 2 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 67 The multiple components of the decoding device 200 shown are Figure 68 The processor b1 and the memory b2 shown are implemented.
[0672] Processor b1 is a circuit that processes information and can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 can also be a processor such as a CPU. In addition, processor b1 can also be a collection of multiple electronic circuits. In addition, for example, processor b1 can also play a role in Figure 67 The functions of the multiple components of the decoding device 200 excluding the components for storing information are as follows.
[0673] Memory b2 is a dedicated or general-purpose memory that stores information used by processor b1 to decode the stream. Memory b2 can be an electronic circuit or connected to processor b1. Alternatively, memory b2 can be included in processor b1. Alternatively, memory b2 can be a collection of multiple electronic circuits. Furthermore, memory b2 can be a magnetic disk or optical disk, or it can be a storage device or recording medium. Furthermore, memory b2 can be either non-volatile or volatile memory.
[0674] For example, the memory b2 can store images or streams. In addition, the memory b2 can also store a program for the processor b1 to decode the stream.
[0675] In addition, for example, memory b2 can also play a role Figure 67The memory b2 can serve as a component for storing information among the multiple components of the decoding device 200 shown in FIG. Figure 67 The functions of the block memory 210 and the frame memory 214 are shown. More specifically, the memory b2 can store reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.).
[0676] In addition, in the decoding device 200, it is not necessary to install Figure 67 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 67 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.
[0677] After describing the overall processing flow of decoding device 200, the components included in decoding device 200 will be described below. Detailed descriptions of components included in decoding device 200 that perform the same processing as components included in encoding device 100 will be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, addition unit 208, block memory 210, frame memory 214, intra-frame prediction unit 216, inter-frame prediction unit 218, prediction control unit 220, and loop filter unit 212 included in decoding device 200 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, addition unit 116, block memory 118, frame memory 122, intra-frame prediction unit 124, inter-frame prediction unit 126, prediction control unit 128, and loop filter unit 120 included in encoding device 100, respectively.
[0678] [Overall decoding process]
[0679] Figure 69 This is a flowchart showing an example of the overall decoding process performed by the decoding device 200 .
[0680] First, the partition determination unit 224 of the decoding device 200 determines a partition pattern for each of a plurality of fixed-size blocks (128×128 pixels) included in a picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This partition pattern is the partition pattern selected by the encoding device 100. The decoding device 200 then performs steps Sp_2 to Sp_6 on each of the plurality of blocks comprising this partition pattern.
[0681] The entropy decoding unit 202 decodes (specifically, performs entropy decoding) the encoded quantization coefficients and prediction parameters of the current block (step Sp_2).
[0682] Next, the inverse quantization unit 204 and the inverse transformation unit 206 perform inverse quantization and inverse transformation on the plurality of quantized coefficients to restore the prediction residual of the current block (step Sp_3).
[0683] Next, the prediction processing unit composed of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a predicted image of the current block (step Sp_4).
[0684] Next, the adding unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the predicted image to the prediction residual (step Sp_5).
[0685] Then, when the reconstructed image is generated, the loop filter unit 212 filters the reconstructed image (step Sp_6).
[0686] Then, the decoding device 200 determines whether decoding of the entire picture is completed (step Sp_7 ). If it is determined that decoding is not completed (No in step Sp_7 ), the processing from step Sp_1 is repeatedly executed.
[0687] Furthermore, the processing of steps Sp_1 to Sp_7 may be performed sequentially by the decoding device 200 , and a plurality of some of these processing may be performed in parallel or in a different order.
[0688] [Division Decision Section]
[0689] Figure 70 : is a diagram showing the relationship between the division determination unit 224 and other components. As an example, the division determination unit 224 may also perform the following processing.
[0690] The partition decision unit 224 collects block information from, for example, the block memory 210 or the frame memory 214 and further obtains parameters from the entropy decoding unit 202. Furthermore, the partition decision unit 224 may determine a partitioning pattern for the fixed-size blocks based on the block information and parameters. Furthermore, the partition decision unit 224 may output information indicating the determined partitioning pattern to the inverse transform unit 206, the intra-frame prediction unit 216, and the inter-frame prediction unit 218. The inverse transform unit 206 may inversely transform the transform coefficients based on the partitioning pattern indicated by the information from the partition decision unit 224. The intra-frame prediction unit 216 and the inter-frame prediction unit 218 may generate a predicted image based on the partitioning pattern indicated by the information from the partition decision unit 224.
[0691] [Entropy decoding unit]
[0692] Figure 71 This is a block diagram showing an example of the functional configuration of the entropy decoding unit 202 .
[0693] The entropy decoding unit 202 entropy decodes the stream to generate quantization coefficients, prediction parameters, and parameters related to the segmentation pattern. CABAC is used, for example, in this entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a multi-valued unit 202c. The binary arithmetic decoding unit 202a performs arithmetic decoding on the stream as a binary signal using the context values derived by the context control unit 202b. Similar to the context control unit 110b of the encoding device 100, the context control unit 202b derives context values, i.e., the probability of occurrence of the binary signal, based on the characteristics of the syntactic elements or the surrounding conditions. The multi-valued unit 202c performs debinarization, converting the binary signal output from the binary arithmetic decoding unit 202a into a multi-valued signal representing the aforementioned quantization coefficients, etc. This debinarization is performed using the binarization method described above.
[0694] The entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the stream to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 (see Figure 1 The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device 100 side.
[0695] [Entropy decoding unit]
[0696] Figure 72 3 is a diagram showing the CABAC flow in the entropy decoding unit 202 .
[0697] Initialization is first performed in the CABAC encoding process within the entropy decoding unit 202. This initialization initializes the binary arithmetic decoding unit 202c and sets the initial context value. The binary arithmetic decoding unit 202c and the multi-valued encoding unit 202c then perform arithmetic decoding and multi-valued encoding on, for example, the coded data of a CTU. At this time, the context control unit 202b updates the context value each time arithmetic decoding is performed. The context control unit 202b then backs off the context value as post-processing. This backed-off context value is used, for example, as the initial context value for the next CTU.
[0698] [Inverse quantization unit]
[0699] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block, which are input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantization unit 204 then outputs the inversely quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0700] Figure 73 This is a block diagram showing an example of the functional configuration of the inverse quantization unit 204 .
[0701] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204 a , a predicted quantization parameter generation unit 204 b , a quantization parameter storage unit 204 d , and an inverse quantization processing unit 204 e .
[0702] Figure 74 This is a flowchart showing an example of inverse quantization performed by the inverse quantization unit 204 .
[0703] As an example, the inverse quantization unit 204 may be based on Figure 74 In the flow shown, inverse quantization is performed on each CU. Specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). If it is determined that inverse quantization is to be performed (yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameter for the current block from the entropy decoder 202 (step Sv_12).
[0704] Next, the predicted quantization parameter generator 204b obtains a quantization parameter for a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). Based on the obtained quantization parameter, the predicted quantization parameter generator 204b generates a predicted quantization parameter for the current block (step Sv_14).
[0705] The quantization parameter generator 204a then adds the differential quantization parameter for the current block obtained from the entropy decoder 202 to the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b (step Sv_15). This addition generates the quantization parameter for the current block. Furthermore, the quantization parameter generator 204a stores the quantization parameter for the current block in the quantization parameter storage unit 204d (step Sv_16).
[0706] Next, the inverse quantization processing unit 204e inversely quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameters generated in step Sv_15 (step Sv_17).
[0707] In addition, the differential quantization parameter can also be decoded at the bit sequence level, picture level, slice level, brick level, or CTU level. In addition, the initial value of the quantization parameter can also be decoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0708] Furthermore, the inverse quantization unit 204 may include a plurality of inverse quantizers, and may inversely quantize the quantized coefficients using an inverse quantization method selected from a plurality of inverse quantization methods.
[0709] [Inverse transformation unit]
[0710] The inverse transform unit 206 restores the prediction residual by performing inverse transform on the transform coefficients input from the inverse quantization unit 204 .
[0711] For example, when the information read from the stream indicates that EMT or AMT is applied (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.
[0712] Furthermore, for example, when the information read from the stream indicates that NSST is applied, the inverse transform unit 206 applies inverse re-transformation to the transform coefficients.
[0713] Figure 75 This is a flowchart showing an example of processing performed by the inverse transformation unit 206.
[0714] For example, the inverse transform unit 206 determines whether information indicating that an orthogonal transform is not to be performed exists in the stream (step St_11). If this information is not present (No in step St_11), the inverse transform unit 206 obtains information indicating the transform type decoded by the entropy decoding unit 202 (step St_12). Next, based on this information, the inverse transform unit 206 determines the transform type to be used for the orthogonal transform in the encoding device 100 (step St_13). The inverse transform unit 206 then performs an inverse orthogonal transform using the determined transform type (step St_14).
[0715] Figure 76 This is a flowchart showing another example of the processing performed by the inverse transformation unit 206.
[0716] For example, the inverse transform unit 206 determines whether the transform size is less than or equal to a predetermined value (step Su_11). If it is determined that the transform size is less than or equal to the predetermined value (yes in step Su_11), the inverse transform unit 206 obtains information from the entropy decoding unit 202 indicating which transform type, among the one or more transform types included in the first transform type group, is used by the encoding device 100 (step Su_12). This information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.
[0717] Based on this information, the inverse transform unit 206 determines the transform type used for the orthogonal transform in the encoding device 100 (step Su_13). The inverse transform unit 206 then performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, if it is determined in step Su_11 that the transform size is not less than a predetermined value (No in step Su_11), the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the second transform type group (step Su_15).
[0718] In addition, as an example, the inverse orthogonal transform performed by the inverse transform unit 206 may be performed for each TU according to Figure 75 or Figure 76 Alternatively, the inverse orthogonal transform may be performed using a predetermined transform type without decoding the information indicating the transform type used in the orthogonal transform. Specifically, the transform type may be DST7 or DCT8, and the inverse transform basis functions corresponding to the transform type may be used in the inverse orthogonal transform.
[0719] [Addition Department]
[0720] The adder 208 reconstructs the current block by adding the prediction residual input from the inverse transform unit 206 to the predicted image input from the prediction control unit 220. In other words, it generates a reconstructed image of the current block. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.
[0721] [Block Memory]
[0722] The block memory 210 is a storage unit for storing blocks in the current picture to be referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the adding unit 208 .
[0723] [Loop filter unit]
[0724] The loop filter unit 212 performs loop filtering on the reconstructed image generated by the adder unit 208 , and outputs the filtered reconstructed image to the frame memory 214 , a display device, or the like.
[0725] When the information indicating the on / off of the ALF read from the stream indicates that the ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0726] Figure 77 2 is a block diagram showing an example of the functional configuration of the loop filter unit 212. The loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding device 100.
[0727] The loop filter unit 212 is, for example, Figure 77 As shown, the loop filter unit 212a includes a deblocking filter processing unit 212a, an SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a applies the above-mentioned deblocking filter processing to the reconstructed image. The SAO processing unit 212b applies the above-mentioned SAO processing to the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 212c applies the above-mentioned ALF processing to the reconstructed image after the SAO processing. In addition, the loop filter unit 212 may not include Figure 77 All the processing units disclosed herein may also include only a portion of the processing units. In addition, the loop filter unit 212 may also be configured in accordance with Figure 77 A structure in which the above-mentioned respective processes are performed in an order different from the processing order disclosed in .
[0728] [Frame Memory]
[0729] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter unit 212.
[0730] [Prediction Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]
[0731] Figure 78 This is a flowchart showing an example of processing performed by the prediction unit of the decoding device 200. Furthermore, as an example, the prediction unit is composed of all or part of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The prediction processing unit includes, for example, the intra prediction unit 216 and the inter prediction unit 218.
[0732] The prediction unit generates a predicted image for the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Prediction signals include, for example, intra-frame prediction signals and inter-frame prediction signals. Specifically, the prediction unit generates a predicted image for the current block using a reconstructed image obtained by generating predicted images for other blocks, restoring prediction residuals, and adding the predicted images. The prediction unit of the decoding device 200 generates the same predicted image as the predicted image generated by the prediction unit of the encoding device 100. In other words, the prediction image generation methods used by these prediction units are common or corresponding to each other.
[0733] The reconstructed image may be, for example, an image of a reference picture or an image including the current block, that is, an image of a decoded block in the current picture (ie, the aforementioned other block). The decoded block in the current picture may be, for example, an adjacent block of the current block.
[0734] Figure 79 This is a flowchart showing another example of processing performed by the prediction unit of the decoding device 200 .
[0735] The prediction unit determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode can be determined based on prediction parameters.
[0736] If the first mode is determined to be a mode for generating a predicted image, the prediction unit generates a predicted image according to the first mode (step Sr_2a). Furthermore, if the second mode is determined to be a mode for generating a predicted image, the prediction unit generates a predicted image according to the second mode (step Sr_2b). Furthermore, if the third mode is determined to be a mode for generating a predicted image, the prediction unit generates a predicted image according to the third mode (step Sr_2c).
[0737] The first, second, and third methods are different methods for generating predicted images, and may be, for example, inter-frame prediction, intra-frame prediction, or other prediction methods. In such prediction methods, the above-mentioned reconstructed image may also be used.
[0738] Figure 80A and Figure 80B This is a flowchart showing another example of processing performed in the prediction unit of the decoding device 200 .
[0739] As an example, the forecasting unit may also Figure 80A and Figure 80B The process shown in the figure is used for prediction processing. Figure 80A and Figure 80BThe intra block copy shown is a mode belonging to inter prediction, in which a block included in the current picture is referred to as a reference image or reference block. That is, in the intra block copy, a picture different from the current picture is not referred to. Figure 80A The PCM mode shown is a mode belonging to intra-frame prediction, and is a mode in which neither transformation nor quantization is performed.
[0740] [Intra-frame prediction unit]
[0741] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the stream, referring to blocks in the current picture stored in the block memory 210, thereby generating a predicted image (i.e., an intra predicted image) for the current block. Specifically, the intra prediction unit 216 performs intra prediction by referring to pixel values (e.g., luminance values and chrominance values) of blocks adjacent to the current block, thereby generating an intra predicted image and outputting the intra predicted image to the prediction control unit 220.
[0742] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0743] Furthermore, when the information read from the stream indicates application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.
[0744] Figure 81 This is a diagram showing an example of processing performed by the intra prediction unit 216 of the decoding device 200 .
[0745] The intra-frame prediction unit 216 first determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 exists (yes in step Sw_11), the intra-frame prediction unit 216 obtains information indicating the intra-frame prediction mode selected in the encoding device 100 in the MPM from the entropy decoding unit 202 (step Sw_12). In addition, this information is decoded by the entropy decoding unit 202 and output to the intra-frame prediction unit 216. Next, the intra-frame prediction unit 216 determines the MPM (step Sw_13). The MPM is composed of, for example, 6 intra-frame prediction modes. Then, the intra-frame prediction unit 216 determines the intra-frame prediction mode indicated by the information obtained in step Sw_12 from the multiple intra-frame prediction modes included in the MPM (step Sw_14).
[0746] On the other hand, if it is determined in step Sw_11 that the stream does not contain an MPM flag indicating 1 (No in step Sw_11), the intra prediction unit 216 obtains information indicating the intra prediction mode selected by the encoding device 100 (step Sw_15). Specifically, the intra prediction unit 216 obtains from the entropy decoding unit 202 information indicating the intra prediction mode selected by the encoding device 100, from one or more intra prediction modes not included in the MPM. This information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. The intra prediction unit 216 then determines the intra prediction mode indicated by the information obtained in step Sw_15 from among the one or more intra prediction modes not included in the MPM (step Sw_17).
[0747] The intra prediction unit 216 generates a predicted image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18 ).
[0748] [Inter-frame prediction unit]
[0749] The inter-frame prediction unit 218 predicts the current block by referring to the reference picture stored in the frame memory 214. Prediction is performed for the current block or a sub-block within the current block. A sub-block is contained within a block and is smaller than a block. The sub-block size can be 4x4 pixels, 8x8 pixels, or other sizes. The sub-block size can also be switched based on units such as slices, tiles, or pictures.
[0750] For example, the inter-frame prediction unit 218 uses the motion information (e.g., MV) read from the stream (e.g., the prediction parameters output from the entropy decoding unit 202) to perform motion compensation, thereby generating an inter-frame prediction image of the current block or sub-block, and outputs the inter-frame prediction image to the prediction control unit 220.
[0751] When the information read from the stream indicates that the OBMC mode is applied, the inter prediction unit 218 generates an inter prediction image using not only the motion information of the current block obtained by the motion search but also the motion information of adjacent blocks.
[0752] If the information decoded from the stream indicates that the FRUC mode is used, the inter-frame prediction unit 218 performs a motion search using the pattern matching method (bidirectional matching or template matching) decoded from the stream to derive motion information. The inter-frame prediction unit 218 then performs motion compensation (prediction) using the derived motion information.
[0753] Furthermore, when the BIO mode is applied, the inter-frame prediction unit 218 derives an MV based on a model assuming constant velocity linear motion. Furthermore, when the information read from the stream indicates that the affine mode is applied, the inter-frame prediction unit 218 derives an MV in sub-block units based on the MVs of multiple adjacent blocks.
[0754] [MV export process]
[0755] Figure 82 This is a flowchart showing an example of MV derivation in the decoding apparatus 200 .
[0756] The inter-frame prediction unit 218, for example, determines whether to decode motion information (e.g., MV). For example, the inter-frame prediction unit 218 may make this determination based on the prediction mode included in the stream or other information included in the stream. If the inter-frame prediction unit 218 determines to decode motion information, it derives the MV of the current block in a mode that decodes the motion information. On the other hand, if the inter-frame prediction unit 218 determines not to decode motion information, it derives the MV in a mode that does not decode motion information.
[0757] Here, MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, modes that decode motion information include the normal inter mode, normal merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Furthermore, motion information may include not only MVs but also predicted MV selection information, which will be described later. Furthermore, modes that do not decode motion information include FRUC mode. The inter prediction unit 218 selects a mode from these multiple modes for deriving the MV of the current block and uses this selected mode to derive the MV of the current block.
[0758] Figure 83 This is a flowchart showing another example of MV derivation in the decoding apparatus 200 .
[0759] The inter-frame prediction unit 218 determines whether to decode the differential MV. For example, the inter-frame prediction unit 218 can make this determination based on the prediction mode included in the stream or other information included in the stream. If the inter-frame prediction unit 218 determines to decode the differential MV, it can derive the MV of the current block in a mode that decodes the differential MV. In this case, for example, the differential MV included in the stream is decoded as prediction parameters.
[0760] On the other hand, if it is determined that the difference MV is not to be decoded, the inter prediction unit 218 derives the MV in a mode in which the difference MV is not to be decoded. In this case, the encoded difference MV is not included in the stream.
[0761] As described above, the MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the normal inter mode and affine mode (specifically, affine inter mode) encode the differential MV. Furthermore, the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode) do not encode the differential MV. The inter prediction unit 218 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0762] [MV Export > Normal Interframe Mode]
[0763] For example, when the information decoded from the stream indicates that the normal inter mode is applied, the inter prediction unit 218 derives an MV in the normal merge mode based on the information decoded from the stream, and performs motion compensation (prediction) using the MV.
[0764] Figure 84 This is a flowchart showing an example of inter prediction performed in the normal inter mode in the decoding apparatus 200 .
[0765] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation on each block. To do this, the inter-frame prediction unit 218 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sg_11). In other words, the inter-frame prediction unit 218 creates a candidate MV list.
[0766] Next, the inter-frame prediction unit 218 extracts N (N is an integer greater than or equal to 2) candidate MVs from the plurality of candidate MVs obtained in step Sg_11 as motion vector predictor candidates (also referred to as predicted MV candidates) in a predetermined order of priority (step Sg_12). Alternatively, the order of priority may be predetermined for each of the N predicted MV candidates.
[0767] Next, the inter prediction unit 218 decodes the predicted MV selection information from the input stream, and uses the decoded predicted MV selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted MV of the current block (step Sg_13).
[0768] Next, the inter prediction unit 218 decodes the difference MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded difference MV to the selected prediction MV (step Sg_14).
[0769] Finally, the inter-frame prediction unit 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are performed for each block. For example, when the processes of steps Sg_11 to Sg_15 are performed for all blocks included in a slice, inter-frame prediction using the normal inter-frame mode for that slice is completed. Alternatively, when the processes of steps Sg_11 to Sg_15 are performed for all blocks included in a picture, inter-frame prediction using the normal inter-frame mode for that picture is completed. Furthermore, the processes of steps Sg_11 to Sg_15 may be performed not for all blocks included in a slice, but for a portion of the blocks, in which case inter-frame prediction using the normal inter-frame mode for that slice is completed. Similarly, when the processes of steps Sg_11 to Sg_15 are performed for a portion of the blocks included in a picture, inter-frame prediction using the normal inter-frame mode for that picture is completed.
[0770] [MV Export > Normal Merge Mode]
[0771] For example, when the information read from the stream indicates application of the normal merge mode, the inter prediction unit 218 derives an MV in the normal merge mode and performs motion compensation (prediction) using the MV.
[0772] Figure 85 This is a flowchart illustrating an example of inter-frame prediction based on the normal merge mode in the decoding apparatus 200 .
[0773] The inter-frame prediction unit 218 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sh_11). In other words, the inter-frame prediction unit 218 creates a candidate MV list.
[0774] Next, the inter-frame prediction unit 218 selects one candidate MV from the multiple candidate MVs obtained in step Sh_11 to derive the MV of the current block (step Sh_12). Specifically, the inter-frame prediction unit 218 obtains MV selection information included as a prediction parameter in the stream, for example, and selects the candidate MV identified by the MV selection information as the MV of the current block.
[0775] Finally, the inter-frame prediction unit 218 performs motion compensation on the current block using the derived MV and the decoded reference picture to generate a predicted image for the current block (step Sh_13). The processing of steps Sh_11 to Sh_13 is performed on each block, for example. For example, when the processing of steps Sh_11 to Sh_13 is performed on all blocks included in a slice, the inter-frame prediction using the normal merge mode for the slice is completed. Alternatively, when the processing of steps Sh_11 to Sh_13 is performed on all blocks included in a picture, the inter-frame prediction using the normal merge mode for the picture is completed. Furthermore, the processing of steps Sh_11 to Sh_13 may be performed on a portion of the blocks rather than all blocks included in the slice, and the inter-frame prediction using the normal merge mode for the slice is completed. Similarly, when the processing of steps Sh_11 to Sh_13 is performed on a portion of the blocks included in the picture, the inter-frame prediction using the normal merge mode for the picture is completed.
[0776] [MV Export > FRUC Mode]
[0777] For example, if the information decoded from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 derives an MV in the FRUC mode and uses this MV for motion compensation (prediction). In this case, the motion information is not signaled by the encoding device 100, but is derived by the decoding device 200. For example, the decoding device 200 may also derive the motion information by performing a motion search. In this case, the decoding device 200 does not use the pixel values of the current block for the motion search.
[0778] Figure 86 This is a flowchart illustrating an example of inter-frame prediction based on the FRUC mode in the decoding apparatus 200 .
[0779] First, the inter-frame prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block and generates a list representing these MVs as candidate MVs (that is, a candidate MV list, which can also be shared with the candidate MV list of the normal merge mode) (step Si_11). Next, the inter-frame prediction unit 218 selects the best candidate MV from the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter-frame prediction unit 218 calculates an evaluation value for each candidate MV included in the candidate MV list and selects one candidate MV as the best candidate MV based on the evaluation value. Then, the inter-frame prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block can be derived by performing pattern matching in the surrounding area of the position in the reference picture corresponding to the selected best candidate MV. Specifically, a search is performed using pattern matching and evaluation values in the reference image for the area surrounding the best candidate MV. If an MV with a good evaluation value exists, the best candidate MV may be updated to that MV and set as the final MV for the current block. Updating to an MV with a better evaluation value may not be performed.
[0780] Finally, the inter-frame prediction unit 218 performs motion compensation on the current block using the derived MV and the decoded reference picture to generate a predicted image for the current block (step Si_15). The processing of steps Si_11 to Si_15 is performed on each block, for example. For example, when the processing of steps Si_11 to Si_15 is performed on all blocks included in a slice, the inter-frame prediction using the FRUC mode for the slice is completed. In addition, when the processing of steps Si_11 to Si_15 is performed on all blocks included in a picture, the inter-frame prediction using the FRUC mode for the picture is completed. The processing can also be performed in sub-block units in the same way as the above-mentioned block unit.
[0781] [MV Export > Affine Merge Mode]
[0782] For example, when the information read from the stream indicates application of the affine merge mode, the inter prediction unit 218 derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0783] Figure 87 This is a flowchart illustrating an example of inter-frame prediction based on the affine merge mode in the decoding device 200 .
[0784] In the affine merge mode, the inter prediction unit 218 first derives the MV of each control point of the current block (step Sk_11). Figure 46AAs shown, the control points are the top left and top right corners of the current block, or as Figure 46B As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block.
[0785] For example, when using Figures 47A to 47C In the case of the MV derivation method shown, as Figure 47A As shown, the inter-frame prediction unit 218 checks the decoded blocks A (left), B (top), C (top right), D (bottom left) and E (top left) in this order to determine the first valid block decoded in affine mode.
[0786] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to derive the MV of the control point. For example, if block A is determined and block A has two control points, Figure 47B As shown, the inter-frame prediction unit 218 calculates the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block by projecting the motion vectors v3 and v4 of the upper left corner and upper right corner of the decoded block including block A onto the current block. Thus, the MV of each control point is derived.
[0787] In addition, if Figure 49A As shown, if block A is determined and block A has two control points, the MV of three control points can also be calculated, or as shown in Figure 49B As shown, block A is determined, and when block A has three control points, MVs of two control points are calculated.
[0788] Furthermore, when the stream includes MV selection information as a prediction parameter, the inter prediction unit 218 may use the MV selection information to derive the MV of each control point of the current block.
[0789] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 218 uses two motion vectors v0 and v1 and the above-mentioned equation (1A), or uses three motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the MV of the sub-block as an affine MV (step Sk_12). Then, the inter-frame prediction unit 218 uses these affine MVs and the decoded reference picture to perform motion compensation on the sub-block (step Sk_13). When the processing of steps Sk_12 and Sk_13 is performed on all sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block is completed. That is, motion compensation is performed on the current block, and a predicted image for the current block is generated.
[0790] In addition, in step Sk_11, the candidate MV list mentioned above may also be generated. The candidate MV list may also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. Multiple MV derivation methods may be Figures 47A to 47C The MV derivation method shown, Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B The MV derivation method shown and any combination of other MV derivation methods.
[0791] Furthermore, the candidate MV list may include candidate MVs of modes other than the affine mode that perform prediction in sub-block units.
[0792] Furthermore, as a candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge pattern with two control points and a candidate MV list including candidate MVs of an affine merge pattern with three control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge pattern with two control points and a candidate MV list including candidate MVs of an affine merge pattern with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of either an affine merge pattern with two control points or an affine merge pattern with three control points may be generated.
[0793] [MV Export > Affine Inter Mode]
[0794] For example, when the information read from the stream indicates application of the affine inter mode, the inter prediction unit 218 derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.
[0795] Figure 88 This is a flowchart illustrating an example of inter prediction based on the affine inter mode in the decoding apparatus 200 .
[0796] In the affine inter mode, first, the inter prediction unit 218 derives the predicted MV (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_11). Figure 46A or Figure 46B As shown, it is the point at the upper left corner, upper right corner or lower left corner of the current block.
[0797] The inter-frame prediction unit 218 obtains the prediction MV selection information included in the stream as a prediction parameter, and uses the MV identified by the prediction MV selection information to derive the prediction MV of each control point of the current block. Figure 48A and Figure 48B In the case of the MV derivation method shown in FIG. 1 , the inter-frame prediction unit 218 selects Figure 48Aor Figure 48B The predicted MV (v0, v1) or (v0, v1, v2) of the control point of the current block is derived from the MV of the block identified by the predicted MV selection information in the decoded blocks near each control point of the current block.
[0798] Next, the inter-frame prediction unit 218 obtains, for example, each differential MV included as a prediction parameter in the stream, and adds the predicted MV of each control point of the current block to the differential MV corresponding to the predicted MV (step Sj_12). Thus, the MV of each control point of the current block is derived.
[0799] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 218 uses two motion vectors v0 and v1 and the above-mentioned equation (1A), or uses three motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the MV of the sub-block as an affine MV (step Sj_13). Then, the inter-frame prediction unit 218 uses these affine MVs and the decoded reference picture to perform motion compensation on the sub-block (step Sj_14). When the processing of steps Sj_13 and Sj_14 is performed on all sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block is completed. That is, motion compensation is performed on the current block, and a predicted image for the current block is generated.
[0800] Furthermore, in step Sj_11, the candidate MV list described above may be generated in the same manner as in step Sk_11.
[0801] [MV Export > Triangle Mode]
[0802] For example, when the information read from the stream indicates application of the triangular mode, the inter prediction unit 218 derives an MV in the triangular mode and performs motion compensation (prediction) using the MV.
[0803] Figure 89 This is a flowchart illustrating an example of inter-frame prediction based on the triangular mode in the decoding apparatus 200 .
[0804] In triangular mode, the inter-frame prediction unit 218 first divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter-frame prediction unit 218 can obtain partition information related to the division into each partition from the stream as a prediction parameter. Furthermore, the inter-frame prediction unit 218 can divide the current block into the first partition and the second partition based on the partition information.
[0805] Next, the inter-frame prediction unit 218 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple decoded blocks located temporally or spatially around the current block (step Sx_12). In other words, the inter-frame prediction unit 218 creates a candidate MV list.
[0806] The inter-frame prediction unit 218 then selects the candidate MV for the first partition and the candidate MV for the second partition as the first and second MVs, respectively, from the multiple candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter-frame prediction unit 218 may also obtain MV selection information from the stream as a prediction parameter to identify the selected candidate MVs. The inter-frame prediction unit 218 then selects the first and second MVs based on the MV selection information.
[0807] Next, the inter-frame prediction unit 218 performs motion compensation using the selected first MV and the decoded reference picture to generate a first predicted image (step Sx_14). Similarly, the inter-frame prediction unit 218 performs motion compensation using the selected second MV and the decoded reference picture to generate a second predicted image (step Sx_15).
[0808] Finally, the inter prediction unit 218 generates a prediction image of the current block by performing weighted addition on the first prediction image and the second prediction image (step Sx_16).
[0809] [Sports Search>DMVR]
[0810] For example, when the information read from the stream indicates application of DMVR, the inter prediction unit 218 performs motion search using DMVR.
[0811] Figure 90 3 is a flowchart showing an example of a motion search based on DMVR in the decoding apparatus 200 .
[0812] The inter-frame prediction unit 218 first derives the MV of the current block in merge mode (step S1_11). Next, the inter-frame prediction unit 218 searches the surrounding area of the reference picture represented by the MV derived in step S1_11 to derive the final MV for the current block (step S1_12). In other words, the MV of the current block is determined using DMVR.
[0813] Figure 91 This is a flowchart showing a detailed example of the motion search based on DMVR in the decoding device 200.
[0814] First, the inter-frame prediction unit 218 Figure 58AIn Step 1 shown in FIG. 1 , the cost of the search position (also called the starting point) represented by the initial MV and the eight search positions located around it are calculated. Then, the inter-frame prediction unit 218 determines whether the cost of the search position other than the starting point is the minimum. Here, if it is determined that the cost of the search position other than the starting point is the minimum, the inter-frame prediction unit 218 moves to the search position with the minimum cost and performs Figure 58A On the other hand, if the cost of the starting point is the smallest, the inter-frame prediction unit 218 skips the process of Step 2. Figure 58A The process of Step 2 shown in the figure is replaced by the process of Step 3.
[0815] exist Figure 58A In Step 2, the inter-frame prediction unit 218 uses the search position moved based on the processing results of Step 1 as a new starting point and performs the same search as in Step 1. Furthermore, the inter-frame prediction unit 218 determines whether the cost of a search position other than the starting point is the lowest. If the cost of a search position other than the starting point is the lowest, the inter-frame prediction unit 218 proceeds to Step 4. On the other hand, if the cost of the starting point is the lowest, the inter-frame prediction unit 218 proceeds to Step 3.
[0816] In Step 4 , the inter prediction unit 218 treats the search position of the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a difference vector.
[0817] exist Figure 58A In Step 3 shown, the inter-frame prediction unit 218 determines the fractional-precision pixel position with the lowest cost based on the costs of the four points above, below, and to the left and right of the starting point in Step 1 or Step 2, and uses this pixel position as the final search position. This fractional-precision pixel position is determined by weighted summing the vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) located above, below, and to the left, using the costs of the four search positions as weights. The inter-frame prediction unit 218 then determines the difference between the position indicated by the initial MV and the final search position as a difference vector.
[0818] [Motion Compensation > BIO / OBMC / LIC]
[0819] For example, when the information read from the stream indicates application of correction to the predicted image, the inter prediction unit 218 corrects the predicted image according to the correction mode when generating the predicted image. Examples of such modes include BIO, OBMC, and LIC.
[0820] Figure 92 This is a flowchart showing an example of generation of a predicted image in the decoding device 200 .
[0821] The inter-frame prediction unit 218 generates a predicted image (step Sm_11 ), and corrects the predicted image using any of the above-mentioned modes (step Sm_12 ).
[0822] Figure 93 This is a flowchart showing another example of generation of a predicted image in the decoding device 200 .
[0823] The inter-frame prediction unit 218 derives the MV of the current block (step Sn_11). Next, the inter-frame prediction unit 218 uses this MV to generate a predicted image (step Sn_12) and determines whether to perform correction processing (step Sn_13). For example, the inter-frame prediction unit 218 obtains prediction parameters contained in the stream and determines whether to perform correction processing based on these prediction parameters. These prediction parameters are, for example, flags indicating whether to apply the aforementioned modes. If correction processing is determined to be performed (yes in step Sn_13), the inter-frame prediction unit 218 corrects the predicted image to generate the final predicted image (step Sn_14). Furthermore, in the LIC, the brightness and color difference of the predicted image can be corrected in step Sn_14. On the other hand, if correction processing is determined not to be performed (no in step Sn_13), the inter-frame prediction unit 218 outputs the predicted image as the final predicted image without correction (step Sn_15).
[0824] [Motion Compensation > OBMC]
[0825] For example, when the information read from the stream indicates the application of OBMC, the inter prediction unit 218 corrects the predicted image in accordance with OBMC when generating the predicted image.
[0826] Figure 94 : is a flowchart showing an example of correction of a predicted image based on OBMC in the decoding device 200. Figure 94 The flowchart shows the use of Figure 62 The process of correcting the predicted image of the current picture and the reference picture is shown.
[0827] First, if Figure 62 As shown, the inter prediction unit 218 uses the MV assigned to the current block to obtain a predicted image (Pred) based on normal motion compensation.
[0828] Next, the inter-frame prediction unit 218 applies (reuses) the MV (MV_L) derived for the decoded left-neighboring block to the current block, obtaining a predicted image (Pred_L). The inter-frame prediction unit 218 then performs a first correction on the predicted image by superimposing the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0829] Similarly, the inter-frame prediction unit 218 applies (reuses) the MV (MV_U) derived for the decoded upper adjacent block to the current block, obtaining a predicted image (Pred_U). The inter-frame prediction unit 218 then performs a second correction on the predicted image by overlaying it with the predicted image (e.g., Pred and Pred_L) that has undergone the first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image for the current block, blended (smoothed) with the boundaries between adjacent blocks.
[0830] [Motion Compensation > BIO]
[0831] For example, when the information read from the stream indicates the application of BIO, the inter prediction unit 218 modifies the predicted image according to BIO when generating the predicted image.
[0832] Figure 95 This is a flowchart showing an example of correction of a predicted image based on BIO in the decoding device 200 .
[0833] like Figure 63 As shown, the inter-frame prediction unit 218 uses two reference pictures (Ref0, Ref1) different from the picture containing the current block (CurPic) to derive two motion vectors (M0, M1). The inter-frame prediction unit 218 then uses these two motion vectors (M0, M1) to derive a predicted image for the current block (step Sy_11). Motion vector M0 is the motion vector (MVx0, MVy0) corresponding to reference picture Ref0, and motion vector M1 is the motion vector (MVx1, MVy1) corresponding to reference picture Ref1.
[0834] Next, the inter-frame prediction unit 218 uses the motion vector M0 and the reference picture L0 to derive the interpolated image I of the current block. 0 In addition, the inter-frame prediction unit 218 uses the motion vector M1 and the reference picture L1 to derive the interpolated image I of the current block. 1 (Step Sy_12). Here, the interpolated image I 0 The interpolated image I is the image contained in the reference image Ref0 derived for the current block. 1 It is the image included in the reference picture Ref1 derived for the current block. 0 and interpolated image I 1 The interpolated image I may be the same size as the current block. 0 and interpolated image I 1 They can also be images larger than the current block. 0 and I 1It may include a predicted image derived by applying a motion vector (M0, M1), a reference picture (L0, L1), and a motion compensation filter.
[0835] In addition, the inter-frame prediction unit 218 uses the interpolation image I 0 and interpolated image I 1 Export the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1 )(Step Sy_13). In addition, the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), the vertical gradient image is (Ix 0 , 1x 1 The inter-frame prediction unit 218 may also derive the gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may be any image that represents the spatial variation of pixel values in the horizontal or vertical direction.
[0836] Next, the inter-frame prediction unit 218 uses the interpolation image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) derives the optical flow (vx, vy) as the velocity vector (step Sy_14). As an example, the sub-block may be a sub-CU of 4x4 pixels.
[0837] Next, the inter-frame prediction unit 218 uses the optical flow (vx, vy) to correct the predicted image of the current block. For example, the inter-frame prediction unit 218 uses the optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_15). The inter-frame prediction unit 218 then uses the correction values to correct the predicted image of the current block (step Sy_16). The correction values may be derived for each pixel, for multiple pixels, or for sub-blocks.
[0838] In addition, the BIO process is not limited to Figure 95 The disclosed processing can be implemented only Figure 95 A part of the disclosed processing may be added to or replaced with a different processing, or may be executed in a different processing order.
[0839] [Motion Compensation > LIC]
[0840] For example, when the information read from the stream indicates application of LIC, the inter prediction unit 218 modifies the predicted image in accordance with LIC when generating the predicted image.
[0841] Figure 96 This is a flowchart showing an example of modification of a predicted image based on LIC in the decoding apparatus 200 .
[0842] First, the inter prediction unit 218 obtains a reference image corresponding to the current block from a decoded reference picture using MV (step Sz_11 ).
[0843] Next, the inter-frame prediction unit 218 extracts information indicating how the luminance value of the current block changes between the reference picture and the current picture (step Sz_12). Figure 66A As shown, this extraction is performed based on the luminance pixel values of the decoded left-adjacent reference region (peripheral reference region) and the decoded upper-adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. The inter-frame prediction unit 218 then calculates the luminance correction parameter using information indicating how the luminance value changes (step Sz_13).
[0844] The inter-frame prediction unit 218 generates a predicted image for the current block by applying the brightness correction parameters to the reference image within the reference picture specified by the MV (step Sz_14). Specifically, the predicted image, which is the reference image within the reference picture specified by the MV, is corrected based on the brightness correction parameters. This correction can be performed to correct brightness or color difference.
[0845] [Prediction Control Department]
[0846] The prediction control unit 220 selects either the intra-frame prediction image or the inter-frame prediction image, and outputs the selected prediction image to the addition unit 208. Generally speaking, the structure, function, and processing of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device 200 side may correspond to the structure, function, and processing of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device 100 side.
[0847] [Overview of the first form]
[0848] Figure 97This diagram represents the syntactic structure of information related to the use of decoding units contained in the HRD (Hypothetical Reference Decoder) encoded in the Buffering Period SEI (Supplemental Enhancement Information). The HRD serves to validate the coded bitstream to comply with the standard. In the first form of this embodiment, the encoding device 100 encodes all information required to parse the HRD-related SEI message into the HRD-related SEI message. Specifically, the HRD-related SEI message refers to the picture timing SEI message, the buffering period SEI message, and the decoding unit information SEI message.
[0849] In particular, the encoding device 100 is as follows Figure 97 As shown, a flag encoding information about the usage of decoding units in the HRD is also encoded in the during-buffering SEI message as bp_decoding_unit_hrd_params_present_flag 300. That is, encoding device 100 also encodes the flag encoding information about the usage of decoding units in the HRD encoded in the sequence parameter set in the during-buffering SEI message. Because encoding device 100 has already encoded this flag or information equivalent to this flag using the sequence parameter set, the information indicated by this flag encoded in the during-buffering SEI message must not conflict with information indicated by a flag corresponding to this flag, which is encoded in the sequence parameter set that applies to the same picture as the picture to which the during-buffering SEI message to which this flag is encoded is applied.
[0850] like Figure 97 As shown in , when CPB parameters for a decoding unit are present in a picture timing SEI message or a decoding unit information SEI message, a flag, decoding_unit_cpb_params_in_pic_timing_sei_flag 400, is encoded in the buffering period SEI message. As a result, all three SEI messages—the picture timing SEI message, the buffering period SEI message, and the decoding unit information SEI message—can be parsed independently from the sequence parameter set. Parsing independently from the sequence parameter set means that decoding apparatus 200 can parse all three SEI messages—the picture timing SEI message, the buffering period SEI message, and the decoding unit information SEI message—without referencing information encoded in the sequence parameter set.
[0851] However, in the configuration of this embodiment, the encoding device 100 encodes at least one of the one or more HRD parameters into the picture timing SEI message, relying solely on the buffering period SEI message among the one or more HRD-related SEI messages. Furthermore, the encoding device 100 encodes at least one of the one or more HRD parameters into the decoding unit information SEI message, relying solely on the buffering period SEI message among the one or more HRD-related SEI messages.
[0852] The primary technical advantage of the first aspect of the present disclosure is that the encoding device 100 can encode a signal so that all HRD-related SEI messages, including picture timing SEI messages, buffering period SEI messages, and decoding unit information SEI messages, can be parsed independently of the sequence parameter set. Thus, the encoding device 100 can encode a signal so that upon receiving an HRD-related SEI message, the information in the SEI message can be immediately parsed, without buffering the signal for subsequent parsing.
[0853] Furthermore, the decoding apparatus 200 can independently parse HRD-related SEI messages, including all picture timing SEI messages, buffering period SEI messages, and decoding unit information SEI messages, from the sequence parameter set. This allows the decoding apparatus 200 to immediately parse the information in the HRD-related SEI message upon receipt, without buffering the signal for subsequent parsing.
[0854] [Actions in the first form]
[0855] Figure 98 This is a flowchart showing an example of a process for checking bitstream consistency using HRD in the first form of the embodiment. Figure 98 In the first aspect of the present disclosure, an example of a process for checking the consistency of a bit stream using HRD is described.
[0856] This process can be applied to a typical bitstream that includes a sequence parameter set, a picture parameter set, an SEI message, and a NAL unit for slice data in the above order. Here, the slice data is preferably one or more slice data attached to one picture.
[0857] The decoding apparatus 200 initially parses all available sequence parameter sets. That is, the decoding apparatus 200 parses the bitstream for available sequence parameter sets (step S100). The decoding apparatus 200 then stores the information encoded in all available sequence parameter sets for future reference.
[0858] The decoding device 200 then parses the buffering period SEI message applied to the buffering period being processed (step S101). The decoding device 200 then stores the parsed information. According to the first aspect of the embodiment of the present disclosure, since the information of the sequence parameter set applied to the buffering period being processed is not necessary for parsing the information in the SEI message, the decoding device 200 can immediately perform the parsing process upon receiving the buffering period SEI message.
[0859] If the current picture is not associated with a buffering period SEI message, the decoding apparatus 200 can skip this step. However, a buffering period SEI message not associated with the current picture is typically transmitted in a picture preceding the current picture. Therefore, in this case, the decoding apparatus 200 can already utilize the information contained in the buffering period SEI message not associated with the current picture.
[0860] The decoding device 200 then parses the picture timing SEI message applied to the current picture (step S102). The decoding device 200 then stores the parsed information. According to the first aspect of the embodiment of the present disclosure, since the parsing process of the picture timing SEI message does not require information about the sequence parameter set applied to the current picture, the decoding device 200 can immediately perform the parsing process upon receiving the picture timing SEI message. The decoding device 200 can then utilize the information in the buffering period SEI message.
[0861] Next, the decoding apparatus 200 determines whether the bitstream being processed includes a decoding unit information SEI message for the picture being processed. In other words, the decoding apparatus 200 determines whether the decoding unit information SEI message is available for the picture being processed (step S103).
[0862] If the decoding device 200 determines that a decoding unit information SEI message is available for the current picture (YES in step S103), the decoding device 200 parses the available decoding unit information SEI message (step S104). According to the first aspect of the present disclosure, since the decoding unit information SEI message parsing process does not require information about the sequence parameter set applied to the current picture, the decoding device 200 can immediately perform the parsing process upon receiving the decoding unit information SEI message. The information in the buffering period SEI message is then available to the decoding device 200.
[0863] When the decoding device 200 determines that the decoding unit information SEI message is not available in the processing target picture (No in step S103 ), the decoding device 200 skips parsing the decoding unit information SEI message.
[0864] Next, in order to determine which picture parameter set and which sequence parameter set the decoding apparatus 200 applies to the current picture, the decoding apparatus 200 analyzes the slice header of at least one slice of the current picture (step S105 ).
[0865] The decoding apparatus 200 then uses the analyzed information to check the consistency of the bitstream being processed (step S106 ). At this time, all analyzed information is used for checking the consistency of the bitstream using the HRD.
[0866] At this time, decoding apparatus 200 uses information such as flags and parameters used in checking the consistency of HRD bitstreams described in Annex C of Versatile Video Coding (Draft 10) as analyzed information to check the consistency of the bitstream being processed.
[0867] Then, the decoding device 200 ends the processing.
[0868] [Modification]
[0869] Figure 99 is a diagram showing the syntax structure of information related to the usage of decoding units included in the HRD encoded in the picture timing SEI. When the encoding device 100 has CPB parameters for decoding units in the picture timing SEI or decoding unit information SEI, the decoding device 200 may Figure 97 The illustrated decoding_unit_cpb_params_in_pic_timing_sei_flag 400 is coded into the picture timing SEI message instead of into the buffering duration SEI message.
[0870] In this case, the decoding device 200 relies on information present in both the buffering period SEI message and the picture timing SEI message to parse the decoding unit SEI message. In other words, the decoding device 200 refers to information present in both the buffering period SEI message and the picture timing SEI message to parse the decoding unit SEI message.
[0871] [Install]
[0872] Figure 100 This is a flowchart showing an example of the operation of the encoding device in the embodiment. Figure 7 The encoding device 100 shown performs Figure 100 Specifically, the processor a1 uses the memory a2 to perform the following operations.
[0873] Processor a1 encodes one or more HRD parameters independently of a sequence parameter set into one or more HRD-related SEI messages, wherein the one or more HRD parameters are one or more parameters of the HRD related to the decoding unit, and the one or more HRD-related SEI messages are one or more SEI messages associated with the HRD. (Step S200)
[0874] Alternatively, one or more HRD-related SEI messages may include a buffering period SEI message, and the processor a1 may encode at least one of the one or more HRD parameters into the buffering period SEI message independently of the sequence parameter set and other HRD-related SEI messages.
[0875] Alternatively, one or more HRD-related SEI messages may include a picture timing SEI message, and the processor a1 may encode at least one of the one or more HRD parameters into the picture timing SEI message depending only on the buffering period SEI message in the one or more HRD-related SEI messages.
[0876] Alternatively, one or more HRD-related SEI messages may include a decoding unit information SEI message, and the processor a1 may encode at least one of the one or more HRD parameters into the decoding unit information SEI message depending only on the buffering period SEI message in the one or more HRD-related SEI messages.
[0877] Alternatively, the processor a1 may encode at least one of the more than one HRD parameters into the sequence parameter set, and encode at least one of the more than one HRD parameters into at least one of the more than one HRD-related SEI messages independently of the sequence parameter set.
[0878] Alternatively, the processor a1 may encode the information on whether the CPB parameters for the decoding unit exist in the picture timing SEI message or the decoding unit information SEI message into the buffering period SEI message or the picture timing SEI message.
[0879] Alternatively, when there are multiple HRD parameters, the processor a1 may encode a parameter indicating whether the multiple HRD parameters are present in the picture timing SEI message or the decoding unit information SEI message, into the buffering period SEI message.
[0880] Here, processor al is a specific example of a circuit.
[0881] Figure 101 This is a flowchart showing an example of the operation of the decoding device in the embodiment. Figure 67 The decoding apparatus 200 shown performs Figure 101Specifically, the processor b1 uses the memory b2 to perform the following operations.
[0882] First, the decoding device 200 decodes one or more HRD parameters independently of the sequence parameters from one or more HRD-associated SEI messages, wherein the one or more HRD parameters are one or more parameters of the HRD related to the decoding unit, and the one or more HRD-associated SEI messages are one or more SEI messages associated with the HRD (step S300).
[0883] Alternatively, one or more HRD-related SEI messages may include a buffering period SEI message, and the processor b1 may decode at least one of the one or more HRD parameters from the buffering period SEI message independently of the sequence parameter set and other HRD-related SEI messages.
[0884] The one or more HRD-related SEI messages may include a picture timing SEI message, and the processor b1 may decode at least one of the one or more HRD parameters from the picture timing SEI message depending only on the buffering period SEI message in the one or more HRD-related SEI messages.
[0885] Alternatively, one or more HRD-related SEI messages may include a decoding unit information SEI message, and the processor b1 may decode at least one of the one or more HRD parameters from the decoding unit information SEI message depending only on the buffering period SEI message in the one or more HRD-related SEI messages.
[0886] Alternatively, the processor b1 may encode at least one of the one or more HRD parameters into a sequence parameter, and decode at least one of the one or more HRD parameters from at least one of the one or more HRD-related SEI messages independently of the sequence parameter.
[0887] Alternatively, the processor b1 decodes the information on whether the CPB parameters for the decoding unit exist in the picture timing SEI message or the decoding unit information SEI message from the buffering period SEI message or the picture timing SEI message.
[0888] Alternatively, the processor b1 may decode a parameter indicating whether multiple HRD parameters for the decoding unit exist from the buffering period SEI message as one of the one or more HRD parameters. Alternatively, if multiple HRD parameters exist, the processor b1 may decode a parameter indicating whether the multiple HRD parameters exist in the picture timing SEI message or the decoding unit information SEI message from the buffering period SEI message.
[0889] Here, the processor b1 is a specific example of a circuit.
[0890] Furthermore, one or more aspects disclosed herein may be combined with at least a portion of other aspects disclosed herein for implementation. Furthermore, a portion of the processing, a portion of the device structure, a portion of the syntax, etc. described in the flowcharts of one or more aspects disclosed herein may be combined with other aspects for implementation.
[0891] It is also possible to implement by combining one or more aspects disclosed herein with at least a portion of other aspects disclosed herein. In addition, it is also possible to implement by combining a portion of the processing, a portion of the device structure, a portion of the syntax, etc. described in the flowchart of one or more aspects disclosed herein with other aspects.
[0892] [Implementation and Application]
[0893] In each of the above embodiments, each functional block or active block can generally be implemented by an MPU (microprocessing unit) and memory, etc. Alternatively, the processing of each functional block can be implemented by a program execution unit such as a processor that reads and executes software (programs) recorded in a recording medium such as a ROM. The software can be distributed. The software can also be recorded in various recording media such as semiconductor memories. In addition, each functional block can also be implemented by hardware (dedicated circuit).
[0894] The processing described in each embodiment can be implemented by centralized processing using a single device (system) or by distributed processing using multiple devices. In addition, the processors that execute the above programs can be single or multiple. In other words, centralized processing can be performed or distributed processing can be performed.
[0895] The aspects of the present disclosure are not limited to the above-described embodiments, and various modifications are possible, which are also included in the scope of the aspects of the present disclosure.
[0896] Furthermore, here, application examples of the moving picture encoding method (image encoding method) or moving picture decoding method (image decoding method) described in each of the above embodiments and various systems implementing these application examples are described. Such a system may be characterized by including an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding and decoding device including both. Other configurations of such a system may be modified as appropriate depending on the circumstances.
[0897] [Example of use]
[0898] Figure 102This diagram shows the overall structure of a content supply system ex100 for implementing content distribution services. The communication service provision area is divided into cells of desired sizes, and in the illustrated example, base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.
[0899] In the content delivery system ex100, various devices, such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115, are connected to the Internet ex101 via an Internet service provider ex102, a communication network ex104, and base stations ex106-ex110. The content delivery system ex100 may also connect some of these devices in combination. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network, short-range wireless, or other means, rather than via base stations ex106-ex110. Furthermore, the streaming server ex103 may be connected to various devices, such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115, via the Internet ex101. Furthermore, the streaming server ex103 may be connected to a terminal, such as a hotspot within an airplane ex117, via a satellite ex116.
[0900] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly to the aircraft ex117 without going through the satellite ex116.
[0901] The camera ex113 is a device such as a digital camera capable of capturing both still and moving images. Furthermore, the smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) compatible with mobile communication systems such as 2G, 3G, 3.9G, 4G, and what will be called 5G in the future.
[0902] The home appliance ex114 is a refrigerator or equipment included in a household fuel cell cogeneration system.
[0903] In the content delivery system ex100, terminals with camera functions are connected to the streaming server ex103 via a base station ex106 or the like, enabling on-site distribution and the like. During on-site distribution, terminals (such as the computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and terminals within an airplane ex117) can perform the encoding processing described in the above embodiments on still images or moving image content captured by the user using the terminal, multiplex the encoded video data with audio data obtained by encoding the corresponding audio, and transmit the resulting data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.
[0904] Meanwhile, the streaming server ex103 streams content data sent by requesting clients. Clients are computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or terminals inside airplanes ex117, all capable of decoding the encoded data. Each device that receives the distributed data can also decode and reproduce the data. In other words, each device can function as an image decoding device according to one aspect of the present disclosure.
[0905] [Distributed Processing]
[0906] Alternatively, the streaming server ex103 can consist of multiple servers or computers, distributing data by distributing processing or recording. For example, the streaming server ex103 can be implemented as a CDN (Content Delivery Network), which distributes content through a network connecting numerous edge servers distributed worldwide. In a CDN, physically close edge servers are dynamically assigned to clients. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Furthermore, in the event of various errors or changes in communication status due to increased traffic, processing can be distributed across multiple edge servers, distribution can be switched to other edge servers, or distribution can be continued by bypassing a faulty portion of the network, thus achieving high-speed and stable delivery.
[0907] In addition, the encoding process of the captured data is not limited to the distributed processing itself, and can be performed by each terminal, on the server side, or shared. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or encoding amount of the image of the frame or scene unit is detected. In addition, in the second cycle, a process is performed to maintain the image quality and improve the encoding efficiency. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, the quality and efficiency of the content can be improved while reducing the processing load in each terminal. In this case, if there is a request for almost real-time reception and decoding, the data completed by the first encoding by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be achieved.
[0908] As another example, cameras like the ex113 extract features from images, compress the feature data as metadata, and transmit it to a server. The server, for example, determines the importance of an object based on the features and switches the quantization precision, performing compression tailored to the image's meaning (or content). Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during recompression in the server. Alternatively, the terminal can perform simple encoding such as VLC (Variable Length Coding), while the server performs more processing-intensive encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding).
[0909] As another example, in stadiums, shopping malls, factories, and other locations, there may be multiple video data sets generated by capturing roughly the same scene using multiple terminals. In such cases, encoding can be distributed using the multiple terminals that captured the images, as well as other terminals and servers that di...
Claims
1. An encoding device, wherein: have: circuits; and a memory connected to the circuit, The circuit is in action, HRD parameters related to the decoding unit and related to the HRD, which is the assumed reference decoder, and SEI, which is the supplemental enhancement information, are encoded independently of the sequence parameter set into a buffering period SEI message, and Other HRD parameters are encoded into the decoding unit information SEI message independently of the sequence parameter set and relying only on the buffering duration SEI.
2. A decoding device, wherein: have: circuits; and a memory connected to the circuit, The circuit is in action, decoding HRD parameters related to the decoding unit and related to the HRD, which is the assumed reference decoder, and SEI, which is the supplemental enhancement information, from SEI messages during buffering independently of the sequence parameter set, and Other HRD parameters are decoded from the decoding unit information SEI message independently of the sequence parameter set and relying only on the buffering duration SEI.
3. A coding method, wherein: HRD parameters related to the decoding unit and related to the HRD, which is the assumed reference decoder, and SEI, which is the supplemental enhancement information, are encoded independently of the sequence parameter set into a buffering period SEI message, and Other HRD parameters are encoded into the decoding unit information SEI message independently of the sequence parameter set and relying only on the buffering duration SEI.
4. A decoding method, wherein: decoding HRD parameters related to the decoding unit and related to the HRD, which is the assumed reference decoder, and SEI, which is the supplemental enhancement information, from SEI messages during buffering independently of the sequence parameter set, and Other HRD parameters are decoded from the decoding unit information SEI message independently of the sequence parameter set and relying only on the buffering duration SEI.
5. A non-transitory computer-readable storage medium storing a bit stream generated by an encoding device, wherein: The bitstream includes HRD parameters, which are related to the decoding unit and the HRD, and are encoded in the buffering period SEI message independently of the sequence parameter set. The bitstream also includes other HRD parameters, which are independent of the sequence parameter set and encoded in the decoding unit information SEI message only depending on the buffering period SEI. The HRD parameters and the other HRD parameters are provided to the decoding device for decoding processing, wherein the HRD is a hypothetical reference decoder and the SEI is supplementary enhancement information.