Coding device, decoding device, coding method, and decoding method
By using BDOF processing and intra prediction image mixing technology in video encoding, the final predicted image is generated, which solves the requirements for improving encoding efficiency and image quality in the prior art, simplifies the processing volume and circuit scale, and improves the processing speed.
Patent Information
- Application Number
- CN202080012103.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-08
- Filing Date
- 2020-02-06
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-02-06
AI Technical Summary
The existing video encoding technology has a need for improvement in encoding efficiency, picture quality, processing volume, circuit scale and processing speed, especially in the selection of appropriate filters, block sizes, motion vectors and reference pictures and other elements or actions.
The first predicted image of the processing target block is generated using the BDOF process in the inter prediction mode, and the process of mixing the second predicted image and the first predicted image in the intra prediction is used as an update process, and these two processes are exclusively applied to generate the final predicted image.
Improve coding efficiency, simplify encoding/decoding processing, reduce processing time and circuit scale, enhance processing speed, and appropriately select key elements in encoding and decoding, improving picture quality.
Smart Images

Figure CN113366852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video coding, and for example, to systems, components, and methods for encoding and decoding moving images. Background Art
[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). This advancement has led to a continuous need for improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data used in various applications.
[0003] Furthermore, Non-Patent Document 1 relates to an example of existing standards related to the above-mentioned video encoding technology.
[0004] Prior art literature
[0005] Non-patent literature
[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention
[0007] Problems to be solved by the invention
[0008] Regarding the coding methods described above, it is expected that new methods will be proposed to improve coding efficiency, improve image quality, reduce processing volume, reduce circuit scale, or appropriately select elements or actions such as filters, blocks, sizes, motion vectors, reference pictures or reference blocks.
[0009] The present invention provides a structure or method that can contribute to one or more of the following: improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, the present invention may include structures or methods that can contribute to benefits other than those described above.
[0010] Means for solving problems
[0011] For example, an encoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit. When the circuit is in operation, in an inter-frame prediction mode, it generates a first prediction image of a processing object block based on a derived motion vector, and generates a final prediction image of the processing object block by applying an update process to the first prediction image. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process. The second process is a process of mixing the second prediction image generated in the intra-frame prediction of the processing object block with the first prediction image. In the application of the update process, the first process and the second process are exclusively applied.
[0012] The implementation of several embodiments of the present invention can improve coding efficiency, simplify coding / decoding processing, speed up coding / decoding processing, and efficiently select appropriate components / actions used in coding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.
[0013] The present invention provides further advantages and effects according to the present invention. These advantages and / or effects are achieved through several embodiments and features described in the present invention and the accompanying drawings, but all advantages and / or effects do not necessarily need to be provided in order to achieve one or more advantages and / or effects.
[0014] Furthermore, these general or specific aspects may also be implemented as a system, a method, an integrated circuit, a computer program, a recording medium, or any combination thereof.
[0015] Effects of the Invention
[0016] The structure or method of one aspect of the present invention can contribute to one or more of, for example, improved coding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. Furthermore, the structure or method of one aspect of the present invention can also contribute to benefits other than those listed above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a block diagram showing the functional structure of the encoding device according to the embodiment.
[0018] Figure 2 This is a flowchart showing an example of the overall encoding process performed by the encoding device.
[0019] Figure 3 This is a conceptual diagram showing an example of block division.
[0020] Figure 4AThis is a conceptual diagram showing an example of the structure of a slice.
[0021] Figure 4B This is a conceptual diagram showing an example of a tile structure.
[0022] Figure 5A This is a table showing the transformation basis functions corresponding to various transformation types.
[0023] Figure 5B This is a conceptual diagram showing an example of SVT (Spatially Varying Transform).
[0024] Figure 6A This is a conceptual diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).
[0025] Figure 6B This is a conceptual diagram showing another example of the shape of the filter used in ALF.
[0026] Figure 6C This is a conceptual diagram showing another example of the shape of the filter used in ALF.
[0027] Figure 7 This is a block diagram showing an example of a detailed configuration of a loop filter unit functioning as a DBF (deblocking filter).
[0028] Figure 8 This is a conceptual diagram showing an example of deblocking filtering having filter characteristics that are symmetric with respect to block boundaries.
[0029] Figure 9 This is a conceptual diagram for explaining the block boundary on which the deblocking filtering process is performed.
[0030] Figure 10 This is a conceptual diagram showing an example of the Bs value.
[0031] Figure 11 This is a flowchart showing an example of processing performed by the prediction processing unit of the encoding device.
[0032] Figure 12 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device.
[0033] Figure 13 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device.
[0034] Figure 14 This is a conceptual diagram showing an example of 67 intra prediction modes in the intra prediction according to the embodiment.
[0035] Figure 15 This is a flowchart showing an example of the flow of basic processing of inter-frame prediction.
[0036] Figure 16 This is a flowchart showing an example of motion vector derivation.
[0037] Figure 17 This is a flowchart showing another example of motion vector derivation.
[0038] Figure 18 This is a flowchart showing another example of motion vector derivation.
[0039] Figure 19 This is a flowchart showing an example of inter prediction based on the normal inter mode.
[0040] Figure 20 is a flowchart illustrating an example of inter-frame prediction based on merge mode.
[0041] Figure 21 This is a conceptual diagram for explaining an example of motion vector derivation processing based on the merge mode.
[0042] Figure 22 This is a flowchart showing an example of FRUC (frame rate up conversion) processing.
[0043] Figure 23 This is a conceptual diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0044] Figure 24 This is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0045] Figure 25A This is a conceptual diagram for explaining an example of derivation of a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks.
[0046] Figure 25B This is a conceptual diagram for explaining an example of derivation of a motion vector in sub-block units in an affine mode having three control points.
[0047] Figure 26A This is a conceptual diagram used to illustrate the affine merge mode.
[0048] Figure 26B This is a conceptual diagram for explaining the affine merge mode with two control points.
[0049] Figure 26C This is a conceptual diagram for explaining the affine merge mode with three control points.
[0050] Figure 27 This is a flowchart showing an example of processing in the affine merge mode.
[0051] Figure 28A This is a conceptual diagram for explaining the affine inter mode with two control points.
[0052] Figure 28B This is a conceptual diagram for explaining the affine inter-frame mode with three control points.
[0053] Figure 29 This is a flowchart showing an example of processing in the affine inter mode.
[0054] Figure 30A This is a conceptual diagram used to illustrate the affine inter-frame mode in which the current block has 3 control points and the adjacent block has 2 control points.
[0055] Figure 30B This is a conceptual diagram used to illustrate the affine inter-frame mode in which the current block has 2 control points and the adjacent block has 3 control points.
[0056] Figure 31A This is a flowchart showing the merge mode including DMVR (decoder motion vector refinement).
[0057] Figure 31B This is a conceptual diagram for explaining an example of DMVR processing.
[0058] Figure 32 This is a flowchart showing an example of generating a predicted image.
[0059] Figure 33 This is a flowchart showing another example of generating a predicted image.
[0060] Figure 34 This is a flowchart showing another example of generating a predicted image.
[0061] Figure 35 This is a flowchart for explaining an example of a predicted image correction process based on an OBMC (overlapped block motion compensation) process.
[0062] Figure 36 This is a conceptual diagram for explaining an example of predicted image correction processing based on OBMC processing.
[0063] Figure 37 This is a conceptual diagram for explaining the generation of predicted images of two triangles.
[0064] Figure 38This is a conceptual diagram for explaining a model assuming constant velocity linear motion.
[0065] Figure 39 This is a conceptual diagram for explaining an example of a method for generating a predicted image using a brightness correction process based on LIC (local illumination compensation) processing.
[0066] Figure 40 This is a block diagram showing an example of installing an encoding device.
[0067] Figure 41 This is a block diagram showing the functional structure of a decoding device according to an embodiment.
[0068] Figure 42 This is a flowchart showing an example of the overall decoding process performed by the decoding device.
[0069] Figure 43 This is a flowchart showing an example of processing performed by the prediction processing unit of the decoding device.
[0070] Figure 44 This is a flowchart showing another example of processing performed by the prediction processing unit of the decoding device.
[0071] Figure 45 This is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in a decoding device.
[0072] Figure 46 This is a block diagram showing an implementation example of a decoding device.
[0073] Figure 47 This is a diagram for explaining the block size when performing prediction processing.
[0074] Figure 48 This is a diagram schematically showing a first example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0075] Figure 49 This is a flowchart showing a second example of prediction processing performed by the decoding device in the second embodiment.
[0076] Figure 50 This is a diagram schematically showing a second example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0077] Figure 51 This is a flowchart showing a third example of prediction processing performed by the decoding device in the second embodiment.
[0078] Figure 52This is a diagram schematically showing a third example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0079] Figure 53 This is a flowchart showing a fourth example of prediction processing performed by the decoding device in the second embodiment.
[0080] Figure 54 This is a diagram schematically showing a fourth example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0081] Figure 55 This is a flowchart showing a fifth example of prediction processing performed by the decoding device in the second embodiment.
[0082] Figure 56 This is a diagram schematically showing a fifth example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0083] Figure 57 This is a flowchart showing a sixth example of prediction processing performed by the decoding device in the second embodiment.
[0084] Figure 58 This is a diagram schematically showing a sixth example of prediction processing performed by the decoding device in the second embodiment as a configuration example of pipeline processing.
[0085] Figure 59 This is a flowchart showing the seventh example of prediction processing performed by the decoding device in the second embodiment.
[0086] Figure 60 This is a diagram schematically showing the seventh example of prediction processing performed by the decoding device in the second embodiment as a structural example of pipeline processing.
[0087] Figure 61 This is a flowchart showing an example of encoding and decoding processing in Implementation 2.
[0088] Figure 62 This is a flowchart showing another example of encoding and decoding processing in Implementation 2.
[0089] Figure 63 This is a block diagram showing the overall structure of a content provision system that implements content distribution services.
[0090] Figure 64 This is a conceptual diagram showing an example of a coding structure in the case of scalable coding.
[0091] Figure 65 This is a conceptual diagram showing an example of a coding structure in the case of scalable coding.
[0092] Figure 66 This is a conceptual diagram showing an example of a display screen of a web page.
[0093] Figure 67 This is a conceptual diagram showing an example of a display screen of a web page.
[0094] Figure 68 This is a block diagram showing an example of a smart phone.
[0095] Figure 69 This is a block diagram showing a configuration example of a smartphone. DETAILED DESCRIPTION
[0096] An encoding device according to one embodiment of the present disclosure includes a circuit and a memory connected to the circuit. When the circuit is in operation, in an inter-frame prediction mode, it generates a first predicted image of a processing target block based on a derived motion vector, and generates a final predicted image of the processing target block by applying an update process to the first predicted image. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process. The second process is a process of mixing the second predicted image generated in the intra-frame prediction of the processing target block with the first predicted image. When applying the update process, the first process and the second process are exclusively applied.
[0097] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first and second processes is applied to the first predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for encoding an image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0098] In addition, it may also be that the above-mentioned circuit determines whether the above-mentioned second processing is applied to the above-mentioned first predicted image in the application of the above-mentioned update processing, and when it is determined that the above-mentioned second processing is applied, the above-mentioned second processing is set as the above-mentioned update processing; when it is determined that the above-mentioned second processing is not applied, the above-mentioned circuit determines whether the above-mentioned first processing is applied to the above-mentioned first predicted image, and when it is determined that the above-mentioned first processing is applied, the above-mentioned first processing is set as the above-mentioned update processing.
[0099] This increases the possibility that the update process can be appropriately selected.
[0100] Alternatively, the circuit may exclusively perform the first processing and the second processing in one stage, wherein the one stage is included in the pipeline processing for encoding an image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final predicted image and the residual image.
[0101] This reduces the number of stages included in pipeline processing compared to, for example, performing the first and second processes at a stage separate from the reconstruction process, thereby increasing the likelihood of suppressing an increase in the circuit scale of the encoding device.
[0102] In addition, a decoding device according to one embodiment of the present invention includes a circuit and a memory connected to the above circuit. When the above circuit is in operation, in the inter-frame prediction mode, it generates a first prediction image of the processing object block based on the derived motion vector, and generates a final prediction image of the processing object block by applying an update process to the above first prediction image. Candidates for the above update process include a first process and a second process. The above first process is a BDOF (bi-directional optical flow) process. The above second process is a process of mixing the second prediction image generated in the intra-frame prediction of the processing object block with the above first prediction image. In the application of the above update process, the above first process and the above second process are exclusively applied.
[0103] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first and second processes is applied to the first predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for decoding a coded image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0104] In addition, it may also be that the above-mentioned circuit determines whether the above-mentioned second processing is applied to the above-mentioned first predicted image in the application of the above-mentioned update processing, and when it is determined that the above-mentioned second processing is applied, the above-mentioned second processing is set as the above-mentioned update processing; when it is determined that the above-mentioned second processing is not applied, the above-mentioned circuit determines whether the above-mentioned first processing is applied to the above-mentioned first predicted image, and when it is determined that the above-mentioned first processing is applied, the above-mentioned first processing is set as the above-mentioned update processing.
[0105] This increases the possibility that the update process can be appropriately selected.
[0106] Alternatively, the circuit may exclusively perform the first processing and the second processing in one stage, wherein the one stage is included in the pipeline processing for decoding an image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final predicted image and the residual image.
[0107] This reduces the number of stages included in pipeline processing compared to, for example, performing the first and second processes at a stage separate from the reconstruction process, thereby increasing the likelihood of suppressing an increase in the circuit scale of the decoding device.
[0108] Furthermore, for example, an encoding device according to one embodiment of the present disclosure includes a partitioning unit, an intra-frame prediction unit, an inter-frame prediction unit, a prediction control unit, a transform unit, a quantization unit, and an entropy encoding unit.
[0109] The partitioning unit partitions a current picture constituting a moving image into a plurality of blocks. The intra-frame prediction unit performs intra-frame prediction to generate an intra-frame prediction image for a processing target block in the current picture using a reference image in the current picture. The inter-frame prediction unit performs inter-frame prediction to generate an inter-frame prediction image for the processing target block using a reference image in a reference picture different from the current picture.
[0110] The prediction control unit controls intra-frame prediction by the intra-frame prediction unit and inter-frame prediction by the inter-frame prediction unit. The transform unit transforms a residual image between a prediction image composed of at least one of the intra-frame prediction image and the inter-frame prediction image and the image of the processing target block to generate a transform coefficient signal for the processing target block. The quantization unit quantizes the transform coefficient signal. The entropy coding unit encodes the quantized transform coefficient signal.
[0111] Here, the inter-frame prediction unit generates a first predicted image for the processing block based on the derived motion vector in inter-frame prediction mode, and applies an update process to the first predicted image to generate a final predicted image for the processing block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a process of blending the second predicted image generated by intra-frame prediction of the processing block with the first predicted image. Furthermore, when applying the update process, the first and second processes are exclusively applied.
[0112] In addition, for example, a decoding device of one embodiment of the present invention is a decoding device that uses a predicted image to decode a moving image, and includes an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra-frame prediction unit, an inter-frame prediction unit, a prediction control unit, and an addition unit (reconstruction unit).
[0113] The entropy decoding unit decodes a quantized transform coefficient signal of a processing target block in a decoding target picture constituting the moving image. The inverse quantization unit inversely quantizes the quantized transform coefficient signal. The inverse transform unit inversely transforms the transform coefficient signal to obtain a residual image of the processing target block.
[0114] The intra prediction unit performs intra prediction to generate an intra prediction image for the processing target block using a reference image in the decoding target picture. The inter prediction unit performs inter prediction to generate an inter prediction image for the processing target block using a reference image in a reference picture different from the decoding target picture. The prediction control unit controls the intra prediction performed by the intra prediction unit and the inter prediction performed by the inter prediction unit.
[0115] The adding unit adds a predicted image formed of at least one of the intra predicted image and the inter predicted image and the residual image to reconstruct an image of the processing target block.
[0116] Here, the inter-frame prediction unit generates a first predicted image for the processing block based on the derived motion vector in inter-frame prediction mode, and applies an update process to the first predicted image to generate a final predicted image for the processing block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a process of blending the second predicted image generated by intra-frame prediction of the processing block with the first predicted image. Furthermore, when applying the update process, the first and second processes are exclusively applied.
[0117] Moreover, these inclusive or specific forms may be implemented by systems, devices, methods, integrated circuits, computer programs, or non-transitory recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0118] The following embodiments are described in detail with reference to the accompanying drawings. The embodiments described below are intended to be inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connections of components, steps, and the relationships and order of steps shown in the following embodiments are merely examples and are not intended to limit the claims.
[0119] The following describes embodiments of encoding and decoding devices. The embodiments are examples of encoding and decoding devices to which the processing and / or structures described in the various aspects of the present invention can be applied. The processing and / or structures can also be implemented in encoding and decoding devices that differ from the embodiments. For example, the processing and / or structures applied to the embodiments may include any of the following.
[0120] (1) Any of the multiple components of the encoding device or decoding device described in the embodiments of the present invention may be replaced by another component described in any of the embodiments of the present invention, or these components may be combined.
[0121] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be modified by adding, replacing, deleting, or other arbitrary changes. For example, any function or process may be replaced by another function or process described in any of the various aspects of the present invention, or these functions or processes may be combined.
[0122] (3) In the method implemented by the encoding device or decoding device of the embodiment, any changes such as addition, replacement, or deletion may be made to a portion of the multiple processes included in the method. For example, any process in the method may be replaced with another process described in one of the various aspects of the present invention, or these processes may be combined.
[0123] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with a component described in any of the aspects of the present invention, may be combined with a component having a portion of the functions described in any of the aspects of the present invention, or may be combined with a component that performs a portion of the processing performed by a component described in any of the aspects of the present invention.
[0124] (5) A component having a portion of the functions of the encoding device or decoding device of the embodiment, or a component implementing a portion of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced with a component described in any of the aspects of the present invention, a component having a portion of the functions described in any of the aspects of the present invention, or a component implementing a portion of the processing described in any of the aspects of the present invention;
[0125] (6) In a method implemented by an encoding device or decoding device according to an embodiment, one of the multiple processes included in the method is replaced by one of the processes described in each aspect of the present invention or a similar process, or a combination of these processes;
[0126] (7) Some of the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the process described in any of the aspects of the present invention.
[0127] (8) The implementation of the processes and / or structures described in the various aspects of the present invention is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or structures may be implemented in a device used for a purpose different from that of the video encoding or decoding disclosed in the embodiments.
[0128] (Implementation Method 1)
[0129] [Encoding device]
[0130] First, the encoding device according to the embodiment will be described. Figure 1 1 is a block diagram showing the functional structure of the encoding device 100 according to the embodiment. The encoding device 100 is a moving picture encoding device that encodes a moving picture in units of blocks.
[0131] like Figure 1 As shown, the encoding device 100 is a device that encodes an image in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126 and a prediction control unit 128.
[0132] The encoding device 100 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes a software program stored in the memory, the processor functions as the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.
[0133] Below, after describing the overall processing flow of the encoding device 100 , each component included in the encoding device 100 will be described.
[0134] [Overall flow of encoding processing]
[0135] Figure 2 3 is a flowchart showing an example of the overall encoding process performed by the encoding device 100 .
[0136] First, the segmentation unit 102 of the encoding device 100 segments each picture included in an input image, which is a moving image, into a plurality of fixed-size blocks (e.g., 128×128 pixels) (step Sa_1). The segmentation unit 102 then selects a segmentation pattern (also called a block shape) for each of the fixed-size blocks (step Sa_2). Specifically, the segmentation unit 102 further segments the fixed-size blocks into a plurality of blocks that conform to the selected segmentation pattern. The encoding device 100 then performs steps Sa_3 through Sa_9 on each of the plurality of blocks (i.e., the encoding target block).
[0137] That is, the prediction processing unit composed of all or part of the intra-frame prediction unit 124, the inter-frame prediction unit 126 and the prediction control unit 128 generates a prediction signal (also called a prediction block) of the encoding target block (also called the current block) (step Sa_3).
[0138] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).
[0139] Next, the transform unit 106 and the quantization unit 108 transform and quantize the difference block to generate a plurality of quantized coefficients (step Sa_5). In addition, a block composed of a plurality of quantized coefficients is also called a coefficient block.
[0140] Next, the entropy coding unit 110 encodes the coefficient block and the prediction parameters related to the generation of the prediction signal (specifically, entropy coding) to generate a coded signal (step Sa_6). In addition, the coded signal is also called a coded bit stream, a compressed bit stream, or a stream.
[0141] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore a plurality of prediction residuals (ie, difference blocks) by performing inverse quantization and inverse transformation on the coefficient block (step Sa_7).
[0142] Next, the adding unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the prediction block to the restored difference block (step Sa_8).
[0143] When the reconstructed image is generated, the loop filter unit 120 filters the reconstructed image as necessary (step Sa_9).
[0144] Then, the encoding device 100 determines whether encoding of the entire picture is completed (step Sa_10 ), and if it is determined that encoding is not completed (No in step Sa_10 ), it repeats the process from step Sa_2 .
[0145] Furthermore, in the above example, encoding device 100 selects a single partitioning pattern for fixed-size blocks and encodes each block according to that partitioning pattern. However, encoding device 100 may also encode each block according to each of a plurality of partitioning patterns. In this case, encoding device 100 may evaluate the cost for each of the plurality of partitioning patterns and, for example, select as the output encoded signal the encoded signal obtained by encoding according to the partitioning pattern with the lowest cost.
[0146] As shown in the figure, the processes of steps Sa_1 to Sa_10 are sequentially performed by the encoding device 100. Alternatively, a part of a plurality of these processes may be performed in parallel, or the order of these processes may be reversed.
[0147] [Division]
[0148] The segmentation unit 102 segments each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the picture into blocks of a fixed size (e.g., 128×128). Other fixed block sizes may also be used. The fixed-size blocks are sometimes called coding tree units (CTUs). Furthermore, the segmentation unit 102 segments each fixed-size block into blocks of a variable size (e.g., 64×64 or less), for example, based on recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. The variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be used as processing units of CUs, PUs, or TUs.
[0149] Figure 3 This is a conceptual diagram showing an example of block division in an implementation manner. Figure 3 In the figure, the solid line represents the block boundary based on quadtree block partitioning, and the dotted line represents the block boundary based on binary tree block partitioning.
[0150] Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into four square 64×64 blocks (quadtree block division).
[0151] The 64×64 block on the upper left is further vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is further vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the 64×64 block on the upper left is divided into two 16×64 blocks 11 and 12 and a 32×64 block 13.
[0152] The 64×64 block in the upper right corner is horizontally partitioned into two rectangular 64×32 blocks 14 and 15 (binary tree block partitioning).
[0153] The 64×64 block on the lower left is divided into four square 32×32 blocks (quadtree block partitioning). The upper left and lower right blocks of the four 32×32 blocks are further partitioned. The upper left 32×32 block is vertically partitioned into two rectangular 16×32 blocks, and the right 16×32 block is further partitioned horizontally into two 16×16 blocks (binary tree block partitioning). The lower right 32×32 block is horizontally partitioned into two 32×16 blocks (binary tree block partitioning). As a result, the lower left 64×64 block is partitioned into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.
[0154] The lower right 64×64 block 23 is not split.
[0155] As above, in Figure 3 In FIG, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. This type of partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.
[0156] In addition, Figure 3 In the example above, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). Division including this ternary tree division is called MBT (multi type tree) division.
[0157] [Structural slices / tiles of the image]
[0158] In order to decode pictures in parallel, pictures may be constructed in slice units or tile units. The picture constructed in slice units or tile units can be constructed by the partitioning unit 102.
[0159] A slice is a basic coding unit that constitutes a picture. A picture is composed of one or more slices. In addition, a slice is composed of one or more consecutive CTUs (Coding Tree Units).
[0160] Figure 4A This is a conceptual diagram showing an example of the structure of a slice. For example, a picture includes 11×8 CTUs and is divided into 4 slices (slices 1 to 4). Slice 1 consists of 16 CTUs, slice 2 consists of 21 CTUs, slice 3 consists of 29 CTUs, and slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any slice. The shape of the slice becomes the shape that divides the picture in the horizontal direction. The boundary of the slice does not need to be the end of the picture, and can be any position in the boundary of the CTU in the picture. The processing order (encoding order or decoding order) of the CTU in the slice is, for example, a raster scan order. In addition, the slice includes header information and encoded data. The header information may also record the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.
[0161] A tile is a unit of rectangular area that constitutes a picture. Each tile may be assigned a number called a TileId in raster scan order.
[0162] Figure 4B This is a conceptual diagram showing an example of a tile structure. For example, a picture includes 11×8 CTUs and is divided into four rectangular area tiles (tiles 1 to 4). When tiles are used, the processing order of CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed in raster scan order. When tiles are used, at least one CTU in each of the multiple tiles is processed in raster scan order. For example, Figure 4B As shown, the processing order of multiple CTUs included in tile 1 is from the left end of the 1st row of tile 1 to the right end of the 1st row of tile 1, and then from the left end of the 2nd row of tile 1 to the right end of the 2nd row of tile 1.
[0163] In addition, one tile may include more than one slice, and one slice may include more than one tile.
[0164] [Subtraction Department]
[0165] The subtraction unit 104 subtracts the prediction signal (prediction samples input from the prediction control unit 128, described below) from the original signal (original samples) in units of blocks input from the segmentation unit 102 and segmented by the segmentation unit 102. Specifically, the subtraction unit 104 calculates a prediction error (also referred to as a residual) for the current block to be coded (hereinafter referred to as the current block). The subtraction unit 104 then outputs the calculated prediction error (residual) to the transformation unit 106.
[0166] The original signal is an input signal to the encoding device 100 and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0167] [Conversion Unit]
[0168] The transform unit 106 transforms the spatial domain prediction error into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction error. The predetermined DCT or DST may also be predetermined.
[0169] Alternatively, the transform unit 106 may adaptively select a transform type from a plurality of transform types and transform the prediction error into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes referred to as an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).
[0170] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transformation basis functions corresponding to the transformation type examples. Figure 5A Where N represents the number of input pixels. The selection of a transform type from among these multiple transform types may depend on, for example, the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode.
[0171] Information indicating whether such EMT or AMT is applied (e.g., an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. However, signaling of this information is not limited to the CU level and may be performed at other levels (e.g., bitstream level, picture level, slice level, tile level, or CTU level).
[0172] In addition, the transform unit 106 may also re-transform the transform coefficients (transform results). Such re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 re-transforms each sub-block (for example, a 4×4 sub-block) contained in the block of transform coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but can also be other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0173] Separable transformation and non-separable transformation can also be applied in the transformation unit 106. Separable transformation refers to a method of performing multiple transformations in each direction corresponding to the number of dimensions of the input. Non-separable transformation refers to a method of treating two or more dimensions as one dimension when the input is multi-dimensional and transforming them together.
[0174] For example, as an example of non-separable transformation, when a 4×4 block is input, it is considered as an array of 16 elements, and the array is transformed using a 16×16 transformation matrix.
[0175] In a further example of the non-separable transformation, a 4×4 input block may be regarded as an array of 16 elements, and then a transformation (Hypercube Givens Transform) may be performed by performing multiple Givens rotations on the array.
[0176] In the transformation in the transformation unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, Figure 5BAs shown, the CU is divided into two equal parts in the horizontal or vertical direction, and only the area on one side is transformed into the frequency area. The type of transformation base can be set for each area, for example, DST7 and DCT8. In this example, only one of the two areas in the CU is transformed, and the other is not transformed, but both areas can also be transformed. In addition, the division method is not limited to bisection, but can be more flexible, such as quartering or encoding the information indicating the division separately, and signaling it in the same way as the CU division. In addition, SVT is sometimes also called SBT (Sub-block Transform).
[0177] [Quantitative Department]
[0178] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112. The predetermined scanning order may also be predetermined.
[0179] The predetermined scanning order is the order used for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order may be defined in ascending order (from low frequency to high frequency) or descending order (from high frequency to low frequency) of frequency.
[0180] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.
[0181] Quantization also sometimes uses a quantization matrix. For example, multiple quantization matrices are sometimes used to correspond to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. Furthermore, quantization refers to digitizing values sampled at specified intervals and assigning them to specified levels. In this technical field, other representations such as rounding, integer scaling, and scaling may also be used. The specified intervals and levels may also be predetermined.
[0182] Methods for using quantization matrices include using a quantization matrix directly set on the encoding device side and using a default quantization matrix (default matrix). By directly setting the quantization matrix on the encoding device side, a quantization matrix suitable for image characteristics can be set. However, this method has the disadvantage of increasing the amount of code due to encoding the quantization matrix.
[0183] On the other hand, there is also a method of performing quantization so that the coefficients of high-frequency components and low-frequency components are the same without using a quantization matrix. This method is equivalent to using a quantization matrix (flat matrix) in which all coefficients have the same value.
[0184] The quantization matrix can be specified by, for example, an SPS (Sequence Parameter Set) or a PPS (Picture Parameter Set). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and PPS are sometimes referred to simply as parameter sets.
[0185] [Entropy coding unit]
[0186] The entropy coding unit 110 generates a coded signal (coded bit stream) based on the quantization coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110 binarizes the quantization coefficients, performs arithmetic coding on the binary signal, and outputs a compressed bit stream or sequence.
[0187] [Inverse quantization unit]
[0188] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. Furthermore, the inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114. The predetermined scanning order may also be predetermined.
[0189] [Inverse transformation unit]
[0190] The inverse transform unit 114 restores the prediction error (residual) by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform performed by the transform unit 106 on the transform coefficients. The inverse transform unit 114 then outputs the restored prediction error to the addition unit 116.
[0191] Furthermore, the restored prediction error generally loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error generally includes a quantization error.
[0192] [Addition Department]
[0193] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a locally decoded block.
[0194] [Block Memory]
[0195] The block memory 118 is a storage unit for storing blocks in a current picture to be coded (referred to as a current picture) to be referenced in intra prediction, for example. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116 .
[0196] [Frame Memory]
[0197] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120 .
[0198] [Loop filter unit]
[0199] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the encoding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).
[0200] In ALF, a least squares error filter is used to remove coding distortion. For example, for each 2×2 sub-block in the current block, one filter is selected from multiple filters based on the direction and activity of the local gradient.
[0201] Specifically, sub-blocks (e.g., 2×2 sub-blocks) are first classified into multiple classes (e.g., 15 or 25 classes). Sub-block classification is performed based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4), a classification value C is calculated (e.g., C = 5D + A). Based on the classification value C, the sub-blocks are then classified into multiple classes.
[0202] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.
[0203] Based on the result of such classification, a filter to be used for the sub-block is determined from among a plurality of filters.
[0204] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 6A to 6C FIG. 1 is a diagram showing a plurality of examples of filter shapes used in ALF. Figure 6Arepresents a 5×5 diamond-shaped filter, Figure 6B represents a 7×7 diamond-shaped filter, Figure 6C Represents a 9×9 diamond-shaped filter. Information representing the filter shape is typically signaled at the picture level. However, signaling of the filter shape information need not be limited to the picture level and can also be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0205] The on / off of ALF can also be determined at the picture level or the CU level. For example, for luminance, whether to use ALF can be determined at the CU level, and for chrominance, whether to use ALF can be determined at the picture level. The information indicating whether ALF is on / off is usually signaled at the picture level or the CU level. In addition, the signaling of the information indicating whether ALF is on / off does not need to be limited to the picture level or the CU level, and can also be at other levels (for example, the sequence level, the slice level, the tile level, or the CTU level).
[0206] The coefficient sets for multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level. However, the signaling of coefficient sets is not limited to the picture level and can also be performed at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0207] [Loop Filter Section > Deblocking Filter]
[0208] In the deblocking filter, the loop filter unit 120 performs filtering processing on block boundaries of the reconstructed image to reduce distortion generated at the block boundaries.
[0209] Figure 7 This is a block diagram showing an example of a detailed configuration of the loop filter unit 120 functioning as a deblocking filter.
[0210] The loop filter unit 120 includes a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 , and 1206 .
[0211] The boundary determination unit 1201 determines whether or not there is a pixel (ie, a target pixel) to be subjected to deblocking filtering near a block boundary, and outputs the determination result to the switch 1202 and the processing determination unit 1208 .
[0212] When the boundary determination unit 1201 determines that the target pixel exists near a block boundary, the switch 1202 outputs the image before filtering to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel does not exist near a block boundary, the switch 1202 outputs the image before filtering to the switch 1206.
[0213] The filter determination unit 1203 determines whether to perform deblocking filtering on the target pixel based on the pixel value of at least one surrounding pixel located around the target pixel, and then outputs the determination result to the switch 1204 and the processing determination unit 1208 .
[0214] When the filter determination unit 1203 determines that the deblocking filtering process is to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that the deblocking filtering process is not to be performed on the target pixel, the switch 1204 outputs the pre-filtering image obtained via the switch 1202 to the switch 1206.
[0215] When the pre-filtered image is obtained via switches 1202 and 1204 , the filter processing unit 1205 performs deblocking filtering on the target pixel using the filter characteristics determined by the filter characteristic determination unit 1207 . The filter processing unit 1205 then outputs the filtered pixel to the switch 1206 .
[0216] According to the control of the processing determination unit 1208 , the switch 1206 selectively outputs pixels that have not been processed by the deblocking filter and pixels that have been processed by the deblocking filter by the filter processing unit 1205 .
[0217] The processing determination unit 1208 controls the switch 1206 based on the determination results of the boundary determination unit 1201 and the filter determination unit 1203. Specifically, when the boundary determination unit 1201 determines that the target pixel is located near a block boundary and the filter determination unit 1203 determines that the target pixel is to be subjected to deblocking filtering, the processing determination unit 1208 outputs the pixel subjected to deblocking filtering from the switch 1206. Otherwise, in other cases, the processing determination unit 1208 outputs the pixel that has not been subjected to deblocking / filtering from the switch 1206. By repeating this pixel outputting process, the filtered image is output from the switch 1206.
[0218] Figure 8 This is a conceptual diagram showing an example of a deblocking filter having a filter characteristic that is symmetric with respect to a block boundary.
[0219] In the deblocking filter process, for example, using pixel values and quantization parameters, one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter, is selected. Figure 8 As shown, when pixels p0 to p2 and pixels q0 to q2 exist across a block boundary, the pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the calculation shown in the following equation, for example.
[0220] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0221] q'1=(p0+q0+q1+q2+2) / 4
[0222] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0223] In the above equations, p0 through p2 and q0 through q2 are the pixel values of pixels p0 through p2 and q0 through q2, respectively. Furthermore, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the side opposite the block boundary. On the right side of each of the above equations, the coefficients multiplied by the pixel values of each pixel used in the deblocking filter process are the filter coefficients.
[0224] Furthermore, during deblocking filtering, clipping can be performed to ensure that the calculated pixel value does not exceed a threshold. In this clipping process, the pixel value calculated based on the above equation is clipped to "the calculated pixel value ± 2 × the threshold" using a threshold determined by the quantization parameter. This prevents excessive smoothing.
[0225] Figure 9 This is a conceptual diagram for explaining the block boundary on which the deblocking filtering process is performed. Figure 10 This is a conceptual diagram showing an example of the Bs value.
[0226] The block boundary for deblocking filtering is, for example, Figure 9 The deblocking filter can be performed in units of 4 rows or 4 columns. First, for Figure 9 The block P and block Q shown are as follows: Figure 10 That determines the Bs (Boundary Strength) value.
[0227] according to Figure 10The Bs value determines whether to perform deblocking filtering with different intensities even at block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering is performed on the color difference signal. When the Bs value is 1 or more and the specified conditions are met, deblocking filtering is performed on the luminance signal. The specified conditions can also be predetermined. In addition, the determination conditions of the Bs value are not limited to Figure 10 The conditions shown can also be determined based on other parameters.
[0228] [Prediction Processing Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]
[0229] Figure 11 12 is a flowchart showing an example of processing performed by the prediction processing unit of the encoding device 100. The prediction processing unit is composed of all or part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0230] The prediction processing unit generates a predicted image for the current block (step Sb_1). This predicted image is also referred to as a prediction signal or a prediction block. Prediction signals include, for example, intra-frame prediction signals and inter-frame prediction signals. Specifically, the prediction processing unit generates a predicted image for the current block using the reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.
[0231] The reconstructed image may be, for example, an image of a reference picture, or an image including the current block, that is, an image of an already coded block within the current picture. The already coded block within the current picture may be, for example, an adjacent block of the current block.
[0232] Figure 12 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device 100.
[0233] The prediction processing unit generates a predicted image using the first method (step Sc_1a), generates a predicted image using the second method (step Sc_1b), and generates a predicted image using the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating predicted images, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may also be used in these prediction methods.
[0234] Next, the prediction processing unit selects any one of the multiple prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). The selection of the prediction image, that is, the selection of the method or mode for obtaining the final prediction image can also be performed by calculating the cost for each prediction image generated and based on the cost. In addition, the selection of the prediction image can be performed based on the parameters used for the encoding process. The encoding device 100 can signal information for determining the selected prediction image, method or mode into a coded signal (also called a coded bit stream). The information can be, for example, a flag, etc. Thus, the decoding device can generate a prediction image in accordance with the method or mode selected in the encoding device 100 based on the information. In addition, in Figure 12 In the example shown, the prediction processing unit selects one of the predicted images after generating them using various methods. However, before generating these predicted images, the prediction processing unit may select a method or mode based on the parameters used in the encoding process described above and generate the predicted images based on that method or mode.
[0235] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction processing unit may select a final prediction image for the current block from prediction images generated according to these prediction methods.
[0236] Figure 13 This is a flowchart showing another example of processing performed by the prediction processing unit of the encoding device 100.
[0237] First, the prediction processing unit generates a predicted image through intra-frame prediction (step Sd_1a), and generates a predicted image through inter-frame prediction (step Sd_1b). In addition, the predicted image generated through intra-frame prediction is also called an intra-frame predicted image, and the predicted image generated through inter-frame prediction is also called an inter-frame predicted image.
[0238] Next, the prediction processing unit evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). Cost can also be used in this evaluation. That is, the prediction processing unit calculates the cost C of each of the intra-frame prediction image and the inter-frame prediction image. The cost C can be calculated by an equation of the RD optimization model, such as C=D+λ×R. In this equation, D is the coding distortion of the predicted image, and is represented by, for example, the sum of the absolute values of the differences between the pixel values of the current block and the pixel values of the predicted image. In addition, R is the amount of generated coding of the predicted image, specifically, the amount of coding required for encoding motion information, etc. for generating the predicted image. In addition, λ is, for example, an undetermined multiplier of Lagrange.
[0239] Then, the prediction processing unit selects the prediction image with the minimum cost C from the intra-frame prediction image and the inter-frame prediction image as the final prediction image of the current block (step Sd_3). In other words, the prediction method or mode for generating the prediction image of the current block is selected.
[0240] [Intra-frame prediction unit]
[0241] The intra prediction unit 124 performs intra prediction (also called intra-screen prediction) on the current block by referring to blocks in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra prediction unit 124 generates the intra-prediction signal by performing intra prediction with reference to samples (e.g., luminance values and chrominance values) of blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 128.
[0242] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes. The plurality of predetermined modes may also be predetermined.
[0243] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.
[0244] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. Furthermore, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 14 This is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent the 33 directions specified by the H.265 / HEVC specification, and the dotted arrows represent the additional 32 directions (2 non-directional prediction modes in Figure 14 (not shown in the figure).
[0245] In various processing examples, the intra prediction of a chrominance block may also refer to the luma block. That is, the chrominance component of the current block may be predicted based on the luma component of the current block. This type of intra prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra prediction mode for a chrominance block that refers to a luma block (e.g., CCLM mode) may also be added as one of the intra prediction modes for a chrominance block.
[0246] The intra-frame prediction unit 124 may also modify the intra-predicted pixel values based on the gradients of reference pixels in the horizontal and vertical directions. Intra-frame prediction with such modification is sometimes called PDPC (position-dependent intraprediction combination). Information indicating whether PDPC is used (e.g., a PDPC flag) is typically signaled at the CU level. However, this information is not necessarily signaled at the CU level and may be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0247] [Inter-frame prediction unit]
[0248] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) on the current block by referring to a reference picture stored in the frame memory 122, which is different from the current picture, to generate a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed on the current block or current sub-block within the current block (e.g., a 4×4 block). For example, the inter-frame prediction unit 126 performs motion estimation on the current block or current sub-block within the reference picture to find the reference block or sub-block that most closely matches the current block or sub-block. Furthermore, the inter-frame prediction unit 126 obtains motion information (e.g., a motion vector) to compensate for motion or changes from the reference block or sub-block to the current block or sub-block. Based on this motion information, the inter-frame prediction unit 126 performs motion compensation (or motion prediction) to generate an inter-frame prediction signal for the current block or sub-block. The inter-frame prediction unit 126 then outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0249] The motion information used in motion compensation can be signaled as an inter-frame prediction signal in various forms. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.
[0250] [Basic process of inter-frame prediction]
[0251] Figure 15 This is a flowchart showing an example of the basic flow of inter-frame prediction.
[0252] The inter prediction unit 126 first generates a predicted image (steps Se_1 to Se_3 ). Next, the subtraction unit 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4 ).
[0253] Here, in generating the predicted image, the inter-frame prediction unit 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). In determining the MV, the inter-frame prediction unit 126 determines the MV by selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The candidate MV is selected, for example, by selecting a candidate MV from a candidate MV list.
[0254] In addition, in the derivation of MV, the inter-frame prediction unit 126 can also select at least one candidate MV.
[0255] At least one candidate MV is further selected from the at least one candidate MV, and the selected at least one candidate MV is determined as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by each of the at least one selected candidate MVs. The operation of searching the region of the reference picture may also be referred to as motion estimation.
[0256] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126 . However, for example, the processing of step Se_1 or step Se_2 may be performed by other components included in the encoding device 100 .
[0257] [Flow of Motion Vector Derivation]
[0258] Figure 16 This is a flowchart showing an example of motion vector derivation.
[0259] The inter-frame prediction unit 126 derives the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information is encoded as prediction parameters and signaled. That is, the encoded motion information is included in the encoded signal (also called the encoded bitstream).
[0260] Alternatively, the inter prediction unit 126 derives the MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the encoded signal.
[0261] Here, the MV derivation modes may include the normal inter mode, merge mode, FRUC mode, and affine mode described later. Among these modes, the modes that encode motion information include the normal inter mode, merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Furthermore, motion information may include not only the MV but also the predicted motion vector selection information described later. Furthermore, modes that do not encode motion information include the FRUC mode. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0262] Figure 17 This is a flowchart showing another example of motion vector derivation.
[0263] The inter-frame prediction unit 126 derives the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. In other words, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.
[0264] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode in which the difference MV is not encoded. In this case, the encoded difference MV is not included in the encoded signal.
[0265] As described above, the MV derivation modes include the normal inter mode, merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the normal inter mode and affine mode (specifically, affine inter mode) encode the difference MV. Furthermore, the FRUC mode, merge mode, and affine mode (specifically, affine merge mode) do not encode the difference MV. The inter prediction unit 126 selects a mode from these multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0266] [Flow of Motion Vector Derivation]
[0267] Figure 18This is a flowchart showing another example of motion vector derivation. There are multiple modes for MV derivation, that is, inter-frame prediction modes, which are roughly divided into modes that encode differential MVs and modes that do not encode differential motion vectors. Modes that do not encode differential MVs include merge mode, FRUC mode, and affine mode (specifically, affine merge mode). The details of these modes will be described later. Simply put, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from the surrounding coded blocks, and the FRUC mode is a mode in which the MV of the current block is derived by searching between coded areas. In addition, the affine mode is a mode in which the motion vectors of each of the multiple sub-blocks constituting the current block are derived as the MV of the current block, assuming an affine transformation.
[0268] Specifically, as shown in the figure, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_2) based on the merge mode. In addition, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_3) based on the FRUC mode. In addition, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_4) based on the affine mode (specifically, the affine merge mode). In addition, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_5) based on the mode for encoding the differential MV (for example, the normal inter-frame mode).
[0269] [MV Export > Normal Interframe Mode]
[0270] Normal inter mode is an inter prediction mode that derives the MV of the current block from a block similar to the image of the current block in an area of a reference picture indicated by a candidate MV. In this normal inter mode, a differential MV is encoded.
[0271] Figure 19 This is a flowchart showing an example of inter prediction based on the normal inter mode.
[0272] First, the inter prediction unit 126 obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter prediction unit 126 creates a candidate MV list.
[0273] Next, the inter-frame prediction unit 126 extracts N (N is an integer greater than or equal to 2) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as motion vector predictor candidates (also referred to as predicted MV candidates) in a predetermined order of priority (step Sg_2). Alternatively, the order of priority may be predetermined for each of the N candidate MVs.
[0274] Next, the inter-frame prediction unit 126 selects one motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor (also called predicted MV) for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor into the stream. The stream is the coded signal or coded bitstream described above.
[0275] Next, the inter-frame prediction unit 126 refers to the coded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference between the derived MV and the predicted motion vector as a differential MV into the stream. The coded reference picture is a picture composed of multiple blocks reconstructed after encoding.
[0276] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). Note that the predicted image is the inter-frame prediction signal described above.
[0277] Furthermore, information indicating the inter prediction mode (in the above example, the normal inter mode) used in generating the predicted image, which is included in the encoded signal, is encoded as, for example, a prediction parameter.
[0278] The candidate MV list can also be used in conjunction with lists used in other modes. Furthermore, processing related to the candidate MV list can be applied to processing related to lists used in other modes. Processing related to the candidate MV list includes, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.
[0279] [MV Export > Merge Mode]
[0280] The merge mode is an inter-frame prediction mode that derives the MV of the current block by selecting a candidate MV from a candidate MV list as the MV of the current block.
[0281] Figure 20 is a flowchart illustrating an example of inter-frame prediction based on merge mode.
[0282] First, the inter prediction unit 126 obtains multiple candidate MVs for the current block based on information about multiple coded block MVs located temporally or spatially around the current block (step Sh_1). In other words, the inter prediction unit 126 creates a candidate MV list.
[0283] Next, the inter prediction unit 126 selects one candidate MV from the plurality of candidate MVs acquired in step Sh_1 to derive the MV of the current block (step Sh_2). At this time, the inter prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream.
[0284] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).
[0285] Furthermore, information indicating the inter prediction mode (in the above example, the merge mode) used for generating the predicted image, which is included in the encoded signal, is encoded as, for example, a prediction parameter.
[0286] Figure 21 This is a conceptual diagram for explaining an example of motion vector derivation processing of the current picture based on the merge mode.
[0287] First, a prediction MV list is generated, which registers candidate prediction MVs. These include: spatially neighboring prediction MVs, which are MVs of multiple coded blocks located in the spatial vicinity of the target block; temporally neighboring prediction MVs, which are MVs of blocks near the target block's position in the coded reference picture; combined prediction MVs, which are generated by combining the MV values of spatially neighboring prediction MVs and temporally neighboring prediction MVs; and zero prediction MVs, which have a value of zero.
[0288] Next, one predicted MV is selected from a plurality of predicted MVs registered in the predicted MV list to determine the MV of the target block.
[0289] Then, the variable length coding unit encodes merge_idx, which is a signal indicating which predicted MV is selected, by describing it in the stream.
[0290] In addition, Figure 21 The predicted MVs registered in the predicted MV list described in the figure are just examples, and may be a number different from the number in the figure, or a structure that does not include some of the types of predicted MVs in the figure, or a structure that adds predicted MVs other than the types of predicted MVs in the figure.
[0291] The final MV may be determined by performing a DMVR (decoder motion vector refinement) process described later using the MV of the target block derived in the merge mode.
[0292] In addition, the candidate for the predicted MV is the candidate MV mentioned above, and the predicted MV list is the candidate MV list mentioned above. In addition, the candidate MV list can also be called the candidate list. In addition, merge_idx is MV selection information.
[0293] [MV Export > FRUC Mode]
[0294] Motion information may be derived on the decoder side instead of being signaled on the encoder side. Furthermore, as described above, the merge mode specified in the H.265 / HEVC specification may be used. Furthermore, motion information may be derived, for example, by performing a motion search on the decoder side. In one embodiment, the decoder side performs a motion search without using the pixel values of the current block.
[0295] Here, a mode for performing motion estimation on the decoding device side is described. This mode for performing motion estimation on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0296] In the form of a flow chart Figure 22An example of FRUC processing is shown in Figure 1. First, a list of multiple candidates (i.e., a candidate MV list, which may also be shared with a merge list) each including a predicted motion vector (MV) is generated by referring to the motion vectors of previously coded blocks that are spatially or temporally adjacent to the current block (step Si_1). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and one candidate is selected based on the evaluation value. Furthermore, a motion vector for the current block is derived based on the motion vector of the selected candidate (step Si_4). Specifically, for example, the selected candidate motion vector (the best candidate MV) is derived as is as the motion vector for the current block. Alternatively, the motion vector for the current block can be derived by performing pattern matching in the surrounding area of the position in the reference picture corresponding to the selected candidate motion vector. Specifically, the surrounding area of the best candidate MV can be searched using pattern matching and evaluation values in the reference picture. If an MV with a better evaluation value is found, the best candidate MV is updated to the above MV and used as the final MV for the current block. A configuration may be adopted in which the process of updating to an MV having a better evaluation value is not performed.
[0297] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).
[0298] Exactly the same processing can be performed even when processing is performed in sub-block units.
[0299] The evaluation value can also be calculated using various methods. For example, the reconstructed image of the region within the reference picture corresponding to the motion vector is compared with the reconstructed image of a predetermined region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). The predetermined region may also be predetermined.
[0300] Then, the difference between the pixel values of the two reconstructed images may be calculated and used for the motion vector evaluation value. Alternatively, the evaluation value may be calculated using other information in addition to the difference value.
[0301] Next, we'll explain a pattern matching example in detail. First, a candidate MV from a candidate MV list (e.g., a merge list) is selected as the starting point for a pattern matching search. For example, pattern matching can involve either a first pattern match or a second pattern match. First and second pattern matching are also known as bilateral matching and template matching, respectively.
[0302] [MV Export > FRUC > Bidirectional Matching]
[0303] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, the region within the other reference pictures along the motion trajectory of the current block is used as the predetermined region for calculating the candidate evaluation value. The predetermined region may also be predetermined.
[0304] Figure 23 This is a conceptual diagram for explaining an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along the motion trajectory. Figure 23 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best-matching pair among pairs of blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by a symmetric MV scaled by the display time interval is derived, and the evaluation value is calculated using the resulting difference. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs, achieving excellent results.
[0305] Assuming a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, mirror-symmetric bidirectional motion vectors are derived in the first pattern matching.
[0306] [MV Export > FRUC > Template Matching]
[0307] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the block above and / or to the left)) and a block in the reference picture. Therefore, in the second pattern matching, blocks adjacent to the current block in the current picture are used as the predetermined area for calculating the candidate evaluation value described above.
[0308] Figure 24This is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 24 As shown, in the second pattern matching, the motion vector of the current block is derived by searching the reference picture (Ref0) for a block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded area adjacent to the left and above, or one of the two, and the reconstructed image at the same position in the coded reference picture (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value can be selected as the best candidate MV from among multiple candidate MVs.
[0309] Such information indicating whether FRUC mode is adopted (e.g., called a FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is adopted (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (first pattern matching or second pattern matching) is signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and may be performed at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0310] [MV Export > Affine Mode]
[0311] Next, the affine mode for deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes referred to as the affine motion compensation prediction mode.
[0312] Figure 25A This is a conceptual diagram for explaining an example of deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks. Figure 25A In the example, the current block consists of 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, according to the following formula (1A), the two motion vectors v0 and v1 can be projected, and the motion vectors (v x , v y ).
[0313]
Formula 1
[0314]
[0315] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predetermined weight coefficient. The predetermined weight coefficient may also be predetermined.
[0316] Information indicating this affine mode (e.g., called an affine flag) can be signaled as a CU-level signal. Furthermore, the signaling of information indicating this affine mode need not be limited to the CU-level, and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0317] Furthermore, such affine modes may include several modes that differ in how the motion vectors for the upper left and upper right control points are derived. For example, affine modes include affine inter (also called affine normal inter) mode and affine merge mode.
[0318] [MV Export > Affine Mode]
[0319] Figure 25B This is a conceptual diagram for explaining an example of derivation of a motion vector for a sub-block unit in an affine mode having three control points. Figure 25B In the example, the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vector of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vector of the adjacent block, and the motion vector v2 of the lower left corner control point of the current block is derived based on the motion vector of the adjacent block. Then, according to the following formula (1B), the three motion vectors v0, v1 and v2 can be projected, and the motion vectors (v x , v y ).
[0320]
Formula 2
[0321]
[0322] Here, x and y represent the horizontal and vertical positions of the sub-block center, respectively, w represents the width of the current block, and h represents the height of the current block.
[0323] Affine modes with different numbers of control points (e.g., 2 and 3) can also be switched and signaled at the CU level. Furthermore, information indicating the number of control points in the affine mode used at the CU level can also be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0324] Furthermore, in such an affine mode with three control points, several modes may be included, each with different methods for deriving motion vectors for the upper left, upper right, and lower left control points. For example, the affine mode includes the affine inter (also called affine normal inter) mode and the affine merge mode.
[0325] [MV Export > Affine Merge Mode]
[0326] Figure 26A 、 Figure 26B and Figure 26C This is a conceptual diagram used to illustrate the affine merge mode.
[0327] In affine merge mode, such as Figure 26A As shown, for example, based on multiple motion vectors corresponding to blocks coded in affine mode among the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, a predicted motion vector for each of the control points of the current block is calculated. Specifically, the already coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in this order to determine the first valid block coded in affine mode. The predicted motion vectors for the control points of the current block are calculated based on the multiple motion vectors corresponding to the determined blocks.
[0328] For example, Figure 26B As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 are derived that are projected onto the positions of the upper left corner and upper right corner of the encoded block including block A. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.
[0329] For example, Figure 26C As shown in FIG2 , when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, a predicted motion vector v0 of the control point at the upper left corner, a predicted motion vector v1 of the control point at the upper right corner, and a predicted motion vector v2 of the control point at the lower left corner of the current block are calculated.
[0330] In addition, in the following Figure 29 This method of deriving a predicted motion vector may also be used in deriving the predicted motion vectors of the control points of the current block in step Sj_1.
[0331] Figure 27 This is a flowchart showing an example of the affine merge mode.
[0332] In the affine merge mode, as shown in the figure, first, the inter-frame prediction unit 126 derives the predicted MV of each control point of the current block (step Sk_1). Figure 25A As shown, it is the top left and top right corner points of the current block, or as Figure 25B As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block.
[0333] That is to say, if Figure 26A As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.
[0334] Then, in the case where block A is determined and block A has 2 control points, as Figure 26B As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner and the motion vector v1 of the control point at the upper right corner of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block including the block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner and the predicted motion vector v1 of the control point at the upper right corner of the current block.
[0335] Alternatively, in the case where block A is determined and block A has 3 control points, as Figure 26C As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner, the motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner, the predicted motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block.
[0336] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. Specifically, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two predicted motion vectors v0 and v1 and the above-mentioned equation (1A), or three predicted motion vectors v0, v1, and v2 and the above-mentioned equation (1B) to calculate the motion vector for that sub-block as an affine MV (step Sk_2). The inter-frame prediction unit 126 then performs motion compensation on that sub-block using these affine MVs and the coded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image for the current block is generated.
[0337] [MV Export > Affine Inter-frame Mode]
[0338] Figure 28A This is a conceptual diagram for explaining the affine inter mode with two control points.
[0339] In this affine inter-frame mode, if Figure 28A As shown, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block.
[0340] Figure 28B This is a conceptual diagram for explaining the affine inter-frame mode with three control points.
[0341] In this affine inter-frame mode, if Figure 28B As shown in FIG. 1 , a motion vector selected from the motion vectors of blocks A, B, and C that have been coded adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of blocks D and E that have been coded adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block. Furthermore, a motion vector selected from the motion vectors of blocks F and G that have been coded adjacent to the current block is used as the predicted motion vector v2 of the control point at the lower left corner of the current block.
[0342] Figure 29 This is a flowchart showing an example of the affine inter mode.
[0343] As shown in the figure, in the affine inter mode, first, the inter prediction unit 126 derives the predicted MV (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). Figure 25A or Figure 25B As shown, the control point is the point at the upper left corner, upper right corner or lower left corner of the current block.
[0344] That is, the inter-frame prediction unit 126 selects Figure 28A or Figure 28B The inter prediction unit 126 derives the predicted motion vector (v0, v1) or (v0, v1, v2) of the control point of the current block based on the motion vector of a block in the coded blocks near each control point of the current block. At this time, the inter prediction unit 126 encodes predicted motion vector selection information for identifying the two selected motion vectors into the stream.
[0345] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select as the predicted motion vector of the control point from the encoded blocks adjacent to the current block by using cost evaluation, etc., and can record a flag indicating which predicted motion vector is selected in the bitstream.
[0346] Next, the inter-frame prediction unit 126 performs a motion search (steps Sj_3 and Sj_4) while updating the predicted motion vector selected or derived in step Sj_1 (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vector of each sub-block corresponding to the predicted motion vector to be updated as an affine MV and calculates it using the above-mentioned equation (1A) or equation (1B) (step Sj_3). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on each sub-block (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference between the determined MV and the predicted motion vector as a differential MV into the stream.
[0347] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0348] [MV Export > Affine Inter-frame Mode]
[0349] When affine modes with different numbers of control points (for example, 2 and 3) are switched at the CU level for signaling, the number of control points may differ between the coded block and the current block. Figure 30A as well as Figure 30B This is a conceptual diagram for explaining a method for deriving a prediction vector for control points when the number of control points in an already coded block and a current block is different.
[0350] For example, Figure 30AAs shown in FIG. 1 , when the current block has three control points, namely, the upper left corner, the upper right corner, and the lower left corner, and the block A adjacent to the left of the current block is coded in an affine mode having two control points, motion vectors v3 and v4 are derived, projected onto the positions of the upper left corner and the upper right corner of the coded block including block A. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated. Furthermore, the predicted motion vector v2 of the control point at the lower left corner is calculated based on the derived motion vectors v0 and v1.
[0351] For example, Figure 30B As shown in FIG, when the current block has two control points, namely, the upper left corner and the upper right corner, and the block A adjacent to the left of the current block is coded in an affine mode having three control points, motion vectors v3, v4, and v5 are derived that are projected onto the positions of the upper left corner, upper right corner, and lower left corner of the coded block including block A. Then, based on the derived motion vectors v3, v4, and v5, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.
[0352] exist Figure 29 This method of deriving a predicted motion vector may also be used in deriving each predicted motion vector of the control point of the current block in step Sj_1.
[0353] [MV Export>DMVR]
[0354] Figure 31A This is a flowchart showing the relationship between the merge mode and DMVR.
[0355] The inter-frame prediction unit 126 derives a motion vector for the current block in merge mode (step S1_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, i.e., a motion search (step S1_2). If it is determined not to perform a motion search (No in step S1_2), the inter-frame prediction unit 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). In other words, in this case, the motion vector for the current block is determined in merge mode.
[0356] On the other hand, if it is determined in step S1_1 that a motion search is to be performed (Yes in step S1_2), the inter-frame prediction unit 126 searches the surrounding area of the reference picture represented by the motion vector derived in step S1_1 to derive a final motion vector for the current block (step S1_3). In other words, in this case, the motion vector of the current block is determined by DMVR.
[0357] Figure 31B This is a conceptual diagram for explaining an example of DMVR processing for determining MV.
[0358] First, the best MVP set for the current block (for example, in merge mode) is set as a candidate MV. Next, based on the candidate MV (L0), reference pixels are determined from the first reference picture (L0), which is the coded picture in the L0 direction. Similarly, based on the candidate MV (L1), reference pixels are determined from the second reference picture (L1), which is the coded picture in the L1 direction. A template is generated by averaging these reference pixels.
[0359] Next, using the template, the surrounding areas of the candidate MVs in the first reference image (L0) and the second reference image (L1) are searched, and the MV with the lowest cost is determined as the final MV. Alternatively, the cost value can be calculated using, for example, the difference between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values.
[0360] Typically, the configuration and operation of the processing described here are basically common in the encoding device and the decoding device described later.
[0361] Even if it is not the processing example described here, any processing can be used as long as it is a processing that can search the vicinity of the candidate MV and derive the final MV.
[0362] [Motion Compensation > BIO / OBMC]
[0363] In motion compensation, there are modes for generating a predicted image and then correcting the predicted image. Examples of such modes include BIO and OBMC, which will be described later.
[0364] Figure 32 This is a flowchart showing an example of generating a predicted image.
[0365] The inter prediction unit 126 generates a predicted image (step Sm_1 ), and corrects the predicted image using, for example, any of the above-described modes (step Sm_2 ).
[0366] Figure 33 This is a flowchart showing another example of generating a predicted image.
[0367] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, if it is determined that a correction process is to be performed (yes in step Sn_3), the inter-frame prediction unit 126 corrects the predicted image to generate a final predicted image (step Sn_4). On the other hand, if it is determined that a correction process is not to be performed (no in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correction (step Sn_5).
[0368] Furthermore, in motion compensation, there is a mode for correcting brightness when generating a predicted image. This mode is, for example, LIC, which will be described later.
[0369] Figure 34 This is a flowchart showing another example of generating a predicted image.
[0370] The inter-frame prediction unit 126 derives the motion vector for the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform brightness correction processing (step So_2). If it is determined that brightness correction processing is to be performed (yes in step So_2), the inter-frame prediction unit 126 generates a predicted image while performing brightness correction (step So_3). In other words, the predicted image is generated using LIC. On the other hand, if it is determined that brightness correction processing is not to be performed (no in step So_2), the inter-frame prediction unit 126 generates a predicted image using standard motion compensation without performing brightness correction (step So_4).
[0371] [Motion Compensation > OBMC]
[0372] Inter-frame prediction signals can be generated using not only the motion information of the current block obtained through motion search, but also the motion information of neighboring blocks. Specifically, inter-frame prediction signals can be generated in sub-block units within the current block by weighted addition of prediction signals based on motion information obtained through motion search (within the reference picture) and prediction signals based on motion information of neighboring blocks (within the current picture). This type of inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0373] In OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) can also be signaled at the sequence level. Furthermore, information indicating whether OBMC mode is applied (e.g., OBMC flag) can also be signaled at the CU level. Furthermore, the signaling level for this information need not be limited to the sequence and CU levels and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0374] An example of the OBMC mode will be described in more detail. Figure 35 and Figure 36 It is a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.
[0375] First, if Figure 36 As shown in FIG, the motion vector (MV) assigned to the processing target (current) block is used to obtain the predicted image (Pred) based on the usual motion compensation. Figure 36 In FIG, the arrow “MV” points to the reference picture and indicates which block the current block of the current picture refers to to obtain the predicted image.
[0376] Next, the motion vector (MV_L) derived for the already coded left-adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The motion vector (MV_L) is represented by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by overlaying the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0377] Similarly, the motion vector (MV_U) derived for the previously coded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The motion vector (MV_U) is represented by an arrow "MV_U" pointing from the current block to the reference picture. The predicted image Pred_U is then overlaid with the predicted image (e.g., Pred and Pred_L) that has undergone the first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image for the current block, with the boundaries with adjacent blocks blended (smoothed).
[0378] In addition, the above example is a two-path correction method using left-adjacent and upper-adjacent blocks, but the correction method can also be a three-path correction method or more than three-path correction method using right-adjacent and / or lower-adjacent blocks.
[0379] Furthermore, the area to be overlapped may not be the entire pixel area of the block, but may be only a partial area near the block boundary.
[0380] In this description, the predicted image correction process using OBMC is described as obtaining a single predicted image Pred by superimposing a single reference picture with the additional predicted images Pred_L and Pred_U. However, when correcting the predicted image based on multiple reference pictures, the same process can be applied to each of the multiple reference pictures. In this case, after performing OBMC image correction based on multiple reference pictures and obtaining a corrected predicted image from each reference picture, the final predicted image is obtained by further superimposing the obtained multiple corrected predicted images.
[0381] In OBMC, the target block unit may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.
[0382] One method for determining whether to apply OBMC processing involves using, for example, obmc_flag, a signal indicating whether OBMC processing should be applied. As a specific example, the encoding device may determine whether the target block belongs to a region with complex motion. If the target block belongs to a region with complex motion, the encoding device sets obmc_flag to a value of 1 and applies OBMC processing to the target block. If the target block does not belong to a region with complex motion, the encoding device sets obmc_flag to a value of 0 and does not apply OBMC processing to the target block. Meanwhile, the decoding device decodes obmc_flag described in a stream (e.g., a compressed sequence) and switches whether to apply OBMC processing based on the value of the flag during decoding.
[0383] In the above example, the inter-frame prediction unit 126 generates a single rectangular predicted image for a rectangular current block. However, the inter-frame prediction unit 126 may generate multiple predicted images having shapes other than a rectangle for the rectangular current block and may combine these multiple predicted images to generate a final rectangular predicted image. A shape other than a rectangle may be, for example, a triangle.
[0384] Figure 37 This is a conceptual diagram for explaining the generation of predicted images of two triangles.
[0385] The inter-frame prediction unit 126 performs motion compensation on the first triangular partition within the current block using the first MV of the first partition to generate a triangular predicted image. Similarly, the inter-frame prediction unit 126 performs motion compensation on the second triangular partition within the current block using the second MV of the second partition to generate a triangular predicted image. The inter-frame prediction unit 126 then combines these predicted images to generate a predicted image with the same rectangular shape as the current block.
[0386] In addition, Figure 37 In the example shown, the first partition and the second partition are each a triangle, but they may also be a trapezoid or may be different shapes from each other. Figure 37 In the example shown, the current block is composed of 2 partitions, but it can also be composed of 3 or more partitions.
[0387] Furthermore, the first and second partitions may overlap. That is, the first and second partitions may contain the same pixel region. In this case, the predicted image in the first and second partitions can be used to generate the predicted image for the current block.
[0388] In addition, although this example shows an example in which predicted images are generated by inter-frame prediction for both of the two partitions, a predicted image may be generated by intra-frame prediction for at least one partition.
[0389] [Motion Compensation > BIO]
[0390] Next, the method for deriving motion vectors will be described. First, a mode for deriving motion vectors based on a model assuming constant-speed linear motion will be described. This mode is sometimes called the BIO (bidirectional optical flow) mode.
[0391] Figure 38 This is a conceptual diagram used to explain a model assuming constant velocity linear motion. Figure 38 In (v x , v y ) represents the velocity vector, τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represent the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represent the motion vector corresponding to the reference picture Ref1.
[0392] At this time, it can also be that the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, using the following optical flow equation (2).
[0393]
Formula 3
[0394]
[0395] Here, I(k) represents the luminance value of reference image k (k = 0, 1) after motion compensation. This optical flow equation states that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Block-level motion vectors obtained from a merge list, etc., may be corrected on a pixel-by-pixel basis based on a combination of this optical flow equation and Hermite interpolation.
[0396] Furthermore, the decoding device may derive motion vectors using a method different from the method based on a model assuming constant velocity linear motion. For example, motion vectors may be derived on a sub-block basis based on motion vectors of a plurality of adjacent blocks.
[0397] [Motion Compensation > LIC]
[0398] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.
[0399] Figure 39 This is a conceptual diagram for explaining an example of a method for generating a predicted image using a brightness correction process based on an LIC process.
[0400] First, MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.
[0401] Next, information is extracted for the current block, indicating how the luminance values vary between the reference picture and the current picture. This extraction is based on the luminance pixel values of the coded left-adjacent reference region (peripheral reference region) and the coded upper-adjacent reference region (peripheral reference region) in the current picture, as well as the luminance pixel values at equivalent positions within the reference picture specified by the derived MV. This information, indicating how the luminance values vary, is then used to calculate the luminance correction parameters.
[0402] The brightness correction parameters are applied to the reference image in the reference picture specified by the MV to perform brightness correction processing, thereby generating a predicted image for the current block.
[0403] in addition, Figure 39 The shape of the peripheral reference area in the figure is an example, and shapes other than these may be used.
[0404] In addition, the process of generating a predicted image based on one reference image is described here, but the same is true when generating a predicted image based on multiple reference images. The predicted image can also be generated after brightness correction processing is performed on the reference images obtained from each reference image in the same way as described above.
[0405] One method for determining whether to use the LIC process involves using, for example, a lic_flag, which serves as a signal indicating whether the LIC process is used. As a specific example, the encoder determines whether the current block belongs to an area where luminance changes. If the current block belongs to an area where luminance changes, the lic_flag is set to a value of 1, and encoding is performed using the LIC process. If the current block does not belong to an area where luminance changes, the lic_flag is set to a value of 0, and encoding is performed without using the LIC process. Alternatively, the decoder can decode the lic_flag described in the stream and switch whether to use the LIC process based on its value for decoding.
[0406] Another method for determining whether to use the LIC process is to determine whether the LIC process was used in surrounding blocks. As a specific example, when the current block is in merge mode, a determination is made as to whether the surrounding coded blocks selected during MV derivation in merge mode were coded using the LIC process. Based on the determination, whether or not to use the LIC process is switched for coding. In this example, the same process also applies to the decoding device.
[0407] use Figure 39 The form of the LIC process (luminance correction process) has been described, and its details will be described below.
[0408] First, the inter prediction unit 126 derives a motion vector for acquiring a reference image corresponding to a current block to be encoded from a reference picture that is an already encoded picture.
[0409] Next, the inter-frame prediction unit 126 uses the luminance pixel values of the coded neighboring reference areas to the left and above, as well as the luminance pixel values at the same position in the reference picture specified by the motion vector, to extract information indicating how the luminance values in the reference picture and the current picture change, and calculates a luminance correction parameter. For example, the luminance pixel value of a pixel in the neighboring reference area in the current picture is set to p0, and the luminance pixel value of the pixel in the neighboring reference area at the same position in the reference picture is set to p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1+B=p0 for multiple pixels in the neighboring reference area as luminance correction parameters.
[0410] Next, the inter-frame prediction unit 126 uses the brightness correction parameters to perform brightness correction on the reference image within the reference picture specified by the motion vector to generate a predicted image for the current block. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after brightness correction is set to p3. The inter-frame prediction unit 126 calculates A×p2+B=p3 for each pixel in the reference image to generate the predicted image after brightness correction.
[0411] also, Figure 39 The shape of the peripheral reference area in is an example, and other shapes can also be used. Figure 39 A portion of the surrounding reference area shown is used. For example, an area including a predetermined number of pixels thinned out from the upper and left adjacent pixels may be used as the surrounding reference area. Furthermore, the surrounding reference area is not limited to an area adjacent to the encoding target block and may also be an area not adjacent to the encoding target block. The predetermined number of pixels may also be predetermined.
[0412] In addition, Figure 39 In the example shown, the peripheral reference region within the reference picture is an area specified by the motion vector of the current picture from among the peripheral reference regions within the current picture, but it may also be an area specified by another motion vector. For example, the other motion vector may be the motion vector of the peripheral reference region within the current picture.
[0413] Here, the operation in the encoding device 100 is described, but typically, the operation in the decoding device 200 is also similar.
[0414] Furthermore, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for all of them.
[0415] Furthermore, the LIC process may be applied in sub-block units. For example, the modification parameters may be derived using the surrounding reference region of the current sub-block and the surrounding reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0416] [Prediction Control Department]
[0417] The prediction control unit 128 selects one of the intra-frame prediction signal (the signal output from the intra-frame prediction unit 124) and the inter-frame prediction signal (the signal output from the inter-frame prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0418] like Figure 1 As shown, in various examples of encoding devices, the prediction control unit 128 may also output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate a coded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may also be used in the decoding device. The decoding device may also receive and decode the coded bitstream, performing the same prediction processing as that performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction parameters may include a prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra-frame prediction unit 124 or the inter-frame prediction unit 126), or any index, flag, or value based on the prediction processing performed by the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128 or indicating the prediction processing.
[0419] [Encoding device installation example]
[0420] Figure 40 1 is a block diagram showing an implementation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. For example, Figure 1The multiple components of the encoding device 100 shown are Figure 40 The processor a1 and the memory a2 shown are implemented.
[0421] Processor a1 is a circuit that processes information and can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. Processor a1 can also be a processor such as a CPU. In addition, processor a1 can also be a collection of multiple electronic circuits. In addition, for example, processor a1 can also play a role. Figure 1 The functions of multiple components of the encoding device 100 shown in FIG.
[0422] Memory a2 is a dedicated or general-purpose memory that stores information used by processor a1 to encode moving images. Memory a2 can be an electronic circuit or connected to processor a1. Alternatively, memory a2 can be included in processor a1. Alternatively, memory a2 can be a collection of multiple electronic circuits. Memory a2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory a2 can be either non-volatile or volatile memory.
[0423] For example, the memory a2 may store a coded moving image or a bit string corresponding to the coded moving image. In addition, the memory a2 may store a program for the processor a1 to encode the moving image.
[0424] In addition, for example, memory a2 can also serve as Figure 1 The memory a2 can be used as a component for storing information among the multiple components of the encoding device 100 shown in FIG. Figure 1 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks and reconstructed pictures.
[0425] In addition, in the encoding device 100, it is not necessary to install Figure 1 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 1 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.
[0426] [Decoding device]
[0427] Next, a decoding device that can decode the coded signal (coded bit stream) output from, for example, the above-described coding device 100 will be described. Figure 412 is a block diagram showing the functional structure of a decoding device 200 according to an embodiment. The decoding device 200 is a moving picture decoding device that decodes a moving picture in units of blocks.
[0428] like Figure 41 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.
[0429] Decoding device 200 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes a software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. Alternatively, decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to entropy decoding unit 202, inverse quantization unit 204, inverse transform unit 206, addition unit 208, loop filter unit 212, intra-frame prediction unit 216, inter-frame prediction unit 218, and prediction control unit 220.
[0430] Below, after describing the overall processing flow of the decoding device 200 , each component included in the decoding device 200 will be described.
[0431] [Overall decoding process]
[0432] Figure 42 This is a flowchart showing an example of the overall decoding process performed by the decoding device 200 .
[0433] First, the entropy decoding unit 202 of the decoding device 200 determines a partitioning pattern for a fixed-size block (e.g., 128×128 pixels) (step Sp_1). This partitioning pattern is the same as the partitioning pattern selected by the encoding device 100. The decoding device 200 then performs steps Sp_2 to Sp_6 on each of the multiple blocks that make up this partitioning pattern.
[0434] That is, the entropy decoding unit 202 decodes (specifically, performs entropy decoding) the encoded quantization coefficients and prediction parameters of the decoding target block (also referred to as the current block) (step Sp_2).
[0435] Next, the inverse quantization unit 204 and the inverse transformation unit 206 perform inverse quantization and inverse transformation on the plurality of quantized coefficients to restore a plurality of prediction residuals (ie, difference blocks) (step Sp_3).
[0436] Next, the prediction processing unit composed of all or part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction signal (also referred to as a prediction block) of the current block (step Sp_4).
[0437] Next, the adding unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the differential block (step Sp_5).
[0438] Then, when the reconstructed image is generated, the loop filter unit 212 filters the reconstructed image (step Sp_6).
[0439] Then, the decoding device 200 determines whether decoding of the entire picture is completed (step Sp_7 ). If it is determined that decoding is not completed (No in step Sp_7 ), the processing from step Sp_1 is repeatedly executed.
[0440] As shown in the figure, the processing of steps Sp_1 to Sp_7 is sequentially performed by the decoding device 200, or a plurality of processes of some of these processes may be performed in parallel, or the order may be reversed.
[0441] [Entropy decoding unit]
[0442] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the coded bit stream to the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220 in the embodiment (see Figure 1 The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.
[0443] [Inverse quantization unit]
[0444] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantization unit 204 then outputs the inversely quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0445] [Inverse transformation unit]
[0446] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients input from the inverse quantization unit 204 .
[0447] For example, when the information read from the coded bitstream indicates that EMT or AMT is adopted (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.
[0448] Furthermore, for example, when the information decoded from the coded bit stream indicates that NSST is adopted, the inverse transform unit 206 applies inverse re-transformation to the transform coefficients.
[0449] [Addition Department]
[0450] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction sample input from the prediction control unit 220 . The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212 .
[0451] [Block Memory]
[0452] The block memory 210 is a storage unit for storing blocks in a decoding target picture (hereinafter referred to as a current picture) to be referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the adding unit 208 .
[0453] [Loop filter unit]
[0454] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device.
[0455] When the ALF on / off information read from the coded bitstream indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0456] [Frame Memory]
[0457] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.
[0458] [Prediction Processing Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]
[0459] Figure 432 is a flowchart showing an example of processing performed by the prediction processing unit of the decoding device 200. The prediction processing unit is composed of all or part of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0460] The prediction processing unit generates a predicted image for the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Prediction signals include, for example, intra-frame prediction signals and inter-frame prediction signals. Specifically, the prediction processing unit generates a predicted image for the current block using the reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.
[0461] The reconstructed image may be, for example, an image of a reference picture, or an image including the current block, that is, an image of a decoded block within the current picture. The decoded block within the current picture may be, for example, an adjacent block of the current block.
[0462] Figure 44 This is a flowchart showing another example of processing performed by the prediction processing unit of the decoding device 200 .
[0463] The prediction processing unit determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode can be determined based on prediction parameters or the like.
[0464] If the first mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the first mode (step Sr_2a). Furthermore, if the second mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the second mode (step Sr_2b). Furthermore, if the third mode is determined to be a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the third mode (step Sr_2c).
[0465] The first, second, and third methods are different methods for generating predicted images, and may be, for example, inter-frame prediction, intra-frame prediction, or other prediction methods. In such prediction methods, the above-mentioned reconstructed image may also be used.
[0466] [Intra-frame prediction unit]
[0467] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the coded bitstream, referring to blocks in the current picture stored in the block memory 210, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 performs intra prediction by referring to samples (e.g., luminance values and chrominance values) of blocks adjacent to the current block, thereby generating an intra prediction signal and outputting the intra prediction signal to the prediction control unit 220.
[0468] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0469] Furthermore, when the information read from the coded bit stream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.
[0470] [Inter-frame prediction unit]
[0471] The inter-frame prediction unit 218 predicts the current block by referring to the reference picture stored in the frame memory 214. Prediction is performed on a per-block basis (e.g., 4×4 blocks) for the current block or a sub-block within the current block. For example, the inter-frame prediction unit 218 performs motion compensation using motion information (e.g., motion vectors) decoded from the coded bitstream (e.g., prediction parameters output from the entropy decoding unit 202). This generates an inter-frame prediction signal for the current block or sub-block and outputs the inter-frame prediction signal to the prediction control unit 220.
[0472] When the information read from the encoded bit stream indicates that the OBMC mode is adopted, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks.
[0473] Furthermore, when the information decoded from the coded bitstream indicates that the FRUC mode is used, the inter-frame prediction unit 218 performs motion estimation using the pattern matching method (bidirectional matching or template matching) decoded from the coded stream to derive motion information. Furthermore, the inter-frame prediction unit 218 performs motion compensation (prediction) using the derived motion information.
[0474] Furthermore, when the BIO mode is used, the inter-frame prediction unit 218 derives a motion vector based on a model assuming constant-speed linear motion. Furthermore, when the information decoded from the coded bitstream indicates that the affine motion compensation prediction mode is used, the inter-frame prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0475] [MV Export > Normal Interframe Mode]
[0476] When the information read from the encoded bit stream indicates that the normal inter mode is applied, the inter prediction unit 218 derives an MV based on the information read from the encoded bit stream, and performs motion compensation (prediction) using the MV.
[0477] Figure 45 This is a flowchart illustrating an example of inter prediction based on the normal inter mode in the decoding apparatus 200 .
[0478] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation on each block. Based on information such as the MVs of multiple decoded blocks temporally and spatially surrounding the current block, the inter-frame prediction unit 218 obtains multiple candidate MVs for the current block (step Ss_1). In other words, the inter-frame prediction unit 218 creates a candidate MV list.
[0479] Next, the inter-frame prediction unit 218 extracts N (N is an integer greater than or equal to 2) candidate MVs from the multiple candidate MVs obtained in step Ss_1 as motion vector predictor candidates (also referred to as predicted MV candidates) in a predetermined priority order (step Ss_2). Alternatively, the priority order may be predetermined for each of the N predicted MV candidates.
[0480] Next, the inter-frame prediction unit 218 decodes the predicted motion vector selection information from the input stream (i.e., the encoded bit stream), and uses the decoded predicted motion vector selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted motion vector (also called predicted MV) of the current block (step Ss_3).
[0481] Next, the inter prediction section 218 decodes the difference MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded difference MV to the selected predicted motion vector (step Ss_4).
[0482] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).
[0483] [Prediction Control Department]
[0484] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as the prediction signal to the adder 208. Generally speaking, the structure, function, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side can correspond to the structure, function, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.
[0485] [Decoding device installation example]
[0486] Figure 46 2 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 41 The components of the decoding device 200 are shown in FIG. Figure 46The processor b1 and memory b2 are shown as being installed.
[0487] Processor b1 is a circuit that processes information and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes the encoded moving image (i.e., the encoded bit stream). Processor b1 can also be a processor such as a CPU. In addition, processor b1 can also be a collection of multiple electronic circuits. In addition, for example, processor b1 can also play a role in Figure 41 The functions of multiple components of the decoding device 200 shown in FIG.
[0488] Memory b2 is a dedicated or general-purpose memory that stores information used by processor b1 to decode the coded bit stream. Memory b2 can be an electronic circuit or connected to processor b1. Alternatively, memory b2 can be included in processor b1. Alternatively, memory b2 can be a collection of multiple electronic circuits. Alternatively, memory b2 can be a magnetic disk or optical disk, or can be a storage device or recording medium. Memory b2 can be either non-volatile or volatile memory.
[0489] For example, the memory b2 may store moving images or coded bit streams. In addition, the memory b2 may also store a program for the processor b1 to decode the coded bit stream.
[0490] In addition, for example, memory b2 can serve as Figure 41 The memory b2 is a component for storing information among the components of the decoding device 200 shown in FIG. Figure 41 The functions of the block memory 210 and the frame memory 214 are shown. More specifically, the memory b2 can store reconstructed blocks and reconstructed pictures.
[0491] In addition, in the decoding device 200, it is not necessary to install Figure 41 All of the multiple components shown above may not perform all of the multiple processes described above. Figure 41 Part of the multiple components shown in the figure may be included in other devices, and part of the multiple processes described above may be executed by other devices.
[0492] [Definition of each term]
[0493] As an example, each term may be defined as follows.
[0494] A picture is an arrangement of multiple luma samples in monochrome format, or an arrangement of multiple luma samples and two corresponding arrangements of multiple color difference samples in 4:2:0, 4:2:2, and 4:4:4 color formats. A picture can be a frame or a field.
[0495] A frame is a composition of a top field that generates a plurality of sample rows 0, 2, 4, ... and a bottom field that generates a plurality of sample rows 1, 3, 5, ... .
[0496] A slice is an integer number of coding tree units contained in an independent slice segment and all subsequent dependent slice segments before the next independent slice segment (if any) in the same access unit.
[0497] A tile is a rectangular area of multiple coding tree blocks within a specific tile column and a specific tile row in a picture. Tiles can still apply loop filters across the edges of the tiles, but can also be rectangular areas of a frame that are intended to be independently decoded and encoded.
[0498] A block is an MxN (N rows and M columns) array of samples, or an MxN array of transform coefficients. A block can also be a square or rectangular area of multiple pixels consisting of multiple matrices of one luma and two chroma.
[0499] A CTU (Coding Tree Unit) can be a coding tree block of multiple luma samples for a picture with a three-sample arrangement, or two corresponding coding tree blocks of multiple chroma samples. Alternatively, a CTU can be a coding tree block of any number of samples in a monochrome picture or a picture encoded using the same syntax used for encoding three separate color planes and multiple samples.
[0500] A super block may be composed of one or two mode information blocks, or may be recursively divided into four 32×32 blocks, or further divided into a 64×64 pixel square block.
[0501] (Implementation Method 2)
[0502] The encoding device 100 in this embodiment has the same structure as that in Embodiment 1. Furthermore, the encoding device 100 in this embodiment has additional functions or functions that replace those in Embodiment 1. Similarly, the decoding device 200 in this embodiment has the same structure as that in Embodiment 1. Furthermore, the decoding device 200 in this embodiment has additional functions or functions that replace those in Embodiment 1.
[0503] The decoding device 200 in this embodiment performs processing using picture units called VPDUs, and performs prediction processing as described in Examples 1 through 7 below. Furthermore, in Examples 1 through 7 below, pipeline processing is used. This pipeline processing includes multiple stages (e.g., Stages 1 through 6), each of which is performed by hardware such as a circuit or processor corresponding to the stage.
[0504] [VPDU]
[0505] Figure 47 This is a diagram for explaining the block size when performing prediction processing.
[0506] The largest unit that makes up a picture or image is a CTU. Prediction processing is essentially performed in CUs, which can be smaller than the CTU size. For example, if the CTU size is 128×128 pixels, the CU size can be 4×4 pixels, 4×8 pixels, 8×4 pixels, 8×8 pixels, etc., or 128×128 pixels.
[0507] Another unit is called a VPDU (Virtual Pipeline Decoding Unit). A VPDU is a fixed unit that can be processed in a single stage when pipeline processing is performed in hardware. Specifically, the VPDU size is often set to the largest of the block sizes to which the transform is applied, for example, a size of 64×64 pixels.
[0508] Generally, when the processing target CU is smaller than the VPDU, it is assumed that the multiple CUs contained in the VPDU are processed collectively in a single pipeline processing stage. On the other hand, when the processing target CU is larger than the VPDU, it is assumed that the CU is split into multiple VPDUs, and each VPDU is processed in a single pipeline processing stage.
[0509] [First example of prediction processing]
[0510] Figure 48 This is a diagram schematically showing a first example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0511] Figure 48 The pipeline process PL1 shown includes Stages 1 to 6.
[0512] In Stage 1, entropy decoding process St1 is performed. That is, in Stage 1, the decoding device 200 performs entropy decoding on the input stream of the decoding object, thereby obtaining the information required for decoding the coded image. In addition, the input stream is the coded bit stream mentioned above, and the entropy decoding process St1 is performed by Figure 41 The entropy decoding unit 202 of the decoding device 200 shown in FIG. Figure 47 The VPDU unit shown in the figure is used instead of the CU unit, CTU unit or a unit larger than them.
[0513] In Stage 2, MV derivation processing St21 and memory transfer processing St22 are performed. Specifically, in MV derivation processing St21, the decoding device 200 derives an MV as a motion vector. Then, in memory transfer processing St22, the decoding device 200 transfers at least a portion of the reference image from a memory such as the frame memory 214 using the derived MV. Generally, the processing of each stage after Stage 2 is performed in the following manner: Figure 47 For example, in each stage after Stage 2, if the VPDU is larger than the CU, processing is performed on each CU until all CUs within the VPDU are processed. In addition, if the VPDU is smaller than the CU, for example, one VPDU within the CU is processed, and the processing result of the one VPDU (e.g., the derived MV) is applied to other VPDUs within the CU.
[0514] In Stage 3, DMVR processing St3 is performed. That is, in Stage 3, the decoding apparatus 200 searches for the vicinity of the MV derived as described above through DMVR processing and modifies the MV.
[0515] In Stage 4, MC processing St4 is performed. That is, in Stage 4, the decoding apparatus 200 performs motion compensation processing using the modified MV.
[0516] In addition, the processing of Stage 2 to 4 above is performed by Figure 41 The inter-frame prediction unit 218 shown is performed.
[0517] In Stage 5, LIC processing St51a, intra-frame prediction processing St52, intra-frame / inter-frame mixing processing St53, mode switching processing St57, inverse quantization / inverse transform processing St54 and reconstruction processing St56 are performed.
[0518] In the LIC process St51a, the decoding device 200 applies the LIC process to the inter-frame prediction image obtained by the motion compensation process to correct the inter-frame prediction image. Figure 41 The inter-frame prediction unit 218 shown is performed.
[0519] In addition, in the intra-frame prediction process St52 performed simultaneously with the LIC process St51a, the decoding device 200 generates an intra-frame prediction image by performing intra-frame prediction. Figure 41 The intra prediction unit 216 shown in FIG.
[0520] In the intra / inter mixing process St53, the decoding device 200 generates an intra / inter mixed prediction image by mixing or superimposing its modified inter prediction image and the intra prediction image using an adder. For example, the average value of the pixel value of the intra prediction image and the pixel value of the inter prediction image can be used as the pixel value of the intra / inter mixed prediction image. In addition, the intra / inter mixing process St53 is also called CIIP (Combined inter merge / intra prediction). In the mode switching process St57, the decoding device 200 selectively switches the final prediction image used in decoding the encoded image based on the intra prediction image, the intra / inter mixed prediction image, and the modified inter prediction image.
[0521] The intra / inter mixing process St53 can be considered as one of the MV derivation modes in inter prediction to mix the intra prediction images after performing the MC process St4 in inter prediction. Figure 16 、 Figure 17 However, it can be included in any category. For example, when MV export mode is classified as Figure 17 In the case of a mode in which the differential MV is encoded or a mode in which the differential MV is not encoded, as long as it is included in one of them, it can be included in any category. Figure 16 The same is true for other categories.
[0522] In addition, the intra-frame / inter-frame mixing processing St53 and the mode switching processing St57 are performed by Figure 41 The prediction is performed by at least one of the inter prediction unit 218, the intra prediction unit 216, and the prediction control unit 220 shown.
[0523] In the inverse quantization / inverse transform process St54, the decoding device 200 inversely quantizes and inversely transforms the input stream to generate a residual image. This residual image is the prediction residual or difference block described above. In the reconstruction process St56, the decoding device 200 adds this residual image to the predicted image switched as described above to generate a reconstructed image. This reconstructed image is used as adjacent reference pixels in the processing of the next block or CU and is therefore fed back to the LIC process St51a and the intra-frame prediction process St52.
[0524] In addition, the inverse quantization / inverse transform processing St54 is performed by Figure 41 The inverse quantization unit 204 and the inverse transform unit 206 are shown, and the reconstruction process St56 is performed by Figure 41 The addition unit 208 shown performs.
[0525] In Stage 6, loop filtering process St6 is performed. That is, in Stage 6, the decoding device 200 generates a decoded image by applying loop filtering such as a deblocking filter to the reconstructed image. In addition, the loop filtering process St6 is performed by Figure 41 This is performed by the loop filter unit 212 shown.
[0526] in addition, Figure 48 The pipeline process PL1 shown is an example, and at least one process included in the pipeline process PL1 may be removed, another process may be added to the pipeline process PL1, or the method of dividing the stages may be changed.
[0527] If used Figure 47 As explained above, when the processing target CU is smaller than the VPDU, it is assumed that multiple CUs included in the VPDU are aggregated and processed in one stage of the pipeline. For example, when the size of the VPDU is 64×64 pixels and the size of all CUs included in it is 8×8 pixels, Figure 48 In one stage of the pipeline processing PL1 shown, 64 CUs are processed.
[0528] Here, in the mode switching process St57 in Stage 5, as a worst-case scenario, mixed intra / inter prediction images may be used in decoding all 64 CUs included in the VPDU. In the intra / inter mixing process St53 to generate these mixed intra / inter prediction images, the decoding device 200 mixes the intra prediction images with the modified inter prediction images. Therefore, the decoding device 200 needs to confirm that both the intra prediction process St52 and the inter prediction process for each CU have completed before mixing the intra prediction images with the modified inter prediction images. The inter prediction process includes the MV derivation process St21, the memory transfer process St22, the DMVR process St3, the MC process St4, and the LIC process St51a.
[0529] That is, in Figure 48 In the pipeline process PL1 shown, the decoding device 200 needs to wait for each CU to complete the intra prediction process St52 or the inter prediction process, whichever has the later completion time. As a result, in the worst case, the waiting time may occur 64 times in one stage of processing.
[0530] This waiting time significantly increases the processing time of Stage 5, and as a result, the decoding of a picture may not be completed within the processing time allocated for decoding. In other words, the decoded image may not be displayed at the scheduled time.
[0531] Therefore, the decoding device 200 in this embodiment may also perform the second to seventh examples of prediction processing shown below.
[0532] [Second example of prediction processing]
[0533] Figure 49 2 is a flowchart showing a second example of prediction processing performed by the decoding device 200 in this embodiment. In addition, this flowchart shows the flow of processing on a CU in the second example of prediction processing.
[0534] The decoding device 200 performs steps S11 to S17a on a CU-by-CU basis. Specifically, the decoding device 200 first derives an MV for the target CU (step S11), and then performs DMVR (step S12). In this DMVR process, the decoding device 200 modifies the MV derived in step S11 by searching its surroundings. The decoding device 200 then performs MC, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image.
[0535] Next, the decoding device 200 determines whether to apply the LIC process to the processing target CU (step S14a). If it is determined that the LIC process is to be applied (yes in step S14a), the decoding device 200 modifies the inter-frame prediction image obtained by the MC process using the LIC process to generate a final prediction image (step S16a). The final prediction image is the prediction image used to generate the reconstructed image and is the prediction image added to the residual image.
[0536] On the other hand, if the decoding device 200 determines that the LIC process is not to be applied (No in step S14a), it further determines whether to apply the intra / inter mixing process (step S15a). If the intra / inter mixing process is to be applied (Yes in step S15a), the decoding device 200 mixes the inter-frame prediction image obtained by the MC process with the intra-frame prediction image obtained by the intra-frame prediction process (step S17a). This mixing generates the intra / inter mixed prediction image as the final prediction image. On the other hand, if the decoding device 200 determines that the intra / inter mixing process is not to be applied (No in step S15a), the inter-frame prediction image obtained by the MC process is used as the final prediction image for generating the reconstructed image.
[0537] In addition, Figure 49 In the flowchart of , the decoding apparatus 200 first determines whether to apply the LIC process and then determines whether to apply the intra / inter mixed process. However, these determinations may be performed in the reverse order.
[0538] in addition, Figure 49 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0539] also, Figure 49 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 49 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0540] also, Figure 49 The determination of whether to apply the LIC process (step S14a) and the determination of whether to apply the intra / inter mixed process (step S15a) in the flowchart can also be performed by parsing flags. That is, the decoding device 200 can also make these determinations by, for example, parsing flags described in the stream. Alternatively, the encoding device 100 can make these determinations by, for example, calculating costs using an RD optimization model. That is, for each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process being determined and the cost of the predicted image obtained by not applying the process. The encoding device 100 then makes the above determinations so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0541] Figure 50 This is a diagram schematically showing a second example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0542] Figure 50 The pipeline process PL2 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0543] This pipeline processes Stage 5 of Stage 1 to Stage 6 of PL2 and Figure 48The illustrated pipeline process PL1, Stage 5, is different. Specifically, in Stage 5 of pipeline process PL2, the intra / inter mixed process St53 is performed without applying the LIC process St51a. In other words, the decoding device 200 does not perform the LIC process St51a on the inter-predicted image generated by the MC process St4, but instead mixes the inter-predicted image with the intra-predicted image to generate an intra / inter mixed predicted image.
[0544] In such a pipeline process PL2, at the time of performing the intra-frame / inter-frame mixing process St53, the generation of the inter-frame prediction image required in the intra-frame / inter-frame mixing process St53 has been completed in Stage 4. Therefore, in the pipeline process PL2, it is not necessary to Figure 48 The inter-frame prediction processing required in the pipeline process PL1 shown in FIG5 is not completed. Therefore, the intra-frame / inter-frame mixing process St53 can be performed immediately upon completion of the intra-frame prediction process St52. As a result, the generation of unnecessary waiting time can be suppressed. In other words, the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is performed can be set to be approximately the same as the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is not performed.
[0545] in addition, Figure 50 The pipeline process PL2 shown is an example, and at least one process included in the pipeline process PL2 may be removed, another process may be added to the pipeline process PL2, or the method of dividing the stages may be changed.
[0546] Thus, in the second example of the prediction process, even if Figure 49 as well as Figure 50 Even when encoding multiple CUs in one VPDU, the processing at all stages can be completed within a pre-specified processing time. This increases the possibility of achieving faster processing while suppressing degradation in encoding performance.
[0547] [Changes in the Second Example of Prediction Processing]
[0548] Furthermore, in the second example of prediction processing, intra / inter mixing is always prohibited when the condition that the LIC process is selected is met. However, prohibition of intra / inter mixing is not limited to this condition. For example, intra / inter mixing can also be prohibited when conditions such as the LIC process is selected and the size of the processing target CU is less than a specific threshold are met. While this increases the processing time of Stage 5 for generating the reconstructed image, this increase can be kept within a certain range, further suppressing degradation in coding performance while increasing the likelihood of suppressing increases in processing time.
[0549] Furthermore, LIC processing may be prohibited if the following conditions are met: application of the intra / inter mixed processing is selected and the size of the CU being processed is less than a specific threshold. Furthermore, the "size of the CU being processed is less than a specific threshold" condition may be replaced with "the number of CUs included in a single VPDU is greater than a specific threshold."
[0550] Furthermore, in the second example of prediction processing, in the combination of the LIC process and the intra / inter mixed process, one process is applied while the other is prohibited. However, this is not limited to this combination; similar application and prohibition can be performed in other combinations. For example, any other combination can be used, as long as it involves an inter-prediction process that requires feedback of the reconstructed image and a process that requires a waiting time in Stage 5 for generating the reconstructed image. The process that requires a waiting time is one that requires waiting for the later of the intra-prediction and inter-prediction processes to complete. Furthermore, the inter-prediction process that requires feedback of the reconstructed image can also be, for example, a filtering process performed on the predicted image obtained through inter-prediction.
[0551] [Third example of prediction processing]
[0552] Figure 51 1 is a flowchart showing a third example of prediction processing performed by the decoding device 200 in this embodiment. In addition, this flowchart shows the flow of processing on a CU in the third example of prediction processing.
[0553] The decoding device 200 performs steps S11 to S13, S14b, S15a, S16b, and S17a on a CU basis. Specifically, the decoding device 200 first derives an MV for the CU being processed (step S11), and then performs DMVR processing (step S12). In this DMVR process, the decoding device 200 modifies the MV derived in step S11 by searching the surrounding area. The decoding device 200 then performs MC processing, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image for the CU being processed.
[0554] Next, the decoding device 200 determines whether to apply filtering (also referred to as prediction image filtering) to the inter-frame prediction image of the processing target CU (step S14b). Here, filtering is a process of multiplying a weight by each of a plurality of pixels to correct or update the pixel value. Furthermore, prediction image filtering is a process of correcting or updating the pixel values included in the prediction image.
[0555] If it is determined that filtering is to be applied (yes in step S14b), the decoding device 200 performs filtering on the inter-frame prediction image obtained by MC processing, thereby generating a filtered inter-frame prediction image as the final prediction image (step S16b). For example, in the filtering process, a matrix composed of 5×5 filter coefficients can also be used. Moreover, in the filtering process, pixels included in the reconstructed image of the CU adjacent to the processing object CU (hereinafter referred to as the adjacent CU) can also be used. For example, among the pixels in the inter-frame prediction image of the processing object CU, the pixels close to the boundary between the processing object CU and the adjacent CU are filtered using the pixels included in the reconstructed image of the adjacent CU.
[0556] On the other hand, if it is determined that filtering is not to be applied (No in step S14b), the decoding device 200 further determines whether to apply intra / inter mixing (step S15a). If it is determined that intra / inter mixing is to be applied (Yes in step S15a), the decoding device 200 mixes the inter-frame prediction image obtained by the MC process with the intra-frame prediction image obtained by the intra-frame prediction process (step S17a). This mixing generates an intra / inter mixed prediction image as the final prediction image. On the other hand, if it is determined that intra / inter mixing is not to be applied (No in step S15a), the decoding device 200 uses the inter-frame prediction image obtained by the MC process as it is for generating the reconstructed image.
[0557] In addition, Figure 51 In the flowchart of , the decoding apparatus 200 first determines whether to apply filtering processing to the inter-frame prediction image, and then determines whether to apply the intra / inter mixed processing, but these determinations may be performed in the reverse order.
[0558] in addition, Figure 51 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0559] also, Figure 51 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 51 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0560] also, Figure 51The determination of whether to apply filtering (step S14b) and the determination of whether to apply intra-frame / inter-frame mixing (step S15a) in the flowchart can also be performed by parsing the flag. That is, the decoding device 200 can also perform these determinations by, for example, parsing the flag described in the stream. In addition, the encoding device 100 can also perform these determinations by, for example, calculating the cost using the RD optimization model. That is, in each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process of the determination object and the cost of the predicted image obtained by not applying the process. And, the encoding device 100 performs the above determination so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0561] Figure 52 This is a diagram schematically showing a third example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0562] Figure 52 The pipeline process PL3 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0563] This pipeline processes Stage 5 in Stage 1 to Stage 6 of PL3 and Figure 50 The Stage 5 of the pipeline process PL2 shown is different. That is, in the Stage 5 of the pipeline process PL3, the prediction image filtering process St51b is performed instead of the LIC process St51a. This prediction image filtering process St51b is the above-mentioned filtering process.
[0564] Specifically, in Stage 5 of pipeline processing PL2, intra / inter mixing processing St53 is performed without applying prediction picture filtering processing St51b. In other words, the decoding device 200 does not perform prediction picture filtering processing St51b on the inter-prediction image generated by MC processing St4, but instead generates an intra / inter mixed prediction image by mixing the inter-prediction image with the intra-prediction image. On the other hand, if prediction picture filtering processing St51b is applied to the inter-prediction image generated by MC processing St4, the decoding device 200 uses the inter-prediction image to which prediction picture filtering processing St51b has been applied as the final prediction image for generating the reconstructed image.
[0565] In such pipeline processing PL3, at the time of performing intra-frame / inter-frame mixing processing St53, the generation of the inter-frame prediction image required in the intra-frame / inter-frame mixing processing St53 has been completed in Stage 4. Therefore, in pipeline processing PL3, it is not necessary to Figure 48The inter-frame prediction processing required in the pipeline process PL1 shown in FIG5 is not completed. Therefore, the intra-frame / inter-frame mixing process St53 can be performed immediately upon completion of the intra-frame prediction process St52. As a result, the generation of unnecessary waiting time can be suppressed. In other words, the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is performed can be set to be approximately the same as the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is not performed.
[0566] in addition, Figure 52 The pipeline process PL3 shown is an example, and at least one process included in the pipeline process PL3 may be removed, another process may be added to the pipeline process PL3, or the method of dividing the stages may be changed.
[0567] In addition, the prediction image filtering process St51b is constructed so that a reconstructed image generated by adding the residual image of the block that has completed processing and the prediction image is fed back, and the above-mentioned reconstructed image is also processed in addition to being input, but is not limited to this. It can also be constructed so that the above-mentioned reconstructed image is not fed back and input, but only the prediction image in the processing object block is filtered.
[0568] In other words, in Figure 52 In Stage 5 of the pipeline processing PL3, the reconstructed image generated by adding the residual image of the processed CU and the predicted image is fed back to the predicted image filtering processing St51b. Then, in the predicted image filtering processing St51b, processing is performed using the reconstructed image and the inter-frame predicted image generated by the MC processing St4. That is, the pixels of the inter-frame predicted image are filtered using the pixels of the reconstructed image. However, the third example of the prediction processing is not limited to such feedback. That is, the reconstructed image may not be fed back to the predicted image filtering processing St51b. In this case, in the predicted image filtering processing St51b, the inter-frame predicted image generated by the MC processing St4 is processed without using the reconstructed image.
[0569] Without this feedback, Stage 5 can be divided into a stage for performing prediction image filtering St51b and a stage for performing several other processes, such as reconstruction St56, which also includes prediction image filtering St51b. However, dividing Stage 5 into two stages may increase the required memory and circuitry due to the addition of a stage. However, in pipeline processing PL3, prediction image filtering St51b is performed in the same stage as several other processes, such as reconstruction St56, thereby increasing the likelihood that an increase in memory and circuitry can be avoided.
[0570] Thus, in the third example of the prediction process, even if Figure 51 as well as Figure 52 Even when encoding multiple CUs in a single VPDU, the illustrated processing can complete all stages of the process within a pre-specified processing time. This increases the likelihood of achieving faster processing speeds while suppressing degradation in encoding performance. Furthermore, as described above, the prediction image filtering process St51b is implemented in the same stage as the reconstruction process St56 and other processes, thereby increasing the likelihood of suppressing increases in the circuit scale of the encoding device 100 and decoding device 200.
[0571] [Example 4 of Prediction Processing]
[0572] Figure 53 4 is a flowchart showing a fourth example of prediction processing performed by the decoding device 200 in this embodiment. This flowchart also shows the flow of processing on a CU in the fourth example of prediction processing.
[0573] The decoding device 200 performs steps S11 to S13, S14c, S15a, S16c, and S17a on a CU basis. Specifically, the decoding device 200 first derives an MV for the CU being processed (step S11), and then performs DMVR processing (step S12). In this DMVR processing, the decoding device 200 modifies the MV derived in step S11 by searching the surrounding area. The decoding device 200 then performs MC processing, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image for the CU being processed.
[0574] Next, the decoding device 200 determines whether to apply BIO processing to the inter-frame prediction image of the processing target CU (step S14c). Here, if it is determined to apply BIO processing (yes in step S14c), the decoding device 200 performs BIO processing on the inter-frame prediction image obtained by MC processing, thereby generating a final prediction image (step S16c). In addition, BIO processing is the above-mentioned BIO mode, also known as BDOF (bi-directional optical flow) processing. This BIO processing is a process of correcting the inter-frame prediction image based on the motion vector derived or corrected based on a model assuming constant speed linear motion. The inter-frame prediction image corrected by the BIO processing is used as the final prediction image to generate the reconstructed image.
[0575] On the other hand, if the BIO process is determined not to be applied (No in step S14c), the decoding device 200 further determines whether to apply the intra / inter mixing process (step S15a). If the intra / inter mixing process is determined to be applied (Yes in step S15a), the decoding device 200 mixes the inter-frame prediction image obtained by the MC process with the intra-frame prediction image obtained by the intra-frame prediction process (step S17a). This mixing generates an intra / inter mixed prediction image as the final prediction image. On the other hand, if the intra / inter mixing process is determined not to be applied (No in step S15a), the decoding device 200 uses the inter-frame prediction image obtained by the MC process as the final prediction image for generating the reconstructed image.
[0576] In addition, Figure 53 In the flowchart of , the decoding apparatus 200 first determines whether to apply the BIO process to the inter-frame prediction image, and then determines whether to apply the intra / inter mixed process. However, these determinations may be performed in the reverse order.
[0577] in addition, Figure 53 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0578] also, Figure 53 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 53 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0579] also, Figure 53 The determination of whether to apply BIO processing (step S14c) and the determination of whether to apply intra-frame / inter-frame mixed processing (step S15a) in the flowchart can also be performed by parsing the flag. That is, the decoding device 200 can also perform these determinations by, for example, parsing the flag described in the stream. In addition, the encoding device 100 can also perform these determinations by, for example, calculating the cost using the RD optimization model. That is, in each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process of the determination object and the cost of the predicted image obtained by not applying the process. And, the encoding device 100 performs the above determination so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0580] Figure 54This is a diagram schematically showing a fourth example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0581] Figure 54 The pipeline process PL4 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0582] This pipeline processes Stage 5 in Stage 1 to Stage 6 of PL4 and Figure 50 The pipeline process PL2 shown in Stage 5 is different. Specifically, in Stage 5 of the pipeline process PL4, the LIC process St51a is replaced with the BIO process St51c. Furthermore, the reconstructed image generated by adding the residual image and the predicted image is not fed back in the BIO process St51c.
[0583] In Stage 5 of pipeline processing PL4, if BIO processing St51c is not applied, intra / inter mixing processing St53 is performed. In other words, the decoding device 200 does not perform BIO processing St51c on the inter-prediction image generated by MC processing St4, but instead mixes the inter-prediction image with the intra-prediction image to generate an intra / inter mixed prediction image. On the other hand, if BIO processing St51c is applied to the inter-prediction image generated by MC processing St4, the decoding device 200 does not perform intra / inter mixing processing St53. Therefore, the decoding device 200 uses the inter-prediction image to which BIO processing St51c has been applied as the final prediction image for reconstructed image generation.
[0584] Specifically, in the fourth example of prediction processing, only one of the BIO processing St51c and the intra / inter mixed processing St53 is exclusively selected, and only the selected one is applied to the inter-frame prediction image generated by the MC processing St4. By applying this selected one, a final prediction image of the processing target CU is generated and used to generate a reconstructed image.
[0585] In such pipeline processing PL4, when the intra-frame / inter-frame mixing processing St53 is performed, the generation of the inter-frame prediction image required in the intra-frame / inter-frame mixing processing St53 has been completed in Stage 4. Therefore, in the pipeline processing PL4, it is not necessary to Figure 48The inter-frame prediction processing required in the pipeline process PL1 shown in FIG5 is not completed. Therefore, the intra-frame / inter-frame mixing process St53 can be performed immediately upon completion of the intra-frame prediction process St52. As a result, the generation of unnecessary waiting time can be suppressed. In other words, the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is performed can be set to be approximately the same as the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is not performed.
[0586] in addition, Figure 54 The pipeline process PL4 shown is an example, and at least one process included in the pipeline process PL4 may be removed, another process may be added to the pipeline process PL4, or the method of dividing the stages may be changed.
[0587] In addition, Figure 54 In Stage 5 of the pipeline processing PL4, the reconstructed image generated by adding the residual image of the processed CU and the predicted image is not fed back to the BIO processing St51c. Therefore, Stage 5 can be divided into a stage in which the BIO processing St51c is performed and a stage in which several processes other than the BIO processing St51c, such as the reconstruction processing St56, are performed. However, if Stage 5 is divided into two stages, the addition of the stage may increase the required memory and circuitry. However, in the pipeline processing PL4, the BIO processing St51c is performed at the same stage as the reconstruction processing St56 and other processes, thereby increasing the possibility of avoiding an increase in memory and circuitry.
[0588] Thus, in the fourth example of prediction processing, Figure 53 as well as Figure 54 The processing shown here allows all stages of processing to be completed within a pre-specified processing time, even when encoding multiple CUs within a single VPDU. This increases the likelihood of achieving faster processing while suppressing degradation in encoding performance. Furthermore, as described above, the BIO process St51c is implemented in the same stage as the reconstruction process St56 and other processes, thereby increasing the likelihood of suppressing increases in the circuit scale of the encoding device 100 and decoding device 200.
[0589] As described above, the decoding device 200 in this embodiment generates a first predicted image for the processing block based on the derived motion vector in inter-frame prediction mode. The decoding device 200 then applies an update process to the first predicted image to generate a final predicted image for the processing block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a process that blends the second predicted image generated by intra-frame prediction of the processing block with the first predicted image. Furthermore, the application of the update process exclusively applies the first and second processes. For example, when generating the final predicted image, if the update process is the first process, the decoding device 200 applies the first process to the first predicted image without applying the second process to the first predicted image to generate the final predicted image. Alternatively, if the update process is the second process, the decoding device 200 applies the second process to the first predicted image without applying the first process to the first predicted image to generate the final predicted image. Note that the first predicted image is an inter-frame predicted image, the second predicted image is an intra-frame predicted image, and the processing target block is, for example, a processing target CU.
[0590] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first and second processes is applied to the first predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for decoding a coded image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0591] Furthermore, when applying the above-described update process, the decoding device 200 determines whether to apply the second process to the first predicted image, and if it is determined that the second process is to be applied, the second process is set as the update process. On the other hand, when it is determined that the second process is not to be applied, the decoding device 200 determines whether to apply the first process to the first predicted image, and if it is determined that the first process is to be applied, the first process is set as the update process.
[0592] This increases the possibility that the update process can be appropriately selected.
[0593] In addition, the decoding device 200 exclusively performs the first process and the second process in one stage, which is included in the pipeline process for decoding the encoded image and is the same stage as the reconstruction process for generating a reconstructed image by adding the generated final prediction image and residual image.
[0594] This reduces the number of pipeline stages compared to, for example, performing the first and second processes at a stage separate from the reconstruction process, thereby increasing the likelihood that the circuit scale of decoding apparatus 200 can be suppressed from increasing.
[0595] Similar to the decoding device 200, the encoding device 100 in this embodiment generates a first predicted image for the current block in inter-frame prediction mode based on the derived motion vector. The encoding device 100 then applies an update process to the first predicted image to generate a final predicted image for the current block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a process that blends the second predicted image generated by intra-frame prediction of the current block with the first predicted image. Furthermore, the first and second processes are applied exclusively when applying the update process. For example, when generating the final predicted image, if the update process is the first process, the encoding device 100 does not apply the second process to the first predicted image, but instead applies the first process to generate the final predicted image. Alternatively, if the update process is the second process, the encoding device 100 does not apply the first process to the first predicted image, but instead applies the second process to generate the final predicted image. Note that the first predicted image is an inter-frame predicted image, the second predicted image is an intra-frame predicted image, and the processing target block is, for example, a processing target CU.
[0596] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first and second processes is applied to the first predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for encoding an image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0597] Furthermore, when applying the above-described update process, the encoding device 100 determines whether to apply the second process to the first predicted image, and if it is determined that the second process is to be applied, the second process is set as the update process. On the other hand, when it is determined that the second process is not to be applied, the encoding device 100 determines whether to apply the first process to the first predicted image, and if it is determined that the first process is to be applied, the first process is set as the update process.
[0598] This increases the possibility that the update process can be appropriately selected.
[0599] In addition, the encoding device 100 exclusively performs the first process and the second process in one stage, which is included in the pipeline processing for encoding the image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final prediction image and residual image.
[0600] This reduces the number of pipeline stages compared to, for example, performing the first and second processes at a stage separate from the reconstruction process. Consequently, it is possible to suppress an increase in the circuit scale of the encoding device 100.
[0601] [Example 5 of Prediction Processing]
[0602] Figure 55 1 is a flowchart showing a fifth example of prediction processing performed by the decoding device 200 in this embodiment. In addition, this flowchart shows the flow of processing for a CU in the fifth example of prediction processing.
[0603] The decoding device 200 performs steps S11 to S13, S14c, S15b, S16c, and S17b on a CU-by-CU basis. Specifically, the decoding device 200 first derives an MV for the CU being processed (step S11), and then performs DMVR processing (step S12). In this DMVR processing, the decoding device 200 modifies the MV derived in step S11 by searching the surroundings. The decoding device 200 then performs MC processing, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image for the CU being processed.
[0604] Next, the decoding device 200 determines whether to apply the BIO process to the inter-frame prediction image of the processing target CU (step S14c). If the BIO process is determined to be applied (yes in step S14c), the decoding device 200 performs the BIO process on the inter-frame prediction image obtained by the MC process to generate a final prediction image (step S16c). The inter-frame prediction image corrected by the BIO process is used as the final prediction image to generate the reconstructed image.
[0605] On the other hand, if it is determined that BIO processing is not to be applied (No in step S14c), the decoding device 200 further determines whether filtering processing (also called prediction image filtering processing) is to be applied (step S15b). Here, if it is determined that filtering processing is to be applied (Yes in step S15b), the decoding device 200 performs filtering processing on the inter-frame prediction image obtained by the MC processing (step S17b). This filtering processing generates a final prediction image used for generating the reconstructed image. On the other hand, if it is determined that filtering processing is not to be applied (No in step S15b), the decoding device 200 uses the inter-frame prediction image obtained by the MC processing as it is as the final prediction image for generating the reconstructed image.
[0606] Here, the filtering process can be similar to the filtering process in the third example of the prediction process described above. As mentioned above, it can also be a process of multiplying weights by multiple pixels to correct or update pixel values. Furthermore, the prediction image filtering process is particularly a process for correcting or updating pixel values included in the prediction image. Alternatively or additionally, the prediction image filtering process can also be a process of correcting or updating pixels in the inter-frame prediction image obtained through the MC process, for example.
[0607] In addition, Figure 55 In the flowchart of , the decoding apparatus 200 first determines whether to apply the BIO process to the inter-frame prediction image, and then determines whether to apply the filtering process. However, these determinations may be performed in the reverse order.
[0608] in addition, Figure 55 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0609] also, Figure 55 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 55 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0610] in addition, Figure 55The determination of whether to apply BIO processing (step S14c) and the determination of whether to apply filtering processing (step S15b) in the flowchart can also be performed by parsing flags. That is, the decoding device 200 can also perform these determinations by, for example, parsing flags described in the stream. In addition, the encoding device 100 can also perform these determinations by, for example, calculating costs using the RD optimization model. That is, in each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process of the determination object and the cost of the predicted image obtained by not applying the process. And, the encoding device 100 performs the above determination so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0611] Figure 56 This is a diagram schematically showing a fifth example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0612] Figure 56 The pipeline process PL5 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0613] This pipeline processes Stage 5 in Stage 1 to Stage 6 of PL5 and Figure 50 The pipeline process PL2 shown in Stage 5 is different. Specifically, in Stage 5 of pipeline process PL5, only one of the BIO process St51c and the prediction image filtering process St51b is applied to the inter-frame prediction image generated by the MC process St4 to generate the final prediction image. The prediction image filtering process St51b is the filtering process described above.
[0614] That is, in Stage 5 of pipeline processing PL5, BIO processing St51c is performed without applying prediction picture filtering St51b, and prediction picture filtering St51b is performed without applying BIO processing St51c. In other words, the decoding device 200 does not perform prediction picture filtering St51b on the inter-frame prediction image generated by MC processing St4, but instead applies BIO processing St51c to generate the final prediction image. Alternatively, the decoding device 200 does not perform BIO processing St51c on the inter-frame prediction image generated by MC processing St4, but instead applies prediction picture filtering St51b to generate the final prediction image.
[0615] Specifically, in the fifth example of prediction processing, only one of the BIO processing St51c and the prediction image filtering processing St51b is selected, and only the selected one of these processing is applied to the inter-frame prediction image generated by the MC processing St4. By applying this selected process, a final prediction image of the processing target CU is generated and used to generate a reconstructed image.
[0616] When the BIO processing St51c and the prediction image filtering processing St51b are applied simultaneously, the processing time or the number of processing cycles of Stage 5 increases because the two processes are performed continuously. Therefore, in this case, the processing of Stage 5 may not be completed within the specified processing time. Figure 56 In the pipeline process PL5 shown, only one of the BIO process St51c and the prediction map filtering process St51b is performed. Therefore, it is possible to suppress the increase in the processing time or the number of processing cycles of Stage 5 and complete the processing of Stage 5 within the specified processing time.
[0617] in addition, Figure 56 The pipeline process PL5 shown is an example, and at least one process included in the pipeline process PL5 may be removed, another process may be added to the pipeline process PL5, or the method of dividing the stages may be changed.
[0618] In addition, the prediction image filtering process St51b is configured to feed back a reconstructed image generated by adding the residual image of the block that has been processed and the prediction image, and the reconstructed image is also processed in addition to being input, but is not limited to this. It can also be configured not to feed back and input the reconstructed image, but to filter only the prediction image in the processing target block.
[0619] In other words, in Figure 56 In Stage 5 of the pipeline processing PL5, the reconstructed image generated by adding the residual image of the processed CU and the predicted image is fed back to the predicted image filtering processing St51b. Then, in the predicted image filtering processing St51b, processing is performed using the reconstructed image and the inter-frame predicted image generated by the MC processing St4. That is, the pixels of the inter-frame predicted image are filtered using the pixels of the reconstructed image. However, the fifth example of the prediction processing is not limited to such feedback. That is, the reconstructed image may not be fed back to the predicted image filtering processing St51b. In this case, in the predicted image filtering processing St51b, the inter-frame predicted image generated by the MC processing St4 is processed without using the reconstructed image.
[0620] Similarly, no reconstructed image feedback is performed in BIO processing St51c. Therefore, Stage 5 can be divided into a stage containing BIO processing St51c and a stage containing several processes other than BIO processing St51c, such as prediction image filtering processing St51b and reconstruction processing St56. However, if Stage 5 is divided into two stages, the additional stages may increase the required memory and circuitry. However, in pipeline processing PL5, prediction image filtering processing St51b and BIO processing St51c are performed in the same stage as several processes such as reconstruction processing St56, thereby increasing the possibility of avoiding an increase in memory and circuitry.
[0621] Thus, in the fifth example of the prediction process, Figure 55 as well as Figure 56 The processing shown here allows all stages of processing to be completed within a pre-specified processing time, even when encoding multiple CUs within a single VPDU. This increases the likelihood of achieving faster processing while suppressing degradation in encoding performance. Furthermore, as described above, the prediction image filtering process St51b and the BIO process St51c are implemented in the same stage as the reconstruction process St56 and other processes, thereby increasing the likelihood of suppressing increases in the circuit scale of the encoding device 100 and the decoding device 200.
[0622] As described above, the decoding device 200 in this embodiment generates a predicted image for the processing block based on the derived motion vector in inter-prediction mode, and applies an update process to the predicted image to generate a final predicted image for the processing block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a filtering process that updates the pixel values of the predicted image. Furthermore, the first and second processes are applied exclusively when applying the update process. For example, when generating the final predicted image, if the update process is the first process, the decoding device 200 does not apply the second process to the predicted image, but instead applies the first process to generate the final predicted image. Alternatively, if the update process is the second process, the decoding device 200 does not apply the first process to the predicted image, but instead applies the second process to generate the final predicted image. Furthermore, the processing block is, for example, the processing target CU. Furthermore, the predicted image generated in inter-prediction mode is an inter-prediction image.
[0623] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first or second process is applied to the predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for decoding a coded image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0624] Furthermore, when applying the update process, the decoding device 200 determines whether to apply the second process to the predicted image, and if it is determined that the second process is to be applied, the second process is set as the update process. On the other hand, when it is determined that the second process is not to be applied, the decoding device 200 determines whether to apply the first process to the predicted image, and if it is determined that the first process is to be applied, the first process is set as the update process.
[0625] This increases the possibility that the update process can be performed appropriately.
[0626] Furthermore, the second process updates the pixel values of the predicted image based on the predicted image. For example, the second process is performed without using the reconstructed image. Alternatively, the second process updates the pixel values of the predicted image based on (1) the predicted image and (2) the reconstructed image of a block different from the processing target block.
[0627] This increases the possibility of updating appropriate pixel values by the second processing.
[0628] In addition, the decoding device 200 exclusively performs the first process and the second process in one stage, which is included in the pipeline processing for decoding the encoded image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final prediction image and residual image.
[0629] This reduces the number of pipeline stages compared to, for example, performing the first and second processes at a stage separate from the reconstruction process, thereby increasing the likelihood that the circuit scale of decoding apparatus 200 can be suppressed from increasing.
[0630] Here, the candidate update processes may also include a third process different from the first and second processes. The third process is a process that mixes the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image generated by inter-frame prediction mode. In this case, the first, second, and third processes are exclusively applied when applying the update process. For example, the third process is the intra / inter mixing process described above.
[0631] This increases the types of update processing that can be applied, and enables, for example, the generation of a predicted image with high accuracy.
[0632] Furthermore, when applying the update process, the decoding device 200 may also (1) determine whether to apply the third process to the predicted image when the inter-frame prediction mode is an inter-frame prediction mode in which the third process can be selected. Then, when the decoding device 200 determines that the third process is to be applied, the third process is set as the update process, and when the decoding device 200 determines that the third process is not to be applied, the first process is determined to be applied to the predicted image. Furthermore, when the decoding device 200 determines that the first process is to be applied, the first process is set as the update process. On the other hand, the decoding device 200 may (2) determine whether to apply the second process to the predicted image when the inter-frame prediction mode is an inter-frame prediction mode in which the third process cannot be selected. Then, when the decoding device 200 determines that the second process is to be applied, the second process is set as the update process, and when the decoding device 200 determines that the second process is not to be applied, the first process is determined to be applied to the predicted image. Furthermore, when the decoding device 200 determines that the first process is to be applied, the first process is set as the update process.
[0633] This increases the possibility that the update process can be performed appropriately.
[0634] In addition, the decoding device 200 exclusively performs the first process, the second process, and the third process in one stage, which is included in the pipeline process for decoding the encoded image and is the same stage as the reconstruction process for generating a reconstructed image by adding the generated final prediction image and residual image.
[0635] This reduces the number of pipeline stages compared to, for example, performing the first, second, and third processes at a stage separate from the reconstruction process. Consequently, it is possible to suppress an increase in the circuit scale of decoding apparatus 200.
[0636] The decoding apparatus 200 also performs filtering processing to update pixel values included in a reconstructed image obtained by adding the generated final prediction image and the residual signal. For example, this filtering processing is loop filtering such as deblocking filtering.
[0637] This makes it possible to remove distortion and the like of the reconstructed image.
[0638] Similar to the decoding device 200, the encoding device 100 of this embodiment generates a predicted image for the processing block in inter-frame prediction mode based on the derived motion vector and applies an update process to the predicted image to generate a final predicted image for the processing block. Candidates for the update process include a first process and a second process. The first process is a BDOF (bi-directional optical flow) process, and the second process is a filtering process that updates the pixel values of the predicted image. Furthermore, when applying the update process, the first and second processes are applied exclusively. For example, when generating the final predicted image, if the update process is the first process, the encoding device 100 does not apply the second process to the predicted image, but instead applies the first process to generate the final predicted image. Alternatively, if the update process is the second process, the encoding device 100 does not apply the first process to the predicted image, but instead applies the second process to generate the final predicted image. Furthermore, if the update process is the second process, the encoding device 100 does not apply the first process to the predicted image, but instead applies the second process to generate the final predicted image. Note that the processing block is, for example, the processing target CU. The predicted image generated in inter-frame prediction mode is an inter-frame predicted image.
[0639] Thus, even if the first process is an update process and the second process is also an update process, no process other than the update process in the first or second process is applied to the predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for encoding an image includes both the first and second processes in a single stage, performing both the first and second processes together can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0640] Furthermore, when applying the update process, the encoding device 100 determines whether to apply the second process to the predicted image, and if it is determined that the second process is to be applied, the second process is set as the update process. On the other hand, when it is determined that the second process is not to be applied, the encoding device 100 determines whether to apply the first process to the predicted image, and if it is determined that the first process is to be applied, the first process is set as the update process.
[0641] This increases the possibility that the update process can be performed appropriately.
[0642] Furthermore, the second process updates the pixel values of the predicted image based on the predicted image. For example, the second process is performed without using the reconstructed image. Alternatively, the second process updates the pixel values of the predicted image based on (1) the predicted image and (2) the reconstructed image of a block different from the processing target block.
[0643] This increases the possibility of updating appropriate pixel values by the second processing.
[0644] In addition, the encoding device 100 exclusively performs the first process and the second process in one stage, which is included in the pipeline processing for encoding the image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final prediction image and residual image.
[0645] This reduces the number of pipeline stages compared to, for example, performing the first and second processes at a stage separate from the reconstruction process. Consequently, it is possible to suppress an increase in the circuit scale of the encoding device 100.
[0646] Here, the candidate update processes may also include a third process different from the first and second processes. The third process is a process that mixes the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image generated by inter-frame prediction mode. In this case, the first, second, and third processes are exclusively applied when applying the update process. For example, the third process may be an intra / inter mixed process.
[0647] This increases the types of update processing that can be applied, and enables, for example, the generation of a predicted image with high accuracy.
[0648] Furthermore, the encoding device 100 may also, in applying the update process, (1) determine whether to apply the third process to the predicted image when the inter-frame prediction mode is an inter-frame prediction mode in which the third process can be selected. Furthermore, if the encoding device 100 determines that the third process is to be applied, the third process is set as the update process, and if the encoding device 100 determines that the third process is not to be applied, the first process is determined to be applied to the predicted image. Furthermore, if the encoding device 100 determines that the first process is to be applied, the first process is set as the update process. Alternatively, the encoding device 100 may (2) determine whether to apply the second process to the predicted image when the inter-frame prediction mode is an inter-frame prediction mode in which the third process cannot be selected. Then, if the encoding device 100 determines that the second process is to be applied, the second process is set as the update process, and if the encoding device 100 determines that the second process is not to be applied, the first process is determined to be applied to the predicted image. Furthermore, if the encoding device 100 determines that the first process is to be applied, the first process is set as the update process.
[0649] This increases the possibility that the update process can be performed appropriately.
[0650] In addition, the encoding device 100 exclusively performs the first process, the second process, and the third process in one stage, which is included in the pipeline processing for encoding the image and is the same stage as the reconstruction processing for generating a reconstructed image by adding the generated final prediction image and residual image.
[0651] This reduces the number of stages included in pipeline processing, compared to, for example, performing the first, second, and third processes at a stage separate from the reconstruction process. Consequently, it is possible to suppress an increase in the circuit scale of encoding device 100.
[0652] Furthermore, the encoding apparatus 100 performs filtering processing to update pixel values included in a reconstructed image obtained by adding the generated final prediction image and the residual signal. For example, this filtering processing is loop filtering such as deblocking filtering.
[0653] This makes it possible to remove distortion and the like of the reconstructed image.
[0654] [Example 6 of Prediction Processing]
[0655] Figure 57 1 is a flowchart showing a sixth example of prediction processing performed by the decoding device 200 in this embodiment. This flowchart also shows the flow of processing on a CU in the sixth example of prediction processing.
[0656] The decoding device 200 performs steps S11 to S13, S14d, S15a, S16d, and S17a on a CU basis. Specifically, the decoding device 200 first derives an MV for the CU being processed (step S11), and then performs DMVR processing (step S12). In this DMVR processing, the decoding device 200 modifies the MV derived in step S11 by searching the surroundings. The decoding device 200 then performs MC processing, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image for the CU being processed.
[0657] Next, the decoding device 200 determines whether to apply the OBMC process to the inter-frame prediction image of the processing target CU (step S14d). If it is determined that the OBMC process is to be applied (yes in step S14d), the decoding device 200 performs the OBMC process on the inter-frame prediction image obtained by the MC process to generate a final prediction image (step S16d).
[0658] On the other hand, if the OBMC process is determined not to be applied (No in step S14d), the decoding device 200 further determines whether to apply intra / inter mixing (step S15a). If the intra / inter mixing process is determined to be applied (Yes in step S15a), the decoding device 200 mixes the inter-frame prediction image obtained by the OBMC process with the intra-frame prediction image obtained by the intra prediction process (step S17a). This mixing generates a final prediction image. On the other hand, if the intra / inter mixing process is determined not to be applied (No in step S15a), the decoding device 200 uses the inter-frame prediction image obtained by the OBMC process as the final prediction image for generating the reconstructed image.
[0659] In addition, Figure 57 In the flowchart of , the decoding apparatus 200 first determines whether to apply the OBMC process to the inter-frame prediction image, and then determines whether to apply the intra / inter mixed process. However, these determinations may be performed in the reverse order.
[0660] in addition, Figure 57 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0661] also, Figure 57 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 57 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0662] also, Figure 57 The determination of whether to apply the OBMC process (step S14d) and the determination of whether to apply the intra-frame / inter-frame mixed process (step S15a) in the flowchart can also be performed by parsing the flag. That is, the decoding device 200 can also perform these determinations by, for example, parsing the flag described in the stream. In addition, the encoding device 100 can also perform these determinations by, for example, calculating the cost using the RD optimization model. That is, in each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process of the determination object and the cost of the predicted image obtained by not applying the process. And, the encoding device 100 performs the above determination so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0663] Figure 58This is a diagram schematically showing a sixth example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0664] Figure 58 The pipeline process PL6 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0665] This pipeline processes Stage 5 in Stage 1 to Stage 6 of PL6 and Figure 50 The pipeline process PL2 in Stage 5 shown is different. Specifically, in the pipeline process PL6 in Stage 5, the OBMC process St51d is performed instead of the LIC process St51a. Furthermore, in the OBMC process St51d, the reconstructed image generated by adding the residual image and the predicted image is not fed back.
[0666] In Stage 5 of pipeline processing PL6, intra / inter mixing processing St53 is performed without applying OBMC processing St51d. In other words, the decoding device 200 does not perform OBMC processing St51d on the inter-prediction image generated by MC processing St4, but instead mixes the inter-prediction image with the intra-prediction image to generate an intra / inter mixed prediction image. On the other hand, if the decoding device 200 applies OBMC processing St51d to the inter-prediction image generated by MC processing St4, the inter-prediction image to which OBMC processing St51d has been applied is used directly as the final prediction image for generating the reconstructed image.
[0667] In such pipeline processing PL6, when the intra-frame / inter-frame mixing processing St53 is performed, the generation of the inter-frame prediction image required in the intra-frame / inter-frame mixing processing St53 has been completed in Stage 4. Therefore, in the pipeline processing PL6, it is not necessary to Figure 48 The inter-frame prediction processing required in the pipeline process PL1 shown in FIG5 is not completed. Therefore, the intra-frame / inter-frame mixing process St53 can be performed immediately upon completion of the intra-frame prediction process St52. As a result, the generation of unnecessary waiting time can be suppressed. In other words, the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is performed can be set to be approximately the same as the processing time of Stage 5 when the intra-frame / inter-frame mixing process St53 is not performed.
[0668] in addition, Figure 58 The pipeline process PL6 shown is an example, and at least one process included in the pipeline process PL6 may be removed, another process may be added to the pipeline process PL6, or the method of dividing the stages may be changed.
[0669] In addition, Figure 58 In Stage 5 of pipeline processing PL6, the reconstructed image generated by adding the residual image of the processed CU to the predicted image is not fed back to OBMC processing St51d. Therefore, Stage 5 can be divided into a stage where OBMC processing St51d is performed and a stage where several processes, such as reconstruction processing St56, are performed in addition to OBMC processing St51d. However, if Stage 5 is divided into two stages, the additional stages may increase the required memory and circuitry. However, in pipeline processing PL6, OBMC processing St51d is performed at the same stage as several processes, such as reconstruction processing St56, thereby increasing the possibility of avoiding an increase in memory and circuitry.
[0670] Thus, in the sixth example of the prediction process, Figure 57 as well as Figure 58 The processing shown here allows all stages of processing to be completed within a pre-specified processing time, even when encoding multiple CUs within a single VPDU. This increases the likelihood of achieving faster processing speeds while suppressing degradation in encoding performance. Furthermore, as described above, the OBMC process St51d is implemented in the same stage as the reconstruction process St56 and other processes, thereby increasing the likelihood of suppressing increases in the circuit scale of the encoding device 100 and the decoding device 200.
[0671] [Example 7 of Prediction Processing]
[0672] Figure 59 2 is a flowchart showing a seventh example of prediction processing performed by the decoding device 200 in this embodiment. In addition, this flowchart shows the flow of processing for a CU in the seventh example of prediction processing.
[0673] The decoding device 200 performs steps S11 to S13, S14d, S15b, S16d, and S17b on a CU-by-CU basis. Specifically, the decoding device 200 first derives an MV for the CU being processed (step S11), and then performs DMVR processing (step S12). In this DMVR processing, the decoding device 200 modifies the MV derived in step S11 by searching the surroundings. The decoding device 200 then performs MC processing, or motion compensation, using the modified MV (step S13). This results in an inter-frame predicted image for the CU being processed.
[0674] Next, the decoding device 200 determines whether to apply the OBMC process to the inter-frame prediction image of the processing target CU (step S14d). If it is determined that the OBMC process is to be applied (yes in step S14d), the decoding device 200 performs the OBMC process on the inter-frame prediction image obtained by the MC process to generate a final prediction image (step S16d).
[0675] On the other hand, if it is determined that OBMC processing is not to be applied (No in step S14d), the decoding device 200 further determines whether filtering processing (also called prediction image filtering processing) should be applied (step S15b). Here, if it is determined that filtering processing should be applied (Yes in step S15b), the decoding device 200 performs filtering processing on the inter-frame prediction image obtained through MC processing (step S17b). This filtering processing generates a final prediction image. On the other hand, if it is determined that filtering processing is not to be applied (No in step S15b), the decoding device 200 uses the inter-frame prediction image obtained through MC processing as it is, as is, the final prediction image for generating the reconstructed image.
[0676] In addition, Figure 59 In the flowchart of , the decoding apparatus 200 first determines whether to apply the OBMC process to the inter-frame prediction image and then determines whether to apply the filtering process. However, these determinations may be performed in the reverse order.
[0677] in addition, Figure 59 The flowchart is an example, and at least one step included in the flowchart may be removed, or other processing or condition determination steps may be added to the flowchart.
[0678] also, Figure 59 The process of the flowchart shown can also be performed by the encoding device 100. That is, the encoding device 100 and the decoding device 200 perform the same prediction process, and there is a difference between encoding an image into a stream and decoding the stream into an image using the prediction process. Figure 59 The flow of the prediction process shown in the flowchart is also basically common in the encoding device 100.
[0679] in addition, Figure 59The determination of whether to apply OBMC processing (step S14d) and the determination of whether to apply filtering processing (step S15b) in the flowchart can also be performed by parsing the flag. That is, the decoding device 200 can also perform these determinations by, for example, parsing the flag described in the stream. In addition, the encoding device 100 can also perform these determinations by, for example, calculating the cost using the RD optimization model. That is, in each of the above determinations, the encoding device 100 calculates the cost of the predicted image obtained by applying the process of the determination object and the cost of the predicted image obtained by not applying the process. And, the encoding device 100 performs the above determination so that the predicted image corresponding to the smaller of these costs is used as the final predicted image.
[0680] Figure 60 This is a diagram schematically showing a seventh example of prediction processing performed by the decoding device 200 in this embodiment as a configuration example of pipeline processing.
[0681] Figure 60 The pipeline process PL7 shown includes Stages 1 to 6. The processes of each stage after Stage 2 are performed in units of VPDU, for example.
[0682] This pipeline processes Stage 5 of Stage 1 to Stage 6 of PL7 and Figure 50 The pipeline process PL2 shown in Stage 5 is different. Specifically, in Stage 5 of pipeline process PL7, only one of the OBMC process St51d and the prediction image filtering process St51b is applied to the inter-frame prediction image generated by the MC process St4 to generate the final prediction image. The prediction image filtering process St51b is the filtering process described above.
[0683] Specifically, in Stage 5 of pipeline processing PL7, OBMC processing St51d is performed without applying prediction image filtering St51b, and prediction image filtering St51b is performed without applying OBMC processing St51d. In other words, the decoding device 200 generates the final prediction image by applying OBMC processing St51d to the inter-frame prediction image generated by MC processing St4, without performing prediction image filtering St51b. Alternatively, the decoding device 200 generates the final prediction image by applying prediction image filtering St51b to the inter-frame prediction image generated by MC processing St4, without performing OBMC processing St51d.
[0684] When the OBMC process St51d and the prediction image filtering process St51b are applied simultaneously, the processing time or the number of processing cycles of Stage 5 increases because the two processes are performed continuously. Therefore, in this case, the processing of Stage 5 may not be completed within the specified processing time. Figure 60 In the pipeline process PL7 shown, only one of the OBMC process St51d and the prediction map filtering process St51b is performed. Therefore, it is possible to suppress the increase in the processing time or the number of processing cycles of Stage 5 and complete the processing of Stage 5 within the specified processing time.
[0685] in addition, Figure 60 The pipeline process PL7 shown is an example, and at least one process included in the pipeline process PL7 may be removed, another process may be added to the pipeline process PL7, or the method of dividing the stages may be changed.
[0686] In addition, Figure 60 In Stage 5 of the pipeline processing PL7, the reconstructed image generated by adding the residual image of the processed CU and the predicted image is fed back to the predicted image filtering processing St51b. Then, in the predicted image filtering processing St51b, processing is performed using the reconstructed image and the inter-frame predicted image generated by the MC processing St4. That is, the pixels around the boundary of the inter-frame predicted image are filtered using the pixels of the reconstructed image. However, the seventh example of the prediction processing is not limited to such feedback. That is, the reconstructed image may not be fed back to the predicted image filtering processing St51b. In this case, in the predicted image filtering processing St51b, the inter-frame predicted image generated by the MC processing St4 is processed without using the reconstructed image.
[0687] Similarly, OBMC processing St51d does not provide feedback on the reconstructed image. Therefore, Stage 5 can be divided into a stage containing OBMC processing St51d and a stage containing several processes other than OBMC processing St51d, such as prediction image filtering St51b and reconstruction processing St56. However, dividing Stage 5 into two stages may increase the required memory and circuitry due to the addition of a stage. However, in pipeline processing PL5, prediction image filtering St51b and OBMC processing St51d are performed in the same stage as reconstruction processing St56 and other processes, thereby increasing the likelihood that an increase in memory and circuitry can be avoided.
[0688] Thus, in the seventh example of the prediction process, Figure 59 as well as Figure 60The processing shown here allows all stages of processing to be completed within a pre-specified processing time, even when encoding multiple CUs within a single VPDU. This increases the likelihood of achieving faster processing speeds while suppressing degradation in encoding performance. Furthermore, as described above, since the prediction image filtering process St51b and the OBMC process St51d are implemented in the same stage as the reconstruction process St56 and other processes, it increases the likelihood of suppressing increases in the circuit scale of the encoding device 100 and the decoding device 200.
[0689] [Modification]
[0690] Furthermore, multiple examples of the prediction processing in Examples 1, 2, 3, 4, 5, 6, and 7 of this embodiment may be combined. For example, the five processes including LIC, BIO, OBMC, prediction image filtering, and intra / inter mixed processing may be implemented in the same stage as the reconstruction process. Furthermore, only one of these five processes may be selected and applied to the inter-frame prediction image. Alternatively, at least two of these five processes may be implemented in the same stage as the reconstruction process. Furthermore, only one of these at least two processes may be selected and applied to the inter-frame prediction image. Furthermore, while only one of the five processes is selected in the above example, only one of six or more processes may be selected and applied to the inter-frame prediction image. In this case, other processes other than LIC, BIO, OBMC, prediction image filtering, and intra / inter mixed processing may also be included in the six or more processes described above. Furthermore, the processes that become candidates for selection are not limited to the LIC process, the BIO process, the OBMC process, the prediction image filtering process, and the intra / inter mixed process, but may be any process.
[0691] [Representative examples of structure and processing]
[0692] Representative examples of the configuration and processing of the encoding apparatus 100 and decoding apparatus 200 described above are described below.
[0693] Figure 61 1 is a flowchart showing an example of encoding and decoding processing. For example, the encoding device 100 and the decoding device 200 each include a circuit and a memory connected to the circuit. The circuit and memory included in the encoding device 100 may also be connected to Figure 40 The processor a1 and memory a2 shown in FIG. 1 correspond to each other, and the circuits and memories included in the decoding device 200 may also correspond to each other. Figure 46 The processor b1 and the memory b2 shown correspond to each other. The circuits of the encoding device 100 and the decoding device 200 perform the following operations.
[0694] Specifically, the circuit operates to perform the processes of steps S1 and S2a as in the fourth example of prediction processing. More specifically, the circuit first generates a first predicted image of the processing target block based on the derived motion vector in inter prediction mode (step S1).
[0695] Next, the circuit generates a final predicted image of the processing object block by applying an update process to the first predicted image (step S2a). Here, candidates for the update process include the first process and the second process. In addition, the first process is the BDOF (bi-directional optical flow) process, that is, the above-mentioned BIO process. In addition, the second process is the above-mentioned intra-frame / inter-frame mixing process. That is, the second process is a process of mixing the second predicted image generated in the intra-frame prediction of the processing object block with the above-mentioned first predicted image. Moreover, in the application of the update process, the first process and the second process are applied exclusively. That is, when the update process is the first process, the circuit does not apply the second process to the first predicted image, but applies the first process, thereby generating the final predicted image. On the other hand, when the update process is the second process, the circuit does not apply the first process to the first predicted image, but applies the second process, thereby generating the final predicted image.
[0696] Thus, even if the first process is an update process and the second process is also an update process, processes other than the update process in the first and second processes are not applied to the inter-frame predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for encoding an image includes the first and second processes in a single stage, performing the first and second processes simultaneously can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0697] Figure 62 1 is a flowchart showing another example of encoding and decoding processing. For example, the encoding device 100 and the decoding device 200 each include a circuit and a memory connected to the circuit. The circuit and memory included in the encoding device 100 may also be connected to Figure 40 The processor a1 and memory a2 shown in FIG. 1 correspond to each other, and the circuits and memories included in the decoding device 200 may also correspond to each other. Figure 46 The processor b1 and the memory b2 shown correspond to each other. The circuits of the encoding device 100 and the decoding device 200 perform the following operations.
[0698] Specifically, the circuit operates to perform the processes of step S1 and step S2b as in the fifth example of prediction processing. More specifically, the circuit first generates a predicted image of the processing target block based on the derived motion vector in the inter prediction mode (step S1).
[0699] Next, the circuit generates a final predicted image of the processing object block by applying an update process to the predicted image (step S2b). Here, candidates for the update process include the first process and the second process. In addition, the first process is a BDOF (bi-directional optical flow) process, that is, the above-mentioned BIO process. In addition, the second process is a filtering process for updating the pixel values of the predicted image. Moreover, in the application of the update process, the first process and the second process are applied exclusively. That is, when the update process is the first process, the circuit does not apply the second process to the predicted image, and generates the final predicted image by applying the first process. On the other hand, when the update process is the second process, the circuit does not apply the first process to the predicted image, and generates the final predicted image by applying the second process.
[0700] Thus, even if the first process is an update process and the second process is also an update process, processes other than the update process in the first and second processes are not applied to the inter-frame predicted image. In other words, the application of processes other than the update process is prohibited. Therefore, for example, if a pipeline process for encoding an image includes the first and second processes in a single stage, performing the first and second processes simultaneously can increase the likelihood of suppressing an increase in the processing time required for that stage. In other words, it can increase the likelihood of suppressing an increase in processing time.
[0701] [Other examples]
[0702] The encoding device 100 and decoding device 200 in each of the above examples can be used as an image encoding device and an image decoding device, respectively, or can be used as a moving picture encoding device and a moving picture decoding device, respectively.
[0703] Furthermore, the encoding device 100 and the decoding device 200 may process (more specifically, encode or decode) only a portion of the plurality of processing elements, while other devices may process (more specifically, encode or decode) other processing elements. Furthermore, the encoding device 100 and the decoding device 200 may include only a portion of the plurality of components, while other devices may include the other components.
[0704] In addition, at least a part of each of the above examples can be used as an encoding method, a decoding method, a prediction method, or other methods.
[0705] In addition, each component is composed of dedicated hardware, but can also be implemented by executing a software program suitable for each component. Each component can be implemented by a program execution unit such as a CPU or a processor reading and executing the software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0706] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to processor a1 or b1, and the storage device corresponds to memory a2 or b2.
[0707] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a storage device. In addition, when the processing circuit includes a program execution unit, the storage device stores a software program executed by the program execution unit.
[0708] Here, the software for realizing the above-mentioned encoding device 100 or decoding device 200 is to execute a computer Figures 47 to 62 The processing procedure is shown.
[0709] In addition, as described above, each component may also be a circuit. These circuits may constitute a single circuit as a whole, or they may be different circuits. In addition, each component may be implemented by a general-purpose processor or a dedicated processor.
[0710] Furthermore, the processing performed by a specific component may be performed by another component. Furthermore, the order in which the processing is performed may be changed, and multiple processing may be performed simultaneously. Furthermore, the encoding and decoding apparatus may include the encoding apparatus 100 and the decoding apparatus 200.
[0711] In addition, the ordinal numbers such as 1 and 2 used in the description may be replaced as appropriate. In addition, new ordinal numbers may be assigned to components, or ordinal numbers may be removed.
[0712] While the embodiments of the encoding device 100 and the decoding device 200 have been described above based on a plurality of examples, the embodiments of the encoding device 100 and the decoding device 200 are not limited to these examples. The embodiments of the encoding device 100 and the decoding device 200 may also include various modifications that can be conceived by those skilled in the art to the respective examples, or embodiments constructed by combining components from different examples, without departing from the spirit of the present invention.
[0713] It is also possible to implement by combining at least a portion of one or more embodiments disclosed herein with at least a portion of other embodiments of the present invention. In addition, it is also possible to implement by combining a portion of the processing, a portion of the structure of the device, a portion of the grammar, etc. described in the flowchart of one or more embodiments disclosed herein with other embodiments.
[0714] [Implementation and Application]
[0715] In each of the above embodiments, each functional block or active block can generally be implemented by an MPU (microprocessing unit) and a memory. In addition, the processing of each functional block can also be implemented by a program execution unit such as a processor that reads and executes software (programs) recorded in a recording medium such as a ROM. The software can be distributed. The software can also be recorded in various recording media such as semiconductor memories. In addition, each functional block can also be implemented by hardware (dedicated circuit). Various combinations of hardware and software can be used.
[0716] The processing described in each embodiment can be implemented by centralized processing using a single device (system) or by distributed processing using multiple devices. In addition, the processors that execute the above programs can be single or multiple. In other words, centralized processing can be performed or distributed processing can be performed.
[0717] The aspects of the present invention are not limited to the above-described embodiments, and various modifications are possible, which are also included in the scope of the aspects of the present invention.
[0718] Furthermore, here, application examples of the moving picture encoding method (image encoding method) or moving picture decoding method (image decoding method) described in each of the above embodiments and various systems implementing these application examples are described. Such a system may be characterized by including an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding and decoding device including both. Other configurations of such a system may be modified as appropriate depending on the circumstances.
[0719] [Example of use]
[0720] Figure 63 This diagram shows the overall structure of a content supply system ex100 for implementing content distribution services. The communication service provision area is divided into cells of desired sizes, and in the illustrated example, base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.
[0721] In the content delivery system ex100, various devices, such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115, are connected to the Internet ex101 via an Internet service provider ex102, a communication network ex104, and base stations ex106-ex110. The content delivery system ex100 may also connect some of these devices in combination. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network, short-range wireless, or other means, rather than via base stations ex106-ex110. Furthermore, the streaming server ex103 may be connected to various devices, such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115, via the Internet ex101. Furthermore, the streaming server ex103 may be connected to a terminal, such as a hotspot within an airplane ex117, via a satellite ex116.
[0722] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly to the aircraft ex117 without going through the satellite ex116.
[0723] The camera ex113 is a device such as a digital camera capable of capturing both still and moving images. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handy-phone System) compatible with mobile communication systems such as 2G, 3G, 3.9G, 4G, and what will be called 5G in the future.
[0724] The home appliance ex114 is a refrigerator or equipment included in a household fuel cell cogeneration system.
[0725] In the content delivery system ex100, terminals with camera functions are connected to the streaming server ex103 via a base station ex106 or the like, enabling on-site distribution and the like. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) can perform the encoding processing described in the above embodiments on still images or moving image content captured by users using these terminals. Furthermore, the terminals can multiplex the encoded video data with the audio data obtained by encoding the corresponding audio, and transmit the resulting data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present invention.
[0726] Meanwhile, the streaming server ex103 streams content data sent by requesting clients. Clients are computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or terminals inside airplanes ex117, all capable of decoding the encoded data. Each device that receives the distributed data can also decode and reproduce the data. In other words, each device can function as an image decoding device according to one aspect of the present invention.
[0727] [Distributed Processing]
[0728] Alternatively, the streaming server ex103 can consist of multiple servers or computers, distributing data by distributing processing or recording. For example, the streaming server ex103 can be implemented as a CDN (Content Delivery Network), which distributes content through a network connecting numerous edge servers distributed worldwide. In a CDN, physically close edge servers can be dynamically assigned to clients. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Furthermore, in the event of various errors or changes in communication status due to increased traffic, processing can be distributed across multiple edge servers, distribution can be switched to other edge servers, or delivery can be continued by bypassing a faulty portion of the network, thus achieving high-speed and stable delivery.
[0729] In addition, the encoding process of the captured data is not limited to the distributed processing itself, and can be performed by each terminal, on the server side, or shared. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or encoding amount of the image of the frame or scene unit is detected. In addition, in the second cycle, a process is performed to maintain the image quality and improve the encoding efficiency. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, the quality and efficiency of the content can be improved while reducing the processing load in each terminal. In this case, if there is a request for almost real-time reception and decoding, the data completed by the first encoding by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be achieved.
[0730] As another example, cameras ex113 and others extract features (features or feature quantities) from images, compress the feature data as metadata, and transmit it to a server. The server, for example, determines the importance of an object based on the feature data and switches the quantization precision, performing compression appropriate to the meaning (or content) of the image. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during recompression in the server. Alternatively, the terminal can perform simple encoding such as VLC (Variable Length Coding), while the server performs more processing-intensive encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding).
[0731] As another example, in stadiums, shopping malls, factories, and other locations, there may be multiple video data sets generated by capturing roughly the same scene using multiple terminals. In such cases, encoding can be distributed using the multiple terminals that captured the images, as well as other terminals and servers that did not capture the images as needed, for example, by allocating the encoding processing to each GOP (Group of Picture) unit, picture unit, or tile unit obtained by dividing the picture. This reduces latency and achieves better real-time performance.
[0732] Because multiple image data sets represent roughly the same scene, the server can manage and / or instruct the image data captured by each terminal to cross-reference each other. Furthermore, the server can receive encoded data from each terminal and change the reference relationship between the multiple data sets, or modify or replace the image itself before re-encoding it. This allows the generation of a stream with improved quality and efficiency for each data set.
[0733] Furthermore, the server may also perform transcoding to change the encoding method of the video data before distributing the video data. For example, the server may convert the encoding method of the MPEG type to the VP type (such as VP9), or convert H.264 to H.265.
[0734] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions of "server" or "terminal" refer to the entity performing the processing. However, some or all of the processing performed by the server can also be performed by the terminal, and some or all of the processing performed by the terminal can also be performed by the server. Furthermore, the same applies to the decoding process.
[0735] [3D, multi-angle]
[0736] There is an increasing trend for combining and utilizing images or videos captured by multiple devices, such as cameras ex113 and / or smartphones ex115, that are roughly synchronized with each other, to capture different scenes or the same scene from different angles. The images captured by each device can be combined based on the relative positional relationship between the devices, or based on areas with consistent feature points contained in the images.
[0737] The server not only encodes two-dimensional moving images but can also encode still images automatically or at user-specified times based on scene analysis of moving images and transmit them to the receiving terminal. Furthermore, if the server can determine the relative positional relationship between the capturing terminals, it can generate not only two-dimensional moving images but also three-dimensional shapes of the scene based on images of the same scene captured from different angles. The server can also separately encode three-dimensional data generated from point clouds, etc., or, based on the results of identifying or tracking people or objects using three-dimensional data, select or reconstruct images captured by multiple terminals to generate images for transmission to the receiving terminal.
[0738] This allows users to arbitrarily select the images corresponding to each camera terminal to enjoy the scene, or to enjoy content that extracts images from a selected viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, along with the video, audio can be collected from multiple angles. The server multiplexes the audio from a specific angle or space with the corresponding video and transmits the multiplexed video and audio.
[0739] Furthermore, content that connects the real and virtual worlds, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular in recent years. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. These can be encoded using techniques such as Multi-View Coding (MVC) to allow for reference between viewpoint images, or encoded as separate streams without reference to each other. When these separate streams are decoded, they can be played back in sync with the user's viewpoint, recreating a virtual three-dimensional space.
[0740] In the case of AR images, the server can also overlay virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or movement of the user's viewpoint. The decoding device obtains or stores the virtual object information and three-dimensional data, generates a two-dimensional image based on the movement of the user's viewpoint, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can also send the user's viewpoint movement to the server in addition to the request for virtual object information. Alternatively, the server can create overlay data based on the three-dimensional data stored on the server, matching the received viewpoint movement, encode the overlay data, and distribute it to the decoding device. In addition, the overlay data typically has an alpha value indicating transparency in addition to RGB. The server sets the alpha value of the portion other than the target generated based on the three-dimensional data to 0, for example, and encodes the portion in a transparent state. Alternatively, the server can set the RGB value of a specified value as the background, as in a chroma key, and generate data with the portion other than the target as the background color. The specified RGB value can also be predetermined.
[0741] Similarly, the decoding process of the distributed data can be performed by the client (for example, the terminal), can be performed on the server side, or can be shared and performed. As an example, a terminal may first send a reception request to the server, and other terminals may receive the content corresponding to the request and perform decoding processing, and send the decoded signal to a device with a display. By distributing the processing regardless of the performance of the communicative terminal itself and selecting appropriate content, data with better image quality can be reproduced. In addition, as another example, large-size image data can also be received by a TV, etc., and a portion of the image, such as tiles, can be decoded and displayed by the viewer's personal terminal. In this way, while sharing the overall image, it is possible to confirm one's own area of responsibility or the area that one wants to confirm in more detail at hand.
[0742] In situations where multiple short-range, medium-range, or long-range wireless communications can be used indoors and outdoors, it may be possible to seamlessly receive content using distribution system standards such as MPEG-DASH. Users can also freely select their own terminals, decoding devices such as displays installed indoors and outdoors, and switch in real time. In addition, it is possible to switch the decoding terminal and the display terminal and perform decoding using their own location information. As a result, it is also possible to map and display information on a part of the wall or ground of a building next to a display device while the user is moving to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as caching the encoded data in a server that can be accessed from the receiving terminal in a short time, or copying the encoded data in the edge server of the content distribution service.
[0743] [Scalable Coding]
[0744] To switch content, use Figure 64 The example illustrates a scalable stream compressed and encoded using the moving picture coding method described in the above embodiments. For the server, multiple streams with the same content but different qualities can be provided as a single stream. Alternatively, a structure can be employed to switch content by leveraging the temporally and spatially scalable nature of streams achieved through layered coding, as shown in the figure. Specifically, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. This allows the decoding side to freely switch between low-resolution and high-resolution content. For example, if a user wants to watch a video viewed on a smartphone ex115 while on the go, then watch it later on a device such as an internet TV at home, the device can simply decode the same stream into different layers, reducing the burden on the server.
[0745] Furthermore, in addition to the hierarchical structure of encoding pictures per layer and implementing enhancement layers above the base layer as described above, the enhancement layers may also include metadata such as statistical information based on the image. Alternatively, the decoding side may generate high-definition content by super-resolutioning the base layer pictures based on this metadata. Super-resolution can improve the signal-to-noise ratio while maintaining and / or increasing resolution. Meta-information includes information for determining linear or nonlinear filter coefficients used in super-resolution processing, as well as information for determining parameter values in filtering, machine learning, or least-squares operations used in super-resolution processing.
[0746] Alternatively, a structure can be provided that divides a picture into tiles according to the meaning of the object in the image. The decoding side decodes only a part of the area by selecting the tile to be decoded. Moreover, by storing the attributes of the object (people, cars, balls, etc.) and the position in the image (coordinate position in the same image, etc.) as meta information, the decoding side can determine the position of the desired object based on the meta information and decide the tile that includes the object. For example, Figure 65 As shown, SEI (supplemental enhancement information) messages in HEVC, which are different from pixel data, can also be used to store meta-information. This meta-information indicates, for example, the position, size, or color of the main object.
[0747] Meta-information can also be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in the image, and by matching this information with the time information of the picture unit, it can determine the picture in which the target appears and the location of the target within the picture.
[0748] [Web page optimization]
[0749] Figure 66 This is a diagram showing an example of a display screen of a web page on the computer ex111 or the like. Figure 67 1 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. Figure 66 and Figure 67 As shown, a web page may contain multiple link images that serve as links to image content. The display format of these images may vary depending on the viewing device. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture associated with each content as a link image until the user explicitly selects a link image, or until the link image approaches the center of the screen, or until the entire link image enters the screen. Alternatively, the display device (decoding device) may display an image such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer, decode the image, and display it.
[0750] When a linked image is selected by the user, the display device, for example, sets the base layer as the top priority and decodes it. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device can also decode to the enhancement layer. Moreover, in order to ensure real-time performance, before selection or when the communication bandwidth is very tight, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of decoding to the start of display of the content) by decoding and displaying only the forward reference pictures (I pictures, P pictures, and B pictures that are only forward referenced). Furthermore, the display device can also forcibly ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward references, and perform normal decoding as the number of pictures received increases over time.
[0751] [Automatic driving]
[0752] Furthermore, when transmitting and receiving still images or video data, such as two-dimensional or three-dimensional map information, for autonomous driving or driving assistance, the receiving terminal may receive weather or construction information as metadata in addition to image data belonging to one or more layers, and decode these metadata by associating them with each other. Furthermore, the metadata may belong to a layer or be multiplexed solely with the image data.
[0753] In this case, since the receiving terminal, such as a car, drone, or airplane, is moving, the receiving terminal transmits its location information, enabling seamless reception and decoding while switching between base stations ex106-ex110. Furthermore, the receiving terminal can dynamically switch the level of metadata received and the level of map information updated based on user preferences, user status, and / or communication band conditions.
[0754] In the content providing system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0755] [Distribution of Personal Content]
[0756] Furthermore, the content delivery system ex100 can deliver not only high-quality, long-duration content provided by video distributors, but also low-quality, short-duration content provided by individuals, either unicast or multicast. Such personal content is expected to increase in the future. To enhance personal content, the server can also perform encoding after editing. This can be achieved, for example, with the following configuration.
[0757] After taking the photos in real time or accumulating them, the server performs recognition processing such as shooting errors, scene search, meaning analysis, and target detection based on the original image data or encoded data. In addition, based on the recognition results, the server manually or automatically corrects focus deviation or hand shaking, deletes less important scenes such as scenes with lower brightness than other pictures or scenes that are not in focus, emphasizes the edges of the target, or changes the color tone. The server encodes the edited data based on the editing results. In addition, it is known that the viewing rate will decrease if the shooting time is too long. The server can also automatically limit not only the less important scenes as mentioned above, but also scenes with less movement based on the image processing results, so that the content is within a specific time range. Alternatively, the server can also generate a summary based on the results of the scene's meaning analysis and encode it.
[0758] There are cases where personal content in its original state may be written with content that infringes copyright, author's personality rights or portrait rights, etc., or there are cases where the scope of sharing exceeds the desired scope, which is inconvenient for individuals. Therefore, for example, the server can also forcibly change the faces of people in the peripheral part of the screen, or the home, etc. to an out-of-focus image for encoding. In addition, the server can also identify whether the face of a person different from the pre-registered person is captured in the image to be encoded, and if so, perform processing such as applying mosaics to the face part. Alternatively, as pre-processing or post-processing for encoding, the user can also specify the person or background area that he wants to process the image from the perspective of copyright, etc. The server can also replace the specified area with another image, or blur the focus, etc. If it is a person, it can track the person in the moving image and replace the image of the person's face.
[0759] The viewing of personal content with a small amount of data has a strong demand for real-time performance, so although it also depends on the bandwidth, the decoding device can also receive, decode, and reproduce the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the playback is repeated more than twice, such as in the case of looped playback, the enhancement layer is also included in the playback of high-definition images. In this way, if the stream is scalable, it can provide an experience in which the moving image is relatively rough when it is not selected or at the beginning of viewing, but the stream gradually becomes smoother and the image becomes better. In addition to scalable coding, the same experience can be provided when the relatively rough stream played back the first time and the second stream encoded with reference to the moving image of the first time are composed of a single stream.
[0760] [Other application examples]
[0761] In addition, these encoding and decoding processes are usually processed in the LSI ex500 of each terminal. Figure 63 ) can be a single chip or a multi-chip configuration. Alternatively, video encoding or decoding software can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) readable by the computer ex111 or the like, and encoding and decoding can be performed using this software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera can be transmitted. In this case, the video data can be encoded by the LSI ex500 in the smartphone ex115.
[0762] Alternatively, the LSIex500 can be configured to download and activate application software. In this case, the terminal first determines whether it is compatible with the content's encoding scheme or has the capability to perform a specific service. If the terminal is not compatible with the content's encoding scheme or does not have the capability to perform a specific service, it can download the codec or application software to retrieve and play the content.
[0763] Furthermore, the content delivery system ex100 is not limited to the content delivery system ex100 via the Internet ex101; at least one of the video encoding devices (image encoding devices) or video decoding devices (image decoding devices) described in the above-mentioned embodiments can also be incorporated into a digital broadcasting system. Since multiplexed data containing multiplexed video and audio is transmitted and received over broadcast radio waves using satellites, the content delivery system ex100 differs from the unicast-friendly structure of the content delivery system ex100 in that it is suitable for multicast. However, the encoding and decoding processes can be applied in the same manner.
[0764] [Hardware structure]
[0765] Figure 68 Is a further detailed representation Figure 63 FIG. 1 shows a diagram of the smartphone ex115. Figure 69 This diagram shows an example of the configuration of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of capturing both video and still images, and a display unit ex458 that displays images captured by the camera unit ex465 and decoded data such as images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio and sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded or decoded data such as captured videos or still images, recorded audio, received videos or still images, and emails, and a slot unit ex464 that serves as an interface with a SIM card ex468 for identifying users and authenticating access to various data, including the network. Alternatively, an external memory card can be used in place of the memory unit ex467.
[0766] The main control unit ex460, which can perform integrated control of the display unit ex458 and the operation unit ex466, is synchronously connected to the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the sound signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via a bus ex470.
[0767] When the user turns on the power button, the power supply circuit unit ex461 activates the smartphone ex115 to be operational and supplies power to various components from the battery pack.
[0768] The smartphone ex115 performs processes such as calls and data communications under the control of the main control unit ex460, which includes a CPU, ROM, and RAM. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. The signal is then subjected to spread spectrum processing by the modulation / demodulation unit ex452. The transmission / reception unit ex451 then performs digital-to-analog conversion and frequency conversion, and the resulting signal is transmitted via the antenna ex450. Furthermore, received data is amplified, subjected to frequency conversion and analog-to-digital conversion, and then subjected to inverse spread spectrum processing by the modulation / demodulation unit ex452. The audio signal processing unit ex454 converts the signal into an analog audio signal, which is then output from the audio output unit ex457. During data communications, text, still images, or video data can be transmitted via the operation input control unit ex462 under the control of the main control unit ex460, based on operations on the main unit's operation unit ex466. Similar transmission and reception processes are performed. In data communication mode, when transmitting video, still images, or both video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and then sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of capturing video or still images by the camera unit ex465, and then sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and audio data in a predetermined format. The modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451 perform modulation and conversion processing, and then transmit the data via the antenna ex450. The predetermined format can also be predetermined.
[0769] When receiving a video file attached to an email or chat tool, or a video file linked to a webpage, the multiplexing / demultiplexing unit ex453 demultiplexes the multiplexed data received via antenna ex450 into a bitstream of video data and a bitstream of audio data. The multiplexing / demultiplexing unit ex453 then supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronous bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in the above embodiments. The video or still image contained in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs the audio through the audio output unit ex457. As live streaming becomes increasingly common, audio reproduction may become socially inappropriate depending on the user's circumstances. Therefore, it is preferable that as an initial value, a configuration is adopted in which only the video data is reproduced without reproducing the audio signal, and the audio is reproduced in synchronization only when the user performs an operation such as clicking on the video data.
[0770] While the smartphone ex115 is used as an example, other possible terminal configurations include transmitting and receiving terminals with both an encoder and a decoder, as well as transmitting terminals with only an encoder and receiving terminals with only a decoder. In the digital broadcasting system, the description assumes the reception and transmission of multiplexed data containing audio data multiplexed with video data. However, multiplexed data can also contain text data associated with the video in addition to audio data. Furthermore, it is also possible to receive or transmit video data itself, rather than multiplexed data.
[0771] While the main control unit ex460, which includes a CPU, has been described as controlling the encoding and decoding processes, many terminals also include GPUs. Therefore, a configuration can also be implemented where a memory shared by the CPU and GPU, or a memory whose addresses are managed in a mutually usable manner, allows the GPU to process larger areas simultaneously. This can shorten encoding time, ensure real-time performance, and achieve low latency. In particular, it is more efficient if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing are performed simultaneously on a per-picture basis, such as on the GPU, rather than on the CPU.
[0772] Industrial applicability
[0773] The present invention can be utilized in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conferencing system, or an electronic mirror.
[0774] Description of Reference Numerals
[0775] 100 Encoding device
[0776] 102 Division
[0777] 104 Subtraction Department
[0778] 106 Transformation Unit
[0779] 108 Quantitative Department
[0780] 110 Entropy Coding Unit
[0781] 112, 204 Inverse Quantization Unit
[0782] 114, 206 Inverse transformation unit
[0783] 116, 208 Addition Department
[0784] 118, 210 block memory
[0785] 120, 212 loop filter unit
[0786] 122, 214 frame memories
[0787] 124, 216 Intra-frame prediction unit
[0788] 126, 218 Inter-frame prediction unit
[0789] 128, 220 Prediction and Control Department
[0790] 200 Decoding Device
[0791] 202 Entropy Decoding Unit
[0792] 1201 Boundary Judgment Department
[0793] 1202, 1204, 1206 switches
[0794] 1203 Filtering and Judgment Unit
[0795] 1205 Filter Processing Unit
[0796] 1207 Filter Characteristics Determination Unit
[0797] 1208 Processing and Judgment Unit
[0798] a1, b1 processors
[0799] a2, b2 memory
Claims
1. An encoding device, wherein: have: circuits; and A memory, connected to the above circuit, The above circuit is in action. In the inter-frame prediction mode, a first predicted image of the processing target block is generated based on the derived motion vector. By applying the update process to the first predicted image, a final predicted image of the processing target block is generated. The candidates for the update process include the first process and the second process. The first processing mentioned above is bidirectional optical flow processing, namely BDOF processing. The second process is a process of mixing the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image. In application of the update process, the first process and the second process are exclusively applied.
2. The encoding device according to claim 1, wherein The above circuit is, In the application of the above update process, Determine whether to apply the second process to the first predicted image. If it is determined that the second process is to be applied, the second process is set as the update process. If it is determined that the second process is not to be applied, it is determined whether the first process is to be applied to the first predicted image. When it is determined that the first process is to be applied, the first process is set as the update process.
3. The encoding device according to claim 1 or 2, wherein: The above-mentioned circuit exclusively performs the above-mentioned first processing and the above-mentioned second processing in one stage, and the above-mentioned one stage is included in the pipeline processing for encoding the image, and is the same stage as the reconstruction processing of generating a reconstructed image by adding the generated final predicted image and residual image.
4. A decoding device, wherein: have: circuits; and A memory, connected to the above circuit, The above circuit is in action. In the inter-frame prediction mode, a first predicted image of the processing target block is generated based on the derived motion vector. By applying the update process to the first predicted image, a final predicted image of the processing target block is generated. The candidates for the update process include the first process and the second process. The first processing mentioned above is bidirectional optical flow processing, namely BDOF processing. The second process is a process of mixing the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image. In application of the update process, the first process and the second process are exclusively applied.
5. The decoding device according to claim 4, wherein: The above circuit is, In the application of the above update process, Determine whether to apply the second process to the first predicted image. If it is determined that the second process is to be applied, the second process is set as the update process. If it is determined that the second process is not to be applied, it is determined whether the first process is to be applied to the first predicted image. When it is determined that the first process is to be applied, the first process is set as the update process.
6. The decoding device according to claim 4 or 5, wherein: The above-mentioned circuit exclusively performs the above-mentioned first processing and the above-mentioned second processing in one stage, and the above-mentioned one stage is included in the pipeline processing for decoding the encoded image, and is the same stage as the reconstruction processing of generating a reconstructed image by adding the generated final predicted image and residual image.
7. A coding method, wherein: In the inter-frame prediction mode, a first predicted image of the processing target block is generated based on the derived motion vector. By applying the update process to the first predicted image, a final predicted image of the processing target block is generated. The candidates for the update process include the first process and the second process. The first processing mentioned above is bidirectional optical flow processing, namely BDOF processing. The second process is a process of mixing the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image. In application of the update process, the first process and the second process are exclusively applied.
8. A decoding method, wherein: In the inter-frame prediction mode, a first predicted image of the processing target block is generated based on the derived motion vector. By applying the update process to the first predicted image, a final predicted image of the processing target block is generated. The candidates for the update process include the first process and the second process. The first processing mentioned above is bidirectional optical flow processing, namely BDOF processing. The second process is a process of mixing the second predicted image generated by intra-frame prediction of the processing target block with the first predicted image. In application of the update process, the first process and the second process are exclusively applied.