Encoding device, decoding device, and storage medium
By adjusting the motion vector in motion image encoding and decoding, the problems of low processing efficiency and large circuit scale in the prior art are solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202510466685.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-27
- Filing Date
- 2019-08-09
- Publication Date
- 2025-06-06
AI Technical Summary
When encoding moving images, the prior art has low processing efficiency, large circuit scale, and large encoding/decoding processing volume, which affects image quality and equipment performance.
Using an encoding device and a decoding device, the motion vector is adjusted to reduce deviation by using a reference motion vector and a differential motion vector in the prediction of processing object blocks, thereby improving coding efficiency and reducing memory access.
Effectively improve processing efficiency, reduce circuit scale and encoding/decoding processing volume, and improve image quality and equipment performance.
Smart Images

Figure CN120111250A_ABST
Abstract
Description
[0001] This application is a division of an invention patent application with an application date of August 9, 2019, application number 201980055826.9, and invention name “Encoding device, decoding device, encoding method and decoding method”. Technical Field
[0002] The present invention relates to an encoding device and the like for encoding a moving image. Background Art
[0003] Conventionally, there is a standard for encoding moving images, called HEVC (High Efficiency Video Coding) H.265 (Non-Patent Document 1).
[0004] Prior art literature
[0005] Non-patent literature
[0006] Non-patent document 1: H.265 (ISO / IEC 23008-2HEVC) / HEVC (High Efficiency Video Coding) Summary of the invention
[0007] Problems to be solved by the invention
[0008] In such encoding methods and decoding methods, it is desired to propose new methods in order to improve processing efficiency, improve image quality, reduce circuit scale, etc.
[0009] The structures or methods disclosed in the embodiments or parts thereof of the present invention can contribute to at least any one of the following aspects, for example, improvement of coding efficiency, reduction of coding / decoding processing volume, reduction of circuit scale, improvement of coding / decoding speed, appropriate selection of components / actions such as filters, blocks, sizes, motion vectors, reference pictures and reference blocks in encoding and decoding, etc.
[0010] In addition, the present invention also includes the disclosure of structures or methods that can provide benefits other than those described above, for example, structures or methods that improve coding efficiency while suppressing an increase in the amount of processing.
[0011] Means for solving problems
[0012] A coding device according to one embodiment of the present invention is a coding device for coding a moving image, the coding device comprising a circuit and a memory connected to the circuit, wherein the prediction mode of a current block to be coded is an affine mode, and in operation, the circuit: obtains the current block from a coding tree unit CTU; derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on a difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; if it is determined that the motion vector difference is greater than the threshold, sets a first value to a second motion vector, and if it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value to the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and encodes the current block using the second motion vector, wherein the threshold is different when the current block is unidirectionally predicted and when the current block is bidirectionally predicted.
[0013] A decoding device according to one aspect of the present invention is a decoding device for decoding a moving image, the decoding device comprising a circuit and a memory connected to the circuit, wherein the prediction mode of a current block to be decoded is an affine mode, and in operation, the circuit: obtains the current block from a coding tree unit CTU; derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on a difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; if it is determined that the motion vector difference is greater than the threshold, sets a first value to a second motion vector, and if it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value to the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and decodes the current block using the second motion vector, wherein the threshold is different when the current block is unidirectionally predicted and when the current block is bidirectionally predicted.
[0014] A computer-readable non-transitory storage medium according to one aspect of the present invention stores a bit stream, wherein a prediction mode of a current block to be decoded is an affine mode, the bit stream includes a coded signal and syntax information, and a decoding device performs, based on the coded signal and the syntax information: obtaining the current block from a coding tree unit (CTU); deriving a reference motion vector for predicting the current block; deriving a first motion vector different from the reference motion vector; deriving a motion vector difference based on a difference between the reference motion vector and the first motion vector; determining whether the motion vector difference is greater than a threshold; setting a first value to a second motion vector if it is determined that the motion vector difference is greater than the threshold, and setting a second value different from the first value to the second motion vector if it is determined that the motion vector difference is not greater than the threshold, the second motion vector being different from the reference motion vector and the first motion vector; and decoding the current block using the second motion vector, wherein the threshold is different when the current block is unidirectionally predicted and when the current block is bidirectionally predicted.
[0015] In addition, these inclusive or specific forms may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-transitory recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0016] Further benefits and advantages provided by the disclosed embodiments will become clear from the description and drawings. These benefits and advantages are sometimes brought about separately according to the features of various embodiments, descriptions and drawings, and it is not necessary to provide all of them in order to obtain more than one benefit or advantage.
[0017] Effects of the Invention
[0018] The present invention can provide an encoding device, a decoding device, an encoding method or a decoding method that can improve processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a block diagram showing the functional structure of the encoding device according to the embodiment.
[0020] Figure 2 This is a flowchart showing an example of the overall encoding process performed by the encoding device.
[0021] Figure 3 This is a diagram showing an example of block division.
[0022] Figure 4A This is a diagram showing an example of the structure of a slice.
[0023] Figure 4BThis is a diagram showing an example of a tile structure.
[0024] Figure 5A This is a table showing the transformation basis functions corresponding to each transformation type.
[0025] Figure 5B It is a diagram representing SVT (Spatially Varying Transform).
[0026] Fig. 6A This is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).
[0027] Figure 6B FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0028] Figure 6C FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0029] Figure 7 This is a block diagram showing an example of a detailed configuration of a loop filter unit that functions as a DBF.
[0030] Figure 8 A diagram showing an example of a deblocking filter having a filter characteristic that is symmetric with respect to a block boundary.
[0031] Fig. 9 This is a diagram for explaining a block boundary on which a deblocking filtering process is performed.
[0032] Fig.10 This is a diagram showing an example of the Bs value.
[0033] Fig.11 This is a diagram showing an example of processing performed by a prediction processing unit of an encoding device.
[0034] Fig.12 This is a diagram showing another example of processing performed by the prediction processing unit of the encoding device.
[0035] Fig.13 This is a diagram showing another example of processing performed by the prediction processing unit of the encoding device.
[0036] Fig.14 This is a diagram showing an example of 67 intra prediction modes in intra prediction.
[0037] Fig.15 This is a flowchart showing the flow of basic processing of inter-frame prediction.
[0038] Fig.16 This is a flowchart showing an example of motion vector derivation.
[0039] Fig.17 This is a flowchart showing another example of motion vector derivation.
[0040] Fig.18 This is a flowchart showing another example of motion vector derivation.
[0041] Fig.19 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.
[0042] Fig. 20 is a flowchart showing an example of inter-frame prediction based on merge mode.
[0043] Fig.21 This is a diagram for explaining an example of motion vector derivation processing based on the merge mode.
[0044] Fig. 22 This is a flowchart showing an example of FRUC (frame rate up conversion).
[0045] Fig.23 This is a diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0046] Fig.24 This is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0047] Fig.25A This is a diagram for explaining an example of derivation of a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks.
[0048] Fig.25B This is a diagram for explaining an example of derivation of a motion vector in sub-block units in an affine mode having three control points.
[0049] Fig.26A This is a conceptual diagram for explaining the affine merge mode.
[0050] Fig.26B A conceptual diagram for explaining the affine merge mode with two control points.
[0051] Fig.26C This is a conceptual diagram for explaining the affine merge mode with three control points.
[0052] Fig. 27 This is a flowchart showing an example of processing in the affine merge mode.
[0053] Fig.28A A diagram for explaining an affine inter-frame mode having two control points.
[0054] Fig.28B A diagram for explaining an affine inter-frame mode having three control points.
[0055] Fig.29 This is a flowchart showing an example of processing in the affine inter mode.
[0056] Fig. 30A It is a diagram for explaining the affine inter mode in which the current block has 3 control points and the adjacent block has 2 control points.
[0057] Fig. 30B It is a diagram for explaining the affine inter mode in which the current block has 2 control points and the adjacent block has 3 control points.
[0058] Fig.31A It is a diagram showing the relationship between merge mode and DMVR (dynamic motion vector refreshing).
[0059] Fig.31B This is a conceptual diagram for explaining an example of DMVR processing.
[0060] Fig.32 This is a flowchart showing an example of generating a predicted image.
[0061] Fig.33 This is a flowchart showing another example of generating a predicted image.
[0062] Fig.34 This is a flowchart showing another example of generating a predicted image.
[0063] Fig.35 This is a flowchart for explaining an example of a predicted image correction process based on an OBMC (overlapped block motion compensation) process.
[0064] Fig.36 This is a conceptual diagram for explaining an example of predicted image correction processing based on OBMC processing.
[0065] Fig.37 This is a diagram for explaining the generation of predicted images of two triangles.
[0066] Fig.38 This is a diagram for explaining a model assuming uniform linear motion.
[0067] Fig.39This is a diagram for explaining an example of a method of generating a predicted image using a brightness correction process based on a LIC (local illumination compensation) process.
[0068] Fig.40 This is a block diagram showing an installation example of an encoding device.
[0069] Fig.41 It is a block diagram showing the functional structure of a decoding device according to an embodiment.
[0070] Fig.42 This is a flowchart showing an example of the overall decoding process performed by the decoding device.
[0071] Fig.43 This is a diagram showing an example of processing performed by a prediction processing unit of a decoding device.
[0072] Fig.44 This is a diagram showing another example of processing performed by the prediction processing unit of the decoding device.
[0073] Fig.45 This is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in a decoding device.
[0074] Fig.46 This is a block diagram showing an implementation example of a decoding device.
[0075] Fig.47 This is a flowchart showing an example of the inter-frame prediction process in the first aspect.
[0076] Fig.48 This is a diagram showing an example of a processing target block.
[0077] Fig.49 It is a diagram showing an example of calculated threshold values.
[0078] Fig.50 It is an overall structural diagram of the content supply system that realizes content distribution services.
[0079] Fig.51 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0080] Fig.52 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0081] Fig.53 This is a diagram showing an example of a display screen of a web page.
[0082] Fig.54 This is a diagram showing an example of a display screen of a web page.
[0083] Fig.55 The diagram shows an example of a smart phone.
[0084] Fig.56 This is a block diagram showing a structural example of a smart phone. DETAILED DESCRIPTION
[0085] (Foundation that forms the basis of the present invention)
[0086] For example, a coding device for coding a moving image suppresses an increase in the amount of processing in the coding of the moving image, and when coding a moving image that has been subjected to a more fragmented prediction process, subtracts the prediction image from the image constituting the moving image, thereby deriving a prediction error. Furthermore, the coding device performs frequency transformation and quantization on the prediction error, and encodes the result as image data. At this time, when the motion prediction process is performed on the motion of the coding object unit such as a block included in the moving image in units of blocks or sub-blocks constituting the block, there is a possibility that the coding efficiency can be improved by controlling the deviation of the motion vector.
[0087] However, if the deviation of motion vectors is not controlled in encoding of blocks included in moving images, the amount of processing increases and the encoding efficiency decreases.
[0088] Therefore, one form of an encoding device of the present invention is an encoding device for encoding a moving image, comprising a circuit and a memory connected to the circuit, wherein the circuit derives a reference motion vector used in prediction of a processing object block during operation, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold, changes the first motion vector if it is determined that the differential motion vector is greater than the threshold, does not change the first motion vector if it is determined that the differential motion vector is not greater than the threshold, and encodes the processing object block using the changed first motion vector or the unchanged first motion vector.
[0089] Thus, the encoding device can adjust the motion vector so that the deviations of the multiple motion vectors in the processing target block divided into multiple sub-blocks converge within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referenced becomes smaller, so the amount of data transferred from the memory can be suppressed within the specified range. Therefore, the encoding device can reduce the amount of memory access in the inter-frame prediction process, so the encoding efficiency is improved.
[0090] For example, the reference motion vector may correspond to a first pixel set in the processing target block, and the first motion vector may correspond to a second pixel set in the processing target block that is different from the first pixel set.
[0091] Thus, the encoding device can appropriately control the deviation of a plurality of motion vectors in the processing target block using the reference motion vector determined based on the first pixel set and the first motion vector determined based on the second pixel set.
[0092] For example, when the circuit determines that the difference motion vector is larger than the threshold value, the circuit may change the first motion vector using a value obtained by clipping the difference motion vector.
[0093] Thus, the encoding device can change the first motion vector so that the deviation of the first motion vector falls within a predetermined range. Therefore, the encoding device can reduce the amount of memory access in the inter-frame prediction process, thereby improving encoding efficiency.
[0094] For example, the prediction mode of the processing target block may be an affine mode.
[0095] This allows the encoding device to predict the processing target block in sub-block units, thereby improving prediction accuracy.
[0096] For example, the threshold may be determined so that the worst-case memory access amount when the processing object block is predicted using the prediction mode is less than the worst-case memory access amount when the processing object block is predicted using a prediction mode other than the prediction mode.
[0097] Thus, the encoding device can smoothly execute the transfer process by limiting the memory access amount for transferring data from the storage area within a predetermined range, thereby improving the processing efficiency.
[0098] For example, the worst-case memory access amount when prediction processing is performed in a prediction mode other than the prediction mode may be the memory access amount when bidirectional prediction processing is performed on the processing object block for each 8×8 pixels in a prediction mode other than the prediction mode.
[0099] At this time, the encoding device generates a predicted image from a past picture in units of blocks when performing prediction processing without using the affine mode, for example. When using a prediction mode other than the affine mode and performing bidirectional prediction in units of 8×8 pixel blocks, the amount of memory accessed is the largest. Therefore, in the affine mode, the threshold of the affine mode is determined so as to converge within the memory access amount. In this way, the encoding device can smoothly perform the transfer processing by limiting the memory access amount for transferring data from the storage area to a predetermined range, thereby improving the processing efficiency.
[0100] For example, the threshold value may be different between a case where the processing target block is unidirectionally predicted and a case where the processing target block is bidirectionally predicted.
[0101] Thus, the number of accesses when the encoding device performs bidirectional prediction on the processing target block increases compared to the case where the processing target block is unidirectionally predicted, so the threshold value can be set to be smaller in the case of bidirectional prediction than in the case of unidirectional prediction. In this way, the more the number of accesses, the smaller the threshold value is set, and the coding efficiency is improved.
[0102] For example, the circuit may determine whether the differential motion vector is greater than the threshold, and determine that the differential motion vector is greater than the threshold when the absolute value of the horizontal component of the differential motion vector is greater than the first value of the threshold or the absolute value of the vertical component of the differential motion vector is greater than the second value of the threshold; and determine that the differential motion vector is not greater than the threshold when the absolute value of the horizontal component of the differential motion vector is not greater than the first value of the threshold and the absolute value of the vertical component of the differential motion vector is not greater than the second value of the threshold.
[0103] Thus, the threshold is represented by a first value as a horizontal component (hereinafter, also referred to as a first threshold) and a second value as a vertical component (hereinafter, also referred to as a second threshold), so the encoding device can determine the deviation of the first motion vector in two dimensions. Furthermore, the encoding device determines whether the deviation of the first motion vector in the processing target block is within a predetermined range, so the first motion vector in the processing target block can be adjusted to an appropriate value. Therefore, the encoding efficiency is improved.
[0104] For example, when the circuit determines that the differential motion vector is greater than the threshold, the circuit may change the first motion vector using a value obtained by limiting the differential motion vector and the reference motion vector, so that the differential motion vector between the changed first motion vector and the reference motion vector is not greater than the threshold.
[0105] At this time, when the deviation of the first motion vector exceeds a predetermined range, the coding device, for example, clips the absolute value of the component exceeding the threshold value in the horizontal component and the vertical component of the differential motion vector to the value of the threshold value. In addition, the coding device may, for example, add the reference motion vector to the value after clipping the differential motion vector as the changed first motion vector. Thus, the deviation of the first motion vector in the processing target block is brought within the predetermined range. Therefore, the coding efficiency is improved.
[0106] For example, the reference motion vector may be an average of a plurality of the first motion vectors in the processing target block.
[0107] Thus, the encoding device can derive, for each of the plurality of first motion vectors in the processing target block, a difference from a reference using the average of all the first motion vectors in the processing target block as a reference.
[0108] For example, the reference motion vector may be one of the plurality of first motion vectors in the processing target block.
[0109] Thus, the encoding device can derive, for each of the plurality of first motion vectors in the processing target block, a difference from the reference using one first motion vector in the processing target block as a reference.
[0110] For example, the threshold may correspond to the size of the processing target block.
[0111] Thus, the memory access amount of the encoding device varies depending on the size of the processing object block, so the threshold value may be reduced as the memory access amount increases. For example, the larger the size of the processing object block, the smaller the threshold value. Therefore, even if the size of the processing object block increases, the encoding device can converge the memory access amount within a specified range, thereby improving the encoding efficiency.
[0112] For example, the threshold may be predetermined and the stream may not be decoded.
[0113] This allows the encoding device to use a predetermined threshold value, and therefore does not need to encode the threshold value in each prediction process, thereby improving encoding efficiency.
[0114] In addition, a decoding device in one form of the present invention is a decoding device for decoding a moving image, comprising a circuit and a memory connected to the circuit, wherein the circuit derives a reference motion vector used in prediction of a processing object block during operation, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold, changes the first motion vector if it is determined that the differential motion vector is greater than the threshold, does not change the first motion vector if it is determined that the differential motion vector is not greater than the threshold, and decodes the processing object block using the changed first motion vector or the unchanged first motion vector.
[0115] Thus, the decoding device can adjust the motion vector so that the deviation of multiple motion vectors in the processing target block divided into multiple sub-blocks converges within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referenced becomes smaller, so the amount of data transferred from the memory can be suppressed within the specified range. Therefore, since the decoding device can reduce the amount of memory access in the inter-frame prediction process, the processing efficiency is improved.
[0116] For example, the reference motion vector may correspond to a first pixel set in the processing target block, and the first motion vector may correspond to a second pixel set in the processing target block that is different from the first pixel set.
[0117] Thus, the decoding device can appropriately control the deviation of a plurality of motion vectors in the processing target block using the reference motion vector determined based on the first pixel set and the first motion vector determined based on the second pixel set.
[0118] For example, when the circuit determines that the difference motion vector is larger than the threshold value, the circuit may change the first motion vector using a value obtained by clipping the difference motion vector.
[0119] Thus, the decoding device can change the first motion vector so that the deviation of the first motion vector falls within a predetermined range. Therefore, the decoding device can reduce the amount of memory access in the inter-frame prediction process, thereby improving the processing efficiency.
[0120] For example, the prediction mode of the processing target block may be an affine mode.
[0121] This allows the decoding device to predict the processing target block in sub-block units, thereby improving prediction accuracy.
[0122] For example, the threshold may be determined so that the worst-case memory access amount when the processing object block is predicted using the prediction mode is less than the worst-case memory access amount when the processing object block is predicted using a prediction mode other than the prediction mode.
[0123] Thus, the decoding device can smoothly execute the transfer process by limiting the memory access amount for transferring data from the storage area within a predetermined range, thereby improving the processing efficiency.
[0124] For example, the worst-case memory access amount when prediction processing is performed in a prediction mode other than the prediction mode may be the memory access amount when bidirectional prediction processing is performed on the processing object block for each 8×8 pixels in a prediction mode other than the prediction mode.
[0125] At this time, for example, when the prediction processing is not performed using the affine mode, the decoding device generates a predicted image from the past picture in units of blocks. When a prediction mode other than the affine mode is used to perform bidirectional prediction in units of 8×8 pixel blocks, the amount of memory accessed is the largest. Therefore, in the affine mode, the threshold of the affine mode is determined so as to converge within the memory access amount. Therefore, the decoding device can smoothly perform the transfer processing by limiting the memory access amount for transferring data from the storage area to a predetermined range, thereby improving the processing efficiency.
[0126] For example, the threshold value may be different between a case where the processing target block is unidirectionally predicted and a case where the processing target block is bidirectionally predicted.
[0127] Thus, the number of accesses when the encoding device performs bidirectional prediction on the processing target block increases compared to the case where the processing target block is unidirectionally predicted, so the threshold value can be set to be smaller in the case of bidirectional prediction than in the case of unidirectional prediction. In this way, the processing efficiency is improved by setting the threshold value smaller as the number of accesses increases.
[0128] For example, the circuit may determine whether the differential motion vector is greater than the threshold, and determine that the differential motion vector is greater than the threshold when the absolute value of the horizontal component of the differential motion vector is greater than the first value of the threshold or the absolute value of the vertical component of the differential motion vector is greater than the second value of the threshold; and determine that the differential motion vector is not greater than the threshold when the absolute value of the horizontal component of the differential motion vector is not greater than the first value of the threshold and the absolute value of the vertical component of the differential motion vector is not greater than the second value of the threshold.
[0129] Thus, the threshold is represented by a first value as a horizontal component (hereinafter, also referred to as a first threshold) and a second value as a vertical component (hereinafter, also referred to as a second threshold), so that the decoding device can determine the deviation of the first motion vector in two dimensions. Furthermore, the decoding device determines whether the deviation of the first motion vector in the processing target block is within a predetermined range, and thus can adjust the first motion vector in the processing target block to an appropriate value. Therefore, the processing efficiency is improved.
[0130] For example, when the circuit determines that the differential motion vector is greater than the threshold, the circuit may change the first motion vector using a value obtained by limiting the differential motion vector and the reference motion vector, so that the differential motion vector between the changed first motion vector and the reference motion vector is not greater than the threshold.
[0131] At this time, when the deviation of the first motion vector exceeds a predetermined range, the decoding device, for example, clips the absolute value of the component exceeding the threshold value in the horizontal component and the vertical component of the differential motion vector to the value of the threshold value. Furthermore, the decoding device may, for example, add the reference motion vector to the value after clipping the differential motion vector, and obtain the value as the changed first motion vector. Thus, the deviation of the first motion vector in the processing target block is brought within the predetermined range. Therefore, the processing efficiency is improved.
[0132] For example, the reference motion vector may be an average of a plurality of the first motion vectors of the processing target block.
[0133] Thus, the decoding device can derive, for each of the plurality of first motion vectors in the processing target block, a difference from a reference using the average of all the first motion vectors in the processing target block as a reference.
[0134] For example, the reference motion vector may be one of the plurality of first motion vectors of the processing target block.
[0135] Thus, the decoding device can derive, for each of the plurality of first motion vectors in the processing target block, a difference from the reference using one first motion vector in the processing target block as a reference.
[0136] For example, the threshold may correspond to the size of the processing target block.
[0137] Thus, the decoding device may have different memory access amounts depending on the size of the processing target block, so the threshold value may be smaller as the memory access amount increases. For example, the larger the size of the processing target block, the smaller the threshold value. Therefore, even if the size of the processing target block increases, the decoding device can keep the memory access amount within a predetermined range, thereby improving processing efficiency.
[0138] For example, the threshold may be predetermined and the data may not be decoded from the stream.
[0139] This allows the decoding device to use a predetermined threshold value, and therefore does not need to decode the threshold value every time a prediction process is performed, thereby improving processing efficiency.
[0140] In addition, one form of a coding method of the present invention is a coding method for encoding a moving image, deriving a reference motion vector used in predicting a processing object block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on the difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector if it is determined that the differential motion vector is greater than the threshold value, and not changing the first motion vector if it is determined that the differential motion vector is not greater than the threshold value, and encoding the processing object block using the changed first motion vector or the unchanged first motion vector.
[0141] Thus, the encoding method can adjust the motion vector so that the deviations of multiple motion vectors in the processing target block divided into multiple sub-blocks converge within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referenced becomes smaller, so the amount of data transferred from the memory can be suppressed within the specified range. Therefore, the encoding method can reduce the amount of memory access in the inter-frame prediction process, so the encoding efficiency is improved.
[0142] In addition, a decoding method according to one embodiment of the present invention is a decoding method for decoding a moving image, deriving a reference motion vector used in predicting a processing object block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on a difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector if it is determined that the differential motion vector is greater than the threshold value, and not changing the first motion vector if it is determined that the differential motion vector is not greater than the threshold value, and decoding the processing object block using the changed first motion vector or the unchanged first motion vector.
[0143] Thus, the decoding method can adjust the motion vector so that the deviations of the plurality of motion vectors in the processing target block divided into the plurality of sub-blocks converge within a predetermined range. Therefore, compared with the case where the deviation of the motion vector is larger than the predetermined range, the number of pixels to be referenced is reduced, so the amount of data transferred from the memory can be suppressed within the predetermined range. Therefore, since the decoding method can reduce the amount of memory access in the inter-frame prediction process, the processing efficiency is improved.
[0144] Moreover, these inclusive or specific forms may also be implemented by systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0145] The following embodiments are described in detail with reference to the accompanying drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, configuration positions and connection forms of components, steps, relationships and sequences of steps, etc. shown in the following embodiments are examples and are not intended to limit the claims.
[0146] The following describes implementations of a coding device and a decoding device. The implementations are examples of coding devices and decoding devices to which the processing and / or structures described in each form of the present invention can be applied. The processing and / or structures can also be implemented in coding devices and decoding devices different from the implementations. For example, with respect to the processing and / or structures applied to the implementations, for example, one of the following can also be performed.
[0147] (1) Any of the multiple components of the encoding device or decoding device of the embodiment described in each aspect of the present invention may be replaced by another component described in any of the aspects of the present invention, or these components may be combined;
[0148] (2) In the coding device or decoding device of the embodiment, any changes such as addition, replacement, deletion, etc. of functions or processes performed by some of the multiple components of the coding device or decoding device may be made. For example, any function or process may be replaced by another function or process described in any of the embodiments of the present invention, or they may be combined;
[0149] (3) In the method implemented by the encoding device or decoding device of the embodiment, any changes such as addition, replacement, deletion, etc. may be made to a part of the multiple processes included in the method. For example, any process in the method may be replaced by another process described in any of the various aspects of the present invention, or they may be combined;
[0150] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with a component described in any of the aspects of the present invention, may be combined with a component having a part of the function described in any of the aspects of the present invention, or may be combined with a component that implements a part of the processing implemented by the component described in any of the aspects of the present invention.
[0151] (5) A component having a part of the functions of the encoding device or decoding device of the embodiment, or a component implementing a part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced with a component described in any of the aspects of the present invention, a component having a part of the functions described in any of the aspects of the present invention, or a component implementing a part of the processing described in any of the aspects of the present invention;
[0152] (6) In a method implemented by an encoding device or a decoding device of an embodiment, one of a plurality of processes included in the method is replaced with one of the processes described in each aspect of the present invention or a similar process, or a combination of these processes;
[0153] (7) Some of the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the process described in any of the aspects of the present invention.
[0154] (8) The implementation of the processing and / or structure described in each aspect of the present invention is not limited to the encoding device or decoding device of the embodiment. For example, the processing and / or structure may also be implemented in a device used for a purpose different from the motion picture encoding or motion picture decoding disclosed in the embodiment.
[0155] (Implementation Method 1)
[0156] [Encoding device]
[0157] First, the encoding device according to this embodiment will be described. Figure 1 1 is a block diagram showing a functional structure of the encoding device 100 according to the present embodiment. The encoding device 100 is a moving picture encoding device that encodes a moving picture in units of blocks.
[0158] like Figure 1 As shown, the encoding device 100 is a device that encodes an image in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126 and a prediction control unit 128.
[0159] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.
[0160] Hereinafter, after describing the overall processing flow of the encoding device 100 , each component included in the encoding device 100 will be described.
[0161] [Overall flow of encoding processing]
[0162] Figure 2 : is a flowchart showing an example of the overall encoding process performed by the encoding device 100.
[0163] First, the segmentation unit 102 of the encoding device 100 segments each picture included in the input image as a moving image into a plurality of fixed-size blocks (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern (also referred to as a block shape) for the fixed-size block (step Sa_2). That is, the segmentation unit 102 further segments the fixed-size block into a plurality of blocks constituting the selected segmentation pattern. Then, the encoding device 100 performs the processing of steps Sa_3 to Sa_9 on each of the plurality of blocks (i.e., the encoding target block).
[0164] That is, the prediction processing unit composed of all or part of the intra-frame prediction unit 124, the inter-frame prediction unit 126 and the prediction control unit 128 generates a prediction signal (also called a prediction block) of the encoding object block (also called the current block) (step Sa_3).
[0165] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa 4).
[0166] Next, the transform unit 106 and the quantization unit 108 transform and quantize the difference block to generate a plurality of quantized coefficients (step Sa_5). In addition, a block composed of a plurality of quantized coefficients is also called a coefficient block.
[0167] Next, the entropy coding unit 110 generates a coded signal by encoding the coefficient block and prediction parameters related to the generation of the prediction signal (specifically, entropy coding) (step Sa_6). In addition, the coded signal is also called a coded bit stream, a compressed bit stream or a stream.
[0168] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore a plurality of prediction residuals (ie, difference blocks) by inverse quantizing and inverse transforming the coefficient blocks (step Sa_7).
[0169] Next, the adding unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the prediction block to the restored difference block (step Sa_8). Thus, a reconstructed image is generated.
[0170] When the reconstructed image is generated, the loop filter unit 120 filters the reconstructed image as necessary (step Sa_9).
[0171] Then, the encoding device 100 determines whether encoding of the entire picture is completed (step Sa_10 ), and when it is determined that encoding is not completed (No in step Sa_10 ), it repeats the processing from step Sa_2 .
[0172] In addition, in the above example, the encoding device 100 selects one partition pattern for a block of a fixed size and encodes each block according to the partition pattern, but each block may be encoded according to each of a plurality of partition patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of partition patterns, and may select, for example, an encoded signal obtained by encoding according to the partition pattern with the minimum cost as the encoded signal to be finally output.
[0173] Furthermore, the processing of steps Sa_1 to Sa_10 may be performed sequentially by the encoding device 100 , a part of a plurality of the processing may be performed in parallel, or the order may be swapped.
[0174] [Division]
[0175] The segmentation unit 102 segments each picture included in the input motion image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the picture into blocks of a fixed size (e.g., 128×128). The fixed-size blocks are sometimes called coding tree units (CTUs). Furthermore, the segmentation unit 102 segments each fixed-size block into blocks of a variable size (e.g., less than 64×64), for example based on recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. The variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and part or all of the blocks in the picture may be used as processing units of CUs, PUs, or TUs.
[0176] Figure 3 FIG. 1 is a diagram showing an example of block division according to the present embodiment. Figure 3 In FIG, the solid line represents the block boundary based on the quadtree block partition, and the dotted line represents the block boundary based on the binary tree block partition.
[0177] Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into four square 64×64 blocks (quadtree block division).
[0178] The upper left 64×64 block is further divided vertically into two rectangular 32×64 blocks, and the left 32×64 block is further divided vertically into two rectangular 16×64 blocks (binary tree block division). As a result, the upper left 64×64 block is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.
[0179] The upper right 64×64 block is horizontally partitioned into two rectangular 64×32 blocks 14 and 15 (binary tree block partitioning).
[0180] The 64×64 block at the bottom left is divided into four square 32×32 blocks (quadtree block division). The upper left block and the lower right block of the four 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the 16×32 block on the right is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the lower left 64×64 block is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.
[0181] The lower right 64×64 block 23 is not split.
[0182] As above, in Figure 3 In FIG. 1 , block 10 is divided into 13 blocks 11 to 23 of variable size based on recursive quad-tree and binary tree block division. This type of division is sometimes called QTBT (quad-tree plus binary tree) division.
[0183] In addition, Figure 3 In the example, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). The division including such ternary tree division is called MBT (multi type tree) division.
[0184] [Structural slices / tiles of the image]
[0185] In order to decode pictures in parallel, pictures may be constructed in slice units or tile units. Pictures constructed in slice units or tile units may be constructed by the partitioning unit 102.
[0186] A slice is a basic coding unit constituting a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more continuous CTUs (Coding Tree Units).
[0187] Figure 4A This is a diagram showing an example of the structure of a slice. For example, a picture includes 11×8 CTUs and is divided into 4 slices (slices 1 to 4). Slice 1 consists of 16 CTUs, slice 2 consists of 21 CTUs, slice 3 consists of 29 CTUs, and slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any slice. The shape of the slice becomes a shape that divides the picture in the horizontal direction. The boundary of the slice does not need to be the end of the picture, but can be any position in the boundary of the CTU in the picture. The processing order (encoding order or decoding order) of the CTU in the slice is, for example, a raster scan order. In addition, the slice includes header information and encoded data. The header information may also record the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.
[0188] A tile is a unit of a rectangular area constituting a picture, and a number called TileId may be assigned to each tile in a raster scan order.
[0189] Figure 4BThis is a diagram showing an example of a tile structure. For example, a picture includes 11×8 CTUs and is divided into four rectangular area tiles (tiles 1 to 4). When tiles are used, the processing order of CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed in raster scan order. When tiles are used, at least one CTU in each of the multiple tiles is processed in raster scan order. For example, Figure 4B As shown, the processing order of multiple CTUs included in tile 1 is from the left end of the 1st column of tile 1 to the right end of the 1st column of tile 1, and then from the left end of the 2nd column of tile 1 to the right end of the 2nd column of tile 1.
[0190] In addition, one tile may include more than one slice, and one slice may include more than one tile.
[0191] [Subtraction Department]
[0192] The subtracting unit 104 subtracts the prediction signal (the prediction sample input from the prediction control unit 128 shown below) from the original signal (original sample) in units of blocks input from the dividing unit 102 and divided by the dividing unit 102. That is, the subtracting unit 104 calculates the prediction error (also referred to as residual) of the encoding target block (hereinafter referred to as the current block). Then, the subtracting unit 104 outputs the calculated prediction error (residual) to the transforming unit 106.
[0193] The original signal is an input signal of the encoding device 100 and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing an image may be referred to as a sample.
[0194] [Conversion Department]
[0195] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0196] In addition, the transform unit 106 may adaptively select a transform type from a plurality of transform types, and transform the prediction error into a transform coefficient using a transform basis function corresponding to the selected transform type. Such a transform is sometimes called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0197] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transformation basis functions corresponding to each transformation type. Figure 5A In , N represents the number of input pixels. The selection of a transform type from among these multiple transform types may depend on the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode, for example.
[0198] Information indicating whether such EMT or AMT is applied (e.g., called an EMT flag or an AMT flag) and information indicating the selected transform type are usually signaled at the CU level. In addition, the signaling of such information is not necessarily limited to the CU level, but may be other levels (e.g., a bit sequence level, a picture level, a slice level, a tile level, or a CTU level).
[0199] In addition, the transformation unit 106 may also re-transform the transformation coefficients (transformation results). Such re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 re-transforms each sub-block (for example, 4×4 sub-blocks) contained in the block of transformation coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transformation matrix used in NSST are usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but can also be other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0200] Separable transformation and Non-Separable transformation can also be applied in the transformation unit 106. Separable transformation refers to a method of performing multiple transformations in each direction equivalent to the number of dimensions of the input. Non-Separable transformation refers to a method of treating two or more dimensions as one dimension and transforming them together when the input is multi-dimensional.
[0201] For example, as an example of Non-Separable transformation, when a 4×4 block is input, it is considered as an array having 16 elements, and the array is transformed using a 16×16 transformation matrix.
[0202] In a further example of the Non-Separable transformation, a 4×4 input block may be regarded as an array of 16 elements, and then a transformation (Hypercube Givens Transform) may be performed by performing a plurality of Givens rotations on the array.
[0203] In the transformation in the transformation unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region in the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, Figure 5B As shown, the CU is divided into two equal parts in the horizontal or vertical direction, and only the area on one side is transformed into the frequency area. The type of transformation base can be set for each area, for example, DST7 and DCT8 are used. In this example, only one of the two areas in the CU is transformed, and the other is not transformed, but both areas can also be transformed. In addition, the division method is not limited to two equal parts, but can be more flexible, such as four equal parts or encoding the information representing the division separately, and signaling it in the same way as the CU division. In addition, SVT is sometimes referred to as SBT (Sub-block Transform).
[0204] [Quantitative Department]
[0205] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on the quantization parameters (QP) corresponding to the scanned transform coefficients. The quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.
[0206] The predetermined scanning order is the order used for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order is defined in ascending order (order from low frequency to high frequency) or descending order (order from high frequency to low frequency) of frequency.
[0207] The quantization parameter (QP) refers to a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0208] In addition, quantization matrices are sometimes used in quantization. For example, multiple quantization matrices are sometimes used corresponding to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as brightness and color difference. In addition, quantization means digitizing values sampled at predetermined intervals in correspondence with predetermined levels. In this technical field, expressions such as rounding, rounding, and scaling are sometimes used.
[0209] As a method of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device side and a method of using a default quantization matrix (default matrix). On the encoding device side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the encoding amount increases due to the encoding of the quantization matrix.
[0210] On the other hand, there is also a method of performing quantization so that the coefficients of high-frequency components and low-frequency components are the same without using a quantization matrix. This method is equivalent to a method of using a quantization matrix (flat matrix) in which all coefficients have the same value.
[0211] The quantization matrix can be specified by, for example, an SPS (Sequence Parameter Set) or a PPS (Picture Parameter Set). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and PPS are sometimes referred to as parameter sets.
[0212] [Entropy coding unit]
[0213] The entropy coding unit 110 generates a coded signal (coded bit stream) based on the quantization coefficient input from the quantization unit 108. Specifically, the entropy coding unit 110 binarizes the quantization coefficient, performs arithmetic coding on the binary signal, and outputs a compressed bit stream or sequence.
[0214] [Inverse Quantization Unit]
[0215] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. The inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114.
[0216] [Inverse transformation unit]
[0217] The inverse transform unit 114 restores the prediction error (residual) by inverse transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform performed by the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0218] In addition, the restored prediction error usually loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error usually includes a quantization error.
[0219] [Addition Department]
[0220] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. Then, the adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoding block.
[0221] [Block Memory]
[0222] The block memory 118 is a storage unit for storing blocks in a picture to be coded (referred to as a current picture) to be referred to in intra prediction, for example. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116 .
[0223] [Frame Memory]
[0224] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120 .
[0225] [Loop filter section]
[0226] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within a coding loop (in-loop filtering), such as deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).
[0227] In ALF, a least square error filter is used to remove coding distortion. For example, for each 2×2 sub-block in the current block, one filter is selected from a plurality of filters based on the direction and activity of the local gradient.
[0228] Specifically, first, the sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C (e.g., C=5D+A) is calculated. And, based on the classification value C, the sub-blocks are classified into multiple classes.
[0229] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (eg, horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantizing the added result.
[0230] Based on the result of such classification, a filter to be used for a sub-block is determined from among a plurality of filters.
[0231] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figure 6A to Figure 6C The diagrams show a plurality of examples of filter shapes used in ALF. Fig. 6A represents a 5×5 diamond-shaped filter, Figure 6B represents a 7×7 diamond shape filter, Figure 6C Represents a 9×9 diamond shape filter. The information representing the shape of the filter is usually signaled at the picture level. In addition, the signaling of the information representing the shape of the filter does not need to be limited to the picture level, and can also be other levels (for example, sequence level, slice level, tile level, CTU level or CU level).
[0232] The on / off of ALF can also be determined at the picture level or CU level. For example, regarding brightness, whether to use ALF can be determined at the CU level, and regarding color difference, whether to use ALF can be determined at the picture level. Information indicating whether ALF is on / off is usually signaled at the picture level or CU level. In addition, the signaling of information indicating whether ALF is on / off does not need to be limited to the picture level or CU level, and can also be other levels (for example, sequence level, slice level, tile level, or CTU level).
[0233] The coefficient sets of the selectable multiple filters (e.g., filters up to 15 or 25) are usually signaled at the picture level. In addition, the signaling of the coefficient set does not need to be limited to the picture level, and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0234] [Loop filter section > Deblocking filter]
[0235] In the deblocking filter, the loop filter unit 120 performs filtering processing on block boundaries of the reconstructed image to reduce distortion generated at the block boundaries.
[0236] Figure 7 1 is a block diagram showing an example of a detailed configuration of the loop filter unit 120 that functions as a deblocking filter.
[0237] The loop filter unit 120 includes a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 , and 1206 .
[0238] The boundary determination unit 1201 determines whether there is a pixel (ie, a target pixel) to be subjected to the deblocking filtering process near the block boundary, and then outputs the determination result to the switch 1202 and the process determination unit 1208 .
[0239] When the boundary determination unit 1201 determines that the target pixel exists near the block boundary, the switch 1202 outputs the image before filtering to the switch 1204. On the other hand, when the boundary determination unit 1201 determines that the target pixel does not exist near the block boundary, the switch 1202 outputs the image before filtering to the switch 1206.
[0240] The filter determination unit 1203 determines whether to perform deblocking filtering on the target pixel based on the pixel value of at least one surrounding pixel located around the target pixel, and then outputs the determination result to the switch 1204 and the processing determination unit 1208 .
[0241] When the filter determination unit 1203 determines that the deblocking filter process is to be performed on the target pixel, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the filtering process unit 1205. On the contrary, when the filter determination unit 1203 determines that the deblocking filter process is not to be performed on the target pixel, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the switch 1206.
[0242] When the image before filtering is obtained via switches 1202 and 1204 , the filter processing unit 1205 performs a deblocking filter process on the target pixel using the filter characteristics determined by the filter characteristic determination unit 1207 . Then, the filter processing unit 1205 outputs the filtered pixel to the switch 1206 .
[0243] According to the control of the processing determination unit 1208 , the switch 1206 selectively outputs pixels that have not been processed by the deblocking filter and pixels that have been processed by the deblocking filter by the filter processing unit 1205 .
[0244] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filter determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filter determination unit 1203 determines that the deblocking filter process is performed on the target pixel, the processing determination unit 1208 outputs the pixel after the deblocking filter process from the switch 1206. In addition, in cases other than the above, the processing determination unit 1208 outputs the pixel that has not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the image after the filtering process is output from the switch 1206.
[0245] Figure 8 This is a diagram showing an example of a deblocking filter having a filter characteristic that is symmetric with respect to a block boundary.
[0246] In the deblocking filter processing, for example, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter, is selected. Figure 8 As shown, when pixels p0 to p2 and pixels q0 to q2 exist across a block boundary, the pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing calculations shown in the following equations.
[0247] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0248] q'1=(p0+q0+q1+q2+2) / 4
[0249] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0250] In the above equations, p0 to p2 and q0 to q2 are pixel values of pixels p0 to p2 and pixels q0 to q2, respectively. In addition, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficient multiplied by the pixel value of each pixel used in the deblocking filter processing is the filter coefficient.
[0251] Furthermore, in the deblocking filter process, a clipping process may be performed in such a way that the pixel value after the operation does not exceed the threshold value. In the clipping process, the pixel value after the operation based on the above formula is clipped to "the pixel value before the operation ±2×threshold value" using a threshold value determined according to the quantization parameter. This can prevent excessive smoothing.
[0252] Fig. 9 This is a diagram for explaining a block boundary on which a deblocking filtering process is performed. Fig.10 This is a diagram showing an example of the Bs value.
[0253] The block boundary for deblocking filtering is, for example, Fig. 9 The PU (Prediction Unit) or TU (Transform Unit) boundary of the 8×8 pixel block shown in FIG. The deblocking filtering process is performed in units of 4 rows or 4 columns. First, for Fig. 9 The block P and block Q shown are as follows: Fig.10 That determines the Bs (Boundary Strength) value.
[0254] according to Fig.10 The Bs value determines whether to perform deblocking filtering of different strengths even at block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering is performed on the color difference signal. When the Bs value is 1 or more and meets the specified conditions, deblocking filtering is performed on the brightness signal. In addition, the determination condition of the Bs value is not limited to Fig.10 The conditions shown may also be determined based on other parameters.
[0255] [Prediction processing unit (intra-frame prediction unit / inter-frame prediction unit / prediction control unit)]
[0256] Fig.11 1 is a diagram showing an example of processing performed by the prediction processing unit of the encoding device 100. The prediction processing unit is composed of all or part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0257] The prediction processing unit generates a prediction image of the current block (step Sb_1). The prediction image is also called a prediction signal or a prediction block. In addition, the prediction signal includes, for example, an intra-frame prediction signal or an inter-frame prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring a difference block, and generating a decoded image block.
[0258] The reconstructed image may be, for example, an image of a reference picture, or may be a picture including the current block, that is, an image of an encoded block in the current picture. The encoded block in the current picture may be, for example, an adjacent block of the current block.
[0259] Fig.12 This is a diagram showing another example of processing performed by the prediction processing unit of the encoding device 100.
[0260] The prediction processing unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating prediction images, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods. In such a prediction method, the above-mentioned reconstructed image may also be used.
[0261] Next, the prediction processing unit selects any one of the multiple prediction images generated in steps Sc_1a, Sc_1b and Sc_1c (step Sc_2). The selection of the prediction image, that is, the selection of the method or mode for obtaining the final prediction image, may also be performed by calculating the cost for each prediction image generated and based on the cost. In addition, the selection of the prediction image may be performed based on the parameters used for the encoding process. The encoding device 100 may signal information for determining the selected prediction image, method or mode as a coded signal (also referred to as a coded bit stream). The information may be, for example, a flag, etc. Thus, the decoding device can generate a prediction image in accordance with the method or mode selected in the encoding device 100 based on the information. In addition, in Fig.12 In the example shown, the prediction processing unit selects any one of the prediction images after generating the prediction images by each method. However, before generating these prediction images, the prediction processing unit can select a method or mode based on the parameters used for the above-mentioned encoding process, and can generate the prediction image according to the method or mode.
[0262] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction processing unit may select a final prediction image for the current block from prediction images generated according to these prediction methods.
[0263] Fig.13 This is a diagram showing another example of processing performed by the prediction processing unit of the encoding device 100.
[0264] First, the prediction processing unit generates a prediction image by intra prediction (step Sd_1a), and generates a prediction image by inter prediction (step Sd_1b). In addition, the prediction image generated by intra prediction is also called intra prediction image, and the prediction image generated by inter prediction is also called inter prediction image.
[0265] Next, the prediction processing unit evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). The cost can also be used in this evaluation. That is, the prediction processing unit calculates the cost C of each of the intra-frame prediction image and the inter-frame prediction image. The cost C is calculated by a formula of the RD optimization model, such as C=D+λ×R. In this formula, D is the coding distortion of the predicted image, and is represented by, for example, the absolute value of the difference between the pixel value of the current block and the pixel value of the predicted image. In addition, R is the amount of code generated for the predicted image, specifically, the amount of code required for encoding motion information, etc. for generating the predicted image. In addition, λ is, for example, an undetermined multiplier of Lagrange.
[0266] Then, the prediction processing unit selects the prediction image with the minimum cost C calculated from the intra prediction image and the inter prediction image as the final prediction image of the current block (step Sd_3). In other words, the prediction method or mode for generating the prediction image of the current block is selected.
[0267] [Intra-frame prediction unit]
[0268] The intra prediction unit 124 performs intra prediction (also referred to as intra-screen prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 124 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., brightness values and color difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0269] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0270] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.
[0271] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Fig.14 This is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions specified by the H.265 / HEVC specification, and the dotted arrows represent the additional 32 directions (2 non-directional prediction modes in Fig.14 (not shown in the figure).
[0272] In various implementation examples, the luminance block may be referenced in the intra prediction of the chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. Such intra prediction is called CCLM (cross-component linear model) prediction. Such an intra prediction mode of the chrominance block with reference to the luminance block (for example, CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0273] The intra prediction unit 124 may also modify the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction. Intra prediction accompanied by such modification is called PDPC (position dependent intraprediction combination). Information indicating whether PDPC is used (e.g., called a PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but may also be other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0274] [Inter-frame prediction unit]
[0275] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-screen prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or the current sub-block (e.g., 4×4 block) in the current block. For example, the inter-frame prediction unit 126 performs motion search (motion estimation) in the reference picture for the current block or the current sub-block, and searches for the reference block or sub-block that is most consistent with the current block or the current sub-block. In addition, the inter-frame prediction unit 126 obtains motion information (e.g., motion vector) for compensating for the motion or change from the reference block or sub-block to the current block or sub-block. The inter-frame prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, thereby generating an inter-frame prediction signal for the current block or sub-block. In addition, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0276] The motion information used in motion compensation is signaled as an inter-frame prediction signal in various forms. For example, a motion vector may be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) may be signaled.
[0277] [Basic process of inter-frame prediction]
[0278] Fig.15This is a flowchart showing the basic process of inter-frame prediction.
[0279] The inter prediction unit 126 first generates a predicted image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4).
[0280] Here, in the generation of the predicted image, the inter-frame prediction unit 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). In addition, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. In addition, in the derivation of the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from the at least one candidate MV and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching the area of the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. In addition, the action of searching the area of the reference picture may also be referred to as motion estimation.
[0281] In the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126 , but processing such as step Se_1 or step Se_2 may be performed by other components included in the encoding device 100 .
[0282] [Flow of deriving motion vectors]
[0283] Fig.16 This is a flowchart showing an example of motion vector derivation.
[0284] The inter-frame prediction unit 126 derives the MV of the current block in a mode of encoding motion information (e.g., MV). In this case, for example, the motion information is encoded as a prediction parameter and is signaled. That is, the encoded motion information is included in the encoded signal (also referred to as an encoded bitstream).
[0285] Alternatively, the inter prediction unit 126 derives the MV in a mode in which the motion information is not encoded. In this case, the motion information is not included in the encoded signal.
[0286] Here, the modes for deriving MV include the normal inter mode, merge mode, FRUC mode, and affine mode described later. Among these modes, the modes for encoding motion information include the normal inter mode, merge mode, and affine mode (specifically, affine inter mode and affine merge mode). In addition, the motion information may include not only MV but also predicted motion vector selection information described later. In addition, the modes for not encoding motion information include the FRUC mode, etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.
[0287] Fig.17 This is a flowchart showing another example of motion vector derivation.
[0288] The inter-frame prediction unit 126 derives the MV of the current block in a mode of encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and is signaled. That is, the encoded differential MV is included in the encoded signal. The differential MV is the difference between the MV of the current block and its predicted MV.
[0289] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode in which the difference MV is not encoded. In this case, the encoded difference MV is not included in the encoded signal.
[0290] Here, as described above, the MV derivation modes include the normal inter mode, merge mode, FRUC mode, and affine mode described later. Among these modes, the modes for encoding the differential MV include the normal inter mode and the affine mode (specifically, the affine inter mode). In addition, the modes for not encoding the differential MV include the FRUC mode, the merge mode, and the affine mode (specifically, the affine merge mode). The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.
[0291] [Flow of deriving motion vectors]
[0292] Fig.18This is a flowchart showing another example of motion vector derivation. There are multiple modes for MV derivation, i.e., inter-frame prediction modes, which are roughly divided into a mode for encoding the differential MV and a mode for not encoding the differential motion vector. The modes for not encoding the differential MV include the merge mode, the FRUC mode, and the affine mode (specifically, the affine merge mode). The details of these modes will be described later. In short, the merge mode is a mode for deriving the MV of the current block by selecting a motion vector from the surrounding coded blocks, and the FRUC mode is a mode for deriving the MV of the current block by searching between coded areas. In addition, the affine mode is a mode for deriving the motion vectors of each of the multiple sub-blocks constituting the current block as the MV of the current block assuming an affine transformation.
[0293] Specifically, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_2) based on the merge mode. In addition, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_3) according to the FRUC mode. In addition, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_4) according to the affine mode (specifically, the affine merge mode). In addition, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector (Sf_5) according to the mode for encoding the differential MV (for example, the normal inter-frame mode).
[0294] [MV export > Normal interframe mode]
[0295] The normal inter mode is an inter prediction mode that derives the MV of the current block by searching for a block similar to the image of the current block in the area of the reference picture indicated by the candidate MV. In addition, in the normal inter mode, the differential MV is encoded.
[0296] Fig.19 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.
[0297] First, the inter prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of coded blocks temporally or spatially located around the current block (step Sg_1). That is, the inter prediction unit 126 creates a candidate MV list.
[0298] Next, the inter-frame prediction unit 126 extracts N (N is an integer greater than or equal to 2) candidate MVs from the multiple candidate MVs obtained in step Sg_1 as predicted motion vector candidates (also referred to as predicted MV candidates) in a predetermined priority order (step Sg_2). In addition, the priority order is predetermined for each of the N candidate MVs.
[0299] Next, the inter-frame prediction unit 126 selects one prediction motion vector candidate from the N prediction motion vector candidates as the prediction motion vector (also referred to as prediction MV) of the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes prediction motion vector selection information for identifying the selected prediction motion vector into the stream. In addition, the stream is the above-mentioned coded signal or coded bit stream.
[0300] Next, the inter-frame prediction unit 126 refers to the coded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference between the derived MV and the predicted motion vector as a difference MV to the stream. In addition, the coded reference picture is a picture composed of multiple blocks reconstructed after encoding.
[0301] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). Note that the predicted image is the above-mentioned inter-frame prediction signal.
[0302] Furthermore, information indicating the inter prediction mode (in the above example, the normal inter mode) used in generating the predicted image contained in the encoded signal is encoded as, for example, a prediction parameter.
[0303] In addition, the candidate MV list can also be used together with the list used in other modes. In addition, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to the candidate MV list is, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.
[0304] [MV Export > Merge Mode]
[0305] The merge mode is an inter-frame prediction mode that derives a candidate MV from a candidate MV list as the MV of the current block.
[0306] Fig. 20 is a flowchart showing an example of inter-frame prediction based on merge mode.
[0307] First, the inter prediction unit 126 obtains a plurality of candidate MVs for the current block based on information on a plurality of coded block MVs located around the current block in time or space (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.
[0308] Next, the inter prediction unit 126 selects one candidate MV from the plurality of candidate MVs acquired in step Sh_1 to derive the MV of the current block (step Sh_2). At this time, the inter prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream.
[0309] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).
[0310] Furthermore, information indicating the inter prediction mode (in the above example, the merge mode) used in generating the predicted image contained in the encoded signal is encoded as, for example, a prediction parameter.
[0311] Fig.21 This is a diagram for explaining an example of motion vector derivation processing of the current picture based on the merge mode.
[0312] First, a prediction MV list is generated in which candidates for prediction MVs are registered. Candidates for prediction MVs include: spatial neighbor prediction MVs, which are MVs of multiple coded blocks located in the spatial periphery of the object block; temporal neighbor prediction MVs, which are MVs of blocks near the position of the object block in the coded reference picture that are projected; combined prediction MVs, which are MVs generated by combining the MV values of the spatial neighbor prediction MV and the temporal neighbor prediction MV; and zero prediction MVs, which are MVs with a value of zero, etc.
[0313] Next, one predicted MV is selected from a plurality of predicted MVs registered in the predicted MV list to determine the MV of the target block.
[0314] Then, in the variable length coding unit, merge_idx, which is a signal indicating which predicted MV is selected, is described in the stream and coded.
[0315] In addition, Fig.21 The predicted MVs registered in the predicted MV list described in the figure are just examples, and may be a number different from that in the figure, or a structure that does not include some types of predicted MVs in the figure, or a structure that adds predicted MVs other than the types of predicted MVs in the figure.
[0316] The final MV may be determined by performing a DMVR (dynamic motion vector refreshing) process described later using the MV of the target block derived by the merge mode.
[0317] In addition, the candidate of the predicted MV is the candidate MV mentioned above, and the predicted MV list is the candidate MV list mentioned above. In addition, the candidate MV list can also be called a candidate list. In addition, merge_idx is MV selection information.
[0318] [MV Export > FRUC Mode]
[0319] The motion information may also be derived on the decoding device side instead of being signaled on the encoding device side. In addition, as described above, the merge mode specified by the H.265 / HEVC specification may also be used. In addition, for example, the motion information may be derived by performing a motion search on the decoding device side. In this case, the motion search is performed on the decoding device side without using the pixel values of the current block.
[0320] Here, a mode for performing motion estimation on the decoding device side is described. The mode for performing motion estimation on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0321] exist Fig. 22An example of FRUC processing is shown in FIG. First, with reference to the motion vectors of the coded blocks that are spatially or temporally adjacent to the current block, a list of multiple candidates each having a predicted motion vector (MV) is generated (that is, a candidate MV list, which can also be shared with a merge list) (step Si_1). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected based on the evaluation value. And, based on the motion vector of the selected candidate, the motion vector for the current block is derived (step Si_4). Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is derived as it is as the motion vector for the current block. In addition, for example, the motion vector for the current block can also be derived by performing pattern matching in the surrounding area of the position in the reference picture corresponding to the motion vector of the selected candidate. That is, the surrounding area of the best candidate MV can also be searched by using pattern matching and evaluation values in the reference picture. When there is an MV with a better evaluation value, the best candidate MV is updated to the above MV and used as the final MV of the current block. A configuration may be adopted in which the process of updating to an MV having a better evaluation value is not performed.
[0322] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).
[0323] Completely the same processing can be performed even when processing is performed in sub-block units.
[0324] The evaluation value may also be calculated by various methods. For example, the reconstructed image of the area in the reference picture corresponding to the motion vector is compared with the reconstructed image of a predetermined area (for example, as shown below, the area may be an area of another reference picture or an area of an adjacent block of the current picture). Then, the difference in pixel values of the two reconstructed images may be calculated and used for the evaluation value of the motion vector. In addition, the evaluation value may be calculated using other information in addition to the difference value.
[0325] Next, pattern matching is described in detail. First, a candidate MV included in a candidate MV list (e.g., a merge list) is selected as the starting point for a search based on pattern matching. As pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are respectively called bilateral matching and template matching.
[0326] [MV Export > FRUC > Bidirectional Matching]
[0327] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures along the motion trajectory of the current block. Therefore, in the first pattern matching, as the prescribed area for calculating the candidate evaluation value, an area in another reference picture along the motion trajectory of the current block is used.
[0328] Fig.23 1 is a diagram for explaining an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along the motion trajectory. Fig.23 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among the pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV with the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs.
[0329] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0330] [MV Export > FRUC > Template Matching]
[0331] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the prescribed area for calculating the candidate evaluation value described above.
[0332] Fig.24 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Fig.24As shown, in the second pattern matching, the motion vector of the current block is derived by searching the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded area of both or one of the left adjacent and upper adjacent areas and the reconstructed image at the same position in the coded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value is selected as the best candidate MV among multiple candidate MVs.
[0333] Such information indicating whether the FRUC mode is adopted (for example, called a FRUC flag) is signaled at the CU level. In addition, when the FRUC mode is adopted (for example, when the FRUC flag is true), information indicating the method of pattern matching that can be adopted (first pattern matching or second pattern matching) is signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and can also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level or sub-block level).
[0334] [MV Export > Affine Mode]
[0335] Next, the affine mode for deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes referred to as an affine motion compensation prediction mode.
[0336] Fig.25A FIG. 1 is a diagram for explaining an example of deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks. Fig.25A In the example, the current block includes 16 4×4 sub-blocks. Here, the motion vector v of the upper left corner control point of the current block is derived based on the motion vectors of the neighboring blocks. 0 Similarly, the motion vector v of the upper right corner control point of the current block is derived based on the motion vector of the adjacent sub-block 1 Then, according to the following formula (1A), the two motion vectors v are projected 0 and v 1 , derive the motion vectors of each sub-block in the current block (v x , v y ).
[0337]
Formula 1
[0338]
[0339] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predefined weight coefficient.
[0340] The information indicating such an affine mode (e.g., called an affine flag) may be signaled as a CU-level signal. In addition, the signaling of the information indicating the affine mode need not be limited to the CU-level, but may be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0341] In addition, such an affine mode may include several modes in which the motion vectors of the upper left and upper right control points are derived in different ways. For example, the affine mode includes two modes: the affine inter-frame (also called affine normal inter-frame) mode and the affine merge mode.
[0342] [MV Export > Affine Mode]
[0343] Fig.25B FIG. 1 is a diagram for explaining an example of derivation of a motion vector for a sub-block unit in an affine mode having three control points. Fig.25B In the example, the current block includes 16 4×4 sub-blocks. Here, the motion vector v of the upper left corner control point of the current block is derived based on the motion vectors of the neighboring blocks. 0 Similarly, the motion vector v of the upper right corner control point of the current block is derived based on the motion vector of the adjacent block. 1 , derive the motion vector v of the lower left corner control point of the current block based on the motion vector of the adjacent block 2 Then, according to the following formula (1B), project the three motion vectors v 0 、v 1 and v 2 , derive the motion vectors of each sub-block in the current block (v x , v y ).
[0344]
Formula 2
[0345]
[0346] Here, x and y represent the horizontal position and vertical position of the center of the sub-block, respectively, w represents the width of the current block, and h represents the height of the current block.
[0347] Affine modes with different numbers of control points (e.g., 2 and 3) may also be switched and signaled at the CU level. In addition, information indicating the number of control points of the affine mode used at the CU level may also be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0348] In addition, in such an affine mode with three control points, several modes with different methods of deriving motion vectors of the upper left, upper right and lower left control points may be included. For example, in the affine mode, there are two modes, namely, the affine inter-frame (also called affine normal inter-frame) mode and the affine merge mode.
[0349] [MV Export > Affine Merge Mode]
[0350] Fig.26A , Fig.26B and Fig.26C This is a conceptual diagram for explaining the affine merge mode.
[0351] In affine merge mode, such as Fig.26A As shown, for example, based on a plurality of motion vectors corresponding to blocks encoded in an affine mode among the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, the predicted motion vector of each of the control points of the current block is calculated. Specifically, the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) are checked in this order to determine the first valid block encoded in the affine mode. The predicted motion vectors of the control points of the current block are calculated based on the plurality of motion vectors corresponding to the determined blocks.
[0352] For example, Fig.26B As shown, when a block A adjacent to the left side of the current block is encoded in an affine mode with two control points, a motion vector v projected to the upper left corner and the upper right corner of the encoded block containing block A is derived. 3 and v 4 Then, according to the derived motion vector v 3 and v 4 , calculate the predicted motion vector v of the control point in the upper left corner of the current block 0 and the predicted motion vector v of the upper right control point 1 .
[0353] For example, Fig.26C As shown, when a block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v projected to the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A are derived. 3 、v 4 and v 5 Then, according to the derived motion vector v 3 、v 4 and v 5 , calculate the predicted motion vector v of the control point in the upper left corner of the current block 0 , the predicted motion vector v of the control point in the upper right corner 1and the predicted motion vector v of the control point at the lower left corner 2 .
[0354] In addition, in the following Fig.29 This predicted motion vector derivation method can also be used in the derivation of the predicted motion vectors of each control point of the current block in step Sj_1.
[0355] Fig. 27 This is a flowchart showing an example of the affine merge mode.
[0356] In the affine merge mode, first, the inter-frame prediction unit 126 derives the predicted MV of each control point of the current block (step Sk1). Fig.25A As shown, it is the upper left and upper right corner points of the current block, or as Fig.25B As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block.
[0357] That is to say, Fig.26A As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.
[0358] Then, when block A is determined and block A has 2 control points, as Fig.26B As shown, the inter-frame prediction unit 126 predicts the motion vectors v at the upper left corner and the upper right corner of the coded block including block A. 3 and v 4 To calculate the motion vector v of the control point in the upper left corner of the current block 0 and the motion vector v of the upper right control point 1 For example, by taking the motion vectors v of the upper left and upper right corners of the coded block 3 and v 4 Projected onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v of the control point at the upper left corner of the current block. 0 and the predicted motion vector v of the upper right control point 1 .
[0359] Alternatively, if block A is determined and block A has 3 control points, Fig.26C As shown, the inter-frame prediction unit 126 predicts the motion vectors v at the upper left corner, upper right corner, and lower left corner of the coded block including block A. 3 、v 4 and v 5 To calculate the motion vector v of the control point in the upper left corner of the current block 0 , the motion vector v of the control point in the upper right corner 1 and the motion vector v of the lower left control point 2For example, by taking the motion vectors v of the upper left, upper right, and lower left corners of the coded block 3 、v 4 and v 5 Projected onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v of the control point at the upper left corner of the current block. 0 , the predicted motion vector v of the control point in the upper right corner 1 and the motion vector v of the lower left control point 2 .
[0360] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 126 uses two prediction motion vectors v for each of the multiple sub-blocks. 0 and v 1 and the above formula (1A), or the three predicted motion vectors v 0 、v 1 and v 2 The motion vector of the sub-block is calculated as the affine MV using the above equation (1B) (step Sk_2). Then, the inter-frame prediction unit 126 performs motion compensation on the sub-block using the affine MV and the encoded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.
[0361] [MV Export > Affine Inter-frame Mode]
[0362] Fig.28A A diagram for explaining an affine inter-frame mode having two control points.
[0363] In this affine inter-frame mode, if Fig.28A As shown, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v of the control point at the upper left corner of the current block. 0 Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v of the control point at the upper right corner of the current block. 1 .
[0364] Fig.28B A diagram for explaining an affine inter-frame mode having three control points.
[0365] In this affine inter-frame mode, if Fig.28B As shown, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v of the control point at the upper left corner of the current block. 0Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v of the control point at the upper right corner of the current block. 1 In addition, a motion vector selected from the motion vectors of the coded blocks F and G adjacent to the current block is used as the predicted motion vector v of the control point at the lower left corner of the current block. 2 .
[0366] Fig.29 This is a flowchart showing an example of the affine inter mode.
[0367] In the affine inter mode, first, the inter prediction unit 126 derives the prediction MV (v 0 , v 1 ) or (v 0 , v 1 , v 2 )(Step Sj_1). Fig.25A or Fig.25B As shown, the control point is the point at the upper left corner, upper right corner, or lower left corner of the current block.
[0368] That is, the inter-frame prediction unit 126 selects Fig.28A or Fig.28B The motion vector of a block in the coded blocks near each control point of the current block shown in FIG. 1 is used to derive the predicted motion vector (v 0 , v 1 ) or (v 0 , v 1 , v 2 ). At this time, the inter-frame prediction unit 126 encodes prediction motion vector selection information for identifying the two selected motion vectors into the stream.
[0369] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select as the predicted motion vector of the control point from the encoded blocks adjacent to the current block by using cost evaluation, etc., and can record a flag indicating which predicted motion vector is selected in the bit stream.
[0370] Next, the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4) while updating the predicted motion vector selected or derived in step Sj_1 (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vector of each sub-block corresponding to the predicted motion vector to be updated as an affine MV, and calculates it using the above-mentioned formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on each sub-block (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted motion vector as a differential MV into a stream.
[0371] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0372] [MV Export > Affine Inter-frame Mode]
[0373] When affine modes with different numbers of control points (for example, 2 and 3) are switched at the CU level for signaling, the number of control points may be different between the coded block and the current block. Fig. 30A as well as Fig. 30B This is a conceptual diagram for explaining a method of deriving a prediction vector of a control point when the number of control points in an already coded block and a current block is different.
[0374] For example, Fig. 30A As shown in FIG. 1 , when the current block has three control points, namely, the upper left corner, the upper right corner, and the lower left corner, and the block A adjacent to the left side of the current block is encoded in an affine mode having two control points, the motion vector v projected to the position of the upper left corner and the upper right corner of the encoded block including the block A is derived. 3 and v 4 Then, according to the derived motion vector v 3 and v 4 , calculate the predicted motion vector v of the control point in the upper left corner of the current block 0 and the predicted motion vector v of the upper right control point 1 In addition, according to the derived motion vector v 0 and v 1 Calculate the predicted motion vector v of the lower left control point 2 .
[0375] For example, Fig. 30BAs shown in FIG. 1 , when the current block has two control points, the upper left corner and the upper right corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with three control points, the motion vector v projected to the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block including the block A is derived. 3 、v 4 and v 5 Then, according to the derived motion vector v 3 、v 4 and v 5 , calculate the predicted motion vector v of the control point in the upper left corner of the current block 0 and the predicted motion vector v of the upper right control point 1 .
[0376] exist Fig.29 This predicted motion vector derivation method can also be used in the derivation of each predicted motion vector of the control point of the current block in step Sj_1.
[0377] [MV Export>DMVR]
[0378] Fig.31A It is a diagram showing the relationship between the merge mode and DMVR.
[0379] The inter-frame prediction unit 126 derives the motion vector of the current block in the merge mode (step S1_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, that is, a motion search (step S1_2). Here, when it is determined that the motion search is not to be performed (No in step S1_2), the inter-frame prediction unit 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). That is, in this case, the motion vector of the current block is determined in the merge mode.
[0380] On the other hand, when it is determined in step S1_1 that motion search is to be performed (Yes in step S1_2), the inter-frame prediction unit 126 derives a final motion vector for the current block by searching the surrounding area of the reference picture represented by the motion vector derived in step S1_1 (step S1_3). That is, in this case, the motion vector of the current block is determined by DMVR.
[0381] Fig.31B This is a conceptual diagram for explaining an example of DMVR processing for determining MV.
[0382] First, the best MVP set for the current block (for example, in merge mode) is set as a candidate MV. Then, according to the candidate MV (L0), the reference pixel is determined from the first reference picture (L0) which is the coded picture in the L0 direction. Similarly, according to the candidate MV (L1), the reference pixel is determined from the second reference picture (L1) which is the coded picture in the L1 direction. The template is generated by taking the average of these reference pixels.
[0383] Next, the above template is used to search the surrounding areas of the candidate MVs of the first reference image (L0) and the second reference image (L1), respectively, and the MV with the minimum cost is determined as the final MV. In addition, the cost value can also be calculated using, for example, the difference between the pixel values of the template and the pixel values of the search area and the candidate MV value.
[0384] Note that the configuration and operation of the processing described here are basically the same in the encoding device and the decoding device described later.
[0385] Even if it is not the process described here, any process may be used as long as it is a process that can search the vicinity of the candidate MV and derive the final MV.
[0386] [Motion compensation > BIO / OBMC]
[0387] In motion compensation, there is a mode of generating a predicted image and correcting the predicted image. Examples of such modes are BIO and OBMC described below.
[0388] Fig.32 This is a flowchart showing an example of generating a predicted image.
[0389] The inter prediction unit 126 generates a predicted image (step Sm_1 ), and corrects the predicted image using any of the above-mentioned modes (step Sm_2 ).
[0390] Fig.33 This is a flowchart showing another example of generating a predicted image.
[0391] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined that the correction process is to be performed (yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting the predicted image (step Sn_4). On the other hand, when it is determined that the correction process is not to be performed (no in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).
[0392] In motion compensation, there is a mode for correcting brightness when generating a predicted image. This mode is, for example, LIC which will be described later.
[0393] Fig.34 This is a flowchart showing another example of generating a predicted image.
[0394] The inter-frame prediction unit 126 derives the motion vector of the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform brightness correction processing (step So_2). Here, when it is determined that brightness correction processing is to be performed (yes in step So_2), the inter-frame prediction unit 126 generates a predicted image while performing brightness correction (step So_3). In other words, the predicted image is generated by LIC. On the other hand, when it is determined that brightness correction processing is not to be performed (no in step So_2), the inter-frame prediction unit 126 generates a predicted image by normal motion compensation without performing brightness correction (step So_4).
[0395] [Motion Compensation > OBMC]
[0396] Not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent blocks can be used to generate the inter-frame prediction signal. Specifically, the inter-frame prediction signal can be generated in sub-block units within the current block by weighted addition of the prediction signal based on the motion information obtained by motion search (within the reference picture) and the prediction signal based on the motion information of the adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0397] In the OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) may also be signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is applied (e.g., OBMC flag) may also be signaled at the CU level. In addition, the signaling level of such information need not be limited to the sequence level and the CU level, but may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0398] The OBMC mode will be described in more detail. Fig.35 and Fig.36 It is a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.
[0399] First, if Fig.36 As shown, the motion vector (MV) assigned to the processing object (current) block is used to obtain the predicted image (Pred) based on the usual motion compensation. Fig.36In FIG. 1 , the arrow “MV” points to the reference picture and indicates which block the current block of the current picture refers to to obtain a predicted image.
[0400] Next, the motion vector (MV_L) derived for the coded left adjacent block is applied (reused) to the coded target block to obtain a predicted image (Pred_L). The motion vector (MV_L) is represented by an arrow "MV_L" pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0401] Similarly, the motion vector (MV_U) derived for the encoded upper adjacent block is applied (reused) to the encoding target block to obtain a predicted image (Pred_U). The motion vector (MV_U) is represented by an arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the predicted image is performed by superimposing the predicted image Pred_U with the predicted image (for example, Pred and Pred_L) that has been corrected for the first time. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block with the boundaries with the adjacent blocks blended (smoothed).
[0402] In addition, the above example is a two-path correction method using left-adjacent and upper-adjacent blocks, but the correction method may also be a three-path correction method or more than three-path correction method using right-adjacent and / or lower-adjacent blocks.
[0403] Furthermore, the region to be overlapped may not be the pixel region of the entire block, but may be only a partial region near the block boundary.
[0404] In addition, the prediction image correction processing of OBMC is described here, which is used to obtain one prediction image Pred by superimposing one reference image with the additional prediction images Pred_L and Pred_U. However, in the case of correcting the prediction image based on multiple reference images, the same processing can be applied to each of the multiple reference images. In this case, after the corrected prediction images are obtained from each reference image by performing image correction of OBMC based on multiple reference images, the final prediction image is obtained by further superimposing the obtained multiple corrected prediction images.
[0405] In OBMC, the unit of the target block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.
[0406] As a method for determining whether to apply the OBMC process, there is a method using, for example, obmc_flag, a signal indicating whether the OBMC process is applied. As a specific example, the encoding device may also determine whether the target block belongs to a region with complex motion. When the target block belongs to a region with complex motion, the encoding device sets the value of obmc_flag to 1 and applies the OBMC process to perform encoding. When the target block does not belong to a region with complex motion, the encoding device sets the value of obmc_flag to 0 and does not apply the OBMC process to perform encoding of the block. On the other hand, in the decoding device, by decoding obmc_flag described in the stream (for example, a compressed sequence), whether to apply the OBMC process is switched according to the value to perform decoding.
[0407] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter-frame prediction unit 126 may generate a plurality of prediction images having shapes different from the rectangle for the rectangular current block, and may generate a final rectangular prediction image by combining these plurality of prediction images. The shape different from the rectangle may be, for example, a triangle.
[0408] Fig.37 This is a diagram for explaining the generation of predicted images of two triangles.
[0409] The inter-frame prediction unit 126 generates a triangular prediction image by performing motion compensation on the first triangular partition in the current block using the first MV of the first partition. Similarly, the inter-frame prediction unit 126 generates a triangular prediction image by performing motion compensation on the second triangular partition in the current block using the second MV of the second partition. Then, the inter-frame prediction unit 126 generates a rectangular prediction image that is the same as the current block by combining these prediction images.
[0410] In addition, Fig.37 In the example shown, the first partition and the second partition are each a triangle, but they may also be a trapezoid or may be different shapes from each other. Fig.37 In the example shown, the current block is composed of 2 partitions, but it can also be composed of 3 or more partitions.
[0411] In addition, the first partition and the second partition may also be repeated. That is, the first partition and the second partition may also include the same pixel region. In this case, the predicted image in the first partition and the predicted image in the second partition may be used to generate the predicted image of the current block.
[0412] In addition, although this example shows an example in which predicted images are generated by inter-frame prediction for both of the two partitions, a predicted image may be generated by intra-frame prediction for at least one partition.
[0413] [Motion compensation>BIO]
[0414] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as a BIO (bidirectional optical flow) mode.
[0415] Fig.38 This is a diagram for explaining a model that assumes uniform linear motion. Fig.38 In (v x , v y ) represents the velocity vector, τ 0 , τ 1 Respectively represent the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ) in time. (MVx 0 , MVy 0 ) indicates the reference image Ref 0 The corresponding motion vector, (MVx 1 , MVy 1 ) indicates the reference image Ref 1 The corresponding motion vector.
[0416] At this time, the velocity vector (v x , v y ) is assumed to be a uniform linear motion, (MVx 0 , MVy 0 ) and (MVx 1 , MVy 1 ) are respectively expressed as (vxτ 0 , vyτ 0 ) and (-vxτ 1 , -vyτ 1 ), the following optical flow equation (2) holds.
[0417]
Formula 3
[0418]
[0419] Here, I(k) represents the brightness value of the reference image k (k=0, 1) after motion compensation. The optical flow equation indicates that the sum of (i) the temporal differential of the brightness value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, the motion vector of the block unit obtained from the merge list or the like may be corrected in pixel units based on a combination of the optical flow equation and Hermite interpolation.
[0420] In addition, the motion vector may be derived on the decoding device side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0421] [Motion Compensation > LIC]
[0422] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.
[0423] Fig.39 This is a diagram for explaining an example of a method of generating a predicted image using a brightness correction process based on an LIC process.
[0424] First, MV is derived from the encoded reference picture to obtain the reference image corresponding to the current block.
[0425] Next, information indicating how the brightness value changes in the reference picture and the current picture is extracted for the current block. This extraction is based on the brightness pixel values of the encoded left adjacent reference area (peripheral reference area) and the encoded upper adjacent reference area (peripheral reference area) in the current picture, and the brightness pixel values at the same position in the reference picture specified by the derived MV. Then, using the information indicating how the brightness value changes, the brightness correction parameter is calculated.
[0426] The brightness correction parameter is applied to perform brightness correction processing on the reference image in the reference picture specified by the MV, thereby generating a predicted image for the current block.
[0427] in addition, Fig.39 The shape of the peripheral reference area described above is just an example, and shapes other than this may also be used.
[0428] In addition, the process of generating a predicted image based on one reference picture is described here, but the same is true when generating a predicted image based on multiple reference pictures. The predicted image can also be generated after brightness correction processing is performed on the reference images obtained from each reference picture in the same way as described above.
[0429] As a method for determining whether to use LIC processing, there is a method of using lic_flag, which is a signal indicating whether to use LIC processing. As a specific example, in the encoding device, it is determined whether the current block belongs to an area where brightness changes. If it belongs to an area where brightness changes, the value 1 is set as lic_flag, and encoding is performed using LIC processing. If it does not belong to an area where brightness changes, the value 0 is set as lic_flag, and encoding is performed without using LIC processing. On the other hand, in the decoding device, it is also possible to decode lic_flag described in the stream and switch whether to use LIC processing according to its value to perform decoding.
[0430] As another method for determining whether to use the LIC process, there is a method of determining whether the LIC process is used in the surrounding blocks. As a specific example, when the current block is in the merge mode, it is determined whether the surrounding coded blocks selected when deriving the MV in the merge mode process are coded using the LIC process, and based on the result, it is switched whether to use the LIC process for coding. In addition, in the case of this example, the same process is also applied to the decoding device side.
[0431] use Fig.39 The LIC process (luminance correction process) has been described above, and its details will be described below.
[0432] First, the inter prediction unit 126 derives a motion vector for acquiring a reference image corresponding to a current block to be encoded from a reference picture that is an already encoded picture.
[0433] Next, the inter-frame prediction unit 126 uses the brightness pixel values of the encoded surrounding reference areas adjacent to the left and above and the brightness pixel values at the same position in the reference picture specified by the motion vector to extract information indicating how the brightness values in the reference picture and the encoding target picture change, and calculates the brightness correction parameter. For example, the brightness pixel value of a certain pixel in the surrounding reference area in the encoding target picture is set to p0, and the brightness pixel value of the pixel in the surrounding reference area in the reference picture at the same position as the pixel is set to p1. The inter-frame prediction unit 126 calculates the coefficients A and B for optimizing A×p1+B=p0 as brightness correction parameters for multiple pixels in the surrounding reference area.
[0434] Next, the inter-frame prediction unit 126 generates a predicted image for the encoding target block by performing a brightness correction process on the reference image in the reference picture specified by the motion vector using the brightness correction parameter. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after the brightness correction process is set to p3. The inter-frame prediction unit 126 generates a predicted image after the brightness correction process by calculating A×p2+B=p3 for each pixel in the reference image.
[0435] also, Fig.39 The shape of the peripheral reference area in is an example, and other shapes may also be used. Fig.39 For example, a region including a predetermined number of pixels thinned out from the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference region. In addition, the peripheral reference region is not limited to the region adjacent to the encoding target block, and may also be a region not adjacent to the encoding target block. Fig.39 In the example shown, the peripheral reference area in the reference picture is an area specified by the motion vector of the encoding target picture from the peripheral reference area in the encoding target picture, but it may also be an area specified by another motion vector. For example, the other motion vector may also be the motion vector of the peripheral reference area in the encoding target picture.
[0436] In addition, although the operation in the encoding device 100 is described here, the operation in the decoding device 200 is also the same.
[0437] The LIC process is not only applied to brightness but also to color difference. In this case, correction parameters may be derived separately for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.
[0438] Furthermore, the LIC process may be applied in sub-block units. For example, the modification parameters may be derived using the surrounding reference area of the current sub-block and the surrounding reference area of the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0439] [Prediction Control Department]
[0440] The prediction control unit 128 selects one of the intra prediction signal (the signal output from the intra prediction unit 124) and the inter prediction signal (the signal output from the inter prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0441] like Figure 1As shown, in various implementation examples, the prediction control unit 128 may also output prediction parameters input to the entropy coding unit 110. The entropy coding unit 110 may generate a coded bit stream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may also be used in a decoding device. The decoding device may also receive the coded bit stream and decode it, performing the same processing as the prediction processing performed in the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction parameters may include a selection prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra-frame prediction unit 124 or the inter-frame prediction unit 126), or any index, flag, or value based on the prediction processing performed in the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128 or indicating the prediction processing.
[0442] [Encoding device installation example]
[0443] Fig.40 1 is a block diagram showing an implementation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. Figure 1 The multiple components of the encoding device 100 shown are Fig.40 The processor a1 and the memory a2 shown are implemented.
[0444] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. Processor a1 can also be a processor such as a CPU. In addition, processor a1 can also be a collection of multiple electronic circuits. In addition, for example, processor a1 can also play a role in Figure 1 The functions of the multiple components of the encoding device 100 other than the components for storing information are shown in the above.
[0445] The memory a2 is a dedicated or general-purpose memory for storing information used by the processor a1 to encode motion images. The memory a2 can be an electronic circuit or connected to the processor a1. In addition, the memory a2 can also be included in the processor a1. In addition, the memory a2 can also be a collection of multiple electronic circuits. In addition, the memory a2 can be a magnetic disk or an optical disk, or can be expressed as a storage or a recording medium. In addition, the memory a2 can be a non-volatile memory or a volatile memory.
[0446] For example, the memory a2 may store the encoded moving image, or may store a bit string corresponding to the encoded moving image. In addition, the memory a2 may store a program for the processor a1 to encode the moving image.
[0447] In addition, for example, memory a2 can also serve Figure 1 The memory a2 can be used as a component for storing information among the multiple components of the encoding device 100 shown in FIG. Figure 1 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks and reconstructed pictures, etc.
[0448] In addition, in the encoding device 100, it is not necessary to install Figure 1 All of the multiple constituent elements shown above may not perform all of the multiple processes described above. Figure 1 Part of the plurality of components shown in the figure may be included in other devices, and part of the plurality of processes described above may be executed by other devices.
[0449] [Decoding device]
[0450] Next, a decoding device that can decode the coded signal (coded bit stream) output from the coding device 100 described above will be described. Fig.41 2 is a block diagram showing a functional structure of a decoding device 200 according to the present embodiment. The decoding device 200 is a moving picture decoding device that decodes a moving picture in units of blocks.
[0451] like Fig.41 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.
[0452] The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when the processor executes a software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. In addition, the decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0453] Hereinafter, after describing the overall processing flow of the decoding device 200 , each component included in the decoding device 200 will be described.
[0454] [Overall flow of decoding processing]
[0455] Fig.42 This is a flowchart showing an example of the overall decoding process performed by the decoding device 200.
[0456] First, the entropy decoding unit 202 of the decoding device 200 determines a partition pattern of a fixed-size block (128×128 pixels) (step Sp_1). This partition pattern is the partition pattern selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 on each of the multiple blocks constituting the partition pattern.
[0457] That is, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the decoding target block (also referred to as the current block) (step Sp_2).
[0458] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore a plurality of prediction residuals (ie, difference blocks) by performing inverse quantization and inverse transformation on the plurality of quantized coefficients (step Sp_3).
[0459] Next, the prediction processing unit composed of all or part of the intra prediction unit 216, the inter prediction unit 218 and the prediction control unit 220 generates a prediction signal (also referred to as a prediction block) of the current block (step Sp_4).
[0460] Next, the adding unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the differential block (step Sp_5).
[0461] Then, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).
[0462] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7 ), and when it is determined that the decoding has not been completed (No in step Sp_7 ), the processing from step Sp_1 is repeatedly executed.
[0463] In addition, the processing of these steps Sp_1 to Sp_7 can be performed sequentially by the decoding device 200, and a plurality of processing of some of these processing can be performed in parallel or in a different order.
[0464] [Entropy decoding unit]
[0465] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. Then, the entropy decoding unit 202 debinarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the coded bit stream to the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220 (see Figure 1 The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.
[0466] [Inverse Quantization Unit]
[0467] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) as the input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. In addition, the inverse quantization unit 204 outputs the inversely quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0468] [Inverse transformation unit]
[0469] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients which are input from the inverse quantization unit 204 .
[0470] For example, when the information read from the encoded bit stream indicates that EMT or AMT is adopted (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.
[0471] Furthermore, for example, when the information read out from the encoded bit stream indicates that NSST is adopted, the inverse transform unit 206 applies an inverse re-transform to the transform coefficients.
[0472] [Addition Department]
[0473] The adding unit 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction sample input from the prediction control unit 220 . The adding unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212 .
[0474] [Block Memory]
[0475] The block memory 210 is a storage unit for storing blocks in a decoding target picture (hereinafter referred to as a current picture) to be referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the adding unit 208 .
[0476] [Loop filter section]
[0477] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208 , and outputs the reconstructed block after filtering to the frame memory 214 , a display device, and the like.
[0478] When the information indicating the on / off of ALF read from the coded bit stream indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0479] [Frame Memory]
[0480] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212 .
[0481] [Prediction processing unit (intra-frame prediction unit / inter-frame prediction unit / prediction control unit)]
[0482] Fig.43 2 is a diagram showing an example of processing performed by the prediction processing unit of the decoding device 200. The prediction processing unit is composed of all or part of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0483] The prediction processing unit generates a prediction image of the current block (step Sq_1). The prediction image is also called a prediction signal or a prediction block. In addition, the prediction signal includes, for example, an intra-frame prediction signal or an inter-frame prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring a difference block, and generating a decoded image block.
[0484] The reconstructed image may be, for example, an image of a reference picture, or may be a picture including the current block, that is, an image of a decoded block in the current picture. The decoded block in the current picture may be, for example, an adjacent block of the current block.
[0485] Fig.44 This is a diagram showing another example of processing performed by the prediction processing unit of the decoding device 200.
[0486] The prediction processing unit determines a method or mode for generating a prediction image (step Sr_1). For example, the method or mode can be determined based on prediction parameters or the like.
[0487] When it is determined that the first mode is a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the first mode (step Sr_2a). In addition, when it is determined that the second mode is a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the second mode (step Sr_2b). In addition, when it is determined that the third mode is a mode for generating a predicted image, the prediction processing unit generates a predicted image according to the third mode (step Sr_2c).
[0488] The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, or other prediction methods. In such a prediction method, the above-mentioned reconstructed image may also be used.
[0489] [Intra-frame prediction unit]
[0490] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the coded bit stream, and refers to the block in the current picture stored in the block memory 210, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 performs intra prediction by referring to samples (e.g., luminance values, color difference values) of blocks adjacent to the current block, thereby generating an intra prediction signal, and outputs the intra prediction signal to the prediction control unit 220.
[0491] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0492] Furthermore, when the information read from the coded bit stream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.
[0493] [Inter-frame prediction unit]
[0494] The inter-frame prediction unit 218 predicts the current block with reference to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or a sub-block (e.g., a 4×4 block) in the current block. For example, the inter-frame prediction unit 218 performs motion compensation using motion information (e.g., a motion vector) read from the coded bit stream (e.g., prediction parameters output from the entropy decoding unit 202), thereby generating an inter-frame prediction signal of the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0495] When the information read from the encoded bit stream indicates that the OBMC mode is adopted, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of the neighboring blocks.
[0496] Furthermore, when the information read from the coded bit stream indicates that the FRUC mode is adopted, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the coded stream, thereby deriving motion information. Furthermore, the inter-frame prediction unit 218 performs motion compensation (prediction) using the derived motion information.
[0497] In addition, when the BIO mode is adopted, the inter-frame prediction unit 218 derives the motion vector based on the model assuming uniform linear motion. In addition, when the information read from the encoded bit stream indicates that the affine motion compensation prediction mode is adopted, the inter-frame prediction unit 218 derives the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks.
[0498] [MV export > Normal interframe mode]
[0499] When the information read from the encoded bit stream indicates that the normal inter mode is applied, the inter prediction section 218 derives an MV based on the information read from the encoded bit stream, and performs motion compensation (prediction) using the MV.
[0500] Fig.45 This is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding device 200 .
[0501] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple decoded blocks that are temporally or spatially around the current block (step Ss_1). In other words, the inter-frame prediction unit 218 creates a candidate MV list.
[0502] Next, the inter-frame prediction unit 218 extracts N (N is an integer greater than or equal to 2) candidate MVs from the multiple candidate MVs obtained in step Ss_1 as predicted motion vector candidates (also referred to as predicted MV candidates) in a predetermined priority order (step Ss_2). In addition, the priority order is predetermined for each of the N predicted MV candidates.
[0503] Next, the inter-frame prediction unit 218 decodes the predicted motion vector selection information from the input stream (i.e., the encoded bit stream), and uses the decoded predicted motion vector selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted motion vector (also called predicted MV) of the current block (step Ss_3).
[0504] Next, the inter prediction section 218 decodes the difference MV from the input stream, and derives the MV of the current block by adding the difference value that is the decoded difference MV to the selected prediction motion vector (step Ss_4).
[0505] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).
[0506] [Prediction Control Department]
[0507] The prediction control unit 220 selects one of the intra prediction signal and the inter prediction signal, and outputs the selected signal as a prediction signal to the adder 208. In general, the structures, functions, and processes of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side may correspond to the structures, functions, and processes of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.
[0508] [Decoding device installation example]
[0509] Fig.46 2 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. Fig.41 The multiple components of the decoding device 200 shown are Fig.46 The processor b1 and the memory b2 are installed as shown.
[0510] Processor b1 is a circuit that performs information processing and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general electronic circuit that decodes the encoded moving image (i.e., the encoded bit stream). Processor b1 can also be a processor such as a CPU. In addition, processor b1 can also be a collection of multiple electronic circuits. In addition, for example, processor b1 can also play a role in Fig.41 The functions of the multiple components of the decoding device 200 other than the components for storing information are shown in the above.
[0511] The memory b2 is a dedicated or general memory for storing information used by the processor b1 to decode the coded bit stream. The memory b2 may be an electronic circuit, or may be connected to the processor b1. In addition, the memory b2 may be included in the processor b1. In addition, the memory b2 may be a collection of multiple electronic circuits. In addition, the memory b2 may be a magnetic disk or an optical disk, or may be a storage device or a recording medium. In addition, the memory b2 may be a non-volatile memory or a volatile memory.
[0512] For example, the memory b2 may store moving images or coded bit streams. In addition, the memory b2 may also store a program for the processor b1 to decode the coded bit stream.
[0513] In addition, for example, memory b2 can serve as Fig.41 The memory b2 can be used as a component for storing information among the multiple components of the decoding device 200 shown in FIG. Fig.41 The functions of the block memory 210 and the frame memory 214 are shown. More specifically, the memory b2 can store reconstructed blocks and reconstructed pictures, etc.
[0514] In addition, in the decoding device 200, it is not necessary to install Fig.41 All of the multiple constituent elements shown above may not perform all of the multiple processes described above. Fig.41 Part of the plurality of components shown in the figure may be included in other devices, and part of the plurality of processes described above may be executed by other devices.
[0515] [Definition of terms]
[0516] As an example, each term may be defined as follows.
[0517] A picture is an arrangement of multiple luma samples in monochrome format, or an arrangement of multiple luma samples and 2 corresponding arrangements of multiple color difference samples in color formats of 4:2:0, 4:2:2, and 4:4:4. A picture can be a frame or a field.
[0518] A frame is a composition of a top field that generates a plurality of sample rows 0, 2, 4, ... and a bottom field that generates a plurality of sample rows 1, 3, 5, ...
[0519] A slice is an integer number of coding tree units contained in 1 independent slice segment and all subsequent dependent slice segments before the next independent slice segment (if any) in the same access unit.
[0520] A tile is a rectangular region of multiple coding tree blocks within a specific tile column and a specific tile row in a picture. A tile can still apply a loop filter across the edge of the tile, but can also be a rectangular region of a frame that is intended to be independently decodable and coded.
[0521] A block is an MxN (N rows and M columns) arrangement of multiple samples, or an MxN arrangement of multiple transform coefficients. A block can also be a square or rectangular area of multiple pixels consisting of multiple matrices of 1 luminance and 2 chrominance.
[0522] A CTU (Coding Tree Unit) can be a coding tree block of multiple luma samples of a picture with 3 sample arrangements, or 2 corresponding coding tree blocks of multiple chrominance samples. Alternatively, a CTU can be a coding tree block of any number of samples in a monochrome picture and a picture encoded using the syntax structure used in the encoding of 3 separate color planes and multiple samples.
[0523] A super block is composed of one or two mode information blocks, or may be recursively divided into four 32×32 blocks, or further divided into a square block of 64×64 pixels.
[0524] [First Form]
[0525] Hereinafter, a coding device 100, a decoding device 200, a coding method, and a decoding method according to a first aspect of the present invention will be described.
[0526] Fig.47 2 is a flowchart showing an example of the inter-frame prediction process in the first aspect. Hereinafter, an example of the inter-frame prediction process in the decoding device 200 will be described.
[0527] The inter-frame prediction unit 218 in the decoding device 200 derives a reference motion vector for predicting a processing target block, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and decodes the processing target block using the changed first motion vector or the unchanged first motion vector. For example, the reference motion vector corresponds to a first pixel set in the processing target block, and the first motion vector corresponds to a second pixel set different from the first pixel set in the processing target block. The inter-frame prediction process will be described in more detail below with reference to the accompanying drawings. In addition, the reference motion vector and the first motion vector described below are examples and are not limited thereto.
[0528] First, in step S1001, the inter-frame prediction unit 218 derives a reference motion vector for a first pixel set in a processing target block. Fig.48 The processing of step S1001 will be described in more detail.
[0529] Fig.48 FIG. 1 is a diagram showing an example of a processing target block. Fig.48 As shown in FIG. 1 , the processing target block may also be composed of a plurality of sub-blocks (sub-block 0 to sub-block 5). In addition, each sub-block may have a different first motion vector (MV 0 ~MV 5 ). An example of the first pixel set may also be sub-block 0. In this case, the reference motion vector is one of the plurality of first motion vectors in the processing target block. In this example, the reference motion vector is MV 0 .
[0530] Another example of the first pixel set may be the entire processing target block. In this example, the reference motion vector is the average of the first motion vectors of all sub-blocks in the processing target block, that is, the reference motion vector is the average of the first motion vectors of all sub-blocks in the processing target block. 0 To MV 5 Average.
[0531] Next, in step S1002, the inter-frame prediction unit 218 derives the first motion vector of the second pixel set in the processing target block. The second pixel set is different from the first pixel set. An example of the second pixel set may also be sub-block 2. In this example, the first motion vector of the second pixel set is MV 2 .
[0532] In addition, another example of the second pixel set may be sub-block 1. In this example, the first motion vector is MV 1 .
[0533] Next, in step S1003, the inter-frame prediction unit 218 derives a differential vector based on the difference between the reference motion vector of the first pixel set and the first motion vector of the second pixel set. An example of the value of the reference motion vector may be (-3, 4). An example of the first motion vector may be (16, 5). Therefore, the differential motion vector between the first motion vector and the reference motion vector becomes (-19, -1).
[0534] Next, in step S1004, the inter-frame prediction unit 218 determines whether the differential motion vector derived in step S1003 is greater than a threshold. The threshold is a set of a first value (hereinafter referred to as the first threshold) and a second value (hereinafter referred to as the second threshold). An example of the threshold may also be (10, 20). An example of the determination process is described below.
[0535] When the absolute value of the horizontal component of the differential motion vector derived in step S1003 is greater than the first threshold value, or the absolute value of the vertical component of the differential motion vector is greater than the second threshold value, the inter-frame prediction unit 218 determines that the differential motion vector is greater than the threshold value (Yes in step S1004). In other words, if any one of the absolute value of the horizontal component and the absolute value of the vertical component of the differential motion vector is greater than the threshold value, the differential motion vector is determined to be greater than the threshold value. For example, when the absolute value of the horizontal component of the differential motion vector (-19, -1) illustrated in step S1003 is compared with the first threshold value, since |-19|>10, it is determined that the differential motion vector is greater than the threshold value. In this case, the inter-frame prediction unit 218 changes the first motion vector. The details of the change are described in step S1006.
[0536] On the other hand, when the absolute value of the horizontal component of the differential motion vector is not greater than the first threshold value and the absolute value of the vertical component of the differential motion vector is not greater than the second threshold value, the inter-frame prediction unit 218 determines that the differential motion vector is not greater than the threshold value (No in step S1004). In other words, when the absolute value of the horizontal component and the absolute value of the vertical component of the differential motion vector are both less than the threshold value, it is determined that the differential motion vector is less than the threshold value. When the inter-frame prediction unit 218 determines that the differential motion vector is not greater than the threshold value (No in step S1004), it does not change the first motion vector (step S1005).
[0537] Next, in step S1006, when it is determined that the differential motion vector is greater than the threshold value (Yes in step S1004), the inter-frame prediction unit 218 changes the first motion vector using the value after clipping the differential motion vector. At this time, the differential motion vector (-19, -1) is clipped to (-10, -1). As described above, since the absolute value of the horizontal component of the differential motion vector |-19| is greater than 10, which is the first threshold value (10, 20), the absolute value of the horizontal component of the differential motion vector is clipped to match the first threshold value. Therefore, the clipped differential motion vector becomes (-10, -1).
[0538] Next, the inter-frame prediction unit 218 changes the first motion vector using the clipped differential motion vector and the reference motion vector. More specifically, the first motion vector may be changed by adding the differential motion vector to the reference motion vector. For example, the changed first motion vector is (-3, 4)-(-10, -1)=(7, 5).
[0539] In step S1007, the inter prediction unit 218 decodes the second pixel set using the changed first motion vector or the unchanged first motion vector. For example, sub-block 2 is decoded using the changed first motion vector (7, 5).
[0540] also, Fig.47 The illustrated process 1000 may also be a process of an encoding device.
[0541] Process 1000 may be applied to all sub-blocks in the processing target block. When process 1000 is applied to all sub-blocks in the processing target block, the prediction mode of the processing target block may also be an affine mode. In addition, when process 1000 is applied to all sub-blocks in the processing target block, the prediction mode of the processing target block may also be an ATMVP (Alternative Temporal Motion Vector Prediction: optional temporal motion vector prediction) mode.
[0542] In addition, the ATMVP mode is an example of a sub-block mode classified as a merge mode. For example, in an encoded reference picture specified by the MV (MV0) of a block adjacent to the lower left of the current block, a temporal MV reference block corresponding to the current block is identified, and for each sub-block in the current block, the MV used for encoding the area corresponding to the sub-block in the temporal MV reference block is identified.
[0543] Furthermore, the first motion vector of another sub-block of the processing target block may be updated using the first motion vector changed in a specific sub-block.
[0544] For example, the MV of sub-block 0 0 Determine (Japanese: Tongding) as the reference motion vector, and set the MV of sub-block 2 2 The first motion vector is determined as the first motion vector, and the changed first motion vector is determined using the process 1000. The changed first motion vector of sub-block 2 is determined by MV 2 '=(V 2x ',V 2y ') indicates that the reference motion vector MV is used. 0 and the changed first motion vector MV 2 ', the changed first motion vector MV of the sub-block i (i=1, 3, 4, 5) other than the sub-block 2 in the processing target block is calculated by the following formula: i '=(V ix ',V iy ').
[0545] V ix '=(V 2x '-V 0x )*POS ix / W-(V 2y '-V 0y )*POS iy / W+V 0x
[0546] V iy '=(V 2y '-V 0y )*POS ix / W+(V 2x '-V 0x )*POS iy / W+V 0y
[0547] Then, in the MV i 'If it is not greater than the threshold, it is used as MV i 'Update MV i , use MV i Decode sub-block i. POS ixand POS iy are the horizontal and vertical positions of sub-block i.
[0548] exist Fig.48 In the example,
[0549] POS 1x =W / 2, POS 1y =0
[0550] POS 3x =0, POS 3y =H
[0551] POS 4x =W / 2, POS 4y =H
[0552] POS 5x =W, POS 5y =H
[0553] W and H are the horizontal and vertical positions of the sub-block 2 (eg, the X and Y coordinates of the lower left corner).
[0554] The threshold value may also correspond to the size of the processing target block. For example, the larger the size of the processing target block, the larger the threshold value.
[0555] In addition, the threshold value may correspond to the number of reference pictures. For example, the greater the number of reference pictures, the smaller the threshold value.
[0556] The threshold value may also be encoded into a header area such as an SPS header, a PPS header, or a slice header.
[0557] The threshold may also be predetermined so that the stream is not decoded.
[0558] The threshold value may be determined so as to limit the worst case memory access amount when a motion compensation process (prediction process) is performed on a processing target block in a predetermined prediction mode to be less than the memory access amount when a bidirectional motion compensation process (prediction process) is performed on the processing target block for every 8×8 pixels in a prediction mode other than the prediction mode. The predetermined prediction mode is, for example, an affine mode. An example of the limitation is described below.
[0559] “Mem_base” indicates the worst case memory access amount when bidirectional prediction motion compensation processing (luminance value) is performed on a processing target block for each 8×8 block in a prediction mode other than a predetermined prediction mode, with one 8×8 block as one unit.
[0560] The worst case memory access amount when motion compensation processing (luminance value) is performed on a processing target block (size: M×N) in a predetermined prediction mode is represented by “Mem_CU”, and the threshold value “Mem_th” therefor is represented as follows using “Mem_base”.
[0561] Mem_th=M×N / (8×8)*Mem_base
[0562] Mem_CU is limited to be not larger than Mem_th. An example of threshold value calculation processing will be described below.
[0563] Assuming that the size of each pixel of the luminance signal is 1 byte and the number of taps of the filter for motion compensation processing is 8, Mem_base=(8+7)*(8+7)*2=450 (byte unit).
[0564] Assuming that the prediction mode of the processing target block is the affine mode and the size of the processing target block is 64×64 pixels, Mem_th=64 / 8*64 / 8*450=28800 (byte unit).
[0565] It is assumed that “H” and “V” represent a threshold value of a first component of the threshold value (first threshold value) and a threshold value of a second component of the threshold value (second threshold value), respectively.
[0566] The worst condition when prediction processing is performed in a normal inter-frame mode other than the affine mode is as follows: the processing object block is divided into 8×8 pixel blocks, and all 8×8 pixel blocks are motion compensated by bidirectional prediction. The range of the memory that can be referenced when the prediction processing is performed in the affine mode is determined in a manner that does not exceed the memory access amount of the worst condition. Assume that the range of the memory that can be referenced in the affine mode is calculated by the following formula [1]. Calculate H and V so that the range of memory access in the affine mode is less than Mem_th (for example, 28800). In addition, here, an example of the case where bidirectional prediction is performed in this affine mode is described.
[0567] Mem_CU=(64+7+2*H)(64+7+2*V)*2≦28800[1]
[0568] Assuming H = V, solve equation [1] and we get H = V = 24.
[0569] The above-described calculation process of the threshold value is also applied to processing target blocks having different sizes and different numbers of reference pictures. Fig.49 is a diagram showing an example of a calculated threshold value. Fig.49 Although an example of H=V is shown in FIG. 1 , H>V or H<V may also be possible.
[0570] like Fig.49 As shown, the threshold value is different depending on whether the processing object block is unidirectionally predicted (the number of reference frames = 1) or bidirectionally predicted (the number of reference frames = 2). In the example of the above formula [1], the size of the processing object block predicted in the affine mode is 64×64. In the case of bidirectional prediction, that is, when referring to two reference pictures, the first threshold value H and the second threshold value V are both 24. For example, in the case of a 64×64 processing object block predicted in the affine mode in the unidirectional prediction, that is, when referring to only one reference picture, the first threshold value H and the second threshold value V are both 49. Fig.49 In , the above formula [1] is used to calculate the thresholds of all block sizes that can be predicted in the affine mode and tabulate them.
[0571] [Technical advantages of the first form]
[0572] In the first aspect of the present invention, the first motion vector derivation process is introduced in the inter-frame prediction process. As described above, the process of selecting an appropriate first motion vector so that the deviation of multiple first motion vectors in the processing target block converges within a predetermined range can reduce the memory bandwidth of the inter-frame prediction process.
[0573] [Replenish]
[0574] The encoding device 100 and the decoding device 200 in this embodiment can be used as an image encoding device and an image decoding device, respectively, or can be used as a moving picture encoding device and a moving picture decoding device, respectively.
[0575] Alternatively, the encoding device 100 and the decoding device 200 may be used as an entropy encoding device and an entropy decoding device, respectively. That is, the encoding device 100 and the decoding device 200 may correspond to only the entropy encoding unit 110 and the entropy decoding unit 202, respectively. Furthermore, other components may also be included in other devices.
[0576] In addition, at least a part of the present embodiment can be used as an encoding method, can be used as a decoding method, can be used as an entropy encoding method, can be used as an entropy decoding method, or can be used as other methods.
[0577] In addition, in the present embodiment, each component is composed of dedicated hardware, but can also be implemented by executing a software program suitable for each component. Each component can be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.
[0578] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to the processor a1 or b1, and the storage device corresponds to the memory a2 or b2.
[0579] The processing circuit includes at least one of dedicated hardware and a program execution unit, and uses the storage device to execute the processing. In addition, when the processing circuit includes the program execution unit, the storage device stores a software program executed by the program execution unit.
[0580] Here, the software that realizes the encoding device 100 or the decoding device 200 according to this embodiment is the following program.
[0581] For example, the program may also cause a computer to execute the following encoding method: the encoding method is an encoding method for encoding a moving image, deriving a reference motion vector used in predicting a processing object block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on a difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector if it is determined that the differential motion vector is greater than the threshold value, not changing the first motion vector if it is determined that the differential motion vector is not greater than the threshold value, and encoding the processing object block using the changed first motion vector or the unchanged first motion vector.
[0582] In addition, for example, the program can also cause a computer to execute the following decoding method: the decoding method is a decoding method for decoding a moving image, deriving a reference motion vector used in predicting a processing object block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on the difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector if it is determined that the differential motion vector is greater than the threshold value, and not changing the first motion vector if it is determined that the differential motion vector is not greater than the threshold value, and decoding the processing object block using the changed first motion vector or the unchanged first motion vector.
[0583] In addition, as described above, each component may also be a circuit. These circuits may constitute one circuit as a whole, or they may be different circuits. In addition, each component may be implemented by a general-purpose processor or a dedicated processor.
[0584] In addition, the processing performed by a specific component may be performed by other components. In addition, the order of performing the processing may be changed, and a plurality of processing may be performed simultaneously. In addition, the encoding and decoding device may include the encoding device 100 and the decoding device 200.
[0585] In addition, the ordinal numbers such as 1 and 2 used in the description may be replaced as appropriate. In addition, ordinal numbers may be newly assigned to components and the like, or ordinal numbers may be removed.
[0586] The above describes the forms of the encoding device 100 and the decoding device 200 based on the embodiment, but the forms of the encoding device 100 and the decoding device 200 are not limited to the embodiment. As long as it does not depart from the gist of the present invention, various modified forms that can be thought of by those skilled in the art are implemented on the present embodiment, and forms constructed by combining constituent elements in different embodiments can also be included in the scope of the forms of the encoding device 100 and the decoding device 200.
[0587] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0588] (Implementation Method 2)
[0589] [Implementation and Application]
[0590] In each of the above embodiments, each functional block or active block can usually be implemented by an MPU (microprocessing unit) and a memory. In addition, the processing of each functional block can also be implemented by a program execution unit such as a processor that reads and executes the software (program) recorded in a recording medium such as a ROM. The software can be distributed. The software can also be recorded in various recording media such as a semiconductor memory. In addition, each functional block can also be implemented by hardware (dedicated circuit).
[0591] The processing described in each embodiment can be implemented by centralized processing using a single device (system), or can be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing can be performed or distributed processing can be performed.
[0592] The aspects of the present invention are not limited to the above-described embodiments, and various modifications can be made, which are also included in the scope of the aspects of the present invention.
[0593] Furthermore, here, an application example of the moving picture encoding method (image encoding method) or the moving picture decoding method (image decoding method) shown in each of the above embodiments and various systems implementing the application example are described. Such a system may also be characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding and decoding device having both. Other structures of such a system can be appropriately changed according to circumstances.
[0594] [Example of use]
[0595] Fig.50 The figure shows the overall structure of a content supply system ex100 for realizing content distribution service. The communication service providing area is divided into desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the example shown in the figure, are respectively installed in each cell.
[0596] In the content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may also connect some of the above devices in combination. In various embodiments, the devices may be directly or indirectly connected to each other via a telephone network or short-range wireless instead of via base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to the devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101. Furthermore, the streaming server ex103 may be connected to a terminal in a hotspot in an airplane ex117 via a satellite ex116.
[0597] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly connected to the aircraft ex117 without going through the satellite ex116.
[0598] The camera ex113 is a device such as a digital camera capable of taking still images and moving images. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that supports mobile communication systems such as 2G, 3G, 3.9G, 4G, and 5G in the future.
[0599] The household appliance ex114 is a refrigerator or a device included in a household fuel cell cogeneration system.
[0600] In the content supply system ex100, a terminal having a camera function is connected to the streaming server ex103 via the base station ex106, etc., so that on-site distribution can be performed. In on-site distribution, a terminal (computer ex111, game machine ex112, camera ex113, home appliance ex114, smartphone ex115, terminal in airplane ex117, etc.) can perform the encoding process described in the above-mentioned embodiments on the still image or moving image content captured by the user using the terminal, multiplex the image data obtained by encoding and the sound data obtained by encoding the sound corresponding to the image, and transmit the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.
[0601] On the other hand, the streaming server ex103 streams the content data sent by the client that requested it. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117 that can decode the coded data. Each device that receives the distributed data decodes the received data and reproduces it. That is, each device can also function as an image decoding device according to one aspect of the present invention.
[0602] [Distributed processing]
[0603] In addition, the streaming media server ex103 can also be a plurality of servers or a plurality of computers, which distributes the data by distributing the processing or recording. For example, the streaming media server ex103 can also be implemented by CDN (Contents Delivery Network), which implements content distribution by connecting many edge servers scattered in the world to the network between the edge servers. In CDN, the edge servers that are physically closer are dynamically allocated according to the client. And, by caching and distributing the content to the edge servers, the delay can be reduced. In addition, when several types of errors occur or when the communication state changes due to the increase of the communication volume, it is possible to distribute the processing by using a plurality of edge servers, or switch the distribution subject to other edge servers, or bypass the part of the network where the failure occurs and continue to distribute, so that high-speed and stable distribution can be achieved.
[0604] In addition, the encoding process of the captured data is not limited to the distributed processing itself, and the encoding process can be performed by each terminal, on the server side, or shared. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or encoding amount of the image of the frame or scene unit is detected. In addition, in the second cycle, processing is performed to maintain the image quality and improve the encoding efficiency. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, the quality and efficiency of the content can be improved while reducing the processing load in each terminal. In this case, if there is a request for receiving and decoding almost in real time, the data completed by the first encoding by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.
[0605] As another example, the camera ex113 or the like extracts feature quantities from an image, compresses data on the feature quantities as metadata, and transmits the data to the server. The server determines the importance of an object based on the feature quantities, switches the quantization accuracy, and performs compression corresponding to the meaning of the image (or the importance of the content). The feature quantity data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression in the server. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a large processing load such as CABAC (context adaptive binary arithmetic coding).
[0606] As another example, in a stadium, shopping mall, or factory, there are multiple video data obtained by shooting roughly the same scene by multiple terminals. In this case, multiple terminals that have shot, and other terminals and servers that have not shot as needed, are used to distribute the encoding processing, such as GOP (Group of Picture) units, picture units, or tile units obtained by dividing pictures, so as to perform distributed processing. This can reduce delays and better achieve real-time performance.
[0607] Since the multiple image data are substantially the same scene, the server can also manage and / or instruct the image data taken by each terminal to refer to each other. In addition, the server can receive the encoded data from each terminal and change the reference relationship between the multiple data, or modify or replace the image itself and re-encode it. In this way, a stream with improved quality and efficiency of each data can be generated.
[0608] Furthermore, the server may perform transcoding to change the encoding method of the video data and distribute the video data. For example, the server may convert the encoding method of the MPEG class to the VP class (such as VP9), or convert H.264 to H.265.
[0609] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following uses "server" or "terminal" as the subject of the process, but part or all of the process performed by the server can also be performed by the terminal, and part or all of the process performed by the terminal can also be performed by the server. In addition, the same is true for the decoding process.
[0610] [3D, multi-angle]
[0611] There are more and more cases where images or videos taken by terminals such as multiple cameras ex113 and / or smartphones ex115 that are roughly synchronized with each other, or the same scene taken from different angles, are combined and used. The images taken by each terminal are combined based on the relative position relationship between the terminals obtained separately, or the area with the same feature points contained in the images.
[0612] The server not only encodes two-dimensional moving images, but also can encode still images automatically or at a user-specified time based on scene analysis of moving images and send them to the receiving terminal. When the server is able to obtain the relative position relationship between the shooting terminals, it can generate not only two-dimensional moving images but also three-dimensional shapes of the scene based on images of the same scene shot from different angles. The server can also encode three-dimensional data generated by point clouds, etc. separately, and can also select or reconstruct images from images shot by multiple terminals based on the results of identifying or tracking people or targets using three-dimensional data to generate images sent to the receiving terminal.
[0613] In this way, the user can arbitrarily select each image corresponding to each shooting terminal to enjoy the scene, and can also enjoy the content of the image of the selected viewpoint cut out from the three-dimensional data reconstructed using multiple images or images. Furthermore, along with the image, the sound can also be collected from multiple different angles, and the server multiplexes the sound from a specific angle or space with the corresponding image and sends the multiplexed image and sound.
[0614] In addition, in recent years, content that establishes correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server produces viewpoint images for the right eye and the left eye respectively, and can be encoded to allow reference between viewpoint images through Multi-View Coding (MVC) or the like, or encoded as different streams without reference to each other. When decoding different streams, they can be reproduced synchronously with each other according to the user's viewpoint to reproduce a virtual three-dimensional space.
[0615] In the case of AR images, the server may overlap the virtual object information in the virtual space with the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device obtains or maintains the virtual object information and three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and produces overlapping data by smoothly connecting them. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to the commission of the virtual object information. Alternatively, the server may produce overlapping data according to the three-dimensional data maintained in the server, matching the received movement of the viewpoint, encoding the overlapping data and distributing it to the decoding device. In addition, the overlapping data has an alpha value indicating the transmittance other than RGB, and the server sets the alpha value of the part other than the target produced according to the three-dimensional data to 0, etc., and encodes it in a transparent state of the part. Alternatively, the server may set the RGB value of a specified value as the background, as in a chroma key, to generate data that sets the part other than the target as the background color.
[0616] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, or it can be performed on the server side, or it can be shared and performed. As an example, a terminal may first send a receiving request to the server, and other terminals may receive the content corresponding to the request and perform decoding processing, and send the decoded signal to a device with a display. By distributing the processing regardless of the performance of the communicative terminal itself and selecting appropriate content, data with better image quality can be reproduced. In addition, as another example, large-size image data can also be received by a TV, etc., and a portion of the image such as tiles after being divided can be decoded and displayed by the personal terminal of the viewer. In this way, while making the overall image shared, it is possible to confirm one's own area of responsibility or the area that wants to be confirmed in more detail at hand.
[0617] In a situation where multiple short-range, medium-range or long-range wireless communications can be used indoors and outdoors, it may be possible to receive content seamlessly using distribution system standards such as MPEG-DASH. The user can also switch in real time while freely selecting the user's terminal, decoding devices or display devices such as displays configured indoors and outdoors. In addition, it is possible to switch the decoding terminal and the display terminal and perform decoding using its own location information, etc. As a result, it is also possible to map and display information on a part of the wall or ground of a building next to a displayable device while the user is moving to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or the encoded data being copied in an edge server of the content distribution service.
[0618] [Scalable Coding]
[0619] To switch content, use Fig.51 1. The scalable stream compressed and coded using the moving picture coding method described in the above embodiments is described as shown. For the server, a plurality of streams having the same content but different qualities may be provided as a single stream, or a structure in which the content is switched by utilizing the characteristics of a temporally / spatially scalable stream coded in layers as shown in the figure. That is, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band, so that the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when a user wants to watch a video that he / she has watched on the smartphone ex115 while on the move, for example, on an Internet TV or the like after returning home, the device only needs to decode the same stream into different layers, thereby reducing the burden on the server side.
[0620] Furthermore, in addition to the hierarchical structure of encoding the picture for each layer and realizing the enhancement layer above the base layer as described above, the enhancement layer may also include meta-information such as statistical information based on the image. Alternatively, the decoding side may generate high-definition content by super-resolving the picture of the base layer based on the meta-information. Super-resolution can improve the SN ratio while maintaining and / or expanding the resolution. The meta-information includes information for determining linear or nonlinear filter coefficients such as those used in the super-resolution processing, or information for determining parameter values in the filter processing, machine learning, or least squares operation used in the super-resolution processing.
[0621] Alternatively, a structure may be provided for dividing a picture into tiles according to the meaning of an object in the image. The decoding side decodes only a part of the area by selecting tiles to be decoded. Furthermore, by storing the attributes of the object (a person, a car, a ball, etc.) and the position in the image (the coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and determine the tile that includes the object. For example, Fig.52 As shown, the meta-information may be stored using a data storage structure different from the pixel data, such as SEI (supplemental enhancement information) messages in HEVC. The meta-information indicates, for example, the position, size, or color of the main object.
[0622] Meta information may also be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in the image, and by matching the information of the picture unit and the time information, it can determine the picture where the target exists and the position of the target in the picture.
[0623] [Web page optimization]
[0624] Fig.53 This is a diagram showing an example of a display screen of a web page on the computer ex111 or the like. Fig.54 1 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. Fig.53 and Fig.54 As shown, there are cases where a web page includes a plurality of link images that are links to image contents, and the way in which they are visible differs depending on the device being browsed. When a plurality of link images are visible on the screen, before a user explicitly selects a link image, or before a link image approaches the center of the screen, or before the entire link image enters the screen, the display device (decoding device) may display a still image or I picture of each content as a link image, may display an image such as a GIF animation using a plurality of still images or I pictures, or may receive only a base layer and decode and display the image.
[0625] When a linked image is selected by the user, the display device sets the base layer as the top priority and decodes it. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device can also decode to the enhancement layer. Moreover, in order to ensure real-time performance, before selection or when the communication band is very tight, the display device can reduce the delay between the decoding time and the display time of the head picture (the delay from the start of decoding of the content to the start of display) by decoding and displaying only the pictures that are referenced in the front (I pictures, P pictures, and B pictures that are only referenced in the front). Furthermore, the display device can also forcibly ignore the reference relationship of the pictures and set all B pictures and P pictures as forward references and roughly decode them, and perform normal decoding as the number of pictures received increases over time.
[0626] [Automatic driving]
[0627] Furthermore, when still images or video data such as two-dimensional or three-dimensional map information are transmitted and received for the purpose of automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta-information in addition to image data belonging to one or more layers, and decode them by establishing a correspondence between them. Furthermore, the meta-information may belong to a layer or may be multiplexed with image data alone.
[0628] In this case, since the car, drone, or airplane including the receiving terminal is moving, the receiving terminal can transmit the location information of the receiving terminal to achieve seamless reception and decoding while switching the base stations ex106 to ex110. In addition, the receiving terminal can dynamically switch the degree to which meta-information is received or the degree to which map information is updated according to the user's selection, the user's condition, and / or the state of the communication band.
[0629] In the content providing system ex100, the client can receive, decode and reproduce the encoded information sent by the user in real time.
[0630] [Distribution of Personal Content]
[0631] Furthermore, in the content supply system ex100, not only high-quality, long-duration contents provided by video distributors but also low-quality, short-duration contents provided by individuals can be unicasted or multicasted. Such personal contents are expected to increase in the future. In order to make personal contents better, the server may perform encoding after editing. This can be achieved, for example, with the following structure.
[0632] After taking pictures in real time or accumulating them, the server performs recognition processing such as shooting errors, scene search, meaning analysis and target detection based on the original image data or encoded data. In addition, based on the recognition results, the server manually or automatically corrects focus deviation or hand shaking, deletes scenes of low importance such as scenes with lower brightness than other pictures or scenes that are not in focus, or emphasizes the edges of the target, or changes the color tone. The server encodes the edited data based on the editing results. In addition, it is known that if the shooting time is too long, the viewing rate will decrease. The server can also automatically limit not only the scenes of low importance as mentioned above, but also the scenes with less movement based on the image processing results according to the shooting time, so as to become content within a specific time range. Alternatively, the server can also generate a summary based on the results of the scene's meaning analysis and encode it.
[0633] Personal content may be written with content that infringes copyright, author's personality rights or portrait rights in its original state, or it may be shared beyond the desired scope, which is inconvenient for individuals. Therefore, for example, the server may forcibly change the faces of people in the peripheral part of the screen or the home to an out-of-focus image for encoding. In addition, the server may also identify whether the face of a person different from the pre-registered person is captured in the encoded image, and if so, apply mosaics to the face part. Alternatively, as a pre-processing or post-processing of the encoding, the user may specify the person or background area that he wants to process the image from the perspective of copyright, etc. The server may also replace the specified area with another image, or blur the focus, etc. If it is a person, it is possible to track the person in the moving image and replace the image of the face part of the person.
[0634] The viewing and listening of personal content with a small amount of data has a strong requirement for real-time performance, so although it also depends on the bandwidth, the decoding device first receives, decodes and reproduces the basic layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the reproduction is repeated more than twice, the enhancement layer is also included to reproduce the high-definition image. In this way, if the stream is scalably encoded, it is possible to provide an experience in which the stream gradually becomes smoother and the image becomes better, although the moving image is relatively rough when it is not selected or at the beginning of viewing. In addition to scalable encoding, the same experience can be provided when the rougher stream reproduced for the first time and the second stream encoded with reference to the moving image of the first time are composed of one stream.
[0635] [Other application examples]
[0636] In addition, these encoding and decoding processes are usually processed in the LSI ex500 of each terminal. LSI (large scale integration circuitry) ex500 (refer to Fig.50 ) may be a single chip or a multi-chip structure. Alternatively, software for encoding or decoding moving images may be loaded into a recording medium (CD-ROM, floppy disk, hard disk, etc.) readable by the computer ex111 or the like, and the encoding or decoding may be performed using the software. Furthermore, if the smartphone ex115 has a camera, moving image data acquired by the camera may be transmitted. In this case, the moving image data is data encoded by the LSI ex500 of the smartphone ex115.
[0637] In addition, LSIex500 may also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. If the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software, and then obtains and reproduces the content.
[0638] Furthermore, the content supply system ex100 is not limited to the content supply system ex100 via the Internet ex101, and at least one of the video encoding device (video encoding device) or video decoding device (video decoding device) in each of the above-mentioned embodiments can be incorporated into a digital broadcasting system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by using broadcasting radio waves such as satellites, the content supply system ex100 is different from the structure that is easy for unicast in that it is suitable for multicast, but the same application can be made to the encoding process and the decoding process.
[0639] [Hardware Structure]
[0640] Fig.55 is a further detailed representation Fig.50 FIG. 115 is a diagram of a smartphone ex115 shown in FIG. Fig.56 1 is a diagram showing a configuration example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of shooting video and still images, and a display unit ex458 for displaying decoded data such as the video shot by the camera unit ex465 and the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing coded data or decoded data such as shot video or still images, recorded sound, received video or still images, and mails, and a slot unit ex464 as an interface unit with a SIM ex468 for identifying a user and authenticating access to various data such as a network. In addition, an external memory may be used instead of the memory unit ex467.
[0641] The main control unit ex460, which performs integrated control over the display unit ex458 and the operation unit ex466, is synchronously connected to the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / separation unit ex453, the sound signal processing unit ex454, the slot unit ex464 and the memory unit ex467 via a bus ex470.
[0642] When the power key is turned on by a user's operation, the power circuit unit ex461 activates the smartphone ex115 to be operable and supplies power to each unit from the battery pack.
[0643] The smartphone ex115 performs processes such as calls and data communications under the control of a main control unit ex460 including a CPU, ROM, and RAM. During a call, the voice signal collected by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, subjected to a spectrum diffusion process by the modulation / demodulation unit ex452, and subjected to a digital-to-analog conversion process and a frequency conversion process by the transmission / reception unit ex451, and the resulting signal is transmitted via the antenna ex450. In addition, the received data is amplified, subjected to a frequency conversion process and an analog-to-digital conversion process, subjected to a spectrum inverse diffusion process by the modulation / demodulation unit ex452, and converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. During data communications, text, still image, or video data is sent to the main control unit ex460 via the operation input control unit ex462 based on the operation of the operation unit ex466 of the main unit. The same transmission and reception processes are performed. In the data communication mode, when transmitting video, still images, or video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of taking a video or still image by the camera unit ex465, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a predetermined manner, and performs modulation and conversion processing on the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.
[0644] When receiving an image attached to an e-mail or a chat tool, or an image linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of image data and a bit stream of audio data, and supplies the encoded image data to the image signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronous bus ex470. The image signal processing unit ex455 decodes the image signal using a moving image decoding method corresponding to the moving image encoding method shown in the above-mentioned embodiments, and displays the image or still image included in the linked moving image file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs the audio from the audio output unit ex457. As real-time streaming media becomes more and more popular, the reproduction of audio may become socially inappropriate depending on the user's situation. Therefore, as the first value, it is preferable to have a configuration in which only the video data is reproduced without reproducing the audio signal, and the audio is reproduced synchronously only when the user performs an operation such as clicking on the video data.
[0645] In the description here, the smartphone ex115 is used as an example. However, as a terminal, three other installation forms can be considered, namely, a transmitting terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, it is assumed that multiplexed data in which audio data is multiplexed with video data is received and transmitted. However, in addition to audio data, character data related to the video may be multiplexed in the multiplexed data. In addition, the video data itself may be received or transmitted instead of the multiplexed data.
[0646] In addition, although the main control unit ex460 including a CPU is assumed to control the encoding or decoding process, various terminals are often equipped with a GPU. Therefore, a structure can also be made to use the performance of the GPU to process a larger area together through a memory shared by the CPU and GPU, or a memory that manages addresses in a way that can be used together. In this way, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset) and transformation / quantization processing are performed together in units such as pictures by the GPU instead of the CPU.
[0647] Industrial Applicability
[0648] The present invention can be utilized in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conference system, or an electronic mirror.
[0649] Description of Reference Numerals
[0650] 100 Encoding device
[0651] 102 Division
[0652] 104 Subtraction Department
[0653] 106 Transformation Department
[0654] 108 Quantitative Department
[0655] 110 Entropy Coding Unit
[0656] 112, 204 Inverse Quantization Unit
[0657] 114, 206 Inverse transformation unit
[0658] 116, 208 Addition Department
[0659] 118, 210 block memory
[0660] 120, 212 loop filter unit
[0661] 122, 214 frame memory
[0662] 124, 216 Intra-frame prediction unit
[0663] 126, 218 Inter-frame prediction unit
[0664] 128, 220 Prediction and Control Department
[0665] 200 Decoding device
[0666] 202 Entropy Decoding Department
[0667] 1000 Processing
[0668] 1201 Boundary Judgment Department
[0669] 1202, 1204, 1206 Switches
[0670] 1203 Filter determination unit
[0671] 1205 Filter Processing Unit
[0672] 1207 Filter characteristic determination unit
[0673] 1208 Processing and Judgment Department
[0674] a1, b1 processors
[0675] a2, b2 memory
Claims
1. A coding device for coding a moving image, in, The encoding device comprises: Circuits; and A memory connected to the circuit, wherein: The prediction mode of the current block to be encoded is affine mode, and In action, the circuit: Obtain the current block from a coding tree unit CTU; deriving a reference motion vector for predicting the current block; deriving a first motion vector different from the reference motion vector; deriving a motion vector difference based on a difference between the reference motion vector and the first motion vector; Determining whether the motion vector difference is greater than a threshold; If it is determined that the motion vector difference is greater than the threshold, a first value is set for the second motion vector, and if it is determined that the motion vector difference is not greater than the threshold, a second value different from the first value is set for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and The current block is encoded using the second motion vector, wherein: The threshold value is different in a case where the current block is unidirectionally predicted and in a case where the current block is bidirectionally predicted.
2. A decoding device for decoding a moving image, in, The decoding device comprises: Circuits; and A memory connected to the circuit, wherein: The prediction mode of the current block to be decoded is affine mode, and In action, the circuit: Obtain the current block from a coding tree unit CTU; deriving a reference motion vector for predicting the current block; deriving a first motion vector different from the reference motion vector; deriving a motion vector difference based on a difference between the reference motion vector and the first motion vector; Determining whether the motion vector difference is greater than a threshold; If it is determined that the motion vector difference is greater than the threshold, a first value is set for the second motion vector, and if it is determined that the motion vector difference is not greater than the threshold, a second value different from the first value is set for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and The current block is decoded using the second motion vector, wherein: The threshold value is different in a case where the current block is unidirectionally predicted and in a case where the current block is bidirectionally predicted.
3. A computer-readable non-transitory storage medium storing a bit stream, in, The prediction mode of the current block to be decoded is affine mode, The bit stream includes an encoded signal and syntax information, and the decoding device performs, according to the encoded signal and the syntax information: Obtain the current block from a coding tree unit CTU; deriving a reference motion vector for predicting the current block; deriving a first motion vector different from the reference motion vector; deriving a motion vector difference based on a difference between the reference motion vector and the first motion vector; Determining whether the motion vector difference is greater than a threshold; If it is determined that the motion vector difference is greater than the threshold, a first value is set for the second motion vector, and if it is determined that the motion vector difference is not greater than the threshold, a second value different from the first value is set for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and The current block is decoded using the second motion vector, wherein: The threshold value is different in a case where the current block is unidirectionally predicted and in a case where the current block is bidirectionally predicted.