Decoding apparatus, encoding apparatus, decoding method, and encoding method
Patent Information
- Application Number
- CN202310966936.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-26
- Filing Date
- 2018-09-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2038-09-20
AI Technical Summary
[0015]本发明能够提供能够抑制处理延迟的解码装置、编码装置、解码方法或编码方法。
Smart Images

Figure CN116886901B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on September 20, 2018, with application number 201880055108.7 and entitled "Encoding device, decoding device, encoding method and decoding method". Technical Field
[0002] This invention relates to encoding apparatus, decoding apparatus, encoding method, and decoding method. Background Technology
[0003] Previously, H.265 existed as a specification for encoding moving images. H.265 is also known as HEVC (High Efficiency Video Coding).
[0004] Existing technical documents
[0005] Non-patent literature
[0006] Non-patent literature 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention
[0007] The problem that the invention aims to solve
[0008] In such encoding and decoding methods, it is desirable to suppress processing latency.
[0009] The purpose of this invention is to provide a decoding device, encoding device, decoding method, or encoding method that can suppress processing delay.
[0010] Methods for solving problems
[0011] A decoding apparatus according to a technical solution of the present invention includes a circuit and a memory. The circuit uses the memory to perform the following processing in inter-frame prediction processing: determining whether the inter-frame prediction mode is a merging mode; when the inter-frame prediction mode is the merging mode, using the motion vector of a previously processed previous block, deriving a first motion vector of the first current block to be processed; deriving a second motion vector of the first current block by performing motion estimation near a position specified by the first motion vector; performing motion compensation using the second motion vector to generate a predicted image of the first current block; and when the first current block is to be processed... When the second current block, processed after the first block, is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. The fourth motion vector of the second current block is derived by performing motion estimation near the position specified by the third motion vector. Motion compensation is performed by using the fourth motion vector to generate a predicted image of the second current block.
[0012] An encoding apparatus according to a technical solution of the present invention includes a circuit and a memory. The circuit uses the memory and performs the following processing in inter-frame prediction processing: determining whether the inter-frame prediction mode is a merging mode; when the inter-frame prediction mode is the merging mode, deriving a first motion vector of the first current block to be processed using the motion vector of a previously processed previous block; deriving a second motion vector of the first current block by performing motion estimation near a position specified by the first motion vector; and generating a prediction map of the first current block by performing motion compensation using the second motion vector. For example, when the second current block to be processed after the first current block is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. The fourth motion vector of the second current block is derived by performing motion estimation near the position specified by the third motion vector. Motion compensation is performed by using the fourth motion vector to generate a predicted image of the second current block.
[0013] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0014] Invention Effects
[0015] The present invention can provide a decoding device, encoding device, decoding method or encoding method that can suppress processing delay. Attached Figure Description
[0016] Figure 1 This is a block diagram showing the functional structure of the encoding device according to Embodiment 1.
[0017] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1.
[0018] Figure 3 It is a table representing the transformation basis functions corresponding to each transformation type.
[0019] Figure 4A This is a diagram showing an example of the shape of the filter used in ALF.
[0020] Figure 4B This is another example of the shape of the filter used in ALF.
[0021] Figure 4C This is another example of the shape of the filter used in ALF.
[0022] Figure 5A This is a diagram representing the 67 intra-prediction modes of intra-frame prediction.
[0023] Figure 5B This is a flowchart illustrating the outline of predictive image correction processing based on OBMC processing.
[0024] Figure 5C This is a conceptual diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.
[0025] Figure 5D This is a diagram representing an example of FRUC.
[0026] Figure 6 It is a diagram used to illustrate pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0027] Figure 7 It is a diagram used to illustrate pattern matching (template matching) between a template in the current image and a block in a reference image.
[0028] Figure 8 It is a diagram used to illustrate a model that assumes uniform linear motion.
[0029] Figure 9A It is a diagram used to illustrate the derivation of the motion vectors of sub-block units based on the motion vectors of multiple adjacent blocks.
[0030] Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.
[0031] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.
[0032] Figure 9D This is a diagram used to illustrate an overview of a predictive image generation method that employs LIC-based brightness correction processing.
[0033] Figure 10 This is a block diagram illustrating the functional structure of the decoding device according to Embodiment 1.
[0034] Figure 11 This is a schematic diagram showing a first example of the assembly line structure of embodiment 1.
[0035] Figure 12 This is a schematic diagram illustrating an example of block segmentation used in the description of the pipeline processing of Implementation 1.
[0036] Figure 13 This is a time diagram illustrating an example of the processing timing in the first example of the pipeline structure related to Embodiment 1.
[0037] Figure 14 This is a flowchart of the inter-frame prediction process in the first example of the pipeline structure of Implementation 1.
[0038] Figure 15 This is a schematic diagram illustrating a second example of the assembly line structure related to Embodiment 1.
[0039] Figure 16 This is a time diagram showing an example of the processing timing in the second example of the pipeline structure related to Embodiment 1.
[0040] Figure 17 This is a flowchart of the inter-frame prediction process in the second example of the pipeline structure of Implementation 1.
[0041] Figure 18 This is a schematic diagram illustrating a third example of the assembly line structure related to Embodiment 1.
[0042] Figure 19 This is a time diagram showing an example of the processing timing in the third example of the pipeline structure related to Embodiment 1.
[0043] Figure 20 This is a flowchart of the inter-frame prediction process in the third example of the pipeline structure of Implementation 1.
[0044] Figure 21 This is a diagram showing an example of a motion vector for reference to Implementation 1.
[0045] Figure 22 This is a diagram showing an example of a motion vector for reference to Implementation 1.
[0046] Figure 23 This is a block diagram showing an example of the installation of an encoding device.
[0047] Figure 24 This is a block diagram illustrating an example of the installation of a decoding device.
[0048] Figure 25 This is a diagram showing the overall structure of a content supply system that enables content distribution services.
[0049] Figure 26 This is a diagram illustrating an example of encoding construction in the case of hierarchical encoding.
[0050] Figure 27 This is a diagram illustrating an example of encoding construction in the case of hierarchical encoding.
[0051] Figure 28 This is an example of a web page display.
[0052] Figure 29 This is an example of a web page display.
[0053] Figure 30 This is a diagram illustrating an example of a smartphone.
[0054] Figure 31 This is a block diagram representing a structural example of a smartphone. Detailed Implementation
[0055] An encoding apparatus according to a technical solution of the present invention includes a circuit and a memory. The circuit uses the memory to, when encoding an object block in an inter-frame prediction mode in which motion search is performed in a decoding apparatus, derive a first motion vector of the object block, store the derived first motion vector in the memory, derive a second motion vector of the object block, and generate a predicted image of the object block by using motion compensation of the second motion vector. In the deriving of the first motion vector, the first motion vector of the object block is derived using the first motion vector of the processed block.
[0056] Therefore, in pipeline control, the decoding device can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed. Thus, compared to exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device can be reduced, thereby reducing processing latency.
[0057] For example, in the derivation of the first motion vector, (i) using the first motion vector of the processed block, a list of predicted motion vectors representing multiple predicted motion vectors is generated, and (ii) the first motion vector of the object block is determined from the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0058] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be a merging mode, and the second motion vector may be derived by performing motion search processing on the periphery of the first motion vector.
[0059] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be FRUC mode, and the second motion vector may be derived by performing motion search processing on the periphery of the first motion vector.
[0060] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be FRUC mode. In the derivation of the second motion vector, (i) the third motion vector is determined based on the plurality of predicted motion vectors shown in the above-mentioned list of predicted motion vectors, and (ii) the second motion vector is derived by performing motion search processing on the periphery of the third motion vector.
[0061] For example, in determining the first motion vector, the first motion vector may be derived based on the average or central value of each of the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0062] For example, in determining the first motion vector, the predicted motion vector shown at the beginning of the predicted motion vector list among the plurality of predicted motion vectors shown in the predicted motion vector list can be determined as the first motion vector.
[0063] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block, and in determining the first motion vector, the first motion vector may be determined based on the candidate predicted motion vector derived using the second motion vector from among the plurality of predicted motion vectors shown in the above-mentioned list of predicted motion vectors.
[0064] Therefore, the encoding device can use a highly reliable second motion vector to determine the first motion vector, thus suppressing the decrease in the reliability of the first motion vector.
[0065] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block. If the processed block and the object block belong to the same image, the predicted motion vector may be derived using the first motion vector of the processed block. If the processed block and the object block belong to different images, the predicted motion vector may be derived using the second motion vector of the processed block.
[0066] Therefore, when the processed block and the object block belong to different images, the encoding device can improve the reliability of predicting motion vectors by using the second motion vector.
[0067] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block, and the first or second motion vector of the processed block may be used in the derivation of the predicted motion vectors may be determined based on the position of the processed block relative to the object block.
[0068] For example, in generating the predicted motion vector list, for the processed blocks that are N orders ahead of the object block in the processing order among the multiple processed blocks belonging to the same image as the object block, and for the processed blocks that are N orders ahead of the object block in the processing order, the first motion vector of the processed block is used to derive the predicted motion vector. For the processed blocks that are N orders ahead of the object block in the processing order, the second motion vector of the processed block is used to derive the predicted motion vector.
[0069] Therefore, by using the second motion vector for the processed block that is N blocks ahead of the processed block in the processing order, the encoding device can improve the reliability of the predicted motion vector.
[0070] For example, N could also be 1.
[0071] For example, the first motion vector mentioned above may be referenced in other processes besides the derivation of the predicted motion vector.
[0072] For example, another possible approach is cyclic filtering.
[0073] For example, the second motion vector described above can also be used in cyclic filtering.
[0074] For example, if the object block is encoded in a low-latency mode, the first motion vector of the object block can be derived using the first motion vector of the processed block in the derivation of the first motion vector.
[0075] Therefore, the encoding device can perform appropriate processing depending on whether a low-latency mode is used.
[0076] For example, the information indicating whether the object block encoding is performed in the low-latency mode described above can be encoded into the sequence header region, image header region, slice header region, or auxiliary information region.
[0077] For example, it is also possible to switch whether to encode the object block in the low-latency mode, depending on the size of the object image containing the object block.
[0078] For example, depending on the processing capability of the decoding device, it is also possible to switch whether to encode the object block in the aforementioned low-latency mode.
[0079] For example, it is also possible to switch whether to encode the object blocks in the low-latency mode described above, based on the file or level information assigned to the encoded object stream.
[0080] A decoding apparatus according to a technical solution of the present invention is a decoding apparatus comprising a circuit and a memory. The circuit uses the memory to, when decoding an object block in an inter-frame prediction mode in which motion search is performed in the decoding apparatus, derive a first motion vector of the object block, store the derived first motion vector in the memory, derive a second motion vector of the object block, and generate a predicted image of the object block by using motion compensation of the second motion vector. In the deriving of the first motion vector, the first motion vector of the object block is derived using the first motion vector of the processed block.
[0081] Therefore, in pipeline control, this decoding device can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed. Thus, compared to exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device can be reduced, thereby reducing processing latency.
[0082] For example, in the derivation of the first motion vector, (i) using the first motion vector of the processed block, a list of predicted motion vectors representing multiple predicted motion vectors is generated, and (ii) the first motion vector of the object block is determined from the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0083] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be a merging mode, and the second motion vector may be derived by performing motion search processing on the periphery of the first motion vector.
[0084] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be FRUC mode, and the second motion vector may be derived by performing motion search processing on the periphery of the first motion vector.
[0085] For example, the inter-frame prediction mode for motion search in the above-mentioned decoding device may be FRUC mode. In the derivation of the second motion vector, (i) the third motion vector is determined based on the plurality of predicted motion vectors shown in the above-mentioned list of predicted motion vectors, and (ii) the second motion vector is derived by performing motion search processing on the periphery of the third motion vector.
[0086] For example, in determining the first motion vector, the first motion vector may be derived based on the average or central value of each of the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0087] For example, in determining the first motion vector, the predicted motion vector shown at the beginning of the predicted motion vector list among the plurality of predicted motion vectors shown in the predicted motion vector list can be determined as the first motion vector.
[0088] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block, and in determining the first motion vector, the first motion vector may be determined based on the candidate predicted motion vector derived using the second motion vector from among the plurality of predicted motion vectors shown in the above-mentioned list of predicted motion vectors.
[0089] Therefore, the decoding device can use a highly reliable second motion vector to determine the first motion vector, thus suppressing the decrease in the reliability of the first motion vector.
[0090] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block. If the processed block and the object block belong to the same image, the predicted motion vector may be derived using the first motion vector of the processed block. If the processed block and the object block belong to different images, the predicted motion vector may be derived using the second motion vector of the processed block.
[0091] Therefore, when the processed block and the object block belong to different images, the decoding device can improve the reliability of predicting motion vectors by using the second motion vector.
[0092] For example, in generating the above-mentioned list of predicted motion vectors, each of the plurality of predicted motion vectors may be derived using the first or second motion vector of the processed block, and the first or second motion vector of the processed block may be used in the derivation of the predicted motion vectors may be determined based on the position of the processed block relative to the object block.
[0093] For example, in generating the predicted motion vector list, for the processed blocks that are N orders ahead of the object block in the processing order among the multiple processed blocks belonging to the same image as the object block, and for the processed blocks that are N orders ahead of the object block in the processing order, the first motion vector of the processed block is used to derive the predicted motion vector. For the processed blocks that are N orders ahead of the object block in the processing order, the second motion vector of the processed block is used to derive the predicted motion vector.
[0094] Therefore, by using the second motion vector for the processed block that is N blocks ahead of the processed block in the processing order, the decoding device can improve the reliability of the predicted motion vector.
[0095] For example, N could also be 1.
[0096] For example, the first motion vector mentioned above may be referenced in other processes besides the derivation of the predicted motion vector.
[0097] For example, another possible approach is cyclic filtering.
[0098] For example, the second motion vector described above can also be used in cyclic filtering.
[0099] For example, when the object block is decoded in a low-latency mode, the first motion vector of the object block can be derived using the first motion vector of the processed block in the derivation of the first motion vector.
[0100] Therefore, the decoding device can perform appropriate processing depending on whether a low-latency mode is used.
[0101] For example, the information indicating whether the object block is decoded in the low-latency mode can be decoded from the sequence header region, image header region, slice header region, or auxiliary information region, and based on the information, it can be determined whether the object block is decoded in the low-latency mode.
[0102] For example, the pipeline configuration of the decoding device may include a first stage that performs the processing of the first motion vector of the object block and a second stage that performs the processing of the second motion vector of the object block, which is different from the first stage. Without waiting for the processing of the second stage of the blocks up to M times in the processing order of the object block to be completed, the processing of the first stage of the object block begins at the time point when the processing of the first stage of the block immediately preceding it in the processing order is completed.
[0103] For example, the pipeline configuration of the decoding device may include a first stage that performs the processing of deriving the first motion vector of the object block and a second stage that performs the processing of deriving the second motion vector of the object block, which is different from the first stage. Without waiting for the processing of the second stage of the blocks up to M times the processing order of the object block to be completed, the processing of the first stage of the object block begins at the time point when the first motion vector of the blocks up to M times the processing order has been derived.
[0104] For example, it could also be that M is 1.
[0105] An encoding method for a technical solution of the present invention involves encoding an object block in an inter-frame prediction mode for motion search in a decoding device, deriving a first motion vector of the object block, storing the derived first motion vector in a memory, deriving a second motion vector of the object block, generating a predicted image of the object block using motion compensation based on the second motion vector, and deriving the first motion vector of the object block using the first motion vector of the processed block.
[0106] Therefore, in pipeline control, the decoding device can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed. Thus, compared to exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device can be reduced, thereby reducing processing latency.
[0107] A decoding method according to a technical solution of the present invention involves, when decoding an object block in an inter-frame prediction mode for motion search in a decoding device, deriving a first motion vector of the object block, storing the derived first motion vector in the memory, deriving a second motion vector of the object block, and generating a predicted image of the object block using motion compensation based on the second motion vector. In the deriving of the first motion vector, the first motion vector of the object block is derived using the first motion vector of the processed block.
[0108] Therefore, in pipeline control, this decoding method can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed, for example, after the export of the first motion vector of the surrounding block is completed in the pipeline control. Therefore, compared with the case of exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device can be reduced, thus reducing the processing delay.
[0109] Moreover, these inclusive or specific technical solutions can be implemented by systems, devices, methods, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0110] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.
[0111] Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangements and connection methods of constituent elements, steps, and sequences of steps shown in the following embodiments are examples and are not intended to limit the scope of the claims. In addition, any constituent elements in the following embodiments that are not described in the independent claim representing the highest-level concept are described as arbitrary constituent elements.
[0112] (Implementation Method 1)
[0113] First, as an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention described later, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention, and the processing and / or structure described in the various embodiments of the present invention can also be implemented in encoding and decoding apparatuses different from Embodiment 1.
[0114] When applying the processing and / or structure described in various aspects of the present invention to Embodiment 1, one of the following may also be performed, for example.
[0115] (1) For the encoding or decoding device of Embodiment 1, the constituent element that corresponds to the constituent element described in each aspect of the present invention is replaced with the constituent element described in each aspect of the present invention.
[0116] (2) For the encoding or decoding device of Embodiment 1, after any modification such as adding, replacing, or deleting any of the constituent elements of the plurality of constituent elements constituting the encoding or decoding device, the constituent elements corresponding to the constituent elements described in each aspect of the present invention are replaced with the constituent elements described in each aspect of the present invention.
[0117] (3) After adding processing to the method implemented by the encoding or decoding device of Embodiment 1, and / or replacing or deleting any of the processing among the multiple processing included in the method, the processing corresponding to the processing described in each aspect of the present invention is replaced with the processing described in each aspect of the present invention.
[0118] (4) A portion of the constituent elements constituting the encoding or decoding apparatus of Embodiment 1 are combined with constituent elements described in various aspects of the present invention, a portion of constituent elements having the functions of constituent elements described in various aspects of the present invention, or a portion of constituent elements implementing the processing performed by constituent elements described in various aspects of the present invention.
[0119] (5) A component having a portion of the functions of a portion of the components constituting the encoding or decoding apparatus of embodiment 1, or a component implementing a portion of the processing performed by a portion of the components constituting the encoding or decoding apparatus of embodiment 1, is combined with the components described in various aspects of the present invention, the components having a portion of the functions of the components described in various aspects of the present invention, or the components implementing a portion of the processing performed by the components described in various aspects of the present invention.
[0120] (6) For the method implemented by the encoding or decoding device of Embodiment 1, the processing that corresponds to the processing described in each aspect of the present invention among the multiple processing included in the method is replaced with the processing described in each aspect of the present invention.
[0121] (7) A portion of the processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1 is combined with the processing described in the various aspects of the present invention.
[0122] Furthermore, the implementation of the processes and / or structures described in the various embodiments of the present invention is not limited to the examples described above. For example, it may be implemented in an apparatus used for a different purpose than the moving image / image encoding apparatus or moving image / image decoding apparatus disclosed in Embodiment 1, or the processes and / or structures described in each embodiment may be implemented individually. In addition, the processes and / or structures described in different embodiments may be combined and implemented.
[0123] [Overview of the encoding device]
[0124] First, an overview of the encoding device for Embodiment 1 will be provided. Figure 1 This is a block diagram illustrating the functional structure of the encoding apparatus 100 according to Embodiment 1. The encoding apparatus 100 is a motion picture / image encoding apparatus that encodes motion pictures / images in block units.
[0125] like Figure 1 As shown, the encoding device 100 is a device for encoding images in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a cyclic filtering unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128.
[0126] The encoding device 100 is implemented, for example, by a general-purpose processor and memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128.
[0127] The following describes the constituent elements included in the encoding device 100.
[0128] [Divider]
[0129] The segmentation unit 102 divides each image contained in the input moving image into multiple blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the image into fixed-size blocks (e.g., 128×128). These fixed-size blocks may be called coding tree units (CTUs). Furthermore, the segmentation unit 102 divides each fixed-size block into variable-size blocks (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block segmentation. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the image may be used as processing units of CUs, PUs, and TUs.
[0130] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1. In Figure 2 In the diagram, solid lines represent block boundaries based on quadtree block partitioning, and dashed lines represent block boundaries based on binary tree block partitioning.
[0131] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into 4 square blocks of 64×64 (quadtree block partitioning).
[0132] The 64×64 block in the upper left corner is then vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is then vertically divided into two rectangular 16×64 blocks (binary tree block partitioning). As a result, the 64×64 block in the upper left corner is divided into two 16×64 blocks (11 and 12) and a 32×64 block (13).
[0133] The 64×64 block in the upper right corner is horizontally divided into two rectangular 64×32 blocks, 14 and 15 (binary tree block division).
[0134] The 64×64 block in the lower left corner is divided into four 32×32 square blocks (quadtree block partitioning). The upper left and lower right blocks of these four 32×32 blocks are further partitioned. The upper left 32×32 block is vertically divided into two 16×32 rectangular blocks, and the right 16×32 block is horizontally divided into two 16×16 blocks (binary tree block partitioning). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block partitioning). As a result, the lower left 64×64 block is divided into 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.
[0135] The 64×64 block 23 in the lower right corner is not divided.
[0136] As described above, in Figure 2 In the example, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quadtree plus binary tree) partitioning.
[0137] In addition, Figure 2 In this context, a block can be divided into 2 or 4 blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these. For example, a block can also be divided into 3 blocks (ternary tree partitioning). Partitioning including such ternary tree partitioning is sometimes referred to as MBT (multi-type tree) partitioning.
[0138] [Subtraction Section]
[0139] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in block units divided by the segmentation unit 102. That is, the subtraction unit 104 calculates the prediction error (also called residual) of the encoded target block (hereinafter referred to as the current block). Furthermore, the subtraction unit 104 outputs the calculated prediction error to the transformation unit 106.
[0140] The original signal is the input signal of the encoding device 100, which is the signal representing the image of each picture that constitutes the moving image (e.g., luminance signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0141] [Transformation Section]
[0142] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain, and outputs the transformation coefficients vectorization unit 108. Specifically, the transformation unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0143] Alternatively, the transform unit 106 can adaptively select a transform type from multiple transform types and use the transform basis function corresponding to the selected transform type to transform the prediction error into transform coefficients. Such a transform is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0144] Several transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 This is a table representing the transformation basis functions corresponding to each transformation type. Figure 3 In this context, N represents the number of input pixels. The choice of transform type from these multiple transform types can depend on the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode.
[0145] Information indicating whether such EMT or AMT is applied (e.g., referred to as the AMT flag) and information indicating the selected transform type are signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0146] Furthermore, the transform unit 106 can also perform a re-transformation on the transform coefficients (transformation results). Such a re-transformation may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (e.g., a 4×4 sub-block) contained in the block of transform coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are signaled at the CU level. In addition, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0147] Here, a separable transformation refers to a method of performing multiple transformations in each direction, which is equivalent to the number of dimensions of the input. A non-separable transformation refers to a method of treating two or more dimensions together as one dimension and transforming them together when the input is multidimensional.
[0148] For example, as one example of a non-separable transformation, one can cite the way that when the input is a 4×4 block, it is treated as a permutation of 16 elements, and the permutation is transformed using a 16×16 transformation matrix.
[0149] Furthermore, the Hypercube Givens Transform, which treats a 4×4 input block as a permutation of 16 elements and then performs multiple Givens rotations on that permutation, is also an example of a non-separable transformation.
[0150] [Quantitative Department]
[0151] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scan order and quantizes the transform coefficients based on the quantization parameters (QP) corresponding to the scanned transform coefficients. Furthermore, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0152] The specified order is the order in which the transform coefficients are quantized / inverse quantized. For example, the specified scan order is defined by ascending frequency (from low frequency to high frequency) or descending frequency (from high frequency to low frequency).
[0153] The quantization parameter is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0154] [Entropy Coding Department]
[0155] The entropy coding unit 110 generates a coded signal (coded bitstream) by performing variable-length coding on the quantization coefficients, which are input from the quantization unit 108. Specifically, the entropy coding unit 110 performs arithmetic coding on the binary signal, for example, by binarizing the quantization coefficients.
[0156] [De-quantization Department]
[0157] The inverse quantization unit 112 performs inverse quantization on the quantization coefficients that are input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantization coefficients of the current block in a predetermined scan order. Furthermore, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0158] [Inverse Transformation Section]
[0159] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients, which are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Furthermore, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.
[0160] Furthermore, the restored prediction error differs from the prediction error calculated by the subtraction unit 104 because information was lost during quantization. In other words, the restored prediction error includes quantization error.
[0161] [Addition Department]
[0162] The addition unit 116 reconstructs the current block by adding the prediction error, which is input from the inverse transform unit 114, to the prediction sample, which is input from the prediction control unit 128. Furthermore, the addition unit 116 outputs the reconstructed block to the block memory 118 and the cyclic filtering unit 120. The reconstructed block may be referred to as a local decoding block.
[0163] [Block Memory]
[0164] Block memory 118 is a storage unit used to save blocks within the encoded object image (hereinafter referred to as the current image) referenced in intra-frame prediction. Specifically, block memory 118 saves the reconstructed blocks output from addition unit 116.
[0165] [Loop Filtering Section]
[0166] The cyclic filtering unit 120 applies cyclic filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Cyclic filtering refers to filtering used within the encoding loop (in-loop filtering), such as deblocking filtering (DF), sample adaptive offset (SAO), and adaptive cyclic filtering (ALF).
[0167] In ALF, a least-squares error filter is used to remove coding distortion. For example, for each 2×2 sub-block within the current block, one filter is selected from multiple filters based on the direction of the gradient and the activity of the locality.
[0168] Specifically, sub-blocks (e.g., 2×2 sub-blocks) are first classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0–2 or 0–4) and the gradient activity value A (e.g., 0–4), a classification value C (e.g., C = 5D + A) is calculated. Then, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).
[0169] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.
[0170] Based on the results of this classification, the filter used for the sub-block is determined from among multiple filters.
[0171] The shape of the filter used in ALF can be, for example, a circular symmetrical shape. Figures 4A to 4C This is a diagram showing several examples of the shapes of filters used in ALF. Figure 4A This indicates a 5×5 diamond-shaped filter. Figure 4B This indicates a 7×7 diamond-shaped filter. Figure 4C This represents a 9×9 diamond-shaped filter. Information representing the filter's shape is signaled at the image level. However, the signaling of the filter's shape information is not limited to the image level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0172] The on / off state of ALF is determined, for example, at the picture level or the CU level. For instance, regarding luminance, the decision to use ALF is made at the CU level, while regarding chromatic aberration, it is made at the picture level. Information indicating the on / off state of ALF is signaled at the picture level or the CU level. However, the signaling of information indicating the on / off state of ALF is not limited to the picture level or the CU level; it can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0173] The coefficient set of a selectable set of filters (e.g., up to 15 or 25 filters) is signaled at the picture level. Furthermore, the signaling of the coefficient set is not limited to the picture level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0174] [Frame Memory]
[0175] The frame memory 122 is a storage unit used to store reference images used in inter-frame prediction, and is also sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the cyclic filtering unit 120.
[0176] Intra-frame prediction unit
[0177] The intra-frame prediction unit 124 performs intra-frame prediction (also called intra-picture prediction) of the current block by referring to the blocks in the current image stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction by referring to samples (e.g., luminance value, chrominance value) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.
[0178] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of predefined intra-prediction modes. The plurality of intra-prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0179] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode as specified by the H.265 / HEVC (High-Efficiency Video Coding) specification (Non-Patent Document 1).
[0180] Multiple directional prediction modes may include, for example, the 33 directional prediction modes specified in the H.265 / HEVC specification. Alternatively, multiple directional prediction modes may also include 32 additional directional prediction modes (a total of 65 directional prediction modes). Figure 5A This diagram represents the 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. Solid arrows indicate the 33 directions specified by the H.265 / HEVC specification, while dashed arrows indicate the additional 32 directions.
[0181] Additionally, in intra-frame prediction of chroma blocks, luma blocks can also be referenced. That is, the chroma components of the current block can be predicted based on the luma components of the current block. Such intra-frame prediction is sometimes referred to as CCLM (cross-component linear model) prediction. This intra-frame prediction mode of chroma blocks referencing luma blocks (e.g., called CCLM mode) can also be added as one of the intra-frame prediction modes for chroma blocks.
[0182] The intra-prediction unit 124 can also correct the intra-predicted pixel values based on the gradient of the reference pixels in the horizontal / vertical directions. Intra-prediction accompanied by such correction is sometimes referred to as PDPC (position-dependent intraprediction combination). Information indicating whether PDPC has been used (e.g., a PDPC flag) is signaled, for example, at the CU level. Furthermore, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0183] [Inter-frame prediction department]
[0184] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) for the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Furthermore, the inter-frame prediction unit 126 uses motion information (e.g., motion vectors) obtained through motion estimation to perform motion compensation, thereby generating the inter-frame prediction signal for the current block or sub-block. Finally, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0185] The motion information used in motion compensation is signaled. A motion vector predictor can also be used in the signaling of motion vectors. That is, the difference between the motion vector and the predicted motion vector can also be signaled.
[0186] Alternatively, the inter-frame prediction signal can be generated using not only the motion information of the current block obtained through motion estimation but also the motion information of neighboring blocks. Specifically, the prediction signal based on the motion information obtained through motion estimation can be weighted and added together with the prediction signal based on the motion information of neighboring blocks, thereby generating the inter-frame prediction signal in sub-block units within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0187] In this OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) is signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is used (e.g., OBMC flag) is signaled at the CU level. However, the signaling level for this information is not limited to the sequence and CU levels; it can also be other levels (e.g., image level, slice level, tile level, CTU level, or sub-block level).
[0188] The OBMC model will be explained in more detail. Figure 5B and Figure 5C This is a flowchart and concept diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.
[0189] First, the predicted image (Pred) obtained through normal motion compensation is obtained using the motion vectors (MV) assigned to the encoded object block.
[0190] Next, the predicted image (Pred_L) is obtained by using the motion vector (MV_L) of the encoded left adjacent block for the encoded object block. The first correction of the predicted image is performed by weighted superposition of the predicted image and Pred_L.
[0191] Similarly, the predicted image (Pred_U) is obtained by using the motion vector (MV_U) of the upper adjacent block of the encoded object block. The predicted image is then corrected a second time by weighting and superimposing the predicted image after the first correction and Pred_U, and this is used as the final predicted image.
[0192] In addition, this describes a two-stage correction method using the left and top adjacent blocks, but it can also be configured to perform more corrections using the right and bottom adjacent blocks than the two-stage method.
[0193] In addition, the area to be overlaid may not be the entire pixel area of the block, but only a part of the area near the block boundary.
[0194] Furthermore, the process of correcting the predicted image based on a single reference image is explained here. However, the same principle applies when correcting the predicted image based on multiple reference images. After obtaining the corrected predicted image based on each reference image, the resulting predicted images are further superimposed to obtain the final predicted image.
[0195] In addition, the processing target block mentioned above can be a prediction block unit or a sub-block unit that further divides the prediction block.
[0196] One method for determining whether to use OBMC processing is to use a signal called obmc_flag. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a motion-complex region. If it does, the obmc_flag is set to 1 and OBMC processing is performed for encoding. If it does not belong to a motion-complex region, the obmc_flag is set to 0, and OBMC processing is not performed for encoding. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag recorded in the stream and switching between using and not using OBMC processing based on its value.
[0197] Alternatively, motion information can be exported at the decoding device side without being signaled. For example, the merging mode specified by the H.265 / HEVC standard can be used. Furthermore, motion information can also be exported by performing motion estimation at the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.
[0198] Here, we will explain the motion estimation mode performed on the decoding device side. This motion estimation mode on the decoding device side may be called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0199] exist Figure 5D The diagram below illustrates an example of FRUC processing. First, referencing the motion vectors of coded blocks spatially or temporally adjacent to the current block, a list of multiple candidates, each with a predicted motion vector, is generated (this list can also be shared with a merge list). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value is calculated for each candidate included in the candidate list, and one candidate is selected based on the evaluation value.
[0200] Furthermore, based on the selected candidate motion vectors, motion vectors for the current block are derived. Specifically, for example, the selected candidate motion vector (best candidate MV) can be derived as is, using it as the motion vector for the current block. Alternatively, for example, motion vectors for the current block can be derived by performing pattern matching in the surrounding region of the position within the reference image corresponding to the selected candidate motion vector. That is, the surrounding region of the best candidate MV can be searched using the same method, and if an MV with a better evaluation value is found, the best candidate MV is updated to the aforementioned MV and used as the final MV for the current block. Alternatively, a structure that does not perform this processing can be implemented.
[0201] The exact same processing can also be performed when processing is done in sub-block units.
[0202] Furthermore, the evaluation value is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region within the reference image corresponding to the motion vector and the specified region. Alternatively, information other than the difference value can be used to calculate the evaluation value.
[0203] As a pattern matching, either pattern matching 1 or pattern matching 2 is used. Pattern matching 1 and pattern matching 2 can be referred to as bilateral matching and template matching, respectively.
[0204] In the first pattern matching, pattern matching is performed between two blocks within two different reference images, along the motion trajectory of the current block. Therefore, in the first pattern matching, the regions within other reference images along the motion trajectory of the current block are used as the defined regions for calculating the candidate evaluation values described above.
[0205] Figure 6 This diagram illustrates an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory. For example... Figure 6 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among two blocks in two different reference images (Ref0, Ref1) along the motion trajectory of the current block. Specifically, for the current block, the difference between the reconstructed image at a specified position in the first encoded reference image (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference image (Ref1) specified by the symmetrical MV scaled by the aforementioned candidate MV over the display time interval is derived, and the obtained difference value is used to calculate an evaluation value. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs.
[0206] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, in the case where the current image is located between the two reference images in time and the temporal distances from the current image to the two reference images are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0207] In the second pattern matching, pattern matching is performed between the template in the current image (the block adjacent to the current block in the current image (e.g., the upper and / or left adjacent block)) and the block in the reference image. Therefore, in the second pattern matching, the block adjacent to the current block in the current image is used as the defined area for calculating the candidate evaluation value as described above.
[0208] Figure 7 This is an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 7 As shown, in the second pattern matching, the motion vector of the current block is derived by searching within the reference image (Ref0) for the block that best matches the block adjacent to the current block (Cur block) within the current image (Cur Pic). Specifically, for the current block, the difference between the reconstructed images of the encoded regions of the left and top adjacent regions or one of them and the reconstructed image at the same position within the encoded reference image (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.
[0209] Information indicating whether FRUC mode is used (e.g., referred to as the FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is used (e.g., when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (e.g., referred to as the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0210] This section explains how to derive motion vector patterns based on a model that assumes uniform linear motion. This pattern can be termed BIO (bi-directional optical flow).
[0211] Figure 8 This diagram is used to illustrate a model that assumes uniform linear motion. In Figure 8 In the middle, (v x v y () represents the velocity vector, and τ0 and τ1 represent the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1), respectively. (MVx0, MVy0) represents the motion vector corresponding to the reference image Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference image Ref1.
[0212] At this time, in the velocity vector (v x v y Under the assumption of constant linear motion of MVx0, MVy0 and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1) respectively, and the following optical flow equation (1) holds.
[0213] [Formula 1]
[0214]
[0215] Here, I (k) This represents the luminance value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation states that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on this optical flow equation combined with Hermite interpolation, the block-unit motion vector obtained from merge lists, etc., is corrected in pixels.
[0216] Alternatively, motion vectors can be derived on the decoding device side using a different method than deriving motion vectors based on a model assuming constant linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0217] Here, we will explain the mode of deriving motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks. This mode is sometimes referred to as the affine motion compensation prediction mode.
[0218] Figure 9A This is a diagram used to illustrate the derivation of sub-block unit motion vectors based on the motion vectors of multiple adjacent blocks. Figure 9AIn this context, the current block comprises 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the top-left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the top-right control point of the current block is derived. Furthermore, using the two motion vectors v0 and v1, the motion vectors (v0, v1, v1) of each sub-block within the current block are derived using the following equation (2). x v y ).
[0219] [Formula 2]
[0220]
[0221] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents the pre-set weight coefficient.
[0222] Such an affine motion compensation prediction mode may also include several modes with different methods for deriving the motion vectors of the upper left and upper right control points. Information representing such an affine motion compensation prediction mode (e.g., affine flags) is signaled at the CU level. Furthermore, the signaling of information representing this affine motion compensation prediction mode is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, CTU level, or sub-block level).
[0223] [Forecasting and Control Department]
[0224] The prediction control unit 128 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0225] This section illustrates an example of exporting motion vectors from an encoded object image using a merge mode. Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.
[0226] First, a list of candidate predicted MVs registered with the predicted MVs is generated. Candidate predicted MVs include: spatially adjacent predicted MVs (MVs) belonging to multiple coded blocks spatially surrounding the coded object block; temporally adjacent predicted MVs (MVs) belonging to blocks whose positions in the coded reference image are projected nearby; combined predicted MVs (MVs) generated by combining the MV values of spatially adjacent and temporally adjacent predicted MVs; and zero predicted MVs (MVs with a value of zero).
[0227] Next, the MV for the encoded object block is determined by selecting one predicted MV from the multiple predicted MVs registered in the predicted MV list.
[0228] Furthermore, in the variable-length coding section, merge_idx, which represents the signal that selected which prediction MV was recorded in the stream and encoded.
[0229] In addition, Figure 9B The predicted MVs registered in the predicted MV list described in the figure are one example. They may also be a number different from the number shown in the figure, or a structure that does not include a part of the predicted MVs in the figure, or a structure that adds predicted MVs other than the predicted MVs in the figure.
[0230] Alternatively, the MV of the encoded object block exported through the merge mode can be used for the DMVR processing described later, thereby determining the final MV.
[0231] Here, an example of using DMVR to determine MV is explained.
[0232] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.
[0233] First, the optimal MVP set for the processing object block is taken as the candidate MV. According to the candidate MV, reference pixels are obtained from the first reference image of the processed image in the L0 direction and the second reference image of the processed image in the L1 direction, respectively. The template is generated by taking the average of each reference pixel.
[0234] Next, using the template described above, the surrounding areas of the candidate music videos (MVs) for the first and second reference images are searched, and the MV with the lowest cost is selected as the final MV. Furthermore, the cost value is calculated using the differences between the pixel values of the template and the pixel values of the search area, as well as the MV value.
[0235] Furthermore, the general outline of the processing described herein is essentially the same in both the encoding and decoding devices.
[0236] In addition, even if it is not the process described here, any other process that can search for the surrounding of candidate MVs and export the final MV can be used.
[0237] Here, the mode of generating predicted images using LIC processing is explained.
[0238] Figure 9D This is a diagram illustrating an overview of a predictive image generation method using LIC-based brightness correction processing.
[0239] First, export the MV used to obtain the reference image corresponding to the encoded object block from the reference image, which is an encoded image.
[0240] Next, for the encoded object block, using the brightness pixel values of the left and top adjacent encoded surrounding reference areas and the brightness pixel values at the same position in the reference image specified by MV, information indicating how the brightness values change in the reference image and the encoded object image is extracted, and brightness correction parameters are calculated.
[0241] By using the aforementioned brightness correction parameters to perform brightness correction processing on the reference image within the reference image specified by MV, a predicted image for the coded object block is generated.
[0242] in addition, Figure 9D The shape of the surrounding reference area mentioned above is one example; other shapes may also be used.
[0243] Furthermore, the process of generating a prediction image based on a single reference image is described here, but the same applies when generating a prediction image based on multiple reference images. The prediction image is generated after performing brightness correction processing on the reference images obtained from each reference image in the same way.
[0244] One method for determining whether to use LIC processing is to use a lic_flag as a signal indicating whether LIC processing is used. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a region where a brightness change has occurred. If it does, the lic_flag is set to 1, and LIC processing is used for encoding. If it does not belong to a region where a brightness change has occurred, the lic_flag is set to 0, and LIC processing is not used for encoding. On the other hand, in a decoding device, decoding is performed by decoding the lic_flag recorded in the stream and switching between using and not using LIC processing based on its value.
[0245] Other methods for determining whether to use LIC processing include checking whether LIC processing was used in surrounding blocks. As a specific example, when the encoded target block is in merge mode, it is determined whether the surrounding encoded blocks selected during the export of the MV in merge mode processing have been encoded using LIC processing. Based on the result, encoding is switched between using LIC processing and other methods. Furthermore, in this example, the decoding process is exactly the same.
[0246] [Overview of the Decoding Device]
[0247] Next, an outline of a decoding apparatus capable of decoding the encoded signal (encoded bit stream) output from the encoding apparatus 100 will be described. Figure 10This is a block diagram illustrating the functional structure of the decoding device 200 according to Embodiment 1. The decoding device 200 is a motion image / image decoding device that decodes motion images / images in block units.
[0248] like Figure 10 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a cyclic filtering unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.
[0249] The decoding device 200 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes the software program stored in the memory, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 can also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0250] The following describes the constituent elements included in the decoding device 200.
[0251] [Entropy Decoding Department]
[0252] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units.
[0253] [De-quantization Department]
[0254] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoded target block (hereinafter referred to as the current block), which is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient of the current block based on the quantization parameter corresponding to that quantization coefficient. Furthermore, the inverse quantization unit 204 outputs the inverse quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0255] [Inverse Transformation Section]
[0256] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients, which are inputs from the inverse quantization unit 204.
[0257] For example, if the information read from the encoded bitstream represents EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information representing the transform type read from the transducer.
[0258] Furthermore, for example, when the information read from the encoded bitstream is represented using NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients.
[0259] [Addition Department]
[0260] The adder 208 reconstructs the current block by adding the prediction error, which is input from the inverse transform 206, to the prediction sample, which is input from the prediction control 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the cyclic filtering 212.
[0261] [Block Memory]
[0262] Block memory 210 is a storage unit used to store blocks within the decoded target image (hereinafter referred to as the current image) that serve as a reference in intra-frame prediction. Specifically, block memory 210 stores the reconstructed blocks output from adder 208.
[0263] [Loop Filtering Section]
[0264] The cyclic filtering unit 212 applies cyclic filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.
[0265] Given that the information indicating the on / off state of the ALF is read from the encoded bitstream, and the ALF is on, one filter is selected from multiple filters based on the direction and activity of the gradient of locality, and the selected filter is applied to the reconstructed block.
[0266] [Frame Memory]
[0267] The frame memory 214 is a storage unit used to store reference images used in inter-frame prediction; it is also sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the cyclic filtering unit 212.
[0268] Intra-frame prediction unit
[0269] The intra-prediction unit 216 performs intra-prediction based on the intra-prediction pattern read from the encoded bitstream, referring to blocks within the current image stored in the block memory 210, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra-prediction unit 216 performs intra-prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, thereby generating an intra-prediction signal, and outputs the intra-prediction signal to the prediction control unit 220.
[0270] In addition, if the intra-prediction mode of the reference luma block is selected in the intra-prediction of the chromatic difference block, the intra-prediction unit 216 can also predict the chromatic difference component of the current block based on the luma component of the current block.
[0271] Furthermore, when the information read from the encoded bitstream represents PDPC, the intra-prediction unit 216 corrects the pixel values after intra-prediction based on the gradient of the reference pixel in the horizontal / vertical direction.
[0272] [Inter-frame prediction department]
[0273] The inter-frame prediction unit 218 refers to a reference image stored in the frame memory 214 and predicts the current block. Prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 218 uses motion information (e.g., motion vectors) read from the coded bitstream to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0274] Furthermore, when the information read from the encoded bitstream is represented in OBMC mode, the inter-frame prediction unit 218 uses not only the motion information of the current block obtained through motion estimation, but also the motion information of adjacent blocks to generate the inter-frame prediction signal.
[0275] Furthermore, when the information read from the coded bitstream is in FRUC mode, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the coded stream, thereby deriving motion information. The inter-frame prediction unit 218 then uses the derived motion information to perform motion compensation.
[0276] Furthermore, when using BIO mode, the inter-frame prediction unit 218 derives motion vectors based on a model assuming constant-velocity linear motion. Additionally, when the information representation read from the encoded bitstream employs affine motion compensation prediction mode, the inter-frame prediction unit 218 derives motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0277] [Forecasting and Control Department]
[0278] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208.
[0279] [Example 1 of inter-frame prediction processing]
[0280] Figure 11 This is a first example of a schematic diagram of the pipeline structure used in the decoding device 200. The pipeline structure includes four stages: a first stage, a second stage, a third stage, and a fourth stage.
[0281] In the first stage, the decoding device 200 performs entropy decoding on the input stream that becomes the object of decoding, thereby obtaining the information required for decoding (S101).
[0282] In the second stage, the decoding device 200 uses the aforementioned information to derive motion vectors (MVs) for inter-frame prediction processing. Specifically, the decoding device 200 first derives one or more candidate predicted motion vectors (hereinafter MVPs) as motion vectors by referring to surrounding decoded blocks (S102). Then, the decoding device 200 performs memory transfer of the reference image based on the derived MVPs (S103).
[0283] Next, when the inter-frame prediction mode is FRUC mode, the decoding device 200 determines the motion vector by performing best MVP determination (S104) and best MVP periphery search (S105). Furthermore, when the inter-frame prediction mode is merge mode, the decoding device 200 performs DMVR processing (S106) to determine the motion vector.
[0284] In the third stage, the decoding device 200 decodes the residual image through inverse quantization and inverse transform processing (S110). Furthermore, if the target block is an intra-frame block, the decoding device 200 decodes the predicted image through intra-frame prediction processing (S108). If the target block is an inter-frame block, the decoding device 200 performs motion compensation processing, etc., using the motion vectors derived in the second stage, thereby decoding the predicted image (S107).
[0285] Next, the decoding device 200 selects one of the prediction image generated by intra-frame prediction processing and the prediction image generated by inter-frame prediction processing (S109), and generates a reconstructed image by adding the residual image and the selected prediction image (S111).
[0286] In the fourth stage, the decoding device 200 generates a decoded image by performing cyclic filtering on the reconstructed image (S112).
[0287] The motion vectors derived in stage 2 are used as peripheral reference motion vectors for deriving the MVP in the decoding process of subsequent blocks, and are therefore fed back as input to the MVP derivation process (S102). To reference the motion vectors belonging to the block immediately preceding it in the processing order, this feedback process needs to be incorporated into one stage. The result is as follows: Figure 11 As shown, the second stage consists of a large number of processes, making it a stage with a long processing time.
[0288] Furthermore, this outline of the pipeline structure is just one example; a portion of the described processes may be removed, undescribed processes may be added, or the method of dividing the stages may be changed.
[0289] Figure 12 This is a schematic diagram illustrating an example of block segmentation used in pipelined processing. Figure 12 The block segmentation example shown depicts two coding tree units. One coding tree unit comprises two coding units, CU0 and CU1, while the other coding tree unit comprises three coding units, CU2, CU3, and CU4.
[0290] Encoding units CU0, CU1, and CU4 have the same size. Encoding units CU2 and CU3 have the same size. The size of each of encoding units CU0, CU1, and CU4 is twice the size of each of encoding units CU2 and CU3.
[0291] Figure 13 It is represented in time series Figure 11 The diagram illustrates the processing timing of each stage of the decoding object block in the first example of the pipeline structure described in the text. Figure 13 express Figure 12 The processing timing of the five decoded object blocks of the encoding units CU0 to CU4 shown is also described. Figure 13 S1 to S4 show Figure 11 The processing time for stages 1 through 4.
[0292] Since the coding units CU0, CU1, and CU4 are twice the size of the coding units CU2 and CU3, the processing time for each stage is also twice as long.
[0293] In addition, such as Figure 11 As explained, because the processing time of the second stage is long, the processing time of the second stage is twice the length of the other stages.
[0294] Processing in each stage begins after the same stage of the block that precedes it in the processing order has finished. For example, the processing of the second stage of coding unit CU1 begins at the time t6 when the second stage of coding unit CU0 ends. At this time, since the processing time of the second stage of coding unit CU0 is twice the length, a waiting time occurs in coding unit CU1 from the time t4 when the processing of the first stage ends to the time t6 when the processing of the second stage begins.
[0295] Thus, a waiting time is always generated at the beginning of the second stage, and this waiting time is accumulated whenever the processing of the decoded object block progresses. Therefore, in the encoding unit CU4, a waiting time is generated from the end time t8 of the first stage processing to the start time t14 of the second stage processing.
[0296] As a result, at the point when decoding of a single image is complete, the processing time, including waiting time, increases to approximately twice the original processing time. Therefore, it may be difficult to complete the processing of all blocks within the processing time allocated to a single image.
[0297] Figure 14 yes Figure 11 The flowchart of inter-frame prediction processing is shown in the first example of the pipeline structure described in the text. The process is repeated according to the prediction block unit, which is the processing unit for inter-frame prediction processing. Figure 14 The processing is shown. Additionally, it is performed in the encoding device 100 and the decoding device 200. Figure 14 The processing is shown below. In addition, the operation of the inter-frame prediction unit 126 included in the encoding device 100 will be mainly described below, but the operation of the inter-frame prediction unit 218 included in the decoding device 200 is the same.
[0298] The inter-frame prediction unit 126 selects an inter-frame prediction mode from multiple modes (normal inter-frame mode, merge mode, FRUC mode, etc.) for use in an object block that is either the target of encoding or decoding. The inter-frame prediction unit 126 derives motion vectors (MV) using the selected inter-frame prediction mode. Specifically, the inter-frame prediction mode information indicates the inter-frame prediction mode used for the object block.
[0299] When using the normal inter-frame mode (normal inter-frame mode in S201), the inter-frame prediction unit 126 obtains multiple predicted motion vectors (MVPs) by referring to the motion vectors of surrounding processed blocks, and creates a list of ordinary MVPs representing the obtained multiple MVPs. The inter-frame prediction unit 126 selects one MVP from the multiple MVPs shown in the created list of ordinary MVPs, and adds the differential motion vector (MVD) to the selected MVP, thereby determining the final motion vector (S202). Specifically, the encoding device 100 generates a differential motion vector based on the motion vector and the MVP, and transmits the generated differential motion vector to the decoding device 200. The decoding device 200 obtains the motion vector by adding the transmitted differential motion vector to the predicted motion vector.
[0300] When using the merge mode (merging mode in S201), the inter-frame prediction unit 126 obtains one or more MVPs by referring to the motion vectors of the surrounding processed blocks, and creates a mixed MVP list representing the obtained one or more MVPs. Next, the inter-frame prediction unit 126 selects one MVP from the created mixed mode MVP list as the best MVP (S203). Then, the inter-frame prediction unit 126 performs DMVR processing to search for the position with the lowest cost value in the surrounding region of the best MVP using the processed images, thereby determining the final motion vector (S204).
[0301] When using FRUC mode (S201: FRUC), the inter-frame prediction unit 126 obtains multiple MVPs by referring to the motion vectors of the surrounding processed blocks, and creates a list of FRUC MVPs representing the obtained multiple MVPs (S205). Next, the inter-frame prediction unit 126 uses a bilateral matching method or a template matching method to derive the best MVP with the minimum cost value from the multiple MVPs shown in the FRUC MVP list (S206). Then, for the surrounding area of the derived best MVP, the inter-frame prediction unit 126 further uses the same processing to search for the position with the minimum cost, and determines the motion vector obtained through the search as the final motion vector (S207).
[0302] The final motion vectors derived in each way are used as peripheral references (MVs) for deriving the MVP of subsequent blocks, and are therefore stored in the peripheral reference MV memory.
[0303] Finally, the inter-frame prediction unit 126 generates a predicted image by performing motion compensation processing using the final motion vector (S208).
[0304] Thus, when processing object blocks using the merge or FRUC modes, significantly more processing is required compared to using other modes before deriving the final motion vectors. Consequently, processing time increases, which becomes… Figure 13 The explanation provided explains the reasons for the increased waiting time in each stage of the pipeline control process.
[0305] Furthermore, the processing flow shown here is an example, and it is also possible to remove a part of the recorded processing or add unrecorded processing.
[0306] Furthermore, in the encoding device 100 and the decoding device 200, the only difference is that the signal required for processing is encoded into a stream or decoded from a stream; the processing flow described herein is essentially the same.
[0307] [Example 2 of inter-frame prediction processing]
[0308] Figure 15 This is the second example of a schematic diagram showing the pipeline structure used in the decoding device 200. Figure 15 In the second example shown, with Figure 11 Unlike the first example described, as the peripheral reference motion vector used in the MVP export of the inter-frame prediction process, the decoding device 200 does not use the final motion vector (the second motion vector) after all the processing related to motion vector export, but uses a temporary motion vector (the first motion vector), which is generated using one or more MVPs obtained in the MVP export process.
[0309] In this way, by using temporary motion vectors as peripheral reference motion vectors for deriving the MVP in subsequent block decoding processes, the length of the feedback loop can be significantly shortened. Thus, in the first example, while satisfying the condition that the feedback loop does not span between stages, the second stage, which in the first example must be a long stage, can be divided into two short stages, the second stage and the third stage.
[0310] Furthermore, this outline of the pipeline structure is just one example; a portion of the described processes may be removed, undescribed processes may be added, or the method of dividing the stages may be changed.
[0311] Figure 16 It is represented in time series Figure 15 The second example of the pipeline structure described in the diagram shows the processing timing of each stage of the decoding object block.
[0312] and Figure 13 Similarly, Figure 16 Indicates to Figure 12The block segmentation example shown illustrates the processing timing of five decoded object blocks from coding unit CU0 to coding unit CU4. In the first example, one long stage, namely stage 2, is divided into two shorter stages, stage 2 and stage 3, in the second example. Furthermore, the lengths of stages 2 and 3 are the same as the other stages.
[0313] Processing in each stage begins after the same stage of the block that precedes it in the processing order has finished. For example, the processing of the second stage of coding unit CU1 begins at time t4 when the second stage of coding unit CU0 ends. At this time, since the processing time of the second stage of coding unit CU0 is the same as the length of other stages, in coding unit CU1, the processing of the second stage can begin without waiting time after the processing of the first stage ends.
[0314] On the other hand, since the block size of the coding unit CU2 is smaller than that of the coding unit CU1, which is immediately preceding in the processing order, although a waiting time occurs between the first and second stages, this waiting time is not accumulated, and there is no waiting time at the time point of coding unit CU4.
[0315] As a result, even at the completion time of decoding a single image, the processing time, including waiting time, is roughly the same as the original processing time, thus increasing the likelihood that all blocks can be processed within the processing time allocated to a single image.
[0316] Figure 17 Is Figure 15 The flowchart of inter-frame prediction processing is shown in the second example of the pipeline structure described in the text. The process is repeated according to the prediction block unit, which is the processing unit for inter-frame prediction processing. Figure 17 The processing is shown. Additionally, it is performed in the encoding device 100 and the decoding device 200. Figure 17 The processing is shown below. In addition, the operation of the inter-frame prediction unit 126 included in the encoding device 100 will be mainly described below, but the operation of the inter-frame prediction unit 218 included in the decoding device 200 is the same.
[0317] Figure 17 The processing shown is the same as in Figure 14 The first example described in the text differs in the following aspects. Figure 17 In the process shown, instead of storing the final motion vectors derived in each way in the peripheral reference motion vector memory, temporary motion vectors derived using one or more MVPs obtained in each way are stored in the peripheral reference motion vector memory.
[0318] Therefore, feedback of motion vectors for deriving the MVP's surrounding references can be performed at an early point in the processing flow. Thus, as in Figure 16 As explained, the possibility of significantly reducing waiting time in stages of pipeline control is increased.
[0319] The following exports the temporary motion vectors stored in the motion vector memory for the surrounding reference.
[0320] (1) In normal inter-frame mode, the inter-frame prediction unit 126 determines the final motion vector based on the usual motion vector derivation process as a temporary motion vector.
[0321] (2) In the merging mode, the inter-frame prediction unit 126 will determine the temporary motion vector from the MVP (best MVP) specified by the merging index among the multiple MVPs shown in the merging MVP list.
[0322] (3) In FRUC mode, the inter-frame prediction unit 126 uses multiple MVPs shown in the FRUC MVP list to derive a temporary motion vector, for example, using any of the methods described below. The inter-frame prediction unit 126 determines the MVP registered at the beginning of the FRUC MVP list as the temporary motion vector. Alternatively, the inter-frame prediction unit 126 scales each of the multiple MVPs shown in the FRUC MVP list according to the time interval of the nearest reference image. For each of the L0 and L1 directions, the inter-frame prediction unit 126 calculates the average or center value of the multiple MVPs obtained by scaling, and determines the resulting motion vector as the temporary motion vector.
[0323] In addition, the inter-frame prediction unit 126 may exclude the MVPs registered by referring to the temporary motion vectors in the multiple MVPs shown in the MVP list by the FRUC, and apply any of the above methods to the remaining MVPs to derive the temporary motion vectors.
[0324] In addition, Figure 17 In this process, temporary motion vectors stored in the peripheral reference motion vector memory are used as peripheral reference motion vectors for deriving the MVP. However, temporary motion vectors can also be used as peripheral reference motion vectors in other processes, such as cyclic filtering. Furthermore, in other processes such as cyclic filtering, the temporary motion vectors may not be used, and the final motion vectors used in motion compensation may be used instead. Specifically, in... Figure 15 The final motion vector is derived in stage 3, as shown. Therefore, this final motion vector can also be used in processing from stage 4 onwards.
[0325] Furthermore, in the structure of the pipeline that assumes this processing flow... Figure 15In the structure shown, a temporary motion vector is fed back as a motion vector for the surrounding reference at the timing immediately following the MVP export process. However, as long as the same information as the temporary motion vector described here can be obtained, a temporary motion vector can also be fed back at other timings.
[0326] Furthermore, this processing flow is just one example; a portion of the recorded processing can be removed, or unrecorded processing can be added. For instance, in merge mode, without implementing the best MVP's perimeter search processing, the temporary motion vector can be the same as the final motion vector.
[0327] Furthermore, in the encoding device 100 and the decoding device 200, the only difference is that the signal required for processing is encoded into a stream or decoded from a stream; the processing flow described herein is essentially the same.
[0328] [Result of the second example of inter-frame prediction processing]
[0329] According to usage Figures 15 to 17 The described structure allows for feedback of motion vectors used to derive the MVP's peripheral references at an early point in the processing flow. This significantly reduces the waiting time in the pipeline control stages described in the first example. Consequently, even with a low-performance decoding device, the likelihood of completing all block processing within the processing time allocated to one image increases.
[0330] [The third example of inter-frame prediction processing]
[0331] Figure 18 This is the third example of a schematic diagram showing the pipeline structure used in the decoding device 200. Figure 18 In the third example shown, with Figure 11 Unlike the first example described, the decoding device 200 uses the motion vector of the peripheral decoded block as the peripheral reference motion vector used in the MVP derivation of the inter-frame prediction process. Instead of using the final motion vector after all the processing related to motion vector derivation, it uses the temporary motion vector before performing the best MVP peripheral search when the inter-frame prediction mode is FRUC mode, and uses the temporary motion vector before performing DMVR processing when the inter-frame prediction mode is merge mode.
[0332] In this way, by using a temporary motion vector as a peripheral reference motion vector for deriving the MVP in the decoding process of subsequent blocks, the length of the feedback loop becomes shorter. Therefore, the peripheral reference motion vector can be fed back at an earlier point in the third stage. Thus, if the start timing of the MVP derivation process for subsequent blocks can be delayed until the temporary motion vector is determined, the second stage, which in the first example must be a long stage, can be divided into two shorter stages: the second stage and the third stage.
[0333] Furthermore, this outline of the pipeline structure is just one example; a portion of the described processes may be removed, undescribed processes may be added, or the method of dividing the stages may be changed.
[0334] Figure 19 Represented in time series Figure 18 The diagram shows the processing timing of each stage of the decoding object block in the third example of the pipeline structure described in the text.
[0335] and Figure 13 Similarly, Figure 19 Indicates to Figure 12 The block segmentation example shown illustrates the processing timing of five decoded object blocks from encoding unit CU0 to encoding unit CU4. In the first example, a long stage, namely stage 2, is divided into two shorter stages, stage 2 and stage 3, in the third example. Here, stage 2 is shorter than the other stages. This is because the processing in stage 2 only involves MVP derivation and reference image memory transfer, resulting in less processing. Furthermore, this is based on the premise that the reference image memory transfer speed is sufficiently fast.
[0336] Processing in each stage begins after the same stage of the block that precedes it in the processing order has finished. However, in the third example, the second stage begins after the timing of the temporary motion vector is determined in the third stage of the block that precedes it in the processing order. Therefore, for example, the processing of the second stage of coding unit CU1 begins at time t4 after the end of the first half of the processing of the third stage of coding unit CU0. At this time, since the processing time of the second stage of coding unit CU0 is sufficiently short compared to the other stages, it is possible to start the processing of the second stage in coding unit CU1 without waiting time after the processing of the first stage has finished.
[0337] On the other hand, for example, the processing of the second stage of the coding unit CU3 begins after the first half of the processing of the third stage of the coding unit CU2 has ended, but since the processing time of the second stage of the coding unit CU2 is not a sufficiently short length compared to other stages, a certain increase in waiting time is generated.
[0338] As a result, and in Figure 16Compared to the second example described, at the time when the coding unit CU4 is processed, a waiting time from time t8 to time t9 is generated between the first and second stages. However, compared to... Figure 13 Compared to the first example described, the waiting time is significantly reduced. Even at the completion time of decoding a single image, the processing time, including the waiting time, is roughly the same as the original processing time. Therefore, the probability of completing the processing of all blocks within the processing time allocated to a single image is also increased.
[0339] Figure 20 yes Figure 18 The flowchart of inter-frame prediction processing is shown in the third example of the pipeline structure described in the text. The process is repeated according to the prediction block unit, which is the processing unit for inter-frame prediction processing. Figure 20 The processing is shown. Additionally, it is performed in the encoding device 100 and the decoding device 200. Figure 17 The processing is shown below. In addition, the operation of the inter-frame prediction unit 126 included in the encoding device 100 will be mainly described below, but the operation of the inter-frame prediction unit 218 included in the decoding device 200 is the same.
[0340] Figure 20 The processing shown is the same as in Figure 14 The first example described in the text differs in the following aspects. Figure 20 In the process shown, instead of using the final motion vectors derived in each way as the motion vectors for the periphery reference, the intermediate values of the motion vectors derived in each way, i.e., the temporary motion vectors, are stored in the motion vector memory for the periphery reference.
[0341] Therefore, feedback of motion vectors for deriving the MVP's surrounding references can be performed at an early point in the processing flow. Thus, as in Figure 19 As explained, the possibility of significantly reducing waiting time in stages of pipeline control is increased.
[0342] The following exports the temporary motion vectors stored in the motion vector memory for the surrounding reference.
[0343] (1) In normal inter-frame mode, the inter-frame prediction unit 126 determines the final motion vector based on the usual motion vector derivation process as a temporary motion vector.
[0344] (2) In the merging mode, the inter-frame prediction unit 126 will determine the temporary motion vector from the MVP (best MVP) specified by the merging index among the multiple MVPs shown in the merging MVP list.
[0345] (3) In FRUC mode, the inter-frame prediction unit 126 will determine the MVP with the lowest cost value (best MVP) from among the multiple MVPs shown in the FRUC MVP list by using bilateral matching or template matching as the temporary motion vector.
[0346] In addition, Figure 20 In this process, the motion vectors stored in the peripheral reference motion vector memory are used as peripheral reference motion vectors for deriving the MVP, but temporary motion vectors can also be used as peripheral reference motion vectors in other processes, such as cyclic filtering. Furthermore, in other processes such as cyclic filtering, the temporary motion vectors can be omitted, and the final motion vectors used in motion compensation can be used instead. Specifically, in... Figure 18 The final motion vector is derived in stage 3, as shown. Therefore, this final motion vector can also be used in processing from stage 4 onwards.
[0347] Furthermore, in the structure of the pipeline that assumes this processing flow... Figure 18 In the structure shown, a temporary motion vector is fed back as a motion vector for the surrounding reference during the timing immediately before the optimal MVP perimeter search processing and the DMVR processing. However, as long as the same information as the temporary motion vector described here can be obtained, the temporary motion vector can also be fed back at other timings.
[0348] Furthermore, this processing flow is just one example; a portion of the recorded processing can be removed, or unrecorded processing can be added. For instance, in merge mode or FRUC mode, the temporary motion vector can be the same as the final motion vector without implementing the best MVP perimeter search processing.
[0349] Furthermore, in the encoding device 100 and the decoding device 200, the only difference is that the signal required for processing is encoded into a stream or decoded from a stream; the processing flow described herein is essentially the same.
[0350] [Result of the third example of inter-frame prediction processing]
[0351] According to usage Figures 18 to 20 The described structure allows for feedback of motion vectors used to derive the MVP's peripheral references at an early point in the processing flow. This significantly reduces the waiting time in the pipeline control stages described in the first example. Consequently, even with a low-performance decoding device, the likelihood of completing all block processing within the processing time allocated to one image increases.
[0352] In addition, with Figure 17Compared to the second example described, in the case of FRUC mode for inter-frame prediction, motion vectors that have undergone the best MVP determination process can be used as peripheral reference motion vectors. Therefore, it is more likely to improve coding efficiency by using motion vectors with higher reference reliability.
[0353] [A peripheral reference motion vector that combines the final motion vector and the temporary motion vector]
[0354] exist Figure 17 The second example described in the text and Figure 20 In the third example described, in addition to storing temporary motion vectors in the motion vector memory for peripheral reference, it is also possible to store them in... Figure 14 The first example described in the text also saves the final motion vector.
[0355] Therefore, in the subsequent MVP export process of the block, the inter-frame prediction unit 126 obtains the final motion vector from the block from which the final motion vector can be obtained as the motion vector for the peripheral reference, and can obtain the temporary motion vector from the block from which the final motion vector cannot be obtained as the motion vector for the peripheral reference.
[0356] Figure 21 and Figure 22 This is a diagram used to illustrate the surrounding blocks referenced in order to derive the object block. Encoding unit CU4 is the object block, which is the surrounding block of the space that encoding units CU0 to CU3 have already processed. Encoding unit CU col is the surrounding block of the same position in other images that have already been processed.
[0357] Figure 21 This diagram illustrates an example where the final motion vector cannot be obtained for the block immediately preceding the object block in the processing order. In this example, the inter-frame prediction unit 126 obtains a temporary motion vector for coding unit CU3 as a peripheral reference motion vector, and obtains the final motion vector for all other blocks as a peripheral reference motion vector. Furthermore, coding unit CU col belongs to a block of an image that has already been processed. Therefore, the inter-frame prediction unit 126 obtains the final motion vector for coding unit CU col as a peripheral reference motion vector.
[0358] Figure 22This diagram illustrates an example where the final motion vector cannot be obtained for the two preceding blocks in the processing order relative to the target block. In this example, the inter-frame prediction unit 126 obtains temporary motion vectors for coding units CU2 and CU3 as peripheral reference motion vectors, and obtains the final motion vectors for all other blocks as peripheral reference motion vectors. Furthermore, coding unit CU col belongs to a block of the image that has already been processed. Therefore, the inter-frame prediction unit 126 obtains the final motion vector for coding unit CU col as a peripheral reference motion vector.
[0359] In this way, the inter-frame prediction unit 126 obtains the final motion vector from the block in the peripheral reference block where the final motion vector can be obtained, and uses it as the peripheral reference motion vector. Compared with the case where all temporary motion vectors are used as the peripheral reference motion vector, the MVP can be derived by referring to a motion vector with higher reliability. As a result, the possibility of improving coding efficiency increases.
[0360] Furthermore, the inter-frame prediction unit 126 can also switch the referenced motion vectors based on whether the boundary of the object block is the boundary of the CTU. For example, when the object block is not adjacent to the upper boundary of the CTU, the inter-frame prediction unit 126 switches the referenced motion vectors based on whether the boundary of the object block is adjacent to the upper boundary of the CTU. Figure 21 and Figure 22 The method described herein determines whether to refer to the final motion vector or a temporary motion vector for blocks adjacent to the upper side of the target block. On the other hand, when the target block is adjacent to the upper boundary of the CTU, the final motion vector is derived for blocks adjacent to the upper side of the target block, therefore the inter-frame prediction unit 126 always refers to the final motion vector. Similarly, when the target block is adjacent to the left boundary of the CTU, the inter-frame prediction unit 126 always refers to the final motion vector for blocks adjacent to the left side of the target block.
[0361] [A combination of cases 1, 2, and 3]
[0362] You can also use Figure 14 The first example described in the text, Figure 17 The second example described in the text and Figure 20 The processing flow for the third combination described in the text.
[0363] As a specific example, when the inter-frame prediction mode is merge mode, the inter-frame prediction unit 126 uses the final motion vector as the peripheral reference motion vector, as in the first example. When the inter-frame prediction mode is FRUC mode, it uses a temporary motion vector before performing the optimal MVP peripheral search as the peripheral reference motion vector, as in the third example. Since the processing in merge mode is less complex than that in FRUC mode, even if the final motion vector is fed back as the peripheral reference motion vector after waiting for it to be determined, the processing of subsequent blocks can be completed within the required processing time. Therefore, it is possible to perform processing without accumulating waiting time in the pipeline control stages. Furthermore, by using a highly reliable motion vector as the peripheral reference motion vector in merge mode, the possibility of improving coding efficiency increases.
[0364] [Switching based on low-latency mode signals]
[0365] The inter-frame prediction unit 126 determines whether to process the processing object stream in low-latency mode. If the processing object stream is processed in low-latency mode, as described in the second or third example, the motion vector is used as a peripheral reference in the MVP export process, and a temporary motion vector is referenced. If the processing object stream is not processed in low-latency mode, as described in the first example, the motion vector is used as a peripheral reference in the MVP export process, and the final motion vector may also be referenced.
[0366] Therefore, in low-latency mode, by shortening the length of the feedback loop used for the motion vector for peripheral reference, the likelihood of significantly reducing stage latency in pipelined control increases. On the other hand, in non-low-latency mode, stage latency in pipelined control occurs, but since the motion vector used for peripheral reference can reference a highly reliable motion vector, the potential for improving coding efficiency increases.
[0367] Encoding device 100 generates information indicating whether processing is performed in low-latency mode and encodes the generated information into a stream. Decoding device 200 decodes the stream and obtains the information, and determines whether processing is performed in low-latency mode based on the obtained information. Furthermore, this information is recorded in the sequence header region, image header region, slice header region, or auxiliary information region of the processed object stream.
[0368] For example, the encoding device 100 can also switch between processing in low-latency mode based on the size of the image to be encoded. For example, when the image size is small, the encoding device 100 processes fewer blocks and has ample processing time, so it is not set to low-latency mode; when the image size is large, the processing number of blocks is large and there is no ample processing time, so it is set to low-latency mode.
[0369] Furthermore, the encoding device 100 can switch between processing in a low-latency mode based on the processing capability of the decoding device 200 at the destination of the stream. For example, if the decoding device 200 has high processing capability, the encoding device 100 can perform a large amount of processing within a certain processing time in the decoding device 200, so it is not set to low-latency mode. On the other hand, if the decoding device 200 has low processing capability, the encoding device 100 cannot perform a large amount of processing within a certain processing time, so it is set to low-latency mode.
[0370] Furthermore, the encoding device 100 can switch between processing in low-latency mode based on the file or rating information assigned to the stream being encoded. For example, if a file and rating with sufficient processing power for the decoding device are assigned to the stream, the encoding device 100 is not set to low-latency mode. On the other hand, if a file and rating with insufficient processing power for the decoding device are assigned to the stream, the encoding device 100 is set to low-latency mode.
[0371] Furthermore, the information encoded to indicate whether the stream is processed in low-latency mode may not be a signal that directly indicates whether the stream is processed in low-latency mode; it may be encoded as a signal with other meanings. For example, it may be configured such that by directly establishing a correspondence between the information indicating whether the stream is processed in low-latency mode and the file and level, it is possible to determine whether the stream is processed in low-latency mode using only the signals indicating the file and level.
[0372] As described above, when the encoding apparatus 100 of this embodiment encodes an object block in an inter-frame prediction mode in which motion search is performed in the decoding apparatus 200 (for example, in...), Figure 17 In S201 (merging mode or FRUC mode), the first motion vector of the object block is exported (S203 or S205), the exported first motion vector is saved in memory, and the second motion vector of the object block is exported (S204 or S207). A predicted image of the object block is generated by using motion compensation of the second motion vector (S208). In the export of the first motion vector (S203 or S205), the encoding device 100 exports the first motion vector of the object block using the first motion vector of the processed block.
[0373] Therefore, in pipeline control, the decoding device 200 can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed. Thus, compared to exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device 200 can be reduced, thereby reducing processing delay.
[0374] For example, in the derivation of the first motion vector, the encoding device 100 (i) uses the first motion vector of the processed block to generate a list of predicted motion vectors representing multiple predicted motion vectors, and (ii) determines the first motion vector of the object block (e.g., from the multiple predicted motion vectors shown in the list of predicted motion vectors) from the multiple predicted motion vectors. Figure 17 (S203 or S205).
[0375] For example, the inter-frame prediction mode for motion search in decoding device 200 is a merging mode, and encoding device 100 derives the second motion vector by performing motion search processing on the periphery of the first motion vector (e.g., Figure 17 (S204).
[0376] For example, the inter-frame prediction mode for motion search in decoding device 200 is FRUC mode, and encoding device 100 derives the second motion vector by performing motion search processing on the periphery of the first motion vector (e.g., Figure 20 (S207).
[0377] For example, the inter-frame prediction mode for motion search in decoding device 200 is FRUC mode, and in the derivation of the second motion vector, encoding device 100 (i) determines the third motion vector (optimal MVP) based on multiple predicted motion vectors shown in the list of predicted motion vectors (e.g., Figure 17 (i) The second motion vector is derived by performing motion search processing on the periphery of the third motion vector (e.g., S206), (ii) the second motion vector is derived by performing motion search processing on the periphery of the third motion vector. Figure 17 (S207).
[0378] For example, in the determination of the first motion vector (e.g.) Figure 17 In S205), the encoding device 100 derives the first motion vector based on the average or central value of each of the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0379] For example, in the determination of the first motion vector (e.g.) Figure 17 In S205), the encoding device 100 determines the first motion vector as the first one among the multiple predicted motion vectors shown in the predicted motion vector list.
[0380] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the encoding device 100 derives each of a plurality of predicted motion vectors using the first or second motion vector of the processed block. The determination of the first motion vector (e.g., Figure 17In step S205, the encoding device 100 determines the first motion vector based on a candidate predicted motion vector derived using the second motion vector from among a plurality of predicted motion vectors shown in the predicted motion vector list. This allows the first motion vector to be determined using the highly reliable second motion vector, thus suppressing any decrease in the reliability of the first motion vector.
[0381] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, when the processed block and the target block belong to the same image, the encoding device 100 derives the predicted motion vector using the first motion vector of the processed block; when the processed block and the target block belong to different images, it derives the predicted motion vector using the second motion vector of the processed block. Therefore, when the processed block and the target block belong to different images, the encoding device 100 can improve the reliability of the predicted motion vector by using the second motion vector.
[0382] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the encoding device 100 determines, based on the position of the processed block for the target block, whether to use the first motion vector of the processed block or the second motion vector of the processed block in the derivation of the predicted motion vector.
[0383] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the encoding device 100 uses the second motion vector of the processed block in the derivation of the predicted motion vector when the processed block belongs to a different processing unit than the processing unit (e.g., CTU) that contains the object block.
[0384] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205), the encoding device 100 uses the first motion vector of a processed block to derive a predicted motion vector for a processed block that is N times earlier than the object block in the processing order among a plurality of processed blocks belonging to the same image as the object block, and for a processed block that is later than the processed block that is N times earlier in the processing order. For a processed block that is earlier in the processing order than the processed block that is N times earlier in the processing order, the encoding device 100 uses the second motion vector of the processed block to derive a predicted motion vector.
[0385] Therefore, by using the second motion vector for the processed block that is N blocks ahead of the processed block in the processing order, the encoding device 100 can improve the reliability of the predicted motion vector.
[0386] For example, N is 1.
[0387] For example, the first motion vector is also referenced in other processes besides deriving the predicted motion vector. For example, other processes include cyclic filtering.
[0388] For example, in cyclic filtering, the second motion vector is used.
[0389] For example, when encoding an object block in a low-latency mode, the encoding device 100 uses the first motion vector of the processed block to derive the first motion vector of the object block during the derivation of the first motion vector.
[0390] Therefore, the encoding device 100 can perform appropriate processing depending on whether a low-latency mode is used.
[0391] For example, the encoding device 100 will encode information indicating whether to encode the object block in a low-latency mode into the sequence header region, image header region, slice header region, or auxiliary information region.
[0392] For example, the encoding device 100 switches whether to encode the object block in a low-latency mode based on the size of the object image containing the object block.
[0393] For example, the encoding device 100 switches whether to encode the object block in a low-latency mode based on the processing capability of the decoding device.
[0394] For example, the encoding device 100 switches whether to encode object blocks in a low-latency mode based on the file or level information assigned to the encoded object stream.
[0395] In this embodiment, the decoding apparatus 200 encodes object blocks in an inter-frame prediction mode in which motion search is performed within the decoding apparatus 200 (e.g., in...). Figure 17 In S201 (either in merge mode or FRUC mode), the first motion vector of the object block is exported (S203 or S205), the exported first motion vector is saved in memory, and the second motion vector of the object block is exported (S204 or S207). A predicted image of the object block is generated by using motion compensation of the second motion vector (S208). In exporting the first motion vector (S203 or S205), the decoding device 200 uses the first motion vector of the processed block to export the first motion vector of the object block.
[0396] Therefore, in pipeline control, the decoding device 200 can begin exporting the first motion vector of the target block without waiting for the export of the second motion vector of the surrounding block to be completed after the export of the first motion vector of the surrounding block is completed. Thus, compared to exporting the first motion vector using the second motion vector of the surrounding block, the waiting time in the pipeline control of the decoding device 200 can be reduced, thereby reducing processing delay.
[0397] For example, in the derivation of the first motion vector, the decoding device 200 (i) uses the first motion vector of the processed block to generate a list of predicted motion vectors representing multiple predicted motion vectors, and (ii) determines the first motion vector of the object block (e.g., from the multiple predicted motion vectors shown in the list of predicted motion vectors) from the multiple predicted motion vectors. Figure 17 (S203 or S205).
[0398] For example, the inter-frame prediction mode for motion search in the decoding device 200 is a merging mode, and in the derivation of the second motion vector, the decoding device 200 derives the second motion vector by performing motion search processing on the periphery of the first motion vector (e.g., Figure 17 (S204).
[0399] For example, the inter-frame prediction mode for motion search in the decoding device 200 is FRUC mode. In the derivation of the second motion vector, the decoding device 200 derives the second motion vector by performing motion search processing on the periphery of the first motion vector (e.g., Figure 20 (S207).
[0400] For example, the inter-frame prediction mode for motion search in the decoding device 200 is the FRUC mode. In the derivation of the second motion vector, the decoding device 200(i) determines the third motion vector (optimal MVP) based on multiple predicted motion vectors shown in the predicted motion vector list (e.g., Figure 17 (i) The second motion vector is derived by performing motion search processing on the periphery of the third motion vector (e.g., S206), (ii) the second motion vector is derived by performing motion search processing on the periphery of the third motion vector. Figure 17 (S207).
[0401] For example, in the determination of the first motion vector (e.g.) Figure 17 In S205), the decoding device 200 derives the first motion vector based on the average or central value of each of the multiple predicted motion vectors shown in the list of predicted motion vectors.
[0402] For example, in the determination of the first motion vector (e.g.) Figure 17 In S205), the decoding device 200 determines the first motion vector as the first one among the multiple predicted motion vectors shown in the predicted motion vector list.
[0403] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the decoding device 200 derives each of a plurality of predicted motion vectors using the first or second motion vector of the processed block. The determination of the first motion vector (e.g., Figure 17In S205), the decoding device 200 determines the first motion vector based on a candidate predicted motion vector derived using the second motion vector from among a plurality of predicted motion vectors shown in the predicted motion vector list. Therefore, the first motion vector can be determined using the highly reliable second motion vector, thus suppressing any decrease in the reliability of the first motion vector.
[0404] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the decoding device 200 derives a predicted motion vector using the first motion vector of the processed block when the processed block and the target block belong to the same image, and derives a predicted motion vector using the second motion vector of the processed block when the processed block and the target block belong to different images. Therefore, by using the second motion vector when the processed block and the target block belong to different images, the decoding device 200 can improve the reliability of the predicted motion vector.
[0405] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the decoding device 200 determines, based on the position of the processed block for the target block, whether to use the first motion vector of the processed block or the second motion vector of the processed block in the derivation of the predicted motion vector.
[0406] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205, the decoding device 200 uses the second motion vector of the processed block in the derivation of the predicted motion vector when the processed block belongs to a different processing unit than the processing unit (e.g., CTU) containing the object block.
[0407] For example, in the generation of the predicted list of motion vectors (e.g.) Figure 17 In S203 or S205), the decoding device 200 uses the first motion vector of a processed block to derive a predicted motion vector for a processed block that is N times earlier than the object block in the processing order and a processed block that is later than the processed block that is N times earlier in the processing order. For a processed block that is earlier in the processing order than the processed block that is N times earlier in the processing order, the decoding device 200 uses the second motion vector of the processed block to derive a predicted motion vector.
[0408] Therefore, by using the second motion vector for the processed block that is earlier in the processing order than the processed block that is N blocks earlier in the processing order, the decoding device 200 can improve the reliability of the predicted motion vector.
[0409] For example, N is 1.
[0410] For example, the first motion vector is also referenced in other processes besides deriving the predicted motion vector. For example, other processes include cyclic filtering.
[0411] For example, in cyclic filtering, the second motion vector is used.
[0412] For example, when decoding an object block in a low-latency mode, the decoding device 200 uses the first motion vector of the processed block to derive the first motion vector of the object block during the deriving of the first motion vector.
[0413] Therefore, the decoding device 200 can perform appropriate processing depending on whether a low-latency mode is used.
[0414] For example, the decoding device 200 decodes information indicating whether the object block is decoded in low-latency mode from the sequence header region, image header region, slice header region, or auxiliary information region, and determines whether the object block is decoded in low-latency mode based on this information.
[0415] For example, such as Figure 15 As shown, the pipelined structure of the decoding device 200 includes a first stage of processing the first motion vector of the derived object block. Figure 15 The second phase) and the second motion vector processing that implements the derived object block, which is different from the first phase. Figure 15 (Phase 3). The decoding device 200 does not wait for the completion of the second phase of processing relative to the block M preceding it in the processing order, but starts the first phase of processing of the object block at the time point when the first phase of processing of the block immediately preceding it in the processing order is completed.
[0416] For example, such as Figure 18 As shown, the pipelined structure of the decoding device 200 includes a first stage of processing the first motion vector of the derived object block. Figure 18 The second phase) and the second motion vector processing that implements the derived object block, which is different from the first phase. Figure 18 (Phase 3). The decoding device 200 does not wait for the completion of Phase 2 processing of the blocks up to M blocks in the processing order of the object block, and starts Phase 1 processing of the object block at the time point when the first motion vector of the block up to M blocks in the processing order is derived.
[0417] For example, M is 1.
[0418] Furthermore, the encoding apparatus 100 of this embodiment includes: a segmentation unit 102 that segments an image into multiple blocks; an intra-frame prediction unit 124 that predicts blocks contained in the image using a reference image contained in the image; an inter-frame prediction unit 126 that predicts blocks contained in the image using a reference block contained in another image different from the image; a cyclic filtering unit 120 that applies filtering to the blocks contained in the image; a transform unit 106 that transforms the prediction error between the prediction signal generated by the intra-frame prediction unit 124 or the inter-frame prediction unit 126 and the original signal, and generates transform coefficients; a quantization unit 108 that quantizes the transform coefficients and generates quantization coefficients; and an entropy coding unit 110 that generates a coded bitstream by performing variable-length coding on the quantization coefficients. When the inter-frame prediction unit 126 encodes an object block in an inter-frame prediction mode that performs motion search in the decoding apparatus 200 (for example, in...), Figure 17 In S201 (either in merge mode or FRUC mode), the first motion vector of the object block is exported (S203 or S205), the exported first motion vector is saved in memory, and the second motion vector of the object block is exported (S204 or S207). A predicted image of the object block is generated by using motion compensation of the second motion vector (S208). In the export of the first motion vector, the encoding device 100 exports the first motion vector of the object block using the first motion vector of the processed block.
[0419] Furthermore, the decoding apparatus 200 of this embodiment includes: a decoding unit (entropy decoding unit 202) that decodes the encoded bitstream and outputs quantization coefficients; an inverse quantization unit 204 that inverse quantizes the quantization coefficients and outputs transform coefficients; an inverse transform unit 206 that inverse transforms the transform coefficients and outputs prediction errors; an intra-frame prediction unit 216 that predicts blocks contained in the image using a reference image contained in the image; an inter-frame prediction unit 218 that predicts blocks contained in the image using reference blocks contained in other images different from the image; and a cyclic filtering unit 212 that applies filtering to the blocks contained in the image. The inter-frame prediction unit 218, when encoding object blocks in an inter-frame prediction mode in which motion search is performed in the decoding apparatus 200 (for example, in...), Figure 17 In S201 (either in merge mode or FRUC mode), the first motion vector of the object block is exported (S203 or S205), the exported first motion vector is saved in memory, and the second motion vector of the object block is exported (S204 or S207). A predicted image of the object block is generated by using motion compensation of the second motion vector (S208). In the export of the first motion vector, the encoding device 100 exports the first motion vector of the object block using the first motion vector of the processed block.
[0420] [Installation example of the encoding device]
[0421] Figure 23 This is a block diagram illustrating an installation example of the encoding device 100 according to Embodiment 1. The encoding device 100 includes circuitry 160 and a memory 162. For example, Figure 1 The multiple components of the encoding device 100 shown are composed of Figure 23 The circuit 160 and memory 162 shown are installed.
[0422] Circuit 160 is an information processing circuit that accesses memory 162. For example, circuit 160 is a dedicated or general-purpose electronic circuit for encoding moving images. Circuit 160 can also be a processor like a CPU. Alternatively, circuit 160 can be an assembly of multiple electronic circuits. Furthermore, for example, circuit 160 can also function as… Figure 1 The functions of multiple components of the encoding device 100 shown, excluding the component for storing information.
[0423] Memory 162 is a dedicated or general-purpose memory used by storage circuit 160 to encode moving images. Memory 162 can be an electronic circuit or connected to circuit 160. Alternatively, memory 162 can be included within circuit 160. Alternatively, memory 162 can be an assembly of multiple electronic circuits. Memory 162 can be a magnetic disk or optical disk, or it can be represented as storage or a recording medium. Memory 162 can be either non-volatile or volatile memory.
[0424] For example, memory 162 can store the encoded motion image, or it can store the bit string corresponding to the encoded motion image. Additionally, memory 162 can also store the program used by circuit 160 to encode the motion image.
[0425] Additionally, for example, memory 162 can also serve as Figure 1 The encoding device 100 shown above has multiple components that serve as components for storing information. Specifically, the memory 162 can function as... Figure 1 The block memory 118 and frame memory 122 shown have the following functions. More specifically, memory 162 can store reconstructed blocks and reconstructed images, etc.
[0426] Alternatively, the encoding device 100 may not require installation of [something]. Figure 1 All of the multiple constituent elements shown above can also be processed without the aforementioned multiple processing steps. Figure 1 A portion of the multiple constituent elements shown may be included in other devices, or a portion of the aforementioned multiple processes may be performed by other devices. Furthermore, in the encoding device 100, an assembly is installed... Figure 1 One of the multiple components shown above can be used to efficiently perform motion compensation by performing one of the multiple processes described above.
[0427] [Installation example of the decoding device]
[0428] Figure 24 This is a block diagram illustrating an installation example of the decoding device 200 according to Embodiment 1. The decoding device 200 includes circuitry 260 and a memory 262. For example, Figure 10 The multiple components of the decoding device 200 shown are transmitted through Figure 24 The circuit 260 and memory 262 shown are installed.
[0429] Circuit 260 is a circuit that performs information processing and can access memory 262. For example, circuit 260 may be a dedicated or general-purpose electronic circuit for decoding moving images. Circuit 260 can also be a processor like a CPU. Alternatively, circuit 260 can be an assembly of multiple electronic circuits. Furthermore, circuit 260 can also, for example, function as... Figure 10 The functions of multiple components of the decoding device 200 shown, excluding the component for storing information.
[0430] Memory 262 is a dedicated or general-purpose memory that stores information used by circuit 260 to decode moving images. Memory 262 can be an electronic circuit or connected to circuit 260. Alternatively, memory 262 can be included within circuit 260. Alternatively, memory 262 can be an assembly of multiple electronic circuits. Alternatively, memory 262 can be a disk or optical disk, or it can take the form of a memory or recording medium. Furthermore, memory 262 can be either non-volatile or volatile memory.
[0431] For example, memory 262 can store bit strings corresponding to the encoded motion image, or it can store motion images corresponding to the decoded bit strings. Additionally, memory 262 can also store a program for circuit 260 to decode the motion image.
[0432] Additionally, for example, memory 262 can serve as Figure 10 The decoding device 200 shown above has multiple components, including the component for storing information. Specifically, the memory 262 can serve as... Figure 10 The block memory 210 and frame memory 214 shown herein serve a specific purpose. More specifically, memory 262 can store reconstructed blocks and reconstructed images, etc.
[0433] In addition, in the decoding device 200, it is not necessary to install Figure 10All of the multiple constituent elements shown above can also be processed without the aforementioned multiple processing steps. Figure 10 A portion of the multiple components shown may be included in other devices, or a portion of the aforementioned multiple processes may be performed by other devices. Then, in the decoding device 200, an installation is performed... Figure 10 One of the multiple components shown above can be used to efficiently perform motion compensation by performing one of the multiple processes described above.
[0434] [Replenish]
[0435] Furthermore, the encoding device 100 and the decoding device 200 in this embodiment can be used as an image encoding device and an image decoding device, respectively, or as a moving image encoding device and a moving image decoding device. Alternatively, the encoding device 100 and the decoding device 200 can be used as inter-frame prediction devices (inter-picture prediction devices).
[0436] That is, the encoding device 100 and the decoding device 200 may correspond only to the inter-frame prediction unit (inter-picture prediction unit) 126 and the inter-frame prediction unit (inter-picture prediction unit) 218, respectively. Furthermore, other components such as the transformation unit 106 and the inverse transformation unit 206 may also be included in other devices.
[0437] In addition, while each component is constructed using dedicated hardware in this embodiment, it can also be implemented by executing software programs suitable for each component. Each component can be implemented by reading and executing software programs recorded on recording media such as hard disks or semiconductor memory using a program execution unit such as a CPU or processor.
[0438] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuitry and a storage device electrically connected to and accessible from the processing circuitry. For example, the processing circuitry corresponds to circuitry 160 or 260, and the storage device corresponds to memory 162 or 262.
[0439] The processing circuit includes at least one of dedicated hardware and a program execution unit, and uses a storage device to perform the processing. Furthermore, if the processing circuit includes a program execution unit, the storage device stores a software program executed by that program execution unit.
[0440] Here, the software for implementing the encoding device 100 or decoding device 200 of this embodiment is the following program.
[0441] Furthermore, as mentioned above, each component can also be a circuit. These circuits can either form a single circuit as a whole or be separate circuits. Additionally, each component can be implemented using either a general-purpose processor or a dedicated processor.
[0442] Alternatively, other components may perform the processing performed by a specific component. Furthermore, the order of processing can be changed, and multiple processes can be performed simultaneously. Alternatively, the encoding / decoding apparatus may include an encoding device 100 and a decoding device 200.
[0443] The above description illustrates the configurations of the encoding device 100 and the decoding device 200 based on the embodiments, but the configurations of the encoding device 100 and the decoding device 200 are not limited to this embodiment. As long as they do not depart from the spirit of the invention, various modifications that can be conceived by those skilled in the art to this embodiment, and configurations constructed by combining the constituent elements of different embodiments, can also be included within the scope of the configurations of the encoding device 100 and the decoding device 200.
[0444] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.
[0445] (Implementation Method 2)
[0446] In the above embodiments, each functional block is typically implemented using an MPU and memory. Furthermore, the processing of each functional block is usually achieved by a program execution unit such as a processor reading and executing software (programs) recorded in a recording medium such as ROM. This software can be distributed via download or by recording it in a recording medium such as semiconductor memory. Alternatively, each functional block can also be implemented in hardware (dedicated circuitry).
[0447] Furthermore, the processing described in each embodiment can be implemented either centrally using a single device (system) or distributedly using multiple devices. Additionally, the processor executing the above-described program can be either single or multiple. That is, it can be either centrally processed or distributed.
[0448] The present invention is not limited to the above embodiments, and various modifications can be made, which are also included within the scope of the present invention.
[0449] Furthermore, examples of applications of the motion picture encoding method (image encoding method) or motion picture decoding method (image decoding method) shown in the above embodiments and systems using them will be described here. The system is characterized by having an image encoding apparatus using the image encoding method, an image decoding apparatus using the image decoding method, and an image encoding / decoding apparatus possessing both. Other structures within the system can be appropriately modified as needed.
[0450] [Usage Example]
[0451] Figure 25 This is a diagram showing the overall structure of the content delivery system ex100 that implements content distribution services. The provision of communication services is divided into desired sizes, and each unit is equipped with base stations ex106, ex107, ex108, ex109, and ex110, which serve as fixed wireless stations.
[0452] In this content delivery system ex100, various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. This content delivery system ex100 can also combine some of the above elements for connection. Alternatively, the devices can be directly or indirectly interconnected via telephone networks or short-range wireless connections without using base stations ex106 to ex110, which are fixed wireless stations. Furthermore, a streaming media server ex103 is connected to the computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via the Internet ex101, etc. Additionally, the streaming media server ex103 is connected to terminals within a hotspot inside an aircraft ex117 via satellite ex116.
[0453] Alternatively, it can replace base stations ex106 to ex110 by using wireless access points or hotspots. Furthermore, the streaming media server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or it can connect directly to the aircraft ex117 without going through the satellite ex116.
[0454] The camera ex113 is a digital camera or similar device capable of capturing still and moving images. Additionally, the smartphone ex115 refers to a smartphone, mobile phone, or PHS (Personal Handyphone System) corresponding to mobile communication systems commonly referred to as 2G, 3G, 3.9G, 4G, and the future 5G.
[0455] Home appliance EX118 refers to refrigerators or other equipment included in a home fuel cell combined heat and power system.
[0456] In the content supply system ex100, terminals with photography capabilities are connected to the streaming media server ex103 via a base station ex106, enabling on-site distribution. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) perform encoding processing on still or moving images captured by the user using these terminals, as described in the above embodiments. The encoded image data and the encoded audio data are multiplexed, and the resulting data is sent to the streaming media server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present invention.
[0457] On the other hand, the streaming media server ex103 distributes the content data sent by requesting clients. Clients are terminals such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or airplanes ex117 capable of decoding the encoded data. Each device receiving the distributed data decodes and reproduces the received data. That is, each device functions as an image decoding device according to one aspect of the present invention.
[0458] [Distributed processing]
[0459] Furthermore, the streaming media server ex103 can also be multiple servers or multiple computers, distributing data through decentralized processing or recording. For example, the streaming media server ex103 can also be implemented by a CDN (Contents Delivery Network), which distributes content by connecting many edge servers scattered around the world. In a CDN, physically nearby edge servers are dynamically allocated based on the client. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Moreover, in the event of an error or a change in communication status due to increased traffic, processing can be distributed across multiple edge servers, the distribution entity can be switched to other edge servers, or the faulty part of the network can be bypassed to continue distribution, thus achieving high-speed and stable distribution.
[0460] Furthermore, beyond the decentralized processing of distribution itself, the encoding processing of the captured data can be performed by each terminal, on the server side, or shared among them. For example, the encoding process typically involves two processing loops. In the first loop, the complexity or code size of the image per frame or scene unit is detected. In the second loop, processing is performed to maintain image quality while improving encoding efficiency. For instance, by having the terminal perform the first encoding processing and the server receiving the content perform the second, the processing load on each terminal can be reduced while improving both content quality and efficiency. In this case, if there is a request for near real-time reception and decoding, the data encoded by the terminal in the first round can be received and reproduced by other terminals, enabling more flexible real-time distribution.
[0461] As another example, cameras like the ex113 extract features from images, compress the data about these features as metadata, and send it to a server. The server, for instance, adjusts the quantization precision based on the features to determine the importance of the target, performing compression that corresponds to the meaning of the image. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, while more demanding encoding methods like CABAC (Context Adaptive Binary Arithmetic Coding) can be used by the server.
[0462] As another example, in stadiums, shopping malls, or factories, there may be multiple images of roughly the same scene captured by multiple terminals. In such cases, the data is distributed and encoded separately using the terminals that captured the images, as well as other terminals and servers that did not capture images, as needed. This can be done by assigning encoding and processing data to different units, such as GOP (Group of Pictures), image units, or tile units obtained by segmenting images. This reduces latency and improves real-time performance.
[0463] Furthermore, since multiple image datasets depict roughly the same scene, the server can manage and / or instruct the data to be cross-referenced between images captured by different terminals. Alternatively, the server can receive encoded data from each terminal and change the reference relationships between the multiple datasets, or modify or replace the images themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each data item.
[0464] In addition, the server can also transcode the image data by changing its encoding method before distributing it. For example, the server can convert MPEG encoding to VP encoding, or H.264 to H.265.
[0465] In this way, encoding processing can be performed by a terminal or one or more servers. Therefore, the terms "server" or "terminal" will be used below to refer to the main body performing the processing, but it is also possible to perform part or all of the processing performed by the server by the terminal, or vice versa. Furthermore, the same applies to decoding processing.
[0466] [3D, Multi-angle]
[0467] In recent years, there has been an increase in the use of images or videos captured by multiple cameras (ex113 and / or smartphones (ex115)) at roughly the same time, capturing different scenes or capturing the same scene from different angles. These images are then merged based on the relative positions of the cameras or regions containing consistent feature points within the images.
[0468] The server not only encodes 2D moving images, but can also automatically or at user-specified times encode still images based on scene analysis of moving images and send them to the receiving terminal. Furthermore, when the server can obtain the relative positions between shooting terminals, it can generate 3D shapes of scenes not only from 2D moving images, but also from images of the same scene captured from different angles. Additionally, the server can separately encode 3D data generated from point clouds, and can also select or reconstruct images from multiple terminals based on the results of identifying or tracking people or targets using 3D data to generate images to be sent to the receiving terminal.
[0469] In this way, users can freely select images corresponding to each shooting terminal to appreciate the scene, and can also appreciate the content of images extracted from any viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, similar to the images, sound can also be collected from multiple different angles, and the server can match it with the images to multiplex and send sound from a specific angle or space with the images.
[0470] Furthermore, in recent years, content that establishes a correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. This can be achieved through Multi-View Coding (MVC) to allow reference between the viewpoint images, or by encoding them as separate streams without reference to each other. During the decoding of these different streams, they can be synchronously reproduced according to the user's viewpoint to recreate a virtual three-dimensional space.
[0471] In the case of AR images, the server can overlay virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device acquires or holds the virtual object information and 3D data, generates a 2D image based on the user's viewpoint movement, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can send the user's viewpoint movement to the server in addition to the virtual object information. The server creates overlay data by matching the received viewpoint movement with the 3D data held on the server, encodes the overlay data, and distributes it to the decoding device. Furthermore, the overlay data has an α value representing transmittance in addition to RGB. The server sets the α value of the portion outside the target created from the 3D data to 0, etc., and encodes the portion in a state of transmittance. Alternatively, the server can set a predetermined RGB value as the background, similar to a chroma key, to generate data where the portion outside the target is set as the background color.
[0472] Similarly, the decoding of distributed data can be performed by the individual terminals acting as clients, on the server side, or distributed among them. For example, one terminal could first send a receive request to the server, other terminals could receive the content corresponding to that request, decode it, and then send the decoded signal to a device with a display. By distributing the processing independently of the capabilities of the communicating terminals and selecting appropriate content, data with better image quality can be reproduced. Furthermore, as another example, large-format image data can be received by a TV, and the viewer's personal terminal can decode and display a segmented area of the image, such as tiles. This allows for the sharing of the overall image while simultaneously allowing the viewer to identify their own area of responsibility or areas they wish to examine in more detail.
[0473] Furthermore, it is envisioned that in the future, with the availability of multiple short-range, medium-range, or long-range wireless communications both indoors and outdoors, content can be seamlessly received while switching appropriate data between connected communications using distribution system standards such as MPEG-DASH. This would allow users to switch in real-time not only using their own terminals but also freely choosing decoding or display devices such as monitors installed indoors or outdoors. Additionally, decoding can be performed by switching between decoding and display terminals based on the user's location information. This would also allow movement towards a destination while displaying map information on a portion of the wall or ground of a building adjacent to a display device. Furthermore, the bit rate of the received data can be switched based on the ease of accessing encoded data to the network, such as when the encoded data is cached on a server accessible only briefly from the receiving terminal or copied to an edge server of the content distribution service.
[0474] [Hyper-level coding]
[0475] Regarding content switching, use Figure 26 The scalable stream, which is compressed using the motion picture coding method described in the above embodiments, will be explained. For the server, there can be multiple streams with the same content but different qualities, or the content structure can be switched using the characteristics of a temporally / spatially scalable stream achieved through layered coding, as shown in the illustration. That is, by having the decoding side determine which layer to decode to based on intrinsic factors such as performance and extrinsic factors such as the state of the communication band, the decoding side can freely switch between decoding low-resolution and high-resolution content. For example, if one wants to watch a follow-up video viewed on a smartphone (ex115) while on the go, and then watch it at home on an internet TV or similar device, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0476] Furthermore, in addition to the hierarchical structure described above, where images are encoded layer by layer and enhancement layers exist above the base layers, the enhancement layer can also contain metadata such as image statistics. The decoding side then uses this metadata to perform super-resolution on the base layer images to generate high-quality content. Super-resolution can be either an increase in the signal-to-noise ratio (SN ratio) at the same resolution or an increase in resolution. The metadata includes information used to determine the linear or nonlinear filtering coefficients used in the super-resolution process, or information determining the parameter values for filtering, machine learning, or least-squares operations used in the super-resolution process.
[0477] Alternatively, the image can be segmented into tiles based on the meaning of targets within it, and the decoding side can decode only a portion of the region by selecting the tiles to be decoded. Furthermore, by storing the target's attributes (people, cars, balls, etc.) and its position within the image (coordinates within the same image, etc.) as metadata, the decoding side can determine the desired target's location based on this metadata and decide which tiles to include that target. For example, ... Figure 27 As shown, metadata is stored using data storage structures different from pixel data, such as SEI messages in HEVC. This metadata may represent, for example, the position, size, or color of the main target.
[0478] Furthermore, metadata can be stored in units consisting of multiple images, such as streams, sequences, or random access units. This allows the decoder to obtain information such as the time a specific person appears in the image, and by matching this information with the image unit information, it can determine the image containing the target and the target's location within that image.
[0479] [Web page optimization]
[0480] Figure 28 This is an example of a web page display screen in a computer such as ex111. Figure 29 This is an example image showing the display screen of a web page in a smartphone such as the ex115. Figure 28 and Figure 29 As shown, in cases where a web page contains multiple linked images that serve as links to image content, their visibility varies depending on the viewing device. When multiple linked images are visible on the screen, before the user explicitly selects a linked image, or before the linked image is near the center of the screen, or before the entire linked image enters the screen, the display device (decoding device) displays still images or I-images of each content as linked images, or displays images like GIF animations using multiple still images or I-images, or only receives the basic layer and decodes and displays the image.
[0481] When a user selects a linked image, the display device prioritizes decoding the base layer. Additionally, if the HTML constituting the webpage contains information indicating tiered content, the display device can also decode up to the enhancement layer. Furthermore, in situations where real-time performance is ensured before selection or when communication bandwidth is extremely limited, the display device can reduce the delay between decoding and displaying the first image (the delay from the start of content decoding to the start of display) by decoding and displaying only the preceding reference images (I images, P images, and B images only used for preceding reference). Alternatively, the display device can forcibly ignore image reference relationships and coarsely decode all B and P images as preceding references, performing normal decoding as more images are received over time.
[0482] [Autonomous Driving]
[0483] Furthermore, when receiving and transmitting still images or video data such as two-dimensional or three-dimensional map information for the purpose of autonomous driving or driving assistance, the receiving terminal can also receive information such as weather or construction as metadata, in addition to image data belonging to more than one layer, and decode them by establishing correspondences. Moreover, metadata can belong to a layer or be multiplexed only with image data.
[0484] In this scenario, since the vehicle, drone, or aircraft containing the receiving terminal is in motion, the receiving terminal can seamlessly receive and decode data by switching between base stations ex106 to ex110 by sending its location information when a request is received. Furthermore, the receiving terminal can dynamically adjust the level of metadata reception or map information updates based on user selection, user status, or the status of the communication frequency band.
[0485] As described above, in the content delivery system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0486] Distribution of personal content
[0487] Furthermore, the content delivery system ex100 not only handles high-quality, long-duration content provided by video distribution providers, but also enables unicast or multicast distribution of low-quality, short-duration content provided by individuals. Moreover, it's conceivable that such personal content will increase in the future. To improve the quality of personal content, the server can also perform encoding processing after editing. This can be achieved, for example, through a structure like the following.
[0488] During or after capturing images, the server performs image processing such as shooting errors, scene search, meaning analysis, and target detection based on the original images or encoded data. Furthermore, based on the recognition results, the server manually or automatically corrects focus deviations or camera shake, deletes scenes of lower importance (e.g., scenes with lower brightness or out-of-focus areas), emphasizes target edges, or adjusts tones. The server then encodes the edited data. Additionally, recognizing that longer shooting times reduce audiovisual quality, the server can automatically limit not only low-importance scenes as described above but also scenes with minimal movement to specific timeframes based on image processing results. Alternatively, the server can generate and encode summaries based on the meaning analysis results of the scenes.
[0489] Furthermore, personal content may contain elements that infringe on copyrights, the author's moral rights, or portrait rights in its original state, or the sharing scope may exceed the intended scope, causing inconvenience to the individual. Therefore, for example, the server could forcibly encode images such as faces of people in the periphery of a scene or a home, replacing them with out-of-focus images. Additionally, the server could identify whether a face different from a pre-registered person is captured within the image to be encoded, and if so, apply processing such as mosaic to the face. Alternatively, as pre- or post-processing for encoding, from a copyright perspective, the user can specify the person or background area to be processed, and the server can replace the specified area with another image or blur the focus. If it is a person, the face image can be replaced while tracking the person in a moving image.
[0490] Furthermore, personal content with small data volumes has strong real-time requirements for audiovisual presentation. Therefore, although bandwidth also plays a role, the decoding device prioritizes receiving, decoding, and reproducing the base layer. The decoding device can also receive enhancement layers during this process, including them in high-quality image reproduction if the playback is looped or repeated more than twice. In this way, if the stream is scalably encoded, it can provide an experience where the motion images are initially coarse but gradually become smoother and the image quality improves. Besides scalable encoding, the same experience can be provided when the first, coarser stream and a second stream encoded based on the first motion image are combined into a single stream.
[0491] [Other Use Cases]
[0492] Furthermore, these encoding or decoding processes are typically handled within the LSIex500 chip present in each terminal. The LSIex500 can be a single chip or a multi-chip structure. Alternatively, software for motion image encoding or decoding can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer such as the ex111, and the encoding and decoding processes can be performed using this software. Furthermore, when the ex115 smartphone has a camera, motion image data acquired by that camera can also be transmitted. This motion image data is encoded using the LSIex500 chip present in the ex115 smartphone.
[0493] Alternatively, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether it corresponds to the content's encoding method or whether it has the capability to execute a specific service. If the terminal does not correspond to the content's encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software, and then retrieves and reproduces the content.
[0494] Furthermore, not limited to the content delivery system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of the above embodiments can also be assembled in a digital broadcasting system. Since multiplexed data that multiplexes images and sound is carried and transmitted and received using radio waves for broadcasting via satellites or the like, it is more suitable for multicasting than the unicast-friendly structure of the content delivery system ex100, but the same applications can be performed for encoding and decoding processing.
[0495] [Hardware Structure]
[0496] Figure 30 This is a diagram representing the EX115 smartphone. Additionally, Figure 31 This diagram illustrates a structural example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with a base station ex110, a camera unit ex465 for capturing images and still images, and a display unit ex458 for displaying decoded data such as images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466, such as a touch panel; a sound output unit ex457, such as a speaker, for outputting sound or audio; a sound input unit ex456, such as a microphone, for inputting sound; a memory unit ex467 capable of storing encoded or decoded data such as captured images or still images, recorded audio, received images or still images, and emails; and a slot unit ex464 serving as an interface with a SIM ex468 used to identify the user and authenticate access to various data sources, such as the network. Alternatively, an external memory may be used instead of the memory unit ex467.
[0497] Furthermore, the main control unit ex460, which performs integrated control of the display unit ex458 and the operation unit ex466, is interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.
[0498] If the power button is turned on by the user, the power circuit section ex461 will supply power to each part from the battery pack, thus activating the smartphone ex115 into an operational state.
[0499] The smart phone ex115 is controlled by a main control unit ex460, which includes a CPU, ROM, and RAM, for call and data communication processing. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. This digital signal is then subjected to spectral diffusion processing by the modulation / demodulation unit ex452, and finally, digital-to-analog conversion and frequency conversion processing by the transmitting / receiving unit ex451 before being transmitted via the antenna ex450. Similarly, received data is amplified and subjected to frequency conversion and analog-to-digital conversion processing. The data is then subjected to inverse spectral diffusion processing by the modulation / demodulation unit ex452, converted into an analog audio signal by the audio signal processing unit ex454, and output from the audio output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 through the operation unit ex466 of the main unit, and are also processed for transmission and reception. In data communication mode, when transmitting images, still images, or images and sound, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded image data to the multiplexing / demultiplexing unit ex453. Furthermore, the sound signal processing unit ex454 encodes the sound signal collected by the sound input unit ex456 during the capture of images, still images, etc., by the camera unit ex465, and sends the encoded sound data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded image data and encoded sound data in a prescribed manner, performs modulation and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmitting / receiving unit ex451, and transmits the data via the antenna ex450.
[0500] When receiving images attached to emails or chat tools, or images linked to web pages, the multiplexing / demultiplexing unit ex453, in order to decode the multiplexed data received via antenna ex450, separates the multiplexed data into a bitstream of image data and a bitstream of audio data. The encoded image data is supplied to the image signal processing unit ex455 via the synchronization bus ex470, and the encoded audio data is supplied to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method described in the above embodiments, and displays the image or still image contained in the linked motion picture file from the display unit ex458 via the display control unit ex459. Furthermore, the audio signal processing unit ex454 decodes the audio signal and outputs audio from the audio output unit ex457. Additionally, since real-time streaming is becoming increasingly common, depending on the user's situation, the reproduction of audio may be unsuitable for certain situations. Therefore, as an initial consideration, a structure that reproduces only the image data and not the audio signal is preferred. It can also reproduce sound synchronously only when the user performs actions such as clicking on the image data.
[0501] Furthermore, while the example given here is the ex115 smartphone, as a terminal, three installation methods can be considered: a transmitting terminal with both an encoder and a decoder, a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. Moreover, in a digital broadcasting system, the example described involves receiving and transmitting multiplexed data, such as audio data, multiplexed within video data. However, in addition to audio data, multiplexed data can also multiplex character data associated with the video, or the video data itself can be received or transmitted without using multiplexed data.
[0502] Furthermore, while the description assumes the CPU's main control unit (ex460) controls the encoding or decoding process, many terminals also possess GPUs. Therefore, a structure can be implemented that utilizes GPU performance to process larger regions simultaneously, using shared memory between the CPU and GPU, or memory that manages addresses in a shared manner. This reduces encoding time, ensures real-time performance, and achieves low latency. In particular, it is even more effective if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing are performed on the GPU at the image level, rather than by the CPU.
[0503] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.
[0504] Industrial availability
[0505] This invention can be used in, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, video conferencing systems, or electronic mirrors.
[0506] Label Explanation
[0507] 100 encoding device
[0508] 102 Division
[0509] 104 Subtraction Section
[0510] 106 Transformer
[0511] Quantitative Department 108
[0512] 110 Entropy Coding Department
[0513] 112, 204 Inverse Quantization Section
[0514] Inverse Transformation Units 114 and 206
[0515] 116, 208 Addition Department
[0516] 118 and 210 memory blocks
[0517] 120, 212 Circular Filter Section
[0518] 122, 214 frame memory
[0519] Intra-frame prediction units 124 and 216
[0520] Inter-frame prediction units 126 and 218
[0521] Predictive Control Department 128, 220
[0522] 160, 260 circuits
[0523] 162, 262 memory
[0524] 200 decoding device
[0525] 202 Entropy Decoding Department
Claims
1. A decoding device, wherein, have: Circuits; and memory, The circuit described above uses the memory described above to perform the following processing in the inter-frame prediction process: Determine whether the inter-frame prediction mode is a merging mode. When the inter-frame prediction mode described above is the merging mode described above. Using the motion vectors of the previously processed previous blocks, derive the first motion vector of the first current block to be processed. By performing motion estimation near the location specified by the first motion vector, the second motion vector of the first current block is derived. Motion compensation is performed using the second motion vector mentioned above to generate the predicted image of the first current block. When the second current block to be processed after the first current block is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image that is different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. By performing motion estimation near the location specified by the third motion vector, the fourth motion vector of the second current block is derived. Motion compensation is performed using the fourth motion vector described above to generate a predicted image of the second current block described above.
2. An encoding device, wherein, have: Circuits; and memory, The circuit described above uses the memory described above to perform the following processing in the inter-frame prediction process: Determine whether the inter-frame prediction mode is a merging mode. When the inter-frame prediction mode described above is the merging mode described above. Using the motion vectors of the previously processed previous blocks, derive the first motion vector of the first current block to be processed. By performing motion estimation near the location specified by the first motion vector, the second motion vector of the first current block is derived. Motion compensation is performed using the second motion vector mentioned above to generate the predicted image of the first current block. When the second current block to be processed after the first current block is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image that is different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. By performing motion estimation near the location specified by the third motion vector, the fourth motion vector of the second current block is derived. Motion compensation is performed using the fourth motion vector described above to generate a predicted image of the second current block described above.
3. A decoding method, wherein, Determine whether the inter-frame prediction mode is a merging mode. When the inter-frame prediction mode described above is the merging mode described above. Using the motion vectors of the previously processed previous blocks, derive the first motion vector of the first current block to be processed. By performing motion estimation near the location specified by the first motion vector, the second motion vector of the first current block is derived. Motion compensation is performed using the second motion vector mentioned above to generate the predicted image of the first current block. When the second current block to be processed after the first current block is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image that is different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. By performing motion estimation near the location specified by the third motion vector, the fourth motion vector of the second current block is derived. Motion compensation is performed using the fourth motion vector described above to generate a predicted image of the second current block described above.
4. An encoding method, wherein, Determine whether the inter-frame prediction mode is a merging mode. When the inter-frame prediction mode described above is the merging mode described above. Using the motion vectors of the previously processed previous blocks, derive the first motion vector of the first current block to be processed. By performing motion estimation near the location specified by the first motion vector, the second motion vector of the first current block is derived. Motion compensation is performed using the second motion vector mentioned above to generate the predicted image of the first current block. When the second current block to be processed after the first current block is contained in the first image containing the first current block, the third motion vector of the second current block is derived using the first motion vector of the first current block. When the second current block is contained in a second image that is different from the first image, the third motion vector of the second current block is derived using the second motion vector of the first current block. By performing motion estimation near the location specified by the third motion vector, the fourth motion vector of the second current block is derived. Motion compensation is performed using the fourth motion vector described above to generate a predicted image of the second current block described above.
Citation Information
Patent Citations
Moving image encoding device, moving image encoding method and moving image encoding program, as well as moving image decoding device, moving image decoding method and moving image decoding program
CN103563386A
Motion information derivation mode determination in video coding
US20160286230A1