Encoding device and encoding method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-04-04
- Publication Date
- 2026-08-14
AI Technical Summary
[0015]本发明能够提供能够实现进一步的改善的编码装置、解码装置、编码方法或解码方法。
Smart Images

Figure CN117097889B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on April 4, 2018, with application number 201880023611.4 and entitled "Encoding device, decoding device, encoding method and decoding method". Technical Field
[0002] This invention relates to encoding devices and encoding methods. Background Technology
[0003] The video coding standard specification known as HEVC (High-Efficiency Video Coding) was standardized by JCT-VC (Joint Collaborative Team on Video Coding).
[0004] Existing technical documents
[0005] Non-patent literature
[0006] Non-patent literature 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention
[0007] The problem that the invention aims to solve
[0008] Further improvements are required in such encoding and decoding techniques.
[0009] Therefore, the object of the present invention is to provide an encoding device, decoding device, encoding method or decoding method that can achieve further improvements.
[0010] Methods used to solve problems
[0011] An encoding apparatus according to a technical solution of the present invention is an encoding apparatus that encodes an object block using motion vectors, characterized in that it comprises: a processor; and a memory; the processor performs the following processing using the memory: deriving a first candidate vector from one or more candidate vectors of one or more adjacent blocks adjacent to the object block; determining a first peripheral region in a first reference image of the object block that includes the position represented by the first candidate vector; calculating the evaluation values of multiple candidate regions included in the first peripheral region; determining a first motion vector of the object block based on the first candidate region; the first candidate region being the candidate region with the smallest evaluation value; and the first peripheral region being included in a first motion search range determined based on the position represented by the first candidate vector.
[0012] The encoding method of one technical solution of the present invention is an encoding method for encoding an object block using motion vectors. The method is characterized by deriving a first candidate vector from one or more candidate vectors of one or more adjacent blocks adjacent to the object block; determining a first peripheral region in a first reference image of the object block that includes the position represented by the first candidate vector; calculating the evaluation values of multiple candidate regions included in the first peripheral region; determining a first motion vector of the object block based on the first candidate region; the first candidate region being the candidate region with the smallest evaluation value; and the first peripheral region being included within a first motion search range determined based on the position represented by the first candidate vector.
[0013] In addition, these inclusive or specific technical solutions can also be implemented by systems, methods, integrated circuits, computer programs or computer-readable CD-ROMs and other recording media, or by any combination of systems, methods, integrated circuits, computer programs and recording media.
[0014] Invention Effects
[0015] The present invention can provide encoding apparatus, decoding apparatus, encoding method or decoding method that can achieve further improvements. Attached Figure Description
[0016] Figure 1 This is a block diagram showing the functional structure of the encoding device according to Embodiment 1.
[0017] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1.
[0018] Figure 3 It is a table representing the transformation basis functions corresponding to each transformation type.
[0019] Figure 4A This is a diagram showing an example of the shape of the filter used in ALF.
[0020] Figure 4B This is another example of the shape of the filter used in ALF.
[0021] Figure 4C This is another example of the shape of the filter used in ALF.
[0022] Figure 5A This is a diagram representing the 67 intra-prediction modes of intra-frame prediction.
[0023] Figure 5B This is a flowchart illustrating the outline of predictive image correction processing based on OBMC processing.
[0024] Figure 5CThis is a conceptual diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.
[0025] Figure 5D This is a diagram representing an example of FRUC.
[0026] Figure 6 It is a diagram used to illustrate pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0027] Figure 7 It is a diagram used to illustrate pattern matching (template matching) between a template in the current image and a block in a reference image.
[0028] Figure 8 It is a diagram used to illustrate a model that assumes uniform linear motion.
[0029] Figure 9A It is a diagram used to illustrate the derivation of the motion vectors of sub-block units based on the motion vectors of multiple adjacent blocks.
[0030] Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.
[0031] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.
[0032] Figure 9D This is a diagram used to illustrate an overview of a predictive image generation method that employs LIC-based brightness correction processing.
[0033] Figure 10 This is a block diagram illustrating the functional structure of the decoding device according to Embodiment 1.
[0034] Figure 11 This is a block diagram showing the internal structure of the intra-frame prediction unit of the coding apparatus according to Embodiment 1.
[0035] Figure 12 This is a diagram showing an example of the location of motion search range information within the bitstream of Implementation Method 1.
[0036] Figure 13 This is a flowchart illustrating the processing of the intra-frame prediction unit of the encoding / decoding apparatus according to Embodiment 1.
[0037] Figure 14 This is a diagram showing an example of a candidate list for Implementation Method 1.
[0038] Figure 15 This is an example of a list of reference images representing Implementation Method 1.
[0039] Figure 16This is a diagram illustrating an example of the motion search range in Implementation Method 1.
[0040] Figure 17 This is a diagram showing an example of the surrounding area of Implementation Method 1.
[0041] Figure 18 This is a block diagram showing the internal structure of the intra-frame prediction unit of the decoding apparatus in Embodiment 1.
[0042] Figure 19 This is a diagram illustrating an example of the motion search range of a variation of Implementation 1, Example 2.
[0043] Figure 20 This is a diagram illustrating an example of the motion search range of a variation of implementation 1, example 4.
[0044] Figure 21 This is a diagram illustrating an example of the motion search range of a variation of implementation 1, example 5.
[0045] Figure 22 This is a diagram illustrating an example of the motion search range of a variation of Implementation 1, Example 6.
[0046] Figure 23 This is a block diagram illustrating the functional structure of the encoding / decoding system according to Implementation Method 1.
[0047] Figure 24 This is a diagram showing the motion search range of variation 9 of implementation method 1.
[0048] Figure 25 This is a diagram showing the overall structure of a content supply system that enables content distribution services.
[0049] Figure 26 This is a diagram illustrating an example of the encoding structure in hierarchical coding.
[0050] Figure 27 This is a diagram illustrating an example of the encoding structure in hierarchical coding.
[0051] Figure 28 This is an example of a web page display.
[0052] Figure 29 This is an example of a web page display.
[0053] Figure 30 This is a diagram illustrating an example of a smartphone.
[0054] Figure 31 This is a block diagram representing a structural example of a smartphone. Detailed Implementation
[0055] (The insights that form the basis of this disclosure)
[0056] In next-generation moving image compression specifications, to reduce the amount of motion information encoded for motion compensation, a motion search mode has been investigated on the decoding device side. In this mode, the decoding device derives the motion vector for the target block by searching (motion search) for regions within a reference image similar to a previously decoded block that are different from the target block. However, considering the anticipated increase in processing load on the decoding device due to motion search, and the increased memory bandwidth required by the decoding device due to data transfer from the reference image, techniques are needed to suppress these increases in processing load and memory bandwidth.
[0057] Therefore, an encoding apparatus according to one aspect of the present invention uses motion vectors to encode a block of objects. The encoding apparatus includes a processor and a memory. The processor uses the memory to derive multiple candidates, each having at least one motion vector, determines a motion search range in a reference image, performs a motion search within the motion search range of the reference image based on the multiple candidates, and encodes information related to the determined motion search range.
[0058] Accordingly, motion search can be performed within the determined motion search range. Therefore, motion search outside the motion search range is unnecessary, thus reducing the processing load for motion search. Furthermore, since reconstructed images outside the motion search range can be avoided from being read from the frame memory, the memory bandwidth requirement for motion search can be reduced.
[0059] In addition, in an encoding apparatus of a technical solution of the present invention, for example, in the above-mentioned motion search, candidates with motion vectors corresponding to positions outside the above-mentioned motion search range are excluded from the above-mentioned plurality of candidates, candidates are selected from the remaining candidates of the above-mentioned plurality of candidates, and the motion vector for the above-mentioned encoded object block is determined based on the selected candidate.
[0060] Therefore, it is possible to select candidates after excluding those with motion vectors corresponding to positions outside the motion search range. Thus, the processing load for candidate selection can be reduced.
[0061] In addition, in an encoding device relating to a technical solution of the present invention, for example, the information related to the motion search range may include information indicating the size of the motion search range.
[0062] Accordingly, information representing the size of the motion search range can be included in the bitstream. Therefore, a motion search range with the same size as the one used in the encoding apparatus can be used in the decoding device. Furthermore, the processing load for determining the size of the motion search range in the decoding device can be reduced.
[0063] Furthermore, in an encoding apparatus relating to one aspect of the present invention, for example, in the derivation of the plurality of candidates, the plurality of candidates may be derived from a plurality of already encoded blocks that are spatially or temporally adjacent to the encoded object block, and the position of the motion search range may be determined based on the average motion vector of the plurality of motion vectors contained in the plurality of candidates. Additionally, in an encoding apparatus relating to one aspect of the present invention, for example, in the derivation of the plurality of candidates, the plurality of candidates may be derived from a plurality of blocks that are spatially or temporally adjacent to the encoded object block, and the position of the motion search range may be determined based on the central motion vector of the plurality of motion vectors contained in the plurality of candidates.
[0064] Accordingly, the location of the motion search range can be determined based on multiple candidates derived from multiple coded blocks adjacent to the coded object block. Therefore, the motion search range can be determined by defining the region suitable for searching motion vectors for the coded object block, thereby improving the accuracy of the motion vectors.
[0065] In addition, in an encoding apparatus relating to one technical solution of the present invention, for example, the position of the aforementioned motion search range may be determined based on the average motion vector of multiple motion vectors used in the encoding of the encoded image.
[0066] Therefore, the location of the motion search range can be determined based on the motion vectors of the encoded image. Even if the encoded object block within the encoded object image changes, the motion vectors of the encoded image do not change, so it is not necessary to determine the motion search range based on the motion vectors of adjacent blocks every time the encoded object block changes. That is, the processing load for determining the motion search range can be reduced.
[0067] In addition, in an encoding apparatus for a technical solution of the present invention, for example, in determining the motion vector for the aforementioned encoded object block, pattern matching is performed in the surrounding area of the position corresponding to the selected candidate motion vector in the aforementioned reference image to find the most matching area in the surrounding area, and the motion vector for the aforementioned encoded object block is determined based on the most matching area.
[0068] Therefore, in addition to candidate motion vectors, motion vectors for encoding object blocks can be determined based on pattern matching of surrounding areas. This further improves the accuracy of motion vectors.
[0069] Furthermore, in an encoding apparatus relating to a technical solution of the present invention, for example, in determining the motion vector for the encoded object block, it may be determined whether the surrounding area is included in the motion search range; if the surrounding area is included in the motion search range, the pattern matching is performed in the surrounding area; if the surrounding area is not included in the motion search range, the pattern matching is performed in a portion of the surrounding area that is included in the motion search range.
[0070] Therefore, even when the surrounding area is not included in the motion search range, pattern matching can be performed in a portion of the motion search range within the surrounding area. This avoids motion searches outside the motion search range, reducing processing load and memory bandwidth requirements.
[0071] In addition, in an encoding apparatus of a technical solution of the present invention, for example, in determining the motion vector for the encoded object block, it may be determined whether the surrounding area is included in the motion search range; if the surrounding area is included in the motion search range, the pattern matching is performed in the surrounding area; if the surrounding area is not included in the motion search range, the motion vector included in the selected candidate is determined as the motion vector for the encoded object block.
[0072] Therefore, pattern matching in the surrounding area can be omitted if the surrounding area is not included in the motion search range. This avoids motion searches outside the motion search range, reducing processing load and memory bandwidth requirements.
[0073] An encoding method for a technical solution of the present invention uses motion vectors to encode a block of objects, derives multiple candidates, each having at least one motion vector, determines a motion search range in a reference image, performs a motion search within the motion search range of the reference image based on the multiple candidates, and encodes information related to the determined motion search range.
[0074] Therefore, it can achieve the same effect as the aforementioned encoding device.
[0075] A decoding apparatus according to a technical solution of the present invention uses motion vectors to decode a block of objects. The decoding apparatus includes a processor and a memory. The processor uses the memory to interpret information related to the motion search range from the bit stream, derives multiple candidates, each having at least one motion vector, determines the motion search range in a reference image based on the information related to the motion search range, and performs a motion search within the motion search range of the reference image based on the multiple candidates.
[0076] Accordingly, motion search can be performed within the determined motion search range. Therefore, motion search outside the motion search range is unnecessary, thus reducing the processing load for motion search. Furthermore, since reconstructed images outside the motion search range can be avoided from being read from the frame memory, the memory bandwidth requirement for motion search can be reduced.
[0077] In addition, in a decoding apparatus of a technical solution of the present invention, for example, in the above-mentioned motion search, candidates with motion vectors corresponding to positions outside the above-mentioned motion search range are excluded from the above-mentioned plurality of candidates, candidates are selected from the remaining candidates of the above-mentioned plurality of candidates, and the motion vector for the above-mentioned decoding object block is determined based on the selected candidate.
[0078] Therefore, it is possible to select candidates after excluding those with motion vectors corresponding to positions outside the motion search range. Thus, the processing load for candidate selection can be reduced.
[0079] Furthermore, in a decoding device relating to one technical solution of the present invention, for example, the information related to the aforementioned motion search range may include information indicating the size of the aforementioned motion search range.
[0080] Accordingly, information representing the size of the motion search range can be included in the bitstream. Therefore, a motion search range with the same size as the one used in the encoding apparatus can be used in the decoding device. Furthermore, the processing load for determining the size of the motion search range in the decoding device can be reduced.
[0081] Furthermore, in a decoding apparatus according to a technical solution of the present invention, for example, in the derivation of the plurality of candidates, the plurality of candidates may be derived from a plurality of already decoded blocks that are spatially or temporally adjacent to the target block being decoded, and the position of the motion search range may be determined based on the average motion vector of the plurality of motion vectors contained in the plurality of candidates. Furthermore, in a decoding apparatus according to a technical solution of the present invention, for example, in the derivation of the plurality of candidates, the plurality of candidates may be derived from a plurality of blocks that are spatially or temporally adjacent to the target block being decoded, and the position of the motion search range may be determined based on the central motion vector of the plurality of motion vectors contained in the plurality of candidates.
[0082] Accordingly, the location of the motion search range can be determined based on multiple candidates derived from multiple decoded blocks adjacent to the decoded object block. Therefore, the motion search range can be determined by defining the region suitable for searching motion vectors for the decoded object block, thus improving the accuracy of the motion vectors.
[0083] In addition, in a decoding apparatus related to one technical solution of the present invention, for example, the position of the aforementioned motion search range may be determined based on the average motion vector of multiple motion vectors used in the decoding of the decoded image.
[0084] Therefore, the location of the motion search range can be determined based on the decoded image. Even if the decoded object block within the decoded object image changes, the motion vector of the decoded image does not change, so it is not necessary to determine the motion search range based on the motion vectors of adjacent blocks every time the decoded object block changes. That is, the processing load for determining the motion search range can be reduced.
[0085] In addition, in a decoding apparatus of a technical solution of the present invention, for example, in determining the motion vector for the aforementioned decoding object block, pattern matching is performed in the surrounding area of the position corresponding to the selected candidate motion vector in the aforementioned reference image to find the most matching area in the surrounding area, and the motion vector for the aforementioned decoding object block is determined based on the most matching area.
[0086] Therefore, in addition to candidate motion vectors, motion vectors for encoding object blocks can be determined based on pattern matching of surrounding areas. This further improves the accuracy of motion vectors.
[0087] Furthermore, in an encoding apparatus relating to a technical solution of the present invention, for example, in determining the motion vector for the aforementioned decoded object block, it may be determined whether the surrounding region is included in the aforementioned motion search range; if the surrounding region is included in the aforementioned motion search range, the aforementioned pattern matching is performed in the surrounding region; if the surrounding region is not included in the aforementioned motion search range, the aforementioned pattern matching is performed in the portion of the surrounding region that is included in the aforementioned motion search range.
[0088] Therefore, even when the surrounding area is not included in the motion search range, pattern matching can be performed in a portion of the motion search range within the surrounding area. This avoids motion searches outside the motion search range, reducing processing load and memory bandwidth requirements.
[0089] Furthermore, in an encoding apparatus relating to a technical solution of the present invention, for example, in determining the motion vector for the aforementioned decoded object block, it may be determined whether the surrounding region is included in the aforementioned motion search range; if the surrounding region is included in the aforementioned motion search range, the aforementioned pattern matching is performed in the aforementioned surrounding region; if the surrounding region is not included in the aforementioned motion search range, the motion vector included in the selected candidate is determined as the motion vector for the aforementioned decoded object block.
[0090] Therefore, pattern matching in the surrounding area can be omitted if the surrounding area is not included in the motion search range. This avoids motion searches outside the motion search range, reducing processing load and memory bandwidth requirements.
[0091] A decoding method for a technical solution of the present invention uses motion vectors to decode a block of objects, interprets information related to the motion search range from the bitstream, derives multiple candidates, each having at least one motion vector, determines the motion search range in a reference image based on the information related to the motion search range, and performs a motion search within the motion search range of the reference image based on the multiple candidates.
[0092] Therefore, it can achieve the same effect as the aforementioned decoding device.
[0093] In addition, these inclusive or specific technical solutions can also be implemented by systems, integrated circuits, computer programs or computer-readable CD-ROMs and other recording media, or by any combination of systems, methods, integrated circuits, computer programs and recording media.
[0094] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.
[0095] Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangements and connection methods of constituent elements, steps, and sequences of steps shown in the following embodiments are examples and are not intended to limit the scope of the claims. In addition, any constituent elements in the following embodiments that are not described in the independent claim representing the highest-level concept are described as arbitrary constituent elements.
[0096] (Implementation Method 1)
[0097] First, as an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention described later, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention, and the processing and / or structure described in the various embodiments of the present invention can also be implemented in encoding and decoding apparatuses different from Embodiment 1.
[0098] When applying the processing and / or structure described in various aspects of the present invention to Embodiment 1, one of the following may also be performed, for example.
[0099] (1) For the encoding or decoding device of Embodiment 1, the constituent element that corresponds to the constituent element described in each aspect of the present invention is replaced with the constituent element described in each aspect of the present invention.
[0100] (2) For the encoding or decoding device of Embodiment 1, after any modification such as adding, replacing, or deleting any of the constituent elements of the plurality of constituent elements constituting the encoding or decoding device, the constituent elements corresponding to the constituent elements described in each aspect of the present invention are replaced with the constituent elements described in each aspect of the present invention.
[0101] (3) After adding processing to the method implemented by the encoding or decoding device of Embodiment 1, and / or replacing or deleting any of the processing among the multiple processing included in the method, the processing corresponding to the processing described in each aspect of the present invention is replaced with the processing described in each aspect of the present invention.
[0102] (4) A portion of the constituent elements constituting the encoding or decoding apparatus of Embodiment 1 are combined with constituent elements described in various aspects of the present invention, a portion of constituent elements having the functions of constituent elements described in various aspects of the present invention, or a portion of constituent elements implementing the processing performed by constituent elements described in various aspects of the present invention.
[0103] (5) A component having a portion of the functions of a portion of the components constituting the encoding or decoding apparatus of embodiment 1, or a component implementing a portion of the processing performed by a portion of the components constituting the encoding or decoding apparatus of embodiment 1, is combined with the components described in various aspects of the present invention, the components having a portion of the functions of the components described in various aspects of the present invention, or the components implementing a portion of the processing performed by the components described in various aspects of the present invention.
[0104] (6) For the method implemented by the encoding or decoding device of Embodiment 1, the processing that corresponds to the processing described in each aspect of the present invention among the multiple processing included in the method is replaced with the processing described in each aspect of the present invention.
[0105] (7) A portion of the processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1 is combined with the processing described in the various aspects of the present invention.
[0106] Furthermore, the implementation of the processes and / or structures described in the various embodiments of the present invention is not limited to the examples described above. For example, it may be implemented in an apparatus used for a different purpose than the moving image / image encoding apparatus or moving image / image decoding apparatus disclosed in Embodiment 1, or the processes and / or structures described in each embodiment may be implemented individually. In addition, the processes and / or structures described in different embodiments may be combined and implemented.
[0107] [Overview of the encoding device]
[0108] First, an overview of the encoding device for Embodiment 1 will be provided. Figure 1 This is a block diagram illustrating the functional structure of the encoding apparatus 100 according to Embodiment 1. The encoding apparatus 100 is a motion picture / image encoding apparatus that encodes motion pictures / images in block units.
[0109] like Figure 1 As shown, the encoding device 100 is a device for encoding images in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a cyclic filtering unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128.
[0110] The encoding device 100 is implemented, for example, by a general-purpose processor and memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128.
[0111] The following describes the constituent elements included in the encoding device 100.
[0112] [Divider]
[0113] The segmentation unit 102 divides each image contained in the input moving image into multiple blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the image into fixed-size blocks (e.g., 128×128). These fixed-size blocks may be called coding tree units (CTUs). Furthermore, the segmentation unit 102 divides each fixed-size block into variable-size blocks (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block segmentation. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the image may be used as processing units of CUs, PUs, and TUs.
[0114] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1. In Figure 2 In the diagram, solid lines represent block boundaries based on quadtree block partitioning, and dashed lines represent block boundaries based on binary tree block partitioning.
[0115] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into 4 square blocks of 64×64 (quadtree block partitioning).
[0116] The 64×64 block in the upper left corner is then vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is then vertically divided into two rectangular 16×64 blocks (binary tree block partitioning). As a result, the 64×64 block in the upper left corner is divided into two 16×64 blocks (11 and 12) and a 32×64 block (13).
[0117] The 64×64 block in the upper right corner is horizontally divided into two rectangular 64×32 blocks, 14 and 15 (binary tree block division).
[0118] The 64×64 block in the lower left corner is divided into four 32×32 square blocks (quadtree block partitioning). The upper left and lower right blocks of these four 32×32 blocks are further partitioned. The upper left 32×32 block is vertically divided into two 16×32 rectangular blocks, and the right 16×32 block is horizontally divided into two 16×16 blocks (binary tree block partitioning). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block partitioning). As a result, the lower left 64×64 block is divided into 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.
[0119] The 64×64 block 23 in the lower right corner is not divided.
[0120] As described above, in Figure 2 In the example, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quadtree plus binary tree) partitioning.
[0121] In addition, Figure 2 In this context, a block can be divided into 2 or 4 blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these. For example, a block can also be divided into 3 blocks (ternary tree partitioning). Partitioning including such ternary tree partitioning is sometimes referred to as MBT (multi-type tree) partitioning.
[0122] [Subtraction Section]
[0123] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in block units divided by the segmentation unit 102. That is, the subtraction unit 104 calculates the prediction error (also called residual) of the encoded target block (hereinafter referred to as the current block). Furthermore, the subtraction unit 104 outputs the calculated prediction error to the transformation unit 106.
[0124] The original signal is the input signal of the encoding device 100, which is the signal representing the image of each picture that constitutes the moving image (e.g., luminance signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0125] [Transformation Section]
[0126] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain, and outputs the transformation coefficients vectorization unit 108. Specifically, the transformation unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0127] Alternatively, the transform unit 106 can adaptively select a transform type from multiple transform types and use the transform basis function corresponding to the selected transform type to transform the prediction error into transform coefficients. Such a transform is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0128] Several transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 This is a table representing the transformation basis functions corresponding to each transformation type. Figure 3 In this context, N represents the number of input pixels. The choice of transform type from these multiple transform types can depend on the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode.
[0129] Information indicating whether such EMT or AMT is applied (e.g., referred to as the AMT flag) and information indicating the selected transform type are signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0130] Furthermore, the transform unit 106 can also perform a re-transformation on the transform coefficients (transformation results). Such a re-transformation may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (e.g., a 4×4 sub-block) contained in the block of transform coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are signaled at the CU level. In addition, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0131] Here, a separable transformation refers to a method of performing multiple transformations in each direction, which is equivalent to the number of dimensions of the input. A non-separable transformation refers to a method of treating two or more dimensions together as one dimension and transforming them together when the input is multidimensional.
[0132] For example, as one example of a non-separable transformation, one can cite the way that when the input is a 4×4 block, it is treated as a permutation of 16 elements, and the permutation is transformed using a 16×16 transformation matrix.
[0133] Furthermore, the Hypercube Givens Transform, which treats a 4×4 input block as a permutation of 16 elements and then performs multiple Givens rotations on that permutation, is also an example of a non-separable transformation.
[0134] [Quantitative Department]
[0135] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scan order and quantizes the transform coefficients based on the quantization parameters (QP) corresponding to the scanned transform coefficients. Furthermore, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0136] The specified order is the order in which the transform coefficients are quantized / inverse quantized. For example, the specified scan order is defined by ascending frequency (from low frequency to high frequency) or descending frequency (from high frequency to low frequency).
[0137] The quantization parameter is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0138] [Entropy Coding Department]
[0139] The entropy coding unit 110 generates a coded signal (coded bitstream) by performing variable-length coding on the quantization coefficients, which are input from the quantization unit 108. Specifically, the entropy coding unit 110 performs arithmetic coding on the binary signal, for example, by binarizing the quantization coefficients.
[0140] [De-quantization Department]
[0141] The inverse quantization unit 112 performs inverse quantization on the quantization coefficients that are input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantization coefficients of the current block in a predetermined scan order. Furthermore, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0142] [Inverse Transformation Section]
[0143] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients, which are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Furthermore, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.
[0144] Furthermore, the restored prediction error differs from the prediction error calculated by the subtraction unit 104 because information was lost during quantization. In other words, the restored prediction error includes quantization error.
[0145] [Addition Department]
[0146] The addition unit 116 reconstructs the current block by adding the prediction error, which is input from the inverse transform unit 114, to the prediction sample, which is input from the prediction control unit 128. Furthermore, the addition unit 116 outputs the reconstructed block to the block memory 118 and the cyclic filtering unit 120. The reconstructed block may be referred to as a local decoding block.
[0147] [Block Memory]
[0148] Block memory 118 is a storage unit used to save blocks within the encoded object image (hereinafter referred to as the current image) referenced in intra-frame prediction. Specifically, block memory 118 saves the reconstructed blocks output from addition unit 116.
[0149] [Loop Filtering Section]
[0150] The cyclic filtering unit 120 applies cyclic filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Cyclic filtering refers to filtering used within the encoding loop (in-loop filtering), such as deblocking filtering (DF), sample adaptive offset (SAO), and adaptive cyclic filtering (ALF).
[0151] In ALF, a least-squares error filter is used to remove coding distortion. For example, for each 2×2 sub-block within the current block, one filter is selected from multiple filters based on the direction of the gradient and the activity of the locality.
[0152] Specifically, sub-blocks (e.g., 2×2 sub-blocks) are first classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0–2 or 0–4) and the gradient activity value A (e.g., 0–4), a classification value C (e.g., C = 5D + A) is calculated. Then, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).
[0153] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.
[0154] Based on the results of this classification, the filter used for the sub-block is determined from among multiple filters.
[0155] The shape of the filter used in ALF can be, for example, a circular symmetrical shape. Figures 4A to 4C This is a diagram showing several examples of the shapes of filters used in ALF. Figure 4A This indicates a 5×5 diamond-shaped filter. Figure 4BThis indicates a 7×7 diamond-shaped filter. Figure 4C This represents a 9×9 diamond-shaped filter. Information representing the filter's shape is signaled at the image level. However, the signaling of the filter's shape information is not limited to the image level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0156] The on / off state of ALF is determined, for example, at the picture level or the CU level. For instance, regarding luminance, the decision to use ALF is made at the CU level, while regarding chromatic aberration, it is made at the picture level. Information indicating the on / off state of ALF is signaled at the picture level or the CU level. However, the signaling of information indicating the on / off state of ALF is not limited to the picture level or the CU level; it can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0157] The coefficient set of a selectable set of filters (e.g., up to 15 or 25 filters) is signaled at the picture level. Furthermore, the signaling of the coefficient set is not limited to the picture level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0158] [Frame Memory]
[0159] The frame memory 122 is a storage unit used to store reference images used in inter-frame prediction, and is also sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the cyclic filtering unit 120.
[0160] Intra-frame prediction unit
[0161] The intra-frame prediction unit 124 performs intra-frame prediction (also called intra-picture prediction) of the current block by referring to the blocks in the current image stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction by referring to samples (e.g., luminance value, chrominance value) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.
[0162] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of predefined intra-prediction modes. The plurality of intra-prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0163] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode as specified by the H.265 / HEVC (High-Efficiency Video Coding) specification (Non-Patent Document 1).
[0164] Multiple directional prediction modes may include, for example, the 33 directional prediction modes specified in the H.265 / HEVC specification. Alternatively, multiple directional prediction modes may also include 32 additional directional prediction modes (a total of 65 directional prediction modes). Figure 5A This diagram represents the 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. Solid arrows indicate the 33 directions specified by the H.265 / HEVC specification, while dashed arrows indicate the additional 32 directions.
[0165] Additionally, in intra-frame prediction of chroma blocks, luma blocks can also be referenced. That is, the chroma components of the current block can be predicted based on the luma components of the current block. Such intra-frame prediction is sometimes referred to as CCLM (cross-component linear model) prediction. This intra-frame prediction mode of chroma blocks referencing luma blocks (e.g., called CCLM mode) can also be added as one of the intra-frame prediction modes for chroma blocks.
[0166] The intra-prediction unit 124 can also correct the intra-predicted pixel values based on the gradient of the reference pixels in the horizontal / vertical directions. Intra-prediction accompanied by such correction is sometimes referred to as PDPC (position-dependent intraprediction combination). Information indicating whether PDPC has been used (e.g., a PDPC flag) is signaled, for example, at the CU level. Furthermore, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0167] [Inter-frame prediction department]
[0168] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) of the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Furthermore, the inter-frame prediction unit 126 uses motion information (e.g., motion vectors) obtained through motion estimation to perform motion compensation, thereby generating the inter-frame prediction signal for the current block or sub-block. Finally, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0169] The motion information used in motion compensation is signaled. A motion vector predictor can also be used in the signaling of motion vectors. That is, the difference between the motion vector and the predicted motion vector can also be signaled.
[0170] Alternatively, the inter-frame prediction signal can be generated using not only the motion information of the current block obtained through motion search, but also the motion information of neighboring blocks. Specifically, the prediction signal based on the motion information obtained through motion search can be weighted and added together with the prediction signal based on the motion information of neighboring blocks, thereby generating the inter-frame prediction signal in sub-block units within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0171] In this OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) is signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is used (e.g., OBMC flag) is signaled at the CU level. However, the signaling level for this information is not limited to the sequence and CU levels; it can also be other levels (e.g., image level, slice level, tile level, CTU level, or sub-block level).
[0172] The OBMC model will be explained in more detail. Figure 5B and Figure 5C This is a flowchart and concept diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.
[0173] First, the predicted image (Pred) obtained through normal motion compensation is obtained using the motion vectors (MV) assigned to the encoded object block.
[0174] Next, the predicted image (Pred_L) is obtained by using the motion vector (MV_L) of the encoded left adjacent block for the encoded object block. The first correction of the predicted image is performed by weighted superposition of the predicted image and Pred_L.
[0175] Similarly, the predicted image (Pred_U) is obtained by using the motion vector (MV_U) of the upper adjacent block of the encoded object block. The predicted image is then corrected a second time by weighting and superimposing the predicted image after the first correction and Pred_U, and this is used as the final predicted image.
[0176] In addition, this describes a two-stage correction method using the left and top adjacent blocks, but it can also be configured to perform more corrections using the right and bottom adjacent blocks than the two-stage method.
[0177] In addition, the area to be overlaid may not be the entire pixel area of the block, but only a part of the area near the block boundary.
[0178] Furthermore, the process of correcting the predicted image based on a single reference image is explained here. However, the same principle applies when correcting the predicted image based on multiple reference images. After obtaining the corrected predicted image based on each reference image, the resulting predicted images are further superimposed to obtain the final predicted image.
[0179] In addition, the processing target block mentioned above can be a prediction block unit or a sub-block unit that further divides the prediction block.
[0180] One method for determining whether to use OBMC processing is to use a signal called obmc_flag. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a motion-complex region. If it does, the obmc_flag is set to 1 and OBMC processing is performed for encoding. If it does not belong to a motion-complex region, the obmc_flag is set to 0, and OBMC processing is not performed for encoding. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag recorded in the stream and switching between using and not using OBMC processing based on its value.
[0181] Alternatively, motion information can be exported at the decoding device side without being signaled. For example, the merging mode specified by the H.265 / HEVC standard can be used. Furthermore, motion information can also be exported by performing a motion search at the decoding device side. In this case, the motion search is performed without using the pixel values of the current block.
[0182] Here, we will explain the motion search mode performed on the decoding device side. This motion search mode on the decoding device side may be called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0183] exist Figure 5DThe diagram below illustrates an example of FRUC processing. First, referencing the motion vectors of coded blocks spatially or temporally adjacent to the current block, a list of multiple candidates, each with a predicted motion vector, is generated (this list can also be shared with a merge list). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value is calculated for each candidate included in the candidate list, and one candidate is selected based on the evaluation value.
[0184] Furthermore, based on the selected candidate motion vectors, motion vectors for the current block are derived. Specifically, for example, the selected candidate motion vector (best candidate MV) can be derived as is, using it as the motion vector for the current block. Alternatively, for example, motion vectors for the current block can be derived by performing pattern matching in the surrounding region of the position within the reference image corresponding to the selected candidate motion vector. That is, the surrounding region of the best candidate MV can be searched using the same method, and if an MV with a better evaluation value is found, the best candidate MV is updated to the aforementioned MV and used as the final MV for the current block. Alternatively, a structure that does not perform this processing can be implemented.
[0185] The exact same processing can also be performed when processing is done in sub-block units.
[0186] Furthermore, the evaluation value is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region within the reference image corresponding to the motion vector and the specified region. Alternatively, information other than the difference value can be used to calculate the evaluation value.
[0187] As a pattern matching, either pattern matching 1 or pattern matching 2 is used. Pattern matching 1 and pattern matching 2 can be referred to as bilateral matching and template matching, respectively.
[0188] In the first pattern matching, pattern matching is performed between two blocks within two different reference images, along the motion trajectory of the current block. Therefore, in the first pattern matching, the regions within other reference images along the motion trajectory of the current block are used as the defined regions for calculating the candidate evaluation values described above.
[0189] Figure 6 This diagram illustrates an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory. For example... Figure 6As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among two blocks in two different reference images (Ref0, Ref1) along the motion trajectory of the current block. Specifically, for the current block, the difference between the reconstructed image at a specified position in the first encoded reference image (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference image (Ref1) specified by the symmetrical MV scaled by the aforementioned candidate MV over the display time interval is derived, and the obtained difference value is used to calculate an evaluation value. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs.
[0190] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, in the case where the current image is located between the two reference images in time and the temporal distances from the current image to the two reference images are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0191] In the second pattern matching, pattern matching is performed between the template in the current image (the block adjacent to the current block in the current image (e.g., the upper and / or left adjacent block)) and the block in the reference image. Therefore, in the second pattern matching, the block adjacent to the current block in the current image is used as the defined area for calculating the candidate evaluation value as described above.
[0192] Figure 7 This is an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 7 As shown, in the second pattern matching, the motion vector of the current block is derived by searching within the reference image (Ref0) for the block that best matches the block adjacent to the current block (Cur block) within the current image (Cur Pic). Specifically, for the current block, the difference between the reconstructed images of the encoded regions of the left and top adjacent regions or one of them and the reconstructed image at the same position within the encoded reference image (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.
[0193] Information indicating whether FRUC mode is used (e.g., referred to as the FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is used (e.g., when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (e.g., referred to as the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0194] This section explains how to derive motion vector patterns based on a model that assumes uniform linear motion. This pattern can be termed BIO (bi-directional optical flow).
[0195] Figure 8 This diagram is used to illustrate a model that assumes uniform linear motion. In Figure 8 In the middle, (v x v y () represents the velocity vector, and τ0 and τ1 represent the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1), respectively. (MVx0, MVy0) represents the motion vector corresponding to the reference image Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference image Ref1.
[0196] At this time, in the velocity vector (v x v y Under the assumption of constant linear motion of MVx0, MVy0 and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1) respectively, and the following optical flow equation (1) holds.
[0197] [Formula 1]
[0198]
[0199] Here, I (k) This represents the luminance value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation states that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on this optical flow equation combined with Hermite interpolation, the block-unit motion vector obtained from merge lists, etc., is corrected in pixels.
[0200] Alternatively, motion vectors can be derived on the decoding device side using a different method than deriving motion vectors based on a model assuming constant linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0201] Here, we will explain the mode of deriving motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks. This mode is sometimes referred to as the affine motion compensation prediction mode.
[0202] Figure 9A This is a diagram used to illustrate the derivation of sub-block unit motion vectors based on the motion vectors of multiple adjacent blocks. Figure 9A In this context, the current block comprises 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the top-left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the top-right control point of the current block is derived. Furthermore, using the two motion vectors v0 and v1, the motion vectors (v0, v1, v1) of each sub-block within the current block are derived using the following equation (2). x v y ).
[0203] [Formula 2]
[0204]
[0205] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents the pre-set weight coefficient.
[0206] Such an affine motion compensation prediction mode may also include several modes with different methods for deriving the motion vectors of the upper left and upper right control points. Information representing such an affine motion compensation prediction mode (e.g., affine flags) is signaled at the CU level. Furthermore, the signaling of information representing this affine motion compensation prediction mode is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, CTU level, or sub-block level).
[0207] [Forecasting and Control Department]
[0208] The prediction control unit 128 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0209] This section illustrates an example of exporting motion vectors from an encoded object image using a merge mode. Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.
[0210] First, a list of candidate predicted MVs registered with the predicted MVs is generated. Candidate predicted MVs include: spatially adjacent predicted MVs (MVs) belonging to multiple coded blocks spatially surrounding the coded object block; temporally adjacent predicted MVs (MVs) belonging to blocks whose positions in the coded reference image are projected nearby; combined predicted MVs (MVs) generated by combining the MV values of spatially adjacent and temporally adjacent predicted MVs; and zero predicted MVs (MVs with a value of zero).
[0211] Next, the MV for the encoded object block is determined by selecting one predicted MV from the multiple predicted MVs registered in the predicted MV list.
[0212] Furthermore, in the variable-length coding section, merge_idx, which represents the signal that selected which prediction MV was recorded in the stream and encoded.
[0213] In addition, Figure 9B The predicted MVs registered in the predicted MV list described in the figure are one example. They may also be a number different from the number shown in the figure, or a structure that does not include a part of the predicted MVs in the figure, or a structure that adds predicted MVs other than the predicted MVs in the figure.
[0214] Alternatively, the MV of the encoded object block exported through the merge mode can be used for the DMVR processing described later to determine the final MV.
[0215] Here, an example of using DMVR to determine MV is explained.
[0216] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.
[0217] First, the optimal MVP set for the processing object block is taken as the candidate MV. According to the candidate MV, reference pixels are obtained from the first reference image of the processed image in the L0 direction and the second reference image of the processed image in the L1 direction, respectively. The template is generated by taking the average of each reference pixel.
[0218] Next, using the template described above, the surrounding areas of the candidate music videos (MVs) for the first and second reference images are searched, and the MV with the lowest cost is selected as the final MV. Furthermore, the cost value is calculated using the differences between the pixel values of the template and the pixel values of the search area, as well as the MV value.
[0219] Furthermore, the general outline of the processing described herein is essentially the same in both the encoding and decoding devices.
[0220] In addition, even if it is not the process described here, any other process that can search for the surrounding of candidate MVs and export the final MV can be used.
[0221] Here, the mode of generating predicted images using LIC processing is explained.
[0222] Figure 9D This is a diagram illustrating the outline of a predictive image generation method using LIC-based brightness correction processing.
[0223] First, export the MV used to obtain the reference image corresponding to the encoded object block from the reference image, which is an encoded image.
[0224] Next, for the encoded object block, using the brightness pixel values of the left and top adjacent encoded surrounding reference areas and the brightness pixel values at the same position in the reference image specified by MV, information indicating how the brightness values change in the reference image and the encoded object image is extracted, and brightness correction parameters are calculated.
[0225] By using the aforementioned brightness correction parameters to perform brightness correction processing on the reference image within the reference image specified by MV, a predicted image for the coded object block is generated.
[0226] in addition, Figure 9D The shape of the surrounding reference area mentioned above is one example; other shapes may also be used.
[0227] Furthermore, the process of generating a prediction image based on a single reference image is described here, but the same applies when generating a prediction image based on multiple reference images. The prediction image is generated after performing brightness correction processing on the reference images obtained from each reference image in the same way.
[0228] One method for determining whether to use LIC processing is to use a lic_flag as a signal indicating whether LIC processing is used. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a region where a brightness change has occurred. If it does, the lic_flag is set to 1, and LIC processing is used for encoding. If it does not belong to a region where a brightness change has occurred, the lic_flag is set to 0, and LIC processing is not used for encoding. On the other hand, in a decoding device, decoding is performed by decoding the lic_flag recorded in the stream and switching between using and not using LIC processing based on its value.
[0229] Other methods for determining whether to use LIC processing include checking whether LIC processing was used in surrounding blocks. As a specific example, when the encoded target block is in merge mode, it is determined whether the surrounding encoded blocks selected during the export of the MV in merge mode processing have been encoded using LIC processing. Based on the result, encoding is switched between using LIC processing and other methods. Furthermore, in this example, the decoding process is exactly the same.
[0230] [Overview of the Decoding Device]
[0231] Next, an outline of a decoding apparatus capable of decoding the encoded signal (encoded bit stream) output from the encoding apparatus 100 will be described. Figure 10 This is a block diagram illustrating the functional structure of the decoding device 200 according to Embodiment 1. The decoding device 200 is a motion image / image decoding device that decodes motion images / images in block units.
[0232] like Figure 10 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a cyclic filtering unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.
[0233] The decoding device 200 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes the software program stored in the memory, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 can also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0234] The following describes the constituent elements included in the decoding device 200.
[0235] [Entropy Decoding Department]
[0236] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units.
[0237] [De-quantization Department]
[0238] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoded target block (hereinafter referred to as the current block), which is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient of the current block based on the quantization parameter corresponding to that quantization coefficient. Furthermore, the inverse quantization unit 204 outputs the inverse quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0239] [Inverse Transformation Section]
[0240] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients, which are inputs from the inverse quantization unit 204.
[0241] For example, if the information read from the encoded bitstream represents EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information representing the transform type read from the transducer.
[0242] Furthermore, for example, when the information read from the encoded bitstream is represented using NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients.
[0243] [Addition Department]
[0244] The adder 208 reconstructs the current block by adding the prediction error, which is input from the inverse transform 206, to the prediction sample, which is input from the prediction control 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the cyclic filtering 212.
[0245] [Block Memory]
[0246] Block memory 210 is a storage unit used to store blocks within the decoded target image (hereinafter referred to as the current image) that serve as a reference in intra-frame prediction. Specifically, block memory 210 stores the reconstructed blocks output from adder 208.
[0247] [Loop Filtering Section]
[0248] The cyclic filtering unit 212 applies cyclic filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.
[0249] Given that the information indicating the on / off state of the ALF is read from the encoded bitstream, and the ALF is on, one filter is selected from multiple filters based on the direction and activity of the gradient of locality, and the selected filter is applied to the reconstructed block.
[0250] [Frame Memory]
[0251] The frame memory 214 is a storage unit used to store reference images used in inter-frame prediction; it is also sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the cyclic filtering unit 212.
[0252] Intra-frame prediction unit
[0253] The intra-prediction unit 216 performs intra-prediction based on the intra-prediction pattern read from the encoded bitstream, referring to blocks within the current image stored in the block memory 210, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra-prediction unit 216 performs intra-prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, thereby generating an intra-prediction signal, and outputs the intra-prediction signal to the prediction control unit 220.
[0254] In addition, if the intra-prediction mode of the reference luma block is selected in the intra-prediction of the chromatic difference block, the intra-prediction unit 216 can also predict the chromatic difference component of the current block based on the luma component of the current block.
[0255] Furthermore, when the information read from the encoded bitstream represents PDPC, the intra-prediction unit 216 corrects the pixel values after intra-prediction based on the gradient of the reference pixel in the horizontal / vertical direction.
[0256] [Inter-frame prediction department]
[0257] The inter-frame prediction unit 218 refers to a reference image stored in the frame memory 214 and predicts the current block. Prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 218 uses motion information (e.g., motion vectors) read from the coded bitstream to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0258] Furthermore, when the information read from the encoded bitstream is represented in OBMC mode, the inter-frame prediction unit 218 uses not only the motion information of the current block obtained through motion search, but also the motion information of neighboring blocks to generate the inter-frame prediction signal.
[0259] Furthermore, when the information read from the encoded bitstream is in FRUC mode, the inter-frame prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) read from the encoded stream, thereby deriving motion information. The inter-frame prediction unit 218 then uses the derived motion information for motion compensation.
[0260] Furthermore, when using BIO mode, the inter-frame prediction unit 218 derives motion vectors based on a model assuming constant-velocity linear motion. Additionally, when the information representation read from the encoded bitstream employs affine motion compensation prediction mode, the inter-frame prediction unit 218 derives motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0261] [Forecasting and Control Department]
[0262] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208.
[0263] [Internal structure of the inter-frame prediction unit of the coding device]
[0264] Next, the internal structure of the inter-frame prediction unit 126 of the encoding apparatus 100 will be described. Specifically, the functional structure of the inter-frame prediction unit 126 of the encoding apparatus 100 for implementing the motion search mode (FRUC mode) on the decoding device side will be described.
[0265] Figure 11 This is a block diagram showing the internal structure of the inter-frame prediction unit 126 of the encoding apparatus 100 in Embodiment 1. The inter-frame prediction unit 126 includes a candidate derivation unit 1261, a range determination unit 1262, a motion search unit 1263, and a motion compensation unit 1264.
[0266] The candidate derivation unit 1261 derives multiple candidates, each having at least one motion vector. These candidates may be referred to as predicted motion vector candidates. Additionally, the motion vectors included in the candidates may be referred to as predicted motion vectors.
[0267] Specifically, the candidate derivation unit 1261 derives multiple candidates based on the motion vectors of coded blocks (hereinafter referred to as neighboring blocks) that are spatially or temporally adjacent to the current block. The motion vectors of neighboring blocks refer to the motion vectors used in the motion compensation of neighboring blocks.
[0268] For example, when referring to two reference images in inter-frame prediction of an adjacent block, the candidate derivation unit 1261 derives a candidate containing two reference image indices and two motion vectors based on two motion vectors corresponding to the two reference images. Furthermore, for example, when referring to one reference image in inter-frame prediction of an adjacent block, the candidate derivation unit 1261 derives a candidate containing one reference image index and one motion vector based on one motion vector corresponding to that one reference image.
[0269] Multiple candidates derived from multiple adjacent blocks are registered in a candidate list. Duplicate candidates can then be removed from the candidate list. Furthermore, if there are empty slots in the candidate list, candidates with fixed-value motion vectors (e.g., zero motion vectors) can be registered. Additionally, this candidate list can be shared with the merge list used in merge mode.
[0270] Spatially adjacent blocks are those contained within the current image, meaning blocks adjacent to the current block. Examples of spatially adjacent blocks include those to the left, top-left, top, or top-right of the current block. Motion vectors derived from spatially adjacent blocks are sometimes referred to as spatial motion vectors.
[0271] Temporally adjacent blocks refer to blocks contained within encoded / decoded images that are different from the current image. The position of a temporally adjacent block within its encoded / decoded image corresponds to its position within the current image. Temporally adjacent blocks can be referred to as co-located blocks. Furthermore, motion vectors derived from temporally adjacent blocks can be called temporal motion vectors.
[0272] The range determination unit 1262 determines the motion search range of the reference image. The motion search range refers to a portion of the reference image within which motion search is permitted.
[0273] The size of the motion search range is determined, for example, based on factors such as memory bandwidth and processing power. Memory bandwidth and processing power can be obtained from levels defined, for example, by standardized specifications. Alternatively, memory bandwidth and processing power can be obtained from the decoding device. The size of the motion search range refers to the size of a portion of the image, and can be represented, for example, by the number of horizontal and vertical pixels representing the distance from the center of the motion search range to its vertical and horizontal edges.
[0274] The location of the motion search range is determined, for example, based on a statistical representative vector of multiple motion vectors from multiple candidates included in the candidate list. In this embodiment, the average motion vector is used as the statistical representative vector. The average motion vector is a motion vector composed of the average of the horizontal values and the average of the vertical values of multiple motion vectors.
[0275] Information related to the determined motion search range (hereinafter referred to as motion search range information) is encoded in the bitstream. The motion search range information includes at least one of information indicating the size of the motion search range and information indicating the position of the motion search range; in this embodiment, it includes only information indicating the size of the motion search range. The position of the motion search range information within the bitstream is not particularly limited. For example, such as... Figure 12As shown, motion search range information can be written into (i) the Video Parameter Set (VPS), (ii) the Sequence Parameter Set (SPS), (iii) the Picture Parameter Set (PPS), (iv) the slice header, or (v) the video system settings. Furthermore, the motion search range information can be entropy-encoded or not.
[0276] The motion search unit 1263 performs a motion search within the motion search range of the reference image. That is, the motion search unit 1263 performs a motion search within the motion search range of the reference image. Specifically, the motion search unit 1263 performs the motion search as follows.
[0277] First, the motion search unit 1263 reads the reconstructed image of the motion search range within the reference image from the frame memory 122. For example, the motion search unit 1263 only reads the reconstructed image of the motion search range in the reference image. Then, the motion search unit 1263 excludes candidates with motion vectors corresponding to positions outside the motion search range of the reference image from the multiple candidates derived by the candidate derivation unit 1261. That is, the motion search unit 1263 deletes candidates with motion vectors indicating positions outside the motion search range from the candidate list.
[0278] Next, the motion search unit 1263 selects candidates from the remaining candidates. That is, the motion search unit 1263 selects candidates from the candidate list that has removed candidates with motion vectors corresponding to positions outside the motion search range.
[0279] The candidate is selected based on its evaluation value. For example, when applying the first pattern matching (bidirectional matching) described above, the evaluation value of each candidate is calculated based on the difference between the reconstructed image of the region in the reference image corresponding to the motion vector of the candidate and the reconstructed image of the region in another reference image along the motion trajectory of the current block. Furthermore, for example, when applying the second pattern matching (template matching), the evaluation value of each candidate is calculated based on the difference between the reconstructed image of the region in the reference image corresponding to the motion vector of each candidate and the reconstructed image of the coded block adjacent to the current block in the current image.
[0280] Finally, the motion search unit 1263 determines the motion vector for the current block based on the selected candidates. Specifically, the motion search unit 1263 performs pattern matching, for example, in the surrounding region of the position within the reference image corresponding to the motion vector included in the selected candidates, and finds the region that best matches within the surrounding region. Then, the motion search unit 1263 determines the motion vector for the current block based on the region that best matches within the surrounding region. Alternatively, for example, the motion search unit 1263 may determine that the motion vector included in the selected candidates is the motion vector for the current block.
[0281] The motion compensation unit 1264 generates the inter-frame prediction signal for the current block by performing motion compensation using the motion vector determined by the motion search unit 1263.
[0282] [Operation of the inter-frame prediction unit of the coding device]
[0283] Next, refer to Figures 13-17 The operation of the inter-frame prediction unit 126 configured as described above will be explained in detail below. The case of performing inter-frame prediction with reference to a single reference image will be explained in the following text.
[0284] Figure 13 This is a flowchart illustrating the processing of the inter-frame prediction unit of the encoding / decoding apparatus in Embodiment 1. Figure 13 In the brackets, the labels indicate the processing of the inter-frame prediction unit of the decoding device.
[0285] First, the candidate derivation unit 1261 derives multiple candidates from adjacent blocks to generate a candidate list (S101). Figure 14 This is a diagram illustrating an example of the candidate list in Implementation Method 1. Here, each candidate has a candidate index, a reference image index, and a motion vector.
[0286] Next, the scope determination unit 1262 selects a reference image from the list of reference images (S102). For example, the scope determination unit 1262 selects reference images in ascending order of their indexes. For example, in Figure 15 In the list of reference images, the range determination unit 1262 initially selects the reference image with index "0".
[0287] The range determination unit 1262 determines the motion search range in the reference image (S103). Here, refer to... Figure 16 The determination of the motion search range is explained.
[0288] Figure 16 This is a diagram illustrating an example of the motion search range 1022 in Implementation Method 1. Figure 16 In the reference image, the corresponding position represents the current block 1000 and adjacent blocks 1001 to 1004 in the current image.
[0289] First, the range determination unit 1262 obtains the motion vectors 1011 to 1014 of multiple adjacent blocks 1001 to 1004 from the candidate list. Then, the range determination unit 1262 scales the motion vectors 1011 to 1014 as needed and calculates the average motion vector 1020 of the motion vectors 1011 to 1014.
[0290] For example, the scope determination section 1262 refers to Figure 14From the candidate list, calculate the average horizontal value of multiple motion vectors "-25 (=((-48)+(-32)+0+(-20)) / 4)" and the average vertical value "6 (=(0+9+12+3) / 4)", thereby calculating the average motion vector (-26, 6).
[0291] Subsequently, the range determination unit 1262 determines the representative position 1021 of the motion search range based on the average motion vector 1020. Here, the center position is used as the representative position 1021. In addition, the representative position 1021 is not limited to the center position, and any position among the vertex positions of the motion search range (e.g., the top left vertex position) can also be used.
[0292] Furthermore, the range determiner 1262 determines the size of the motion search range based on factors such as memory bandwidth and processing power. For example, the range determination unit 1262 determines the number of horizontal pixels and the number of vertical pixels that represent the size of the motion search range.
[0293] Based on the representative position 1021 and size of the motion search range determined in this way, the range determination unit 1262 determines the motion search range 1022.
[0294] Here, return Figure 13 The flowchart is explained. The motion search unit 1263 excludes candidates from the candidate list that have motion vectors corresponding to positions outside the motion search range (S104). For example, in... Figure 16 In the process, the motion search unit 1263 excludes candidates with motion vectors 1012 and 1013 that indicate positions outside the motion search range from the candidate list.
[0295] The motion search unit 1263 calculates the evaluation values of the remaining candidates in the candidate list (S105). For example, the motion search unit 1263 calculates the difference between the reconstructed image (template) of the neighboring block in the current image and the reconstructed image of the region in the reference image corresponding to the candidate motion vector as the evaluation value (template matching). In this case, the region in the reference image corresponding to the candidate motion vector is the region of the neighboring block in the reference image that has undergone motion compensation using the candidate motion vector. In the evaluation value calculated in this way, the lower the value, the higher the evaluation. Alternatively, the evaluation value can be the reciprocal of the difference value. In this case, the higher the evaluation value, the higher the evaluation.
[0296] The motion search unit 1263 selects a candidate from the candidate list based on the evaluation value (S106). For example, the motion search unit 1263 selects the candidate with the lowest evaluation value.
[0297] The motion search unit 1263 determines the surrounding region of the region corresponding to the selected candidate motion vector (S107). For example, when selecting... Figure 16In the case of a motion vector of 1014, such as Figure 17 As shown, the motion search unit 1263 determines the surrounding area 1023 of the current block's region that has undergone motion compensation using motion vector 1014 in the reference image.
[0298] For example, the size of the peripheral region 1023 can be predefined in the standard specification. Specifically, the size of the peripheral region 1023 can be a fixed size such as 8×8 pixels, 16×16 pixels, or 32×32 pixels, which can be predefined. Alternatively, the size of the peripheral region 1023 can be determined based on processing capabilities. In this case, information related to the size of the peripheral region 1023 can be written into the bitstream. Alternatively, the size of the peripheral region 1023 can be taken into account to determine the number of horizontal and vertical pixels representing the motion search range, and then written into the bitstream.
[0299] The motion search unit 1263 determines whether the determined surrounding area is included in the motion search range (S108). That is, the motion search unit 1263 determines whether the entire surrounding area is included in the motion search range.
[0300] Here, when the surrounding area is included in the motion search range (as in S108), the motion search unit 1263 performs pattern matching within the surrounding area (S109). As a result, the motion search unit 1263 obtains the evaluation value of the region within the reference image that best matches the reconstructed image of the adjacent block within the surrounding area.
[0301] On the other hand, if the surrounding area is not included in the motion search range (no in S108), the motion search unit 1263 performs pattern matching in the portion of the surrounding area that is included in the motion search range (S110). That is, the motion search unit 1263 does not perform pattern matching in the portion of the surrounding area that is not included in the motion search range.
[0302] The scope determination unit 1262 determines whether there is an unselected reference image within the reference image (S111). Here, if there is an unselected reference image (as in S111), the selection of the reference image is returned (S102).
[0303] On the other hand, when there are no unselected reference images (No in S111), the motion search unit 1263 determines the motion vector for the current image based on the evaluation value (S112). That is, the motion search unit 1263 determines the candidate motion vector with the highest evaluation among multiple reference images as the motion vector for the current image.
[0304] [Internal structure of the inter-frame prediction unit of the decoding device]
[0305] Next, the internal structure of the inter-frame prediction unit 218 of the decoding device 200 will be described. Specifically, the functional structure of the inter-frame prediction unit 218 of the decoding device 200 for implementing the motion search mode (FRUC mode) on the decoding device side will be described.
[0306] Figure 18 This is a block diagram showing the internal structure of the inter-frame prediction unit 218 of the decoding apparatus 200 according to Embodiment 1. The inter-frame prediction unit 218 includes a candidate derivation unit 2181, a range determination unit 2182, a motion search unit 2183, and a motion compensation unit 2184.
[0307] Similar to the candidate derivation unit 1261 of the encoding device 100, the candidate derivation unit 2181 derives a plurality of candidates, each having at least one motion vector. Specifically, the candidate derivation unit 2181 derives a plurality of candidates based on the motion vectors of adjacent blocks in space and / or time.
[0308] The range determination unit 2182 determines the motion search range in the reference image. Specifically, the range determination unit 2182 first obtains motion search range information interpreted from the bitstream. Then, the range determination unit 2182 determines the size of the motion search range based on the motion search range information. Furthermore, similar to the range determination unit 1262 of the encoding device 100, the range determination unit 2182 determines the position of the motion search range. Thus, the motion search range in the reference image is determined.
[0309] The motion search unit 2183 performs a motion search within the motion search range of the reference image. Specifically, the motion search unit 2183 first reads the reconstructed image of the motion search range within the reference image from the frame memory 214. For example, the motion search unit 2183 only reads the reconstructed image of the motion search range in the reference image. Then, similarly to the motion search unit 1263 of the encoding device 100, the motion search unit 2183 performs a motion search within the motion search range and determines the motion vector for the current block.
[0310] The motion compensation unit 2184 performs motion compensation by using the motion vector determined by the motion search unit 2183, and generates the inter-frame prediction signal for the current block.
[0311] [Operation of the inter-frame prediction unit of the decoding device]
[0312] Next, refer to Figure 13 The operation of the inter-frame prediction unit 218 configured as described above will be explained. The processing of the inter-frame prediction unit 218 is the same as that of the inter-frame prediction unit 126 of the coding apparatus 100, except that step S103 replaces step S203. Step S203 will be explained below.
[0313] The range determination unit 2182 determines the motion search range in the reference image (S203). At this time, the range determination unit 2182 determines the size of the motion search range based on the motion search range information interpreted from the bitstream. In addition, similar to the range determination unit 1262 of the encoding device 100, the range determination unit 2182 determines the position of the motion search range based on multiple candidates included in the candidate list.
[0314] [Effects, etc.]
[0315] As described above, according to the inter-frame prediction unit 126 of the encoding apparatus 100 and the inter-frame prediction unit 218 of the decoding apparatus 200 of this embodiment, candidate selection can be performed after excluding candidates with motion vectors corresponding to positions outside the motion search range. Therefore, the processing load for candidate selection can be reduced. Furthermore, reconstructed images outside the motion search range can be avoided from being read from the frame memory, thus reducing the memory bandwidth used for motion search.
[0316] Furthermore, according to the encoding apparatus 100 and decoding apparatus 200 of this embodiment, information related to the motion search range can be written into a bitstream, and information related to the motion search range can be interpreted from the bitstream. Therefore, the same motion search range used in the encoding apparatus 100 can also be used in the decoding apparatus 200. Furthermore, the processing load for determining the motion search range in the decoding apparatus 200 can be reduced.
[0317] Furthermore, according to the encoding apparatus 100 and decoding apparatus 200 of this embodiment, information representing the size of the motion search range can be included in the bitstream. Therefore, a motion search range having the same size as the motion search range used in the encoding apparatus 100 can also be used in the decoding apparatus 200. This further reduces the processing load in the decoding apparatus 200 for determining the size of the motion search range.
[0318] Furthermore, according to the inter-frame prediction unit 126 of the encoding apparatus 100 and the inter-frame prediction unit 218 of the decoding apparatus 200 of this embodiment, the position of the motion search range can be determined based on the average motion vectors obtained from multiple candidate blocks derived from multiple blocks adjacent to the current block. Therefore, the region suitable for searching motion vectors for the current block can be determined as the motion search range, thereby improving the accuracy of the motion vectors.
[0319] Furthermore, according to the inter-frame prediction unit 126 of the encoding apparatus 100 and the inter-frame prediction unit 218 of the decoding apparatus 200 of this embodiment, in addition to candidate motion vectors, the motion vector for the current block can also be determined based on pattern matching in the surrounding area. Therefore, the accuracy of the motion vector can be further improved.
[0320] Furthermore, according to the inter-frame prediction unit 126 of the encoding apparatus 100 and the inter-frame prediction unit 218 of the decoding apparatus 200 of this embodiment, pattern matching can be performed in a portion of the motion search range in the surrounding area when the surrounding area is not included in the motion search range. Therefore, motion search outside the motion search range can be avoided, reducing processing load and memory bandwidth requirements.
[0321] (Modification 1 of Implementation Method 1)
[0322] In the above embodiment 1, the position of the motion search range is determined based on the average motion vector of multiple motion vectors among multiple candidates included in the candidate list. However, in this modified example, it is determined based on the central motion vector of multiple motion vectors among multiple candidates included in the candidate list.
[0323] According to the range determination units 1262 and 2182 of this modified example, multiple motion vectors included in the candidate list are obtained. Then, the range determination units 1262 and 2182 calculate the central motion vector of the multiple obtained motion vectors. The central motion vector is a motion vector composed of the median of the horizontal values and the median of the vertical values of the multiple motion vectors.
[0324] Scope determination section 1262, 2182 (see example) Figure 14 The candidate list is calculated by calculating the median of the horizontal values of multiple motion vectors, "-26 (=((-32)+(-20)) / 2)" and the median of the vertical values, "6 (=(9+3) / 2)", and then calculating the central motion vector (-26, 6).
[0325] Next, the range determination units 1262 and 2182 determine the representative position of the motion search range based on the calculated central motion vector.
[0326] As described above, according to the range determination units 1262 and 2182 of this modified example, the position of the motion search range can be determined based on the central motion vectors derived from multiple candidate central motion vectors of multiple blocks adjacent to the current block. Therefore, the region suitable for searching the motion vectors of the current block can be determined as the motion search range, thereby improving the accuracy of the motion vectors.
[0327] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0328] (Modification 2 of Implementation Method 1)
[0329] Next, a variation of Embodiment 1, Example 2, will be described. In this variation, instead of the average motion vector, the position of the motion search range is determined based on the minimum motion vector. Hereinafter, this variation will be described focusing on the differences from Embodiment 1 described above.
[0330] According to the range determination units 1262 and 2182 of this variant example, referencing the candidate list, multiple motion vectors included in the multiple candidates are obtained. Then, the range determination units 1262 and 2182 select the motion vector with the smallest size (i.e., the smallest motion vector) from the multiple obtained motion vectors.
[0331] Scope determination section 1262, 2182 (see example) Figure 14 From the candidate list, select the motion vector (0, 8) with the smallest size among the candidates with candidate index "2".
[0332] Next, the range determination units 1262 and 2182 determine the representative position of the motion search range based on the selected minimum motion vector.
[0333] Figure 19 This is a diagram illustrating an example of the motion search range in a variation of Embodiment 1, Example 2. Figure 19 In this process, the range determination units 1262 and 2182 select the motion vector 1013 with the smallest magnitude among the motion vectors 1011 to 1014 of the adjacent blocks as the minimum motion vector 1030. Next, the range determination units 1262 and 2182 determine the representative position 1031 of the motion search range based on the minimum motion vector 1030. Based on the determined representative position 1031, the range determination units 1262 and 2182 determine the motion search range 1032.
[0334] As described above, according to the range determination units 1262 and 2182 of this modified example, the position of the motion search range can be determined based on the minimum motion vector obtained from multiple candidate blocks derived from multiple blocks adjacent to the current block. Therefore, the region close to the current block can be determined as the motion search range, and the accuracy of the motion vector can be improved.
[0335] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0336] (Modification 3 of Implementation Method 1)
[0337] Next, a variation of Embodiment 1, Example 3, will be described. In this variation, instead of the average motion vector, the position of the motion search range is determined based on the motion vector of an encoded / decoded image that is different from the current image. Hereinafter, this variation will be described focusing on the differences from Embodiment 1 described above.
[0338] Regarding the range determination units 1262 and 2182 of this variant example, they refer to the list of reference images and select a reference image that is an encoded / decoded image different from the current image. For example, the range determination units 1262 and 2182 select the reference image with the minimum reference image index. Alternatively, the range determination units 1262 and 2182 may also select the reference image that is closest to the current image in the output order.
[0339] Next, the range determination units 1262 and 2182 acquire multiple motion vectors used in the encoding / decoding of multiple blocks contained in the selected reference image. Then, the range determination units 1262 and 2182 calculate the average motion vector of the acquired multiple motion vectors.
[0340] Then, the range determination units 1262 and 2182 determine the representative position of the motion search range based on the calculated average motion vector.
[0341] As described above, according to the range determination units 1262 and 2182 of this variant, even if the current block within the current image changes, the motion vector of the encoded / decoded image does not change. Therefore, it is not necessary to determine the motion search range based on the motion vectors of adjacent blocks each time the current block changes. That is, the processing load for determining the motion search range can be reduced.
[0342] Furthermore, here, the representative position of the motion search range is determined based on the average motion vector of the selected reference image, but is not limited to this. For example, the central motion vector can be used instead of the average motion vector. Additionally, for example, the motion vector of the co-located block can be used instead of the average motion vector.
[0343] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0344] (Modification 4 of Implementation Method 1)
[0345] Next, a variation of embodiment 1, number 4, will be described. In this variation, the reference image is divided into multiple regions, and based on the divided regions, multiple motion vectors contained in multiple candidates are grouped. Then, the position of the motion search range is determined based on the group containing the most motion vectors.
[0346] The following will focus on the differences from Embodiment 1 described above. Figure 20 This variation will be explained. Figure 20 This is a diagram illustrating an example of the motion search range in Variation 4 of Implementation Method 1.
[0347] The scope determination parts 1262 and 2182 of this variation divide the reference image into regions. For example, as... Figure 20 As shown, the range determination units 1262 and 2182 divide the reference image into four regions (region 1 to region 4) based on the position of the current image.
[0348] The range determination units 1262 and 2182 group multiple motion vectors of adjacent blocks based on multiple regions. For example, in Figure 20 In the process, the range determination units 1262 and 2182 divide the multiple motion vectors 1011 to 1014 into a first group containing the motion vector 1013 corresponding to the first region and a second group containing the motion vectors 1011, 1012 and 1014 corresponding to the second region.
[0349] The range determination units 1262 and 2182 determine the location of the motion search range based on the group containing the most motion vectors. For example, in Figure 20 In this process, the range determination units 1262 and 2182 determine the representative position 1041 of the motion search range based on the average motion vector 1040 of the motion vectors 1011, 1012, and 1014 included in the second group. Alternatively, the central motion vector or the minimum motion vector may be used instead of the average motion vector.
[0350] As described above, according to the range determination units 1262 and 2182 of this modified example, the area suitable for searching the motion vector of the current block can be determined as the motion search range, thereby improving the accuracy of the motion vector.
[0351] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0352] (Modification 5 of Implementation Method 1)
[0353] Next, a variation 5 of Embodiment 1 will be described. In this variation, the difference from Embodiment 1 is the correction of the position of the motion search range. Hereinafter, focusing on the differences from Embodiment 1, refer to... Figure 21 This variation will be explained. Figure 21This is a diagram illustrating an example of the motion search range in Variation 5 of Implementation Method 1.
[0354] Regarding the range determination units 1262 and 2182 of this modification, for example, they correct the position of the motion search range determined based on the average motion vector. Specifically, the range determination units 1262 and 2182 first temporarily determine the motion search range based on the average motion vector of the multiple motion vectors included in the multiple candidates. For example, as Figure 21 As shown, the range determination units 1262 and 2182 temporarily determine the motion search range of 1050.
[0355] Here, the range determination units 1262 and 2182 determine whether the position corresponding to the zero motion vector is included in the temporarily determined motion search range. That is, the range determination units 1262 and 2182 determine whether the reference position of the current block in the reference image (e.g., the upper left corner) is included in the temporarily determined motion search range 1050. For example, in Figure 21 In the process, the range determination units 1262 and 2182 determine whether the temporarily determined motion search range 1050 includes the position 1051 corresponding to the zero motion vector.
[0356] Here, if the position corresponding to the zero motion vector is not included in the temporarily determined motion search range, the range determination units 1262 and 2182 modify the position of the temporarily determined motion search range so that the motion search range includes the position corresponding to the zero motion vector. For example, in Figure 21 Since the temporarily determined motion search range 1050 does not include the position 1051 corresponding to the zero motion vector, the range determination units 1262 and 2182 modify the motion search range 1050 to a motion search range 1052. As a result, the modified motion search range 1052 includes the position 1051 corresponding to the zero motion vector.
[0357] On the other hand, if the position corresponding to the zero motion vector is included in the temporarily determined motion search range, the range determination units 1262 and 2182 directly determine the temporarily determined motion search range as the motion search range. That is, the range determination units 1262 and 2182 do not modify the position of the motion search range.
[0358] As described above, according to the range determination units 1262 and 2182 of this modified example, the area suitable for searching the motion vector of the current block can be determined as the motion search range, thereby improving the accuracy of the motion vector.
[0359] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0360] (Modification 6 of Implementation Method 1)
[0361] Next, a variation of embodiment 1, number 6, will be described. In variation 5, the position of the motion search range is corrected in a manner that includes the position corresponding to the zero motion vector. However, in this variation, the position of the motion search range is corrected in a manner that includes the position corresponding to the motion vector of one of the multiple adjacent blocks.
[0362] The following is for reference Figure 22 This variation will be explained. Figure 22 This is a diagram illustrating an example of the motion search range in Variation 6 of Implementation Method 1.
[0363] First, the range determination units 1262 and 2182, like in Modification 5, temporarily determine the motion search range based on, for example, the average motion vector. For example, as... Figure 22 As shown, the range determination units 1262 and 2182 temporarily determine the motion search range of 1050.
[0364] Here, the range determination units 1262 and 2182 determine whether the position corresponding to the motion vector of one of the multiple adjacent blocks is included in the temporarily determined motion search range. For example, in Figure 22 In the process, the range determination units 1262 and 2182 determine whether the temporarily determined motion search range 1050 includes the position 1053 corresponding to the motion vector 1011 of the adjacent block 1001. As one of a plurality of adjacent blocks, a predetermined adjacent block can be used, for example, the left adjacent block or the upper adjacent block can be used.
[0365] Here, if the position corresponding to the motion vector of one of the multiple adjacent blocks is not included in the temporarily determined motion search range, the range determination units 1262 and 2182 correct the position of the temporarily determined motion search range so that the motion search range includes the position corresponding to the motion vector. For example, in Figure 22 In the process, the temporarily determined motion search range 1050 does not include the position 1053 corresponding to the motion vector 1011 of the adjacent block 1001, so the range determination units 1262 and 2182 modify the motion search range 1050 to a motion search range 1054. As a result, the position 1053 is included in the modified motion search range 1054.
[0366] On the other hand, if the position corresponding to the motion vector of one of the multiple adjacent blocks is included in the temporarily determined motion search range, the range determination units 1262 and 2182 directly determine the temporarily determined motion search range as the motion search range. That is, the range determination units 1262 and 2182 do not modify the position of the motion search range.
[0367] As described above, according to the range determination units 1262 and 2182 of this modified example, the area suitable for searching the motion vector of the current block can be determined as the motion search range, thereby improving the accuracy of the motion vector.
[0368] This embodiment can be implemented in combination with at least a portion of other technical solutions in this invention. Furthermore, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc., can also be implemented in combination with other embodiments.
[0369] (Modification 7 of Implementation Method 1)
[0370] Next, a variation 7 of Embodiment 1 will be described. In this variation, the difference from Embodiment 1 is that the bitstream does not contain information related to the motion search range. Hereinafter, refer to... Figure 23 This modified example will be described focusing on the differences from Embodiment 1 described above.
[0371] Figure 23 This is a block diagram illustrating the functional structure of the encoding / decoding system 300 in variation 7 of implementation method 1. For example... Figure 23 As shown, the encoding / decoding system 300 includes an encoding system 310 and a decoding system 320.
[0372] The encoding system 310 encodes the input motion image and outputs a bit stream. The encoding system 310 includes a communication device 311, an encoding device 312, and an output buffer 313.
[0373] The communication device 311 exchanges capability information with the decoding system 320 via a communication network (not shown) or the like, and generates motion search range information based on this capability information. Specifically, the communication device 311 sends encoded capability information to the decoding system 320 and receives decoded capability information from the decoding system 320. The encoded capability information includes information such as the processing capability and memory bandwidth of the motion search in the encoding system 310. The decoded capability information includes information such as the processing capability and memory bandwidth used for the motion search in the decoding system 320.
[0374] Encoding device 312 encodes the input motion image and outputs the bit stream to output buffer 313. At this time, encoding device 312 performs essentially the same processing as encoding device 100 in embodiment 1, except that it determines the size of the motion search range based on motion search range information obtained from communication device 311.
[0375] The output buffer 313 is a so-called buffer memory that temporarily stores the bit stream input from the encoding device 312 and outputs the stored bit stream to the decoding system 320 via a communication network or the like.
[0376] The decoding system 320 decodes the bitstream input from the encoding system 310 and outputs the motion image to a display (not shown). The decoding system 320 includes a communication device 321, a decoding device 322, and an input buffer 323.
[0377] Similar to the communication device 311 of the encoding system 310, the communication device 321 exchanges capability information with the encoding system 310 via a communication network or the like, and generates motion search range information based on this capability information. Specifically, the communication device 311 sends decoding capability information to the encoding system 310 and receives encoding capability information from the encoding system 310.
[0378] The decoding device 322 decodes the bitstream input from the input buffer 323 and outputs the motion image to a display or the like. At this time, the decoding device 322 performs essentially the same processing as the decoding device 200 in Embodiment 1, except that it determines the motion search range based on motion search range information obtained from the communication device 321. Furthermore, if the motion search range determined based on the motion search range information obtained from the communication device 321 exceeds the motion search range that the decoding device 322 can handle, a message indicating that decoding is not possible can be sent to the communication device 321.
[0379] The input buffer 323 is a so-called buffer memory that temporarily stores the bit stream input from the encoding system 310 and outputs the stored bit stream to the decoding device 322.
[0380] As described above, according to the encoding / decoding system 300 of this variant, even if the bitstream does not contain information related to the motion search range, motion searching can be performed using the same motion search range in both the encoding device 312 and the decoding device 322. Therefore, the amount of code used for the motion search range can be reduced. Furthermore, since the processing in the range determination unit 1262 for determining the number of horizontal and vertical pixels representing the size of the motion search range is not required, the processing workload can be reduced.
[0381] (Variation 8 of Implementation Method 1)
[0382] Furthermore, in Embodiment 1 described above, all of the multiple reference images included in the reference image list are selected sequentially, but it is not necessary to select all reference images. In this variation, an example where the number of selected reference images is limited will be described.
[0383] Regarding the range determination unit 1262 of the encoding device 100 in this variant example, similarly to the size of the motion search range, it determines the number of reference images allowed in the motion search in the FRUC mode (hereinafter referred to as the reference image allowance number) based on factors such as memory bandwidth and processing power. Information related to the determined reference image allowance number (hereinafter referred to as the reference image allowance number information) is written into the bitstream.
[0384] Furthermore, the range determination unit 2182 of the decoding device 200 in this modified example determines the number of allowed reference images based on the number of allowed reference images interpreted from the bitstream.
[0385] Furthermore, there are no particular restrictions on the position within the bitstream where the reference image's allowed number of information can be written. For example, with... Figure 12 Similarly, the motion search range information shown can be written into VPS, SPS, PPS, slice header, or video system settings parameters, as can the number of images allowed.
[0386] Based on the number of reference images allowed determined in this way, the number of reference images used in the FRUC mode is limited. Specifically, the range determination units 1262 and 2182, for example, in... Figure 13 In step S111, it is determined whether there are any unselected reference images and whether the number of selected reference images is less than the allowed number of reference images. Here, if there are no unselected reference images, or if the number of selected reference images exceeds the allowed number of reference images (as in S111), the process proceeds to step S112. Therefore, selecting reference images from the reference image list in quantities exceeding the allowed number of reference images is prohibited.
[0387] In this case, Figure 13 In step S111, the range determination units 1262 and 2182 can select reference images in ascending order of reference image index values or in order of temporal proximity to the current image. In this case, reference images with smaller reference image index values or those temporally close to the current image are preferentially selected from the reference image list. Furthermore, the temporal distance between the current image and the reference images can be determined based on the POC (Picture Order Count).
[0388] As described above, according to the scope determination parts 1262 and 2182 of this modified example, the number of reference images used in motion search can be limited to below the allowed number of reference images. Therefore, the processing load for motion search can be reduced.
[0389] Additionally, for example, in the case of performing time-scalable encoding / decoding, the range determination units 1262 and 2182 may also limit the number of reference images contained in a lower level than the level of the current image represented by the time identifier, based on the number of allowed reference images.
[0390] (Variation 9 of Implementation Method 1)
[0391] Next, a variation of embodiment 1, number 9, will be described. In this variation, a method for determining the size of the motion search range when referring to multiple reference images in inter-frame prediction will be explained.
[0392] When multiple reference images are referenced in inter-frame prediction, the size of the motion search range can depend on the number of reference images referenced in the inter-frame prediction, in addition to memory bandwidth and processing power. Specifically, the range determination units 1262 and 2182 first determine the total size of the multiple motion search ranges in the multiple reference images referenced in the inter-frame prediction based on memory bandwidth and processing power. Then, the range determination units 1262 and 2182 determine the size of the motion search range of each reference image based on the number of multiple reference images and the determined total size. That is, the range determination unit 1262 determines the size of the motion search range in each reference image such that the sum of the sizes of the multiple motion search ranges in the multiple reference images is consistent with the total size of the multiple motion search ranges determined based on memory bandwidth and processing power.
[0393] Reference Figure 24 The specific explanation outlines the motion search range for each reference image used in this decision. Figure 24 This is a diagram showing the motion search range in variation 9 of implementation method 1. Figure 24 (a) represents an example of the motion search range in a prediction (double prediction) that references two reference images. Figure 24 (b) represents an example of the motion search range in the prediction based on four reference images.
[0394] exist Figure 24 In (a), motion search ranges F20 and B20 are determined for the front reference image 0 and the rear reference image 0, respectively. Pattern matching (template matching or bidirectional matching) is performed within these motion search ranges F20 and B20.
[0395] exist Figure 24In (b), motion search ranges F40, F41, B40, and B41 are determined for the front reference image 0, the front reference image 1, the rear reference image 0, and the rear reference image 1, respectively. Therefore, pattern matching is performed within these motion search ranges F40, F41, B40, and B41.
[0396] Here, the combined dimensions of motion search ranges F20 and B20 are approximately the same as the combined dimensions of motion search ranges F40, F41, B40, and B41. That is, the size of the motion search range in each reference image is determined based on the number of reference images used in inter-frame prediction.
[0397] As described above, according to the range determination units 1262 and 2182 of this modified example, the size of the motion search range in each reference image can be determined based on the number of reference images referenced in inter-frame prediction. Therefore, the total size of the region where motion search is performed can be controlled, and the processing load and memory bandwidth requirements can be reduced more effectively.
[0398] (Other variations of Implementation Method 1)
[0399] The encoding and decoding apparatus of one or more technical solutions of the present invention have been described above based on embodiments and modifications, but the present invention is not limited to these embodiments and modifications. As long as they do not depart from the spirit of the present invention, various modifications conceived by those skilled in the art can be implemented in this embodiment or modification, and technical solutions constructed by combining the constituent elements of different modifications can also be included within the scope of one or more technical solutions of the present invention.
[0400] For example, in the above embodiments and variations, motion search in the FRUC mode is performed in variable-sized block units called coding units (CUs), prediction units (PUs), or transform units (TUs), but is not limited thereto. Motion search in the FRUC mode can also be performed in sub-block units obtained by further dividing the variable-sized blocks. In this case, the vectors used to determine the location of the motion search range (e.g., average vector, center vector, etc.) can be performed in picture units, block units, or sub-block units.
[0401] Furthermore, in the above embodiments and their variations, the size of the motion search range is determined based on factors such as processing power and memory bandwidth, but is not limited thereto. For example, the size of the motion search range can be determined based on the type of reference image. For instance, the range determination unit 1262 can determine the size of the motion search range as a first size when the reference image is a B image, and determine the size of the motion search range as a second size larger than the first size when the reference image is a P image.
[0402] Additionally, for example, in the above embodiments and variations, if the position corresponding to the motion vector included in the candidate is not included in the motion search range, the candidate is excluded from the candidate list, but this is not a limitation. For example, if part or all of the surrounding area of the position corresponding to the motion vector included in the candidate is not included in the motion search range, the candidate may be excluded from the candidate list.
[0403] Furthermore, for example, in the above embodiments and variations, pattern matching is performed in the surrounding region of the position corresponding to the motion vector included in the selected candidate, but this is not a limitation. For example, pattern matching of the surrounding region may not be performed. In this case, the motion vector included in the candidate can be directly determined as the motion vector for the current block.
[0404] Furthermore, for example, in the above embodiments and variations, the candidates excluded from the candidate list have motion vectors corresponding to positions outside the motion search range, but are not limited to this. For example, if the pixel used for interpolation when performing motion compensation with fractional pixel precision using the motion vectors included in the candidate is not included in the motion search range, the candidate may be excluded from the candidate list. That is, it may also be determined whether to exclude a candidate based on the position of the pixel used for fractional pixel interpolation. In addition, for example, when applying BIO or OBMC, candidates utilizing pixels outside the motion search range in BIO or OBMC may be excluded from the candidate list. Furthermore, for example, the candidate with the smallest reference image index among multiple candidates may be retained, while other candidates may be excluded.
[0405] Furthermore, while the application of a mode for defining the motion search range in a reference image has been consistently described in the above embodiments and variations, it is not limited to this. For example, the mode can be selected to be applied / not applied at the video, sequence, image, slice, or block level. In this case, tag information indicating whether the mode is applied can be included in the bitstream. The position of the tag information in the bitstream does not need to be specifically defined. For example, the tag information can also be included in... Figure 12 The location shown is the same as the location of the motion search range information.
[0406] Furthermore, for example, the scaling of motion vectors was not described in detail in the above embodiments and variations. However, for example, the motion vectors of each candidate can be scaled based on a reference image used as a reference. Specifically, a reference image with a different index than the reference image index of the encoding and decoding results can be used as a reference to scale the motion vectors of each candidate. For example, a reference image with reference image index "0" can be used as the reference image. Alternatively, for example, the reference image that is closest to the current image in the output order can also be used as the reference image.
[0407] Furthermore, unlike the inter-frame prediction in the above-described embodiments and variations, when searching for blocks identical to the current block by referring to the area above or to the left of the current block in the current image (e.g., the case of inner block copying), the motion search range can be limited in the same way as in the above-described embodiments and variations.
[0408] Furthermore, in the above embodiments and their variations, information defining the correspondence between the feature quantity or type of the current block or current image and the sizes of multiple motion search ranges can be predetermined. Referring to this information, the size of the motion search range corresponding to the feature quantity or type of the current block or current image is determined. For example, size (number of pixels) can be used as a feature quantity, and a prediction mode (e.g., single prediction, double prediction, etc.) can be used as a type.
[0409] (Implementation Method 2)
[0410] In the above embodiments, each functional block is typically implemented using an MPU and memory. Furthermore, the processing of each functional block is usually achieved by a program execution unit such as a processor reading and executing software (programs) recorded in a recording medium such as ROM. This software can be distributed via download or by recording it in a recording medium such as semiconductor memory. Alternatively, each functional block can also be implemented in hardware (dedicated circuitry).
[0411] Furthermore, the processing described in each embodiment can be implemented either centrally using a single device (system) or distributedly using multiple devices. Additionally, the processor executing the above-described program can be either single or multiple. That is, it can be either centrally processed or distributed.
[0412] The present invention is not limited to the above embodiments, and various modifications can be made, which are also included within the scope of the present invention.
[0413] Furthermore, examples of applications of the motion picture encoding method (image encoding method) or motion picture decoding method (image decoding method) shown in the above embodiments and systems using them will be described here. The system is characterized by having an image encoding apparatus using the image encoding method, an image decoding apparatus using the image decoding method, and an image encoding / decoding apparatus possessing both. Other structures within the system can be appropriately modified as needed.
[0414] [Usage Example]
[0415] Figure 25This is a diagram showing the overall structure of the content delivery system ex100 that implements content distribution services. The communication service provision is divided into desired sizes, and each unit has base stations ex106, ex107, ex108, ex109, and ex110, which serve as fixed wireless stations.
[0416] In this content delivery system ex100, various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. This content delivery system ex100 can also combine some of the above elements for connection. Alternatively, the devices can be directly or indirectly interconnected via telephone networks or short-range wireless connections without using base stations ex106 to ex110, which are fixed wireless stations. Furthermore, a streaming media server ex103 is connected to the computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via the Internet ex101, etc. Additionally, the streaming media server ex103 is connected to terminals within a hotspot inside an aircraft ex117 via satellite ex116.
[0417] Alternatively, it can replace base stations ex106 to ex110 by using wireless access points or hotspots. Furthermore, the streaming media server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or it can connect directly to the aircraft ex117 without going through the satellite ex116.
[0418] The camera ex113 is a digital camera or similar device capable of capturing still and moving images. Additionally, the smartphone ex115 refers to a smartphone, mobile phone, or PHS (Personal Handyphone System) corresponding to mobile communication systems commonly referred to as 2G, 3G, 3.9G, 4G, and the future 5G.
[0419] Home appliance EX118 refers to refrigerators or other equipment included in a home fuel cell combined heat and power system.
[0420] In the content supply system ex100, terminals with photography capabilities are connected to the streaming media server ex103 via a base station ex106, enabling on-site distribution. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) perform encoding processing on still or moving images captured by the user using these terminals, as described in the above embodiments. The encoded image data and the encoded audio data are multiplexed, and the resulting data is sent to the streaming media server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present invention.
[0421] On the other hand, the streaming media server ex103 distributes the content data sent by requesting clients. Clients are terminals such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or airplanes ex117 capable of decoding the encoded data. Each device receiving the distributed data decodes and reproduces the received data. That is, each device functions as an image decoding device according to one aspect of the present invention.
[0422] [Distributed processing]
[0423] Furthermore, the streaming media server ex103 can also be multiple servers or multiple computers, distributing data through decentralized processing or recording. For example, the streaming media server ex103 can also be implemented by a CDN (Contents Delivery Network), which distributes content by connecting many edge servers scattered around the world. In a CDN, physically nearby edge servers are dynamically allocated based on the client. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Moreover, in the event of an error or a change in communication status due to increased traffic, processing can be distributed across multiple edge servers, the distribution entity can be switched to other edge servers, or the faulty part of the network can be bypassed to continue distribution, thus achieving high-speed and stable distribution.
[0424] Furthermore, beyond the decentralized processing of distribution itself, the encoding processing of the captured data can be performed by each terminal, on the server side, or shared among them. For example, the encoding process typically involves two processing loops. In the first loop, the complexity or code size of the image per frame or scene unit is detected. In the second loop, processing is performed to maintain image quality while improving encoding efficiency. For instance, by having the terminal perform the first encoding processing and the server receiving the content perform the second, the processing load on each terminal can be reduced while improving both content quality and efficiency. In this case, if there is a request for near real-time reception and decoding, the data encoded by the terminal in the first round can be received and reproduced by other terminals, enabling more flexible real-time distribution.
[0425] As another example, cameras like the ex113 extract features from images, compress the data about these features as metadata, and send it to a server. The server, for instance, adjusts the quantization precision based on the features to determine the importance of the target, performing compression that corresponds to the meaning of the image. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, while more demanding encoding methods like CABAC (Context Adaptive Binary Arithmetic Coding) can be used by the server.
[0426] As another example, in stadiums, shopping malls, or factories, there may be multiple images of roughly the same scene captured by multiple terminals. In such cases, the data is distributed and encoded separately using the terminals that captured the images, as well as other terminals and servers that did not capture images, as needed. This can be done by assigning encoding and processing data to different units, such as GOP (Group of Pictures), image units, or tile units obtained by segmenting images. This reduces latency and improves real-time performance.
[0427] Furthermore, since multiple image datasets depict roughly the same scene, the server can manage and / or instruct the data to be cross-referenced between images captured by different terminals. Alternatively, the server can receive encoded data from each terminal and change the reference relationships between the multiple datasets, or modify or replace the images themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each data item.
[0428] In addition, the server can also transcode the image data by changing its encoding method before distributing it. For example, the server can convert MPEG encoding to VP encoding, or H.264 to H.265.
[0429] In this way, encoding processing can be performed by a terminal or one or more servers. Therefore, the terms "server" or "terminal" will be used below to refer to the main body performing the processing, but it is also possible to perform part or all of the processing performed by the server by the terminal, or vice versa. Furthermore, the same applies to decoding processing.
[0430] [3D, Multi-angle]
[0431] In recent years, there has been an increase in the use of images or videos captured by multiple cameras (ex113 and / or smartphones (ex115)) at roughly the same time, capturing different scenes or capturing the same scene from different angles. These images are then merged based on the relative positions of the cameras or regions containing consistent feature points within the images.
[0432] The server not only encodes 2D moving images, but can also automatically or at user-specified times encode still images based on scene analysis of moving images and send them to the receiving terminal. Furthermore, when the server can obtain the relative positions between shooting terminals, it can generate 3D shapes of scenes not only from 2D moving images, but also from images of the same scene captured from different angles. Additionally, the server can separately encode 3D data generated from point clouds, and can also select or reconstruct images from multiple terminals based on the results of identifying or tracking people or targets using 3D data to generate images to be sent to the receiving terminal.
[0433] In this way, users can freely select images corresponding to each shooting terminal to appreciate the scene, and can also appreciate the content of images extracted from any viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, similar to the images, sound can also be collected from multiple different angles, and the server can match it with the images to multiplex and send sound from a specific angle or space with the images.
[0434] Furthermore, in recent years, content that establishes a correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. This can be achieved through Multi-View Coding (MVC) to allow reference between the viewpoint images, or by encoding them as separate streams without reference to each other. During the decoding of these different streams, they can be synchronously reproduced according to the user's viewpoint to recreate a virtual three-dimensional space.
[0435] In the case of AR images, the server can overlay virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device acquires or holds the virtual object information and 3D data, generates a 2D image based on the user's viewpoint movement, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can send the user's viewpoint movement to the server in addition to the virtual object information. The server creates overlay data by matching the received viewpoint movement with the 3D data held on the server, encodes the overlay data, and distributes it to the decoding device. Furthermore, the overlay data has an α value representing transmittance in addition to RGB. The server sets the α value of the portion outside the target created from the 3D data to 0, etc., and encodes the portion in a state of transmittance. Alternatively, the server can set a predetermined RGB value as the background, similar to a chroma key, to generate data where the portion outside the target is set as the background color.
[0436] Similarly, the decoding of distributed data can be performed by the individual terminals acting as clients, on the server side, or distributed among them. For example, one terminal could first send a receive request to the server, other terminals could receive the content corresponding to that request, decode it, and then send the decoded signal to a device with a display. By distributing the processing independently of the capabilities of the communicating terminals and selecting appropriate content, data with better image quality can be reproduced. Furthermore, as another example, large-format image data can be received by a TV, and the viewer's personal terminal can decode and display a segmented area of the image, such as tiles. This allows for the sharing of the overall image while simultaneously allowing the viewer to identify their own area of responsibility or areas they wish to examine in more detail.
[0437] Furthermore, it is envisioned that in the future, with the availability of multiple short-range, medium-range, or long-range wireless communications both indoors and outdoors, content can be seamlessly received while switching appropriate data between connected communications using distribution system standards such as MPEG-DASH. This would allow users to switch in real-time not only using their own terminals but also freely choosing decoding or display devices such as monitors installed indoors or outdoors. Additionally, decoding can be performed by switching between decoding and display terminals based on the user's location information. This would also allow movement towards a destination while displaying map information on a portion of the wall or ground of a building adjacent to a display device. Furthermore, the bit rate of the received data can be switched based on the ease of accessing encoded data to the network, such as when the encoded data is cached on a server accessible only briefly from the receiving terminal or copied to an edge server of the content distribution service.
[0438] [Hyper-level coding]
[0439] Regarding content switching, use Figure 26 The scalable stream, which is compressed using the motion picture coding method described in the above embodiments, will be explained. For the server, there can be multiple streams with the same content but different qualities, or the content structure can be switched using the characteristics of a temporally / spatially scalable stream achieved through layered coding, as shown in the illustration. That is, by having the decoding side determine which layer to decode to based on intrinsic factors such as performance and extrinsic factors such as the state of the communication band, the decoding side can freely switch between decoding low-resolution and high-resolution content. For example, if one wants to watch a follow-up video viewed on a smartphone (ex115) while on the go, and then watch it at home on an internet TV or similar device, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0440] Furthermore, in addition to the hierarchical structure described above, where images are encoded layer by layer and enhancement layers exist above the base layers, the enhancement layer can also contain metadata such as image statistics. The decoding side then uses this metadata to perform super-resolution on the base layer images to generate high-quality content. Super-resolution can be either an increase in the signal-to-noise ratio (SN ratio) at the same resolution or an increase in resolution. The metadata includes information used to determine the linear or nonlinear filtering coefficients used in the super-resolution process, or information determining the parameter values for filtering, machine learning, or least-squares operations used in the super-resolution process.
[0441] Alternatively, the image can be segmented into tiles based on the meaning of targets within it, and the decoding side can decode only a portion of the region by selecting the tiles to be decoded. Furthermore, by storing the target's attributes (people, cars, balls, etc.) and its position within the image (coordinates within the same image, etc.) as metadata, the decoding side can determine the desired target's location based on this metadata and decide which tiles to include that target. For example, ... Figure 27 As shown, metadata is stored using data storage structures different from pixel data, such as SEI messages in HEVC. This metadata may represent, for example, the position, size, or color of the main target.
[0442] Furthermore, metadata can be stored in units consisting of multiple images, such as streams, sequences, or random access units. This allows the decoder to obtain information such as the time a specific person appears in the image, and by matching this information with the image unit information, it can determine the image containing the target and the target's location within that image.
[0443] [Web page optimization]
[0444] Figure 28 This is an example of a web page display screen in a computer such as ex111. Figure 29 This is an example image showing the display screen of a web page in a smartphone such as the ex115. Figure 28 and Figure 29 As shown, in cases where a web page contains multiple linked images that serve as links to image content, their visibility varies depending on the viewing device. When multiple linked images are visible on the screen, before the user explicitly selects a linked image, or before the linked image is near the center of the screen, or before the entire linked image enters the screen, the display device (decoding device) displays still images or I-images of each content as linked images, or displays images like GIF animations using multiple still images or I-images, or only receives the basic layer and decodes and displays the image.
[0445] When a user selects a linked image, the display device prioritizes decoding the base layer. Additionally, if the HTML constituting the webpage contains information indicating tiered content, the display device can also decode up to the enhancement layer. Furthermore, in situations where real-time performance is ensured before selection or when communication bandwidth is extremely limited, the display device can reduce the delay between decoding and displaying the first image (the delay from the start of content decoding to the start of display) by decoding and displaying only the preceding reference images (I images, P images, and B images only used for preceding reference). Alternatively, the display device can forcibly ignore image reference relationships and coarsely decode all B and P images as preceding references, performing normal decoding as more images are received over time.
[0446] [Autonomous Driving]
[0447] Furthermore, when receiving and transmitting still images or video data such as two-dimensional or three-dimensional map information for the purpose of autonomous driving or driving assistance, the receiving terminal can also receive information such as weather or construction as metadata, in addition to image data belonging to more than one layer, and decode them by establishing correspondences. Moreover, metadata can belong to a layer or be multiplexed only with image data.
[0448] In this scenario, since the vehicle, drone, or aircraft containing the receiving terminal is in motion, the receiving terminal can seamlessly receive and decode data by switching between base stations ex106 to ex110 by sending its location information when a request is received. Furthermore, the receiving terminal can dynamically adjust the level of metadata reception or map information updates based on user selection, user status, or the status of the communication frequency band.
[0449] As described above, in the content delivery system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0450] Distribution of personal content
[0451] Furthermore, the content delivery system ex100 not only handles high-quality, long-duration content provided by video distribution providers, but also enables unicast or multicast distribution of low-quality, short-duration content provided by individuals. Moreover, it's conceivable that such personal content will increase in the future. To improve the quality of personal content, the server can also perform encoding processing after editing. This can be achieved, for example, through a structure like the following.
[0452] During or after capturing images, the server performs image processing such as error detection, scene search, meaning analysis, and target detection based on the original images or encoded data. Furthermore, based on the recognition results, the server manually or automatically corrects focus deviations or camera shake, deletes less important scenes (such as those with lower brightness or out-of-focus areas), emphasizes target edges, or adjusts color tones. The server then encodes the edited data. Additionally, recognizing that longer shooting times can decrease audiovisual quality, the server can automatically crop not only less important scenes as described above but also scenes with minimal movement based on image processing results, to create content within a specific timeframe. Alternatively, the server can generate and encode summaries based on the meaning analysis results of the scenes.
[0453] Furthermore, personal content may contain elements that infringe on copyrights, the author's moral rights, or portrait rights in its original state, or the sharing scope may exceed the intended scope, causing inconvenience to the individual. Therefore, for example, the server could forcibly encode images such as faces of people in the periphery of a scene or a home, replacing them with out-of-focus images. Additionally, the server could identify whether a face different from a pre-registered person is captured within the image to be encoded, and if so, apply processing such as mosaic to the face. Alternatively, as pre- or post-processing for encoding, from a copyright perspective, the user can specify the person or background area to be processed, and the server can replace the specified area with another image or blur the focus. If it is a person, the face image can be replaced while tracking the person in a moving image.
[0454] Furthermore, personal content with small data volumes has strong real-time requirements for audiovisual presentation. Therefore, although bandwidth also plays a role, the decoding device prioritizes receiving, decoding, and reproducing the base layer. The decoding device can also receive enhancement layers during this process, including them in high-quality image reproduction if the playback is looped or repeated more than twice. In this way, if the stream is scalably encoded, it can provide an experience where the motion images are initially coarse but gradually become smoother and the image quality improves. Besides scalable encoding, the same experience can be provided when the first, coarser stream and a second stream encoded based on the first motion image are combined into a single stream.
[0455] [Other Use Cases]
[0456] Furthermore, these encoding or decoding processes are typically handled within the LSIex500 chip present in each terminal. The LSIex500 can be a single chip or a multi-chip structure. Alternatively, software for motion image encoding or decoding can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer such as the ex111, and the encoding and decoding processes can be performed using this software. Furthermore, when the ex115 smartphone has a camera, motion image data acquired by that camera can also be transmitted. This motion image data is encoded using the LSIex500 chip present in the ex115 smartphone.
[0457] Alternatively, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether it corresponds to the content's encoding method or whether it has the capability to execute a specific service. If the terminal does not correspond to the content's encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software, and then retrieves and reproduces the content.
[0458] Furthermore, not limited to the content delivery system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of the above embodiments can also be assembled in a digital broadcasting system. Since multiplexed data that multiplexes images and sound is carried and transmitted and received using radio waves for broadcasting via satellites or the like, it is more suitable for multicasting than the unicast-friendly structure of the content delivery system ex100, but the same applications can be performed for encoding and decoding processing.
[0459] [Hardware Structure]
[0460] Figure 30 This is a diagram representing the EX115 smartphone. Additionally, Figure 31 This diagram illustrates a structural example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with a base station ex110, a camera unit ex465 for capturing images and still images, and a display unit ex458 for displaying decoded data such as images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466, such as a touch panel; a sound output unit ex457, such as a speaker, for outputting sound or audio; a sound input unit ex456, such as a microphone, for inputting sound; a memory unit ex467 capable of storing encoded or decoded data such as captured images or still images, recorded audio, received images or still images, and emails; and a slot unit ex464 serving as an interface with a SIM ex468 used to identify the user and authenticate access to various data sources, such as the network. Alternatively, an external memory may be used instead of the memory unit ex467.
[0461] Furthermore, the main control unit ex460, which performs integrated control of the display unit ex458 and the operation unit ex466, is interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.
[0462] If the power button is turned on by the user, the power circuit section ex461 will supply power to each part from the battery pack, thus activating the smartphone ex115 into an operational state.
[0463] The smart phone ex115 is controlled by a main control unit ex460, which includes a CPU, ROM, and RAM, for call and data communication processing. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. This digital signal is then subjected to spectral diffusion processing by the modulation / demodulation unit ex452, and finally, digital-to-analog conversion and frequency conversion processing by the transmitting / receiving unit ex451 before being transmitted via the antenna ex450. Similarly, received data is amplified and subjected to frequency conversion and analog-to-digital conversion processing. The data is then subjected to inverse spectral diffusion processing by the modulation / demodulation unit ex452, converted into an analog audio signal by the audio signal processing unit ex454, and output from the audio output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 through the operation unit ex466 of the main unit, and are also processed for transmission and reception. In data communication mode, when transmitting images, still images, or images and sound, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded image data to the multiplexing / demultiplexing unit ex453. Furthermore, the sound signal processing unit ex454 encodes the sound signal collected by the sound input unit ex456 during the capture of images, still images, etc., by the camera unit ex465, and sends the encoded sound data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded image data and encoded sound data in a prescribed manner, performs modulation and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmitting / receiving unit ex451, and transmits the data via the antenna ex450.
[0464] When receiving images attached to emails or chat tools, or images linked to web pages, the multiplexing / demultiplexing unit ex453, in order to decode the multiplexed data received via antenna ex450, separates the multiplexed data into a bitstream of image data and a bitstream of audio data. The encoded image data is supplied to the image signal processing unit ex455 via the synchronization bus ex470, and the encoded audio data is supplied to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method described in the above embodiments, and displays the image or still image contained in the linked motion picture file from the display unit ex458 via the display control unit ex459. Furthermore, the audio signal processing unit ex454 decodes the audio signal and outputs audio from the audio output unit ex457. Additionally, since real-time streaming is becoming increasingly common, depending on the user's situation, the reproduction of audio may be unsuitable for certain situations. Therefore, as an initial consideration, a structure that reproduces only the image data and not the audio signal is preferred. It can also reproduce sound synchronously only when the user performs actions such as clicking on the image data.
[0465] Furthermore, while the example given here is the ex115 smartphone, as a terminal, three installation methods can be considered: a transmitting terminal with both an encoder and a decoder, a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. Moreover, in a digital broadcasting system, the example described involves receiving and transmitting multiplexed data, such as audio data, multiplexed within video data. However, in addition to audio data, multiplexed data can also multiplex character data associated with the video, or the video data itself can be received or transmitted without using multiplexed data.
[0466] Furthermore, while the description assumes the CPU's main control unit (ex460) controls the encoding or decoding process, many terminals also possess GPUs. Therefore, a structure can be implemented that utilizes GPU performance to process larger regions simultaneously, using shared memory between the CPU and GPU, or memory that manages addresses in a shared manner. This reduces encoding time, ensures real-time performance, and achieves low latency. In particular, it is even more efficient to perform motion search, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing on the GPU, image-by-image, rather than using the CPU.
[0467] Industrial availability
[0468] This invention can be applied to, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, or digital video cameras.
[0469] Label Explanation
[0470] 100, 312 encoding devices
[0471] 102 Division
[0472] 104 Subtraction Section
[0473] 106 Transformer
[0474] Quantitative Department 108
[0475] 110 Entropy Coding Department
[0476] 112, 204 Inverse Quantization Section
[0477] Inverse Transformation Units 114 and 206
[0478] 116, 208 Addition Department
[0479] 118 and 210 memory blocks
[0480] 120, 212 Circular Filter Section
[0481] 122, 214 frame memory
[0482] Intra-frame prediction units 124 and 216
[0483] Inter-frame prediction units 126 and 218
[0484] Predictive Control Department 128, 220
[0485] 200, 322 decoding devices
[0486] 202 Entropy Decoding Department
[0487] 300 encoding and decoding system
[0488] 310 Coding System
[0489] 311, 321 communication device
[0490] 313 Output Buffer
[0491] 320 decoding system
[0492] 323 Input Buffer
[0493] Candidate Derivatives 1261 and 2181
[0494] 1262, 2182 Scope Determination Department
[0495] 1263, 2183 Sports Search Department
[0496] 1264, 2184 Sports Compensation Department
Claims
1. An encoding apparatus for encoding a block of objects using motion vectors, characterized in that, have: Processor; and Memory; The processor described above uses the memory described above to perform the following processes: The first candidate vector is derived from one or more candidate vectors of one or more adjacent blocks that are adjacent to the above-mentioned encoded object block. In the first reference image of the aforementioned encoded object block, a first surrounding region containing the position represented by the aforementioned first candidate vector is determined. Calculate the evaluation values of multiple candidate regions contained in the first surrounding area mentioned above. The first motion vector of the coded object block is determined based on the first candidate region, which is the candidate region with the smallest evaluation value. The aforementioned first surrounding region is included within the first motion search range determined based on the position represented by the aforementioned first candidate vector.
2. An encoding method for encoding object blocks using motion vectors, characterized in that, The first candidate vector is derived from one or more candidate vectors of one or more adjacent blocks that are adjacent to the above-mentioned encoded object block. In the first reference image of the aforementioned encoded object block, a first surrounding region containing the position represented by the aforementioned first candidate vector is determined. Calculate the evaluation values of multiple candidate regions contained in the first surrounding area mentioned above. The first motion vector of the coded object block is determined based on the first candidate region, which is the candidate region with the smallest evaluation value. The aforementioned first surrounding region is included within the first motion search range determined based on the position represented by the aforementioned first candidate vector.
Citation Information
Patent Citations
Moving image encoding device, moving image decoding device, moving image encoding method, moving image decoding method, moving image encoding program, moving image decoding program, moving image processing system and moving image processing method
CN102177716A
Motion vector derivation in video coding
US20160286229A1