Symbolization device and decoding device

JP7686858B2Active Publication Date: 2025-06-02PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024153039
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-05-27
Filing Date
2024-09-05
Publication Date
2025-06-02
Estimated Expiration
2037-05-23

AI Technical Summary

Technical Problem

Conventional encoding and decoding devices struggle to handle images captured by non-rectilinear lenses effectively, leading to inefficiencies due to distortion and overlapping areas created by image stitching and correction processes.

Method used

The encoding device employs a stitching process to combine multiple images, identifies empty areas, and performs inter-screen prediction to replace pixel values, while the decoding device adapts to these processes using parameters to ensure efficient handling and encoding/decoding of distorted images.

Benefits of technology

This approach enhances encoding efficiency by reducing duplication and improving the handling of distorted images captured by non-rectilinear lenses, ensuring effective encoding and decoding of stitched images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000046_0000
    Figure 00000046_0000
  • Figure 00000049_0000
    Figure 00000049_0000
  • Figure 00000049_0001
    Figure 00000049_0001
Patent Text Reader

Abstract

To provide an encoder that can properly handle an image to be encoded or decoded.SOLUTION: An encoder 1500 comprises a processing circuit and memories 1505, 1506 connected to the processing circuit. By using the memories 1505, 1506, the processing circuit performs connection processing of connecting a plurality of images to each other to create a connected image, acquires a parameter for specifying a space area in the connected image generated in the connection processing, performs inter-screen prediction processing on the connected image, and writes the parameter in a bit stream. The inter-screen prediction processing includes padding processing of connecting the values of pixels in the space area and replacing the values with the value of another area that is not the space area. The value of the another area is the value of a pixel closest from the space area. The inter-screen prediction processing is performed on an image block basis.SELECTED DRAWING: Figure 37
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an apparatus and method for encoding an image, and an apparatus and method for decoding the encoded image. [Background technology]

[0002] Currently, HEVC has been established as a standard for image coding (see, for example, Non-Patent Document 1). However, the transmission and storage of next-generation videos (e.g., 360-degree videos) will require coding efficiency that exceeds current coding performance. In addition, several studies and experiments have been conducted on the compression of moving images captured by wide-angle lenses such as non-rectilinear lenses. In these studies, the image samples are manipulated to eliminate distortion aberration, thereby making the image to be processed linear before encoding. For this purpose, image processing techniques are generally used. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] H.265(ISO / IEC 23008-2 HEVC(High Efficiency Video Coding)) Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional encoding devices and decoding devices have a problem in that they are unable to appropriately handle images to be encoded or decoded.

[0005] Therefore, the present disclosure provides an encoding device and the like that can appropriately handle images to be encoded or decoded. [Means for solving the problem]

[0006] An encoding device according to one aspect of the present disclosure includes a processing circuit and a memory connected to the processing circuit, the processing circuit uses the memory to perform a stitching process to stitch together a plurality of images to generate a stitched image, obtains parameters that identify an empty area in the stitched image that is generated by the stitching process, performs inter-screen prediction processing on the stitched image, and writes the parameters to a bitstream, the inter-screen prediction processing including a padding process that replaces values ​​of pixels in the empty area with values ​​of another area in the stitched image that is not the empty area, the value of the other area being the value of the pixel closest to the empty area, and the inter-screen prediction processing is performed on an image block basis.

[0007] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. Effect of the Invention

[0008] The encoding device of the present disclosure can appropriately handle images to be encoded or decoded. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of a coding device according to the first embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of block division according to the first embodiment. [Diagram 3] FIG. 3 is a table showing the transform basis functions corresponding to each transform type. [Figure 4A] FIG. 4A is a diagram showing an example of the shape of a filter used in ALF. [Figure 4B] FIG. 4B is a diagram showing another example of the shape of the filter used in the ALF. [Figure 4C]FIG. 4C is a diagram showing another example of the shape of the filter used in the ALF. [Diagram 5] FIG. 5 is a diagram showing 67 intra prediction modes in intra prediction. [Figure 6] FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 8] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. [Figure 9] FIG. 9 is a diagram for explaining derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 10] FIG. 10 is a block diagram showing a functional configuration of a decoding device according to the first embodiment. As shown in FIG. [Figure 11] FIG. 11 is a flowchart showing an example of a video encoding process according to the second embodiment. [Figure 12] FIG. 12 is a diagram showing possible positions of a header in a bitstream in the second embodiment where parameters are written. [Figure 13] FIG. 13 is a diagram showing a captured image and a processed image that has been subjected to image correction processing in the second embodiment. [Figure 14] FIG. 14 is a diagram showing a stitched image generated by stitching a plurality of images together by the stitching process according to the second embodiment. [Figure 15] FIG. 15 is a diagram showing an arrangement of a plurality of cameras and a stitched image including a free area generated by stitching together images captured by the cameras in the second embodiment. [Figure 16] FIG. 16 is a flowchart showing the inter prediction process or motion compensation according to the second embodiment. [Figure 17]FIG. 17 is a diagram showing an example of barrel distortion caused by a non-rectilinear lens or a fisheye lens in the second embodiment. In FIG. [Figure 18] FIG. 18 is a flowchart showing a modification of the inter prediction process or motion compensation process in the second embodiment. [Figure 19] FIG. 19 is a flowchart showing the image reconstruction process in the second embodiment. [Figure 20] FIG. 20 is a flowchart showing a modified example of the image reconstruction process in the second embodiment. [Figure 21] FIG. 21 is a diagram showing an example of partial encoding processing or partial decoding processing for a spliced ​​image in the second embodiment. [Figure 22] FIG. 22 is a diagram showing another example of the partial encoding process or partial decoding process for a spliced ​​image in the second embodiment. [Diagram 23] FIG. 23 is a block diagram of an encoding device according to the second embodiment. [Figure 24] FIG. 24 is a flowchart showing an example of the video decoding process according to the second embodiment. [Diagram 25] FIG. 25 is a block diagram of a decoding device according to the second embodiment. In FIG. [Figure 26] FIG. 26 is a flowchart showing an example of a video encoding process according to the third embodiment. [Figure 27] FIG. 27 is a flowchart showing an example of the joining process according to the third embodiment. [Figure 28] FIG. 28 is a block diagram of an encoding device according to the third embodiment. [Figure 29] FIG. 29 is a flowchart showing an example of a video decoding process according to the third embodiment. [Diagram 30] FIG. 30 is a block diagram of a decoding device according to the third embodiment. [Diagram 31] FIG. 31 is a flowchart showing an example of a video encoding process according to the fourth embodiment. [Diagram 32]FIG. 32 is a flowchart showing the intra prediction process according to the fourth embodiment. [Diagram 33] FIG. 33 is a flowchart showing the motion vector prediction process according to the fourth embodiment. [Diagram 34] FIG. 34 is a block diagram of an encoding device according to the fourth embodiment. [Diagram 35] FIG. 35 is a flowchart showing an example of a video decoding process according to the fourth embodiment. [Diagram 36] FIG. 36 is a block diagram of a decoding device according to the fourth embodiment. In FIG. [Figure 37] FIG. 37 is a block diagram of an encoding device according to one embodiment of the present disclosure. [Figure 38] FIG. 38 is a block diagram of a decoding device according to one embodiment of the present disclosure. [Figure 39] FIG. 39 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Diagram 40] FIG. 40 is a diagram showing an example of a coding structure in scalable coding. [Diagram 41] FIG. 41 is a diagram showing an example of a coding structure in scalable coding. [Diagram 42] FIG. 42 is a diagram showing an example of a display screen of a web page. [Diagram 43] FIG. 43 is a diagram showing an example of a display screen of a web page. [Diagram 44] FIG. 44 is a diagram illustrating an example of a smartphone. [Diagram 45] FIG. 45 is a block diagram showing an example of the configuration of a smartphone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, the embodiment will be described in detail with reference to the drawings.

[0011] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.

[0012] (Embodiment 1) [Outline of the encoding device] First, a description will be given of an overview of a coding device according to embodiment 1. Fig. 1 is a block diagram showing a functional configuration of a coding device 100 according to embodiment 1. The coding device 100 is a video / image coding device that codes a video / image on a block-by-block basis.

[0013] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0014] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The encoding device 100 may also be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0015] Each component included in the encoding device 100 will be described below.

[0016] [Divided part] The division unit 102 divides each picture included in the input video into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides a picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks may be called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in a picture may be the processing units of CUs, PUs, and TUs.

[0017] Fig. 2 is a diagram showing an example of block division according to embodiment 1. In Fig. 2, solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0018] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0019] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0020] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14, 15 (binary tree block division).

[0021] The bottom left 64x64 block is divided into four square 32x32 blocks (quadtree block division). Of the four 32x32 blocks, the top left and bottom right blocks are further divided. The top left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block division). The bottom right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block division). As a result, the bottom left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0022] The bottom right 64x64 block 23 is not split.

[0023] 2, the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0024] In Fig. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0025] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (prediction sample) from an original signal (original sample) for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also called a residual error) of a block to be coded (hereinafter, referred to as a current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0026] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0027] [Conversion section] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0028] The transform unit 106 may adaptively select a transform type from among a plurality of transform types, and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform may be called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0029] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transform basis functions corresponding to each transform type. In Figure 3, N indicates the number of input pixels. The selection of the transform type from among the multiple transform types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0030] Such information indicating whether EMT or AMT is applied (e.g., called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0031] Furthermore, the transform unit 106 may retransform the transform coefficients (transformation results). Such retransformation may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs retransformation for each subblock (e.g., 4x4 subblock) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether or not to apply NSST and information regarding a transform matrix used in NSST are signaled at a CU level. Note that signaling of these pieces of information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0032] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter, referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0033] The predetermined order is an order for quantization / dequantization of the transform coefficients. For example, the predetermined scanning order is defined as ascending (low to high) or descending (high to low) frequency order.

[0034] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0035] [Entropy coding part] The entropy coding unit 110 generates a coded signal (coded bit stream) by variable-length coding the quantized coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110, for example, binarizes the quantized coefficients and arithmetically codes the binary signal.

[0036] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0037] [Inverse conversion section] The inverse transform unit 114 restores the prediction error by inverse transforming the transform coefficients that are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients that corresponds to the transform performed by the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0038] Note that the restored prediction error does not match the prediction error calculated by the subtraction unit 104 because information has been lost due to quantization. That is, the restored prediction error includes a quantization error.

[0039] [Addition section] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transformer 114 and the prediction signal input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0040] [Block memory] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be coded (hereinafter, referred to as a current picture). Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0041] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0042] In ALF, a least squared error filter is applied to remove coding artifacts. For example, for each 2x2 sub-block in the current block, one filter is selected from among multiple filters based on local gradient direction and activity.

[0043] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (e.g., 15 or 25 classes).

[0044] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived, for example, by adding gradients in multiple directions and quantizing the sum.

[0045] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0046] The shape of the filter used in the ALF is, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing a number of examples of the shape of the filter used in the ALF. FIG. 4A shows a 5×5 diamond-shaped filter, FIG. 4B shows a 7×7 diamond-shaped filter, and FIG. 4C shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, and may be at other levels (for example, the sequence level, slice level, tile level, CTU level, or CU level).

[0047] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, whether or not to apply ALF is determined for luminance at the CU level, and whether or not to apply ALF is determined for chrominance at the picture level. Information indicating whether or not to apply ALF is signaled at the picture level or the CU level. Note that the signaling of information indicating whether or not to apply ALF is not limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0048] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Note that the signaling of the coefficient sets does not need to be limited to the picture level, but may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0049] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0050] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called intra-screen prediction) of the current block with reference to a block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0051] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0052] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / High-Efficiency Video Coding (HEVC) standard (Non-Patent Document 1).

[0053] The multiple directional prediction modes include, for example, 33 prediction modes defined in the H.265 / HEVC standard. The multiple directional prediction modes may include 32 prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 5 is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0054] In addition, in the intra prediction of the chrominance block, the luminance block may be referenced. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. Such intra prediction may be called CCLM (cross-component linear model) prediction. An intra prediction mode of the chrominance block that refers to such a luminance block (for example, called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0055] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on the gradient of reference pixels in the horizontal / vertical directions. Intra prediction with such correction may be called position dependent intra prediction combination (PDPC). Information indicating whether or not PDPC is applied (e.g., called a PDPC flag) is signaled, for example, at a CU level. Note that the signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0056] [Inter prediction section] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. The inter prediction is performed in units of the current block or a sub-block (e.g., 4x4 block) in the current block. For example, the inter prediction unit 126 performs motion estimation in the reference picture for the current block or the sub-block. Then, the inter prediction unit 126 generates an inter prediction signal of the current block or the sub-block by performing motion compensation using motion information (e.g., a motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0057] The motion information used for the motion compensation is signaled. For the signaling of the motion vector, a motion vector predictor may be used, i.e. the difference between the motion vector and the motion vector predictor may be signaled.

[0058] In addition, the inter prediction signal may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent block. Specifically, the inter prediction signal may be generated for each sub-block in the current block by performing weighted addition of the prediction signal based on the motion information obtained by motion search and the prediction signal based on the motion information of the adjacent block. Such inter prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0059] In such an OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (e.g., called OBMC flag) is signaled at the CU level. Note that the signaling level of these pieces of information does not need to be limited to the sequence level and CU level, and may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0060] In addition, the motion information may be derived on the decoding device side without being signaled. For example, a merge mode defined in the H.265 / HEVC standard may be used. Also, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation is performed without using pixel values ​​of the current block.

[0061] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a flame rate up-conversion (FRUC) mode.

[0062] First, one of the candidates included in the merge list is selected as a starting position for the search by pattern matching. As the pattern matching, a first pattern matching or a second pattern matching is used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0063] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures along the motion trajectory of the current block.

[0064] Fig. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in Fig. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of blocks that best match among two blocks along the motion trajectory of a current block (Cur block) and in two different reference pictures (Ref0, Ref1).

[0065] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures in time and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives bidirectional motion vectors that are mirror-symmetric.

[0066] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (eg, an upper and / or left adjacent block)) and a block in the reference picture.

[0067] Fig. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in Fig. 7, in the second pattern matching, a motion vector of a current block is derived by searching a block in a reference picture (Ref0) that best matches a block adjacent to a current block (Cur block) in a current picture (Cur Pic).

[0068] Information indicating whether such a FRUC mode is applied (e.g., when the FRUC flag is true) is signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (e.g., when the FRUC mode flag is true) is signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0069] Note that the motion information may be derived on the decoding device side using a method other than the motion estimation. For example, a correction amount of a motion vector may be calculated on a pixel-by-pixel basis using neighboring pixel values ​​based on a model assuming uniform linear motion.

[0070] Here, a mode in which a motion vector is derived based on a model assuming uniform linear motion will be described. This mode is sometimes called a BIO (bi-directional optical flow) mode.

[0071] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. x ,v y) denotes a velocity vector, and τ0 and τ1 denote the temporal distance between the current picture (CurPic) and two reference pictures (Ref0, Ref1), respectively. (MVx0,MVy0) denotes a motion vector corresponding to the reference picture Ref0, and (MVx1,MVy1) denotes a motion vector corresponding to the reference picture Ref1.

[0072] At this time, the velocity vector (v x ,v y Under the assumption of uniform linear motion of the object, (MVx0,MVy0) and (MVx1,MVy1) are respectively (v x τ0,v y τ0) and (-v x τ1,-v y τ1), and the following optical flow equation (1) holds:

[0073]

number

[0074] Here, I (k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like are corrected pixel by pixel.

[0075] Note that the decoding device may derive a motion vector using a method other than the method based on a model assuming uniform linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0076] Here, a mode in which a motion vector is derived for each sub-block based on the motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.

[0077] FIG. 9 is a diagram for explaining derivation of a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 9, the current block includes 16 4x4 sub-blocks. Here, a motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and a motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vectors (v x ,v y ) is derived.

[0078]

number

[0079] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting factor.

[0080] Such affine motion compensation prediction mode may include several modes with different methods of deriving the motion vectors of the upper left and upper right corner control points. Information indicating such affine motion compensation prediction mode (e.g., called affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0081] [Predictive control unit] The prediction control unit 128 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.

[0082] [Overview of the Decryption Device] Next, a description will be given of an overview of a decoding device capable of decoding the coded signal (coded bit stream) output from the above coding device 100. Fig. 10 is a block diagram showing a functional configuration of a decoding device 200 according to the first embodiment. The decoding device 200 is a video / image decoding device that decodes a video / image on a block-by-block basis.

[0083] As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0084] The decoding device 200 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The decoding device 200 may also be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0085] Each component included in the decoding device 200 will be described below.

[0086] [Entropy Decoding Part] The entropy decoding unit 202 entropy decodes the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. The entropy decoding unit 202 then debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantized coefficients to the inverse quantization unit 204 on a block-by-block basis.

[0087] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of a block to be decoded (hereinafter, referred to as a current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0088] [Inverse conversion section] The inverse transform unit 206 restores the prediction error by inverse transforming the transform coefficients input from the inverse quantization unit 204 .

[0089] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0090] Also, for example, if the information decoded from the coded bitstream indicates that NSST is to be applied, the inverse transform unit 206 re-transforms the transformed transform coefficients (transformation results).

[0091] [Addition section] The adder 208 reconstructs the current block by adding the prediction error, which is an input from the inverse transformer 206, and the prediction signal, which is an input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0092] [Block memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter, referred to as a current picture). Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0093] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0094] If the information indicating ALF on / off read from the encoded bitstream indicates ALF on, one filter is selected from among multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0095] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0096] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on an intra prediction mode interpreted from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0097] Note that, when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0098] Furthermore, when information interpreted from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects pixel values ​​after intra prediction based on the gradients of reference pixels in the horizontal / vertical directions.

[0099] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) in the current block. For example, the inter prediction unit 126 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the encoded bitstream, and outputs the inter prediction signal to the prediction control unit 128.

[0100] In addition, when the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks.

[0101] Also, if the information interpreted from the encoded bitstream indicates that the FRUC mode is applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the encoded bitstream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0102] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when information interpreted from the encoded bitstream indicates that an affine motion compensation prediction mode is applied, the inter prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0103] [Predictive control unit] The prediction control unit 220 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the addition unit 208 as a prediction signal.

[0104] (Embodiment 2) Next, a part of the processing performed in the encoding device 100 and the decoding device 200 configured as above will be specifically described with reference to the drawings. Note that it will be apparent to those skilled in the art that the embodiments described below may be combined in order to further enhance the benefits of the present disclosure.

[0105] The encoding device and decoding device in this embodiment can be used for encoding and decoding any multimedia data, and more specifically, can be used for encoding and decoding images captured by a non-rectilinear (e.g., fisheye) camera.

[0106] Here, in the above-mentioned prior art, the same video coding tool is used to compress the processed images and the images captured directly by the rectilinear lens. There is no video coding tool in the prior art that is specifically customized to compress this type of processed images in a different way.

[0107] Typically, images are first captured by multiple cameras, and then the images captured by the multiple cameras are stitched together to create a larger image, known as a 360-degree image. To make the image more comfortable to view on a flat display or to make objects in the image easier to detect using machine learning techniques, image transformation processes may be performed before the image is encoded, including "defishing" or straightening the image. However, this image transformation process usually involves interpolating image samples, which results in duplication of information held in the image. Also, the stitching and image transformation processes may create empty areas in the image, which are usually filled with default pixel values ​​(e.g., black pixels). Such problems caused by the stitching and image transformation processes reduce the coding efficiency of the coding process.

[0108] To solve these problems, the present embodiment uses an adaptive video encoding tool and an adaptive video decoding tool as the customized video encoding tool and video decoding tool. To improve the encoding efficiency, the adaptive video encoding tool can adapt to the image transformation or image stitching process used to process the images prior to the encoder. The present disclosure can reduce any duplication caused by these processes by adapting the adaptive video encoding tool to such processes during the encoding process. The adaptive video decoding tool is similar to the adaptive video encoding tool.

[0109] In this embodiment, the information of the image transformation process and / or image stitching process is used to adapt the video encoding tool and the video decoding tool, so that the video encoding tool and the video decoding tool can be applied to different types of processed images, and therefore, the compression efficiency can be improved in this embodiment.

[0110] [Encoding process] A method of performing video coding on an image captured using a non-rectilinear lens according to the second embodiment of the present disclosure shown in Fig. 11 will be described. Note that the non-rectilinear lens is a wide-angle lens or an example thereof.

[0111] FIG. 11 is a flowchart showing an example of the video encoding process according to the present embodiment.

[0112] In step S101, the encoding device writes a set of parameters into a header. Figure 12 shows possible locations of said header in a compressed video bitstream. The written parameters (i.e., camera image parameters in Figure 12) include one or more parameters related to image enhancement processing. For example, such parameters are written into a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a video system setup parameter set, as shown in Figure 12. That is, the parameters written in this embodiment may be written into any header of the bitstream, or into SEI (Supplemental Enhancement Information). Note that the image enhancement processing corresponds to the above-mentioned image conversion processing.

[0113] <Example of image correction processing parameters> As shown in FIG. 13, the captured image may be distorted due to the characteristics of the lens used during capture of the image. An image correction process was used to linearly correct the captured image. By linearly correcting the captured image, a rectangular image is generated. The written parameters include parameters for identifying or describing the image correction process used. The parameters used in the image correction process include, as an example, parameters constituting a mapping table for mapping pixels of the input image to the intended output pixel values ​​of the image correction process. These parameters may include weight parameters for one or more interpolation processes, or / and position parameters that specify the location of the input and output pixels of the picture. As one possible implementation of the image correction process, the mapping table for the image correction process may be used for all pixels in the corrected image.

[0114] Other examples of parameters used to describe the image correction process include a selection parameter that selects one of multiple predefined correction algorithms, a direction parameter that selects one of multiple predefined directions of the correction algorithm, and / or a calibration parameter that calibrates or fine-tunes the correction algorithm. For example, if there are multiple predefined correction algorithms (e.g., different algorithms are used for different types of lenses), the selection parameter is used to select one of these predefined algorithms. For example, if there are two or more directions in which the correction algorithm can be applied (e.g., the image correction process can be performed horizontally, vertically, or in either direction), the direction parameter selects one of these predefined directions. If the image correction process can be calibrated, the calibration parameter allows the image correction process to be adjusted to suit different types of lenses.

[0115] <Example of parameters for splicing process> The written parameters may further include one or more parameters related to the stitching process. As shown in Figures 14 and 15, the image input to the encoding device may be the result of a stitching process that combines multiple images from different cameras. The written parameters include parameters that provide information about the stitching process, such as the number of cameras, the distortion center or the principal axis of each camera, and the distortion level. Another example of parameters describing the stitching process includes parameters that identify the location of the stitched image that is generated by overlapping pixels from multiple images. Each of these images may contain pixels that may appear in other images, since there may be overlapping areas in the camera angles. In the stitching process, these overlapping pixels are processed and reduced to generate the stitched image.

[0116] Another example of a parameter describing the stitching process includes a parameter specifying the layout of the stitched images. For example, different 360-degree image formats, such as equirectangular projection, cubic 3x2 layout, and cubic 4x3 layout, have different image arrangements in the stitched image. Note that a 3x2 layout is a layout of 6 images arranged in 3 columns and 2 rows, and a 4x3 layout is a layout of 12 images arranged in 4 columns and 3 rows. The above parameter, the arrangement parameter, is used to specify the continuity of images in a direction based on the arrangement of the images. During the motion compensation process, pixels from other images or views can be used for the inter-prediction process, and these images or views are specified by the arrangement parameter. Some images or pixels in images may also need to be rotated to ensure continuity.

[0117] Other examples of parameters include camera and lens parameters (e.g., focal length, principal point, scale factor, type of image sensor used in the camera, etc.) Yet other examples of parameters include physical information about the placement of the camera (e.g., camera position, camera angle, etc.).

[0118] Next, in step S102, the encoding device encodes the image by adaptive video encoding tools based on these written parameters. The adaptive video encoding tools include an inter prediction process. The set of adaptive video encoding tools may further include an image reconstruction process.

[0119] <Distortion correction in inter-frame prediction> FIG. 16 is a flowchart showing an inter-frame prediction process that is applied when an image is identified as being captured using a non-rectilinear lens, when an image is identified as being linearly processed, or when an image is identified as being stitched together from one or more images. As shown in FIG. 16, in step S1901, the encoding device determines a position in an image to be a distortion center or a principal point based on parameters written in a header. FIG. 17 shows an example of a distortion aberration caused by a fisheye lens. Note that a fisheye lens is an example of a wide-angle lens. As the image moves away from the distortion center, the magnification decreases along the focal axis. Therefore, in step S1902, the encoding device can correct the distortion by wrapping pixels in the image to straighten the image based on the distortion center, or can undo the correction that was made. That is, the encoding device performs an image correction process (i.e., wrapping process) on a block of the distorted image to be encoded. Finally, in step S1903, the encoding device can perform block prediction to derive a block of predicted samples based on pixels of the wrapped image. In this embodiment, the wrapping process or wrapping is a process of arranging or rearranging pixels, blocks or images. The coding device may also return the predicted block, which is a predicted block, to the original distorted state before the image correction process is performed, and use the distorted predicted block as the predicted image of the distorted target block. In addition, the predicted image and the target block correspond to the predicted signal and the current block in the first embodiment.

[0120] Another example of an adapted inter prediction process includes adapted motion vector processing. The resolution of the motion vector is lower for image blocks far from the distortion center than for image blocks close to the distortion center. For example, image blocks far from the distortion center may have a motion vector precision up to half pixel precision. Meanwhile, image blocks close to the distortion center may have a high motion vector precision up to 1 / 8 pixel precision. Since the adapted motion vector precision differs based on the image block position, the precision of the motion vector coded in the bitstream may be adaptive depending on the end and / or start position of the motion vector. That is, the coding device may use parameters to vary the precision of the motion vector depending on the position of the block.

[0121] Another example of an adaptive inter prediction process includes an adaptive motion compensation process, in which pixels from different views may be used to predict image samples from a target view based on the alignment parameters written in the header. For example, different 360-degree image formats, such as equirectangular projection, cubic 3x2 layout, cubic 4x3 layout, etc., have different image alignments in the stitched image. The alignment parameters are used to specify image continuity in a certain direction based on the image alignment. During the motion compensation process, pixels from other images or other views can be used for the inter prediction process, and these images or views are specified by the alignment parameters. Some images or pixels in images may also need to be rotated to ensure continuity.

[0122] That is, the encoding device may perform a process to ensure continuity. For example, when encoding the spliced ​​image shown in FIG. 15, the encoding device may perform a wrapping process based on the parameters. Specifically, of the five images (i.e., images A to D and a top view) included in the spliced ​​image, the top view is a 180-degree image, and images A to D are 90-degree images. Therefore, the space displayed in the top view is continuous with the spaces displayed in each of images A to D, and the space displayed in image A is continuous with the space displayed in image B. However, in the spliced ​​image, the top view is not continuous with images A, C, and D, and image A is not continuous with image B. Therefore, the encoding device performs the above-mentioned wrapping process to improve the encoding efficiency. That is, the encoding device rearranges each image included in the spliced ​​image. For example, the encoding device rearranges each image so that image A and image B are continuous. As a result, the objects displayed separately in images A and B are continuous, and the encoding efficiency can be improved. The wrapping process, which is a process of rearranging or arranging each image in this manner, is also called frame packing.

[0123] <Padding in inter-prediction> FIG. 18 is a flow chart showing a variation of the inter prediction process that is applied when the image is identified as captured using a non-rectilinear lens, when the image is identified as linearly processed, or when the image is identified as stitched together from two or more images. As shown in FIG. 18, the encoding device identifies regions of the image as free regions in step S2001 based on parameters written in the header. These free regions are regions of the image that do not contain pixels of the captured image, and are typically replaced with a predetermined pixel value (e.g., black pixels). FIG. 13 shows an example of these regions in an image. FIG. 15 shows another example of these regions in the case of stitching together multiple images. The encoding device then pads the pixels in these identified regions with values ​​from other non-free regions of the image during motion compensation in step S2002 of FIG. 18. The padded values ​​may be values ​​from the nearest pixel in the non-free region or values ​​from the nearest pixel, depending on the physical three-dimensional space. Finally, in step S2003, the encoding apparatus performs block prediction to generate a block of predicted samples based on the padded values.

[0124] <Distortion correction in image reconstruction> FIG. 19 is a flowchart showing an image reconstruction process that is applied when an image is identified as being captured using a non-rectilinear lens, when an image is identified as being linearly processed, or when an image is identified as being stitched together from two or more images. As shown in FIG. 19, the encoding device determines a position in an image as a distortion center or a principal point based on a parameter written in a header in step S1801. FIG. 17 shows an example of a distortion aberration caused by a fisheye lens. As the focal axis moves away from the distortion center, the magnification decreases along the focal axis. Therefore, in step S1802, the encoding device may perform a wrapping process on a reconstructed pixel in an image based on the distortion center to correct the distortion or to undo the correction made to make the image linear. For example, the encoding device generates a reconstructed image by adding an image of a prediction error generated by an inverse transform and a predicted image. At this time, the encoding device performs a wrapping process to make each of the image of the prediction error and the predicted image linear.

[0125] Finally, in step S1803, the encoding device stores in memory the blocks of the image reconstructed based on the pixels of the wrapped image.

[0126] <Replacement of pixel values ​​in image reconstruction> FIG. 20 shows a variation of the image reconstruction process that is applied when the image is identified as being captured using a non-rectilinear lens, or when the image is identified as being processed linearly, or when the image is identified as being stitched together from one or more images. As shown in FIG. 20, based on the parameters written in the header, in step S2101, the encoding device identifies regions of the image as free regions. These free regions are regions of the image that do not contain pixels of the captured image and are typically replaced with a predetermined pixel value (e.g., black pixels). FIG. 13 shows an example of such regions in an image. FIG. 15 shows another example of such regions in the case of stitching together multiple images. Next, in step S2102, the encoding device reconstructs blocks of image samples.

[0127] Also, in step S2103, the encoding device replaces the reconstructed pixels in these identified regions with predetermined pixel values.

[0128] <Omission of encoding process> Another possible variant of the adaptive video coding tool in step S102 of Fig. 11 is that the coding device may skip the coding process of the image, i.e., based on the written parameters of the layout arrangement of the image and the information about the active viewing area based on the user's eye line or head direction, the coding device may skip the coding process of the image, i.e., perform a partial coding process.

[0129] FIG. 21 shows an example of the viewing angle or head orientation of a user with respect to different views captured by different cameras. As shown in the figure, the user's viewing angle is within the image captured by the camera from view 1 only. In this example, images from other views do not need to be encoded because they are outside the user's viewing angle. Therefore, the encoding or transmission process for these images can be omitted to reduce the encoding complexity or to reduce the transmission bit rate of the compressed images. In another possible example shown, since view 5 and view 2 are physically close to the active view 1, images from view 5 and view 2 are also encoded and transmitted. These images are not currently displayed to the viewer or user, but will be displayed to the viewer or user when the viewer turns his or her head. These images are used to improve the user's viewing experience when the viewer turns his or her head.

[0130] FIG. 22 shows another example of gaze angles or head orientations for different views captured by different cameras of a user. Here, the active gaze area is within the image from view 2. Thus, the image from view 2 is coded and displayed to the user. Here, the encoding device predicts the estimated range of the viewer's head movement in the near future and defines a larger area as the possible gaze area of ​​the future frame. The encoding device also codes images from views (other than view 2) that are not within the target's active gaze area but are within the larger future gaze area, and transmits them to enable the viewer to draw the views faster. That is, not only the image from view 2, but also the images from the top view and view 1, which at least partially overlap the possible gaze area shown in FIG. 22, are coded and transmitted. Images from the remaining views (view 3, view 4, and bottom view) are not coded, and the coding process for these images is omitted.

[0131] [Encoding device] FIG. 23 is a block diagram showing a configuration of a coding device that codes moving pictures in this embodiment.

[0132] The encoding device 900 is a device for encoding an input video image for each block in order to generate an output bitstream, and corresponds to the encoding device 100 of Embodiment 1. As shown in Fig. 23, the encoding device 900 includes a transform unit 901, a quantization unit 902, an inverse quantization unit 903, an inverse transform unit 904, a block memory 905, a frame memory 906, an intra prediction unit 907, an inter prediction unit 908, a subtraction unit 921, an addition unit 922, an entropy encoding unit 909, and a parameter derivation unit 910.

[0133] An image of the input video sequence (i.e., a current block) is input to the subtraction unit 921, and the subtracted value is output to the conversion unit 901. That is, the subtraction unit 921 calculates a prediction error by subtracting a prediction image from the current block. The conversion unit 901 converts the subtracted value (i.e., a prediction error) into a frequency coefficient, and outputs the obtained frequency coefficient to the quantization unit 902. The quantization unit 902 quantizes the input frequency coefficient, and outputs the obtained quantized value to the inverse quantization unit 903 and the entropy coding unit 909.

[0134] The inverse quantization unit 903 inverse quantizes the sample values ​​(i.e., quantized values) output from the quantization unit 902, and outputs frequency coefficients to the inverse transformation unit 904. The inverse transformation unit 904 performs inverse frequency transformation to transform the frequency coefficients into sample values ​​of the image, i.e., pixel values, and outputs the obtained sample values ​​to the addition unit 922.

[0135] The parameter derivation unit 910 derives parameters related to image correction processing, or parameters related to a camera, or parameters related to a splicing process from an image, and outputs the parameters to the inter prediction unit 908, the adder 922, and the entropy coding unit 909. For example, the input video may include these parameters, and in this case, the parameter derivation unit 910 extracts and outputs the parameters included in the video. Alternatively, the input video may include parameters that are the base for deriving these parameters. In this case, the parameter derivation unit 910 extracts the base parameters included in the video, converts the extracted base parameters into the above-mentioned parameters, and outputs them.

[0136] The adder 922 adds the sample values ​​output from the inverse transformer 904 to the pixel values ​​of the predicted image output from the intra prediction unit 907 or the inter prediction unit 908. That is, the adder 922 performs an image reconstruction process to generate a reconstructed image. The adder 922 outputs the resulting sum to the block memory 905 or the frame memory 906 for further prediction.

[0137] The intra prediction unit 907 performs intra-screen prediction. That is, the intra prediction unit 907 estimates the image of the processing target block using a reconstructed image included in the same picture as the processing target block picture stored in the block memory 905. The inter prediction unit 908 performs inter-screen prediction. That is, the inter prediction unit 908 estimates the image of the processing target block using a reconstructed image included in a picture different from the processing target block picture stored in the frame memory 906.

[0138] Here, in this embodiment, the inter prediction unit 908 and the adder unit 922 adapt the process based on the parameters derived by the parameter derivation unit 910. That is, the inter prediction unit 908 and the adder unit 922 perform the process according to the flowcharts shown in Figs. 16, 18, 19, and 20 as the process by the adaptive video coding tool described above.

[0139] The entropy coding unit 909 codes the quantized value output from the quantization unit 902 and the parameters derived by the parameter derivation unit 910, and outputs a bitstream. That is, the entropy coding unit 909 writes the parameters into the header of the bitstream.

[0140] [Decryption process] FIG. 24 is a flowchart showing an example of the video decoding process according to the present embodiment.

[0141] In step S201, the decoding device parses a set of parameters from a header, the possible locations of which are shown in Figure 12 in a compressed video bitstream. The parsed parameters include one or more parameters related to image enhancement processes.

[0142] <Example of image correction processing parameters> As shown in FIG. 13, the captured image may be distorted due to the characteristics of the lens used during capture of the image. An image correction process was used to linearly correct the captured image. The interpreted parameters include parameters that identify or describe the image correction process used. Examples of parameters used in the image correction process include parameters that constitute a mapping table for mapping pixels of the input image to the intended output pixel values ​​of the image correction process. These parameters may include weight parameters for one or more interpolation processes, or / and location parameters that identify the location of the input and output pixels of the picture. In one possible implementation of the image correction process, the mapping table for the image correction process may be used for all pixels in the corrected image.

[0143] Other examples of parameters used to describe the image correction process include a selection parameter for selecting one of a plurality of predefined correction algorithms, a direction parameter for selecting one of a plurality of predefined directions of the correction algorithm, and / or a calibration parameter for calibrating or fine-tuning the correction algorithm. For example, if there are a plurality of predefined correction algorithms (e.g., different algorithms are used for different types of lenses), the selection parameter is used to select one of these predefined algorithms. For example, if there are two or more directions in which the correction algorithm can be applied (e.g., the image correction process can be performed horizontally, vertically, or in either direction), the direction parameter selects one of these predefined directions. For example, if the image correction process can be calibrated, the calibration parameter can adjust the image correction process to suit different types of lenses.

[0144] <Example of parameters for splicing process> The interpreted parameters may further include one or more parameters related to the stitching process. As shown in Figures 14 and 15, the encoded image input to the decoding device may be the result of a stitching process that combines multiple images from different cameras. The interpreted parameters include parameters that provide information about the stitching process, such as the number of cameras, the distortion center or principal axis of each camera, and the distortion level. Another example of a parameter that describes the stitching process is a parameter that specifies the location of a stitched image that is generated from overlapping pixels from multiple images. Each of these images may contain pixels that may appear in other images because there may be overlapping areas in the camera angles. In the stitching process, these overlapping pixels are processed and reduced to generate the stitched image.

[0145] Another example of a parameter describing the stitching process includes a parameter specifying the layout of the stitched images. For example, different formats of 360-degree images, such as equirectangular projection, cubic 3x2 layout, or cubic 4x3 layout, have different arrangements of images in the stitched image. The above parameter, the arrangement parameter, is used to specify the continuity of images in a direction based on the arrangement of the images. During the motion compensation process, pixels from other images or views can be used for the inter prediction process, and these images or views are specified by the arrangement parameter. Some images or pixels in images may also need to be rotated to ensure continuity.

[0146] Other examples of parameters include camera and lens parameters (e.g., focal length, principal point, scale factor, type of image sensor used in the camera, etc.) Yet other examples of parameters include physical information about the placement of the camera (e.g., camera position, camera angle, etc.).

[0147] Next, in step S202, the decoding device decodes the image by an adaptive video decoding tool based on the interpreted parameters. The adaptive video decoding tool includes an inter-frame prediction process. The set of adaptive video decoding tools may also include an image reconstruction process. Note that the video decoding tool or the adaptive video decoding tool is the same as or corresponds to the above-mentioned video encoding tool or the adaptive video encoding tool.

[0148] <Distortion correction in inter-frame prediction> FIG. 16 is a flowchart showing an inter-frame prediction process that is applied when an image is identified as being captured using a non-rectilinear lens, when an image is identified as being linearly processed, or when an image is identified as being stitched together from one or more images. As shown in FIG. 16, in step S1901, the decoding device determines that a certain position in the image is a distortion center or a principal point based on the parameters written in the header. FIG. 17 shows an example of a distortion aberration caused by a fisheye lens. As the focal axis moves away from the distortion center, the magnification decreases along the focal axis. Therefore, in step S1902, the decoding device may perform a wrapping process on pixels in the image based on the distortion center to correct the distortion or to undo the correction made to straighten the image. That is, the decoding device performs an image correction process (i.e., a wrapping process) on a block of the distorted image that is the target of the decoding process. Finally, in step S1903, the decoding device may perform a block prediction to derive a block of predicted samples based on the pixels of the image that has been subjected to the wrapping process. The decoding device may also restore the predicted block, which is a prediction block, to its original distorted state before the image correction process was performed, and use the distorted prediction block as a predicted image for the distorted block to be processed.

[0149] Another example of an adapted inter prediction process includes adapted motion vector processing. The resolution of the motion vector is lower for image blocks far from the distortion center than for image blocks close to the distortion center. For example, image blocks far from the distortion center may have a motion vector precision up to half pixel precision. Meanwhile, image blocks close to the distortion center may have a high motion vector precision up to 1 / 8 pixel precision. Since the adapted motion vector precision differs based on the image block position, the motion vector precision coded in the bitstream may be adaptive depending on the end and / or start position of the motion vector. That is, the decoding device may use parameters to vary the precision of the motion vector depending on the position of the block.

[0150] Another example of an adaptive inter-prediction process includes an adaptive motion compensation process, in which pixels from different views may be used to predict image samples from a target view based on the alignment parameters written in the header. For example, different 360-degree image formats, such as equirectangular projection, cubic 3x2 layout, cubic 4x3 layout, etc., have different image alignments in the stitched image. The alignment parameters are used to specify image continuity in a certain direction based on the image alignment. During the motion compensation process, pixels from other images or other views may be used for the inter-prediction process, and these images or views are specified by the alignment parameters. Some images or pixels in images may also need to be rotated to ensure continuity.

[0151] That is, the decoding device may perform a process to ensure continuity. For example, when encoding the spliced ​​image shown in FIG. 15, the decoding device may perform wrapping processing based on the parameters. Specifically, the decoding device rearranges each image so that image A and image B are continuous, similar to the above-mentioned encoding device. This makes it possible to make the objects displayed separately in image A and image B continuous, thereby improving the encoding efficiency.

[0152] <Padding in inter-prediction> FIG. 18 is a flow chart showing a variation of the inter prediction process that is applied when an image is identified as captured using a non-rectilinear lens, when an image is identified as being linearly processed, or when an image is identified as being stitched together from two or more images. As shown in FIG. 18, the decoding device identifies regions of the image as free regions in step S2001 based on parameters read from the header. These free regions are regions of the image that do not contain pixels of the captured image and are generally replaced with a predetermined pixel value (e.g., black pixels). FIG. 13 shows an example of these regions in an image. FIG. 15 shows another example of these regions in the case of stitching together multiple images. The decoding device then pads pixels in these identified regions with values ​​from other non-free regions of the image during motion compensation in step S2002 of FIG. 18. The padded values ​​may be values ​​from the nearest pixel in the non-free region or the nearest pixel depending on the physical three-dimensional space. Finally, in step S2003, the decoder performs block prediction to generate a block of predicted samples based on the padded values.

[0153] <Distortion correction in image reconstruction> FIG. 19 is a flowchart showing an image reconstruction process that is applied when an image is identified as being captured using a non-rectilinear lens, when an image is identified as being linearly processed, or when an image is identified as being stitched together from two or more images. As shown in FIG. 19, the decoding device determines a position in the image as a distortion center or a principal point in step S1801 based on parameters read from the header. FIG. 17 shows an example of distortion aberration caused by a fisheye lens. As the focal axis moves away from the distortion center, the magnification decreases along the focal axis. Therefore, in step S1802, the decoding device may perform a wrapping process on the reconstructed pixels in the image based on the distortion center to correct the distortion or to undo the correction made to make the image linear. For example, the decoding device generates a reconstructed image by adding an image of a prediction error generated by an inverse transform and a predicted image. At this time, the decoding device performs a wrapping process to make each of the image of the prediction error and the predicted image linear.

[0154] Finally, in step S1803, the decoding device stores in memory blocks of an image reconstructed based on the pixels of the wrapped image.

[0155] <Replacement of pixel values ​​in image reconstruction> FIG. 20 shows a variation of the image reconstruction process that is applied when the image is identified as having been captured using a non-rectilinear lens, or when the image is identified as having been processed linearly, or when the image is identified as having been stitched together from one or more images. As shown in FIG. 20, based on the parameters read from the header, in step S2001, the decoder identifies regions of the image as free regions. These free regions are regions of the image that do not contain pixels of the captured image and are typically replaced with a predetermined pixel value (e.g., black pixels). FIG. 13 shows an example of these regions in an image. FIG. 15 shows another example of these regions in the case of stitching together multiple images. Next, in step S2102, the decoder reconstructs blocks of image samples.

[0156] Also, in step S2103, the decoding device replaces the reconstructed pixels in these identified regions with predetermined pixel values.

[0157] <Omission of decryption process> In step S202 of Fig. 24, another possible variant of the adaptive video decoding tool for an image may skip the decoding process of the image, i.e., based on the interpreted parameters of the layout arrangement of the image and the information about the active viewing area based on the user's eye gaze or head direction, the decoding device may skip the decoding process of the image, i.e., the decoding device performs a partial decoding process.

[0158] FIG. 21 shows an example of the viewing angle or head orientation of a user with respect to different views captured by different cameras. As shown in the figure, the user's viewing angle is within the image captured by the camera from view 1 only. In this example, images from other views do not need to be decoded because they are outside the user's viewing angle. Therefore, the decoding or display process for these images can be omitted to reduce the decoding complexity or to reduce the transmission bit rate of the compressed images. In another possible example shown, images from view 5 and view 2 are also decoded because they are physically close to the active view 1. These images are not currently displayed to the viewer or user, but will be displayed to the viewer or user when the viewer changes his or her head orientation. These images are displayed as soon as possible to improve the user's viewing experience when the user changes his or her head orientation by reducing the time to decode and display the views according to the user's head movement.

[0159] FIG. 22 shows another example of gaze angles or head orientations for different views captured by different cameras of a user. Here, the active gaze area is within the image from view 2. Thus, the image from view 2 is decoded and displayed to the user. Here, the decoding device predicts the estimated range of the viewer's head movement in the near future and defines a larger area as the possible gaze area of ​​the future frame. The decoding device also decodes images from views (other than view 2) that are not within the target active gaze area but are within the larger future gaze area. That is, not only the image from view 2, but also the images from the top view and view 1 that overlap at least partially with the possible gaze area shown in FIG. 22 are decoded. This allows the images to be displayed in a way that allows the viewer to draw the views faster. Images from the remaining views (view 3, view 4, and the view below) are not decoded, and the decoding process for these images is omitted.

[0160] [Decryption device] FIG. 25 is a block diagram showing a configuration of a decoding device that decodes moving images in this embodiment.

[0161] The decoding device 1000 is a device for decoding an input coded video (i.e., an input bit stream) for each block in order to generate a decoded video, and corresponds to the decoding device 200 of the first embodiment. As shown in FIG. 25 , the decoding device 1000 includes an entropy decoding unit 1001, an inverse quantization unit 1002, an inverse transform unit 1003, a block memory 1004, a frame memory 1005, an adder 1022, an intra prediction unit 1006, and an inter prediction unit 1007.

[0162] An input bitstream is input to an entropy decoding unit 1001. Thereafter, the entropy decoding unit 1001 performs entropy decoding on the input bitstream, and outputs a value obtained by the entropy decoding (i.e., a quantized value) to an inverse quantization unit 1002. The entropy decoding unit 1001 further decodes parameters from the input bitstream, and outputs the parameters to an inter prediction unit 1007 and an adder unit 1022.

[0163] The inverse quantization unit 1002 inversely quantizes the value obtained by the entropy decoding, and outputs the frequency coefficient to the inverse transform unit 1003. The inverse transform unit 1003 performs an inverse frequency transform on the frequency coefficient to convert the frequency coefficient into a sample value (i.e., a pixel value), and outputs the obtained pixel value to the adder unit 1022. The adder unit 1022 adds the obtained pixel value to the pixel value of the predicted image output from the intra prediction unit 1006 or the inter prediction unit 1007. That is, the adder unit 1022 performs an image reconstruction process to generate a reconstructed image. The adder unit 1022 outputs the value obtained by the addition (i.e., a decoded image) to a display, and outputs the obtained value to the block memory 1004 or the frame memory 1005 for further prediction.

[0164] The intra prediction unit 1006 performs intra-screen prediction. That is, the intra prediction unit 1006 estimates an image of the block to be processed using a reconstructed image included in the same picture as the picture of the block to be processed stored in the block memory 1004. The inter prediction unit 1007 performs inter-screen prediction. That is, the inter prediction unit 1007 estimates an image of the block to be processed using a reconstructed image included in a picture different from the picture of the block to be processed stored in the frame memory 1005.

[0165] Here, in this embodiment, the inter prediction unit 1007 and the adder 1022 adapt the process based on the interpreted parameters. That is, the inter prediction unit 1007 and the adder 1022 perform the process according to the flowcharts shown in Figs. 16, 18, 19, and 20 as the process by the adaptive video decoding tool described above.

[0166] (Embodiment 3) [Encoding process] A method for performing a video coding process on an image captured using a non-rectilinear lens according to the third embodiment of this disclosure shown in FIG. 26 will be described.

[0167] FIG. 26 is a flowchart showing an example of the video encoding process according to the present embodiment.

[0168] In step S301, the encoding device writes a set of parameters into a header. Figure 12 shows possible locations of said header in a compressed video bitstream. The written parameters include one or more parameters relating to the camera position. The written parameters may also include one or more parameters relating to the camera angle or instructions on how to stitch multiple images together.

[0169] Other examples of parameters include camera and lens parameters (e.g., focal length, principal point, scale factor, type of image sensor used in the camera, etc.) Further examples of parameters include physical information regarding the placement of the camera (e.g., camera position, camera angle, etc.).

[0170] In this embodiment, the above parameters written in the header are also called camera parameters or stitching parameters.

[0171] Figure 15 shows an example of a method for stitching together images from two or more cameras. Figure 14 shows another example of a method for stitching together images from two or more cameras.

[0172] Next, in step S302, the encoding device encodes the image. In step S302, the encoding process may be adapted based on the stitched image. For example, the encoding device may refer to the larger stitched image as a reference image in the motion compensation process, instead of an image of the same size as the decoded image (i.e., an unstitched image).

[0173] Finally, in step S303, the encoding device stitches the first image, which is the image encoded and reconstructed in step S302, with the second image based on the written parameters to create a larger image. The stitched image may be used for prediction of future frames (i.e., inter-frame prediction or motion compensation).

[0174] FIG. 27 is a flowchart showing a splicing process in which the parameters written in the header are used. In step S2401, the encoding device determines camera parameters or splicing parameters from the parameters written for the target image. Similarly, in step S2402, the encoding device determines camera parameters or splicing parameters for other images from the parameters written for the other images. Finally, in step S2403, the encoding device uses these determined parameters to splice images together to create a larger image. These determined parameters are written in the header. Note that the encoding device may perform a wrapping process or frame packing to arrange or rearrange multiple images so that the encoding efficiency is further improved.

[0175] [Encoding device] FIG. 28 is a block diagram showing a configuration of a coding device that codes a moving image in this embodiment.

[0176] The encoding device 1100 is a device for encoding an input video image for each block in order to generate an output bitstream, and corresponds to the encoding device 100 of the first embodiment. As shown in FIG. 28, the encoding device 1100 includes a transform unit 1101, a quantization unit 1102, an inverse quantization unit 1103, an inverse transform unit 1104, a block memory 1105, a frame memory 1106, an intra prediction unit 1107, an inter prediction unit 1108, a subtraction unit 1121, an addition unit 1122, an entropy encoding unit 1109, a parameter derivation unit 1110, and an image splicing unit 1111.

[0177] An image of the input video sequence (i.e., a current block) is input to the subtraction unit 1121, and the subtracted value is output to the transformation unit 1101. That is, the subtraction unit 1121 calculates a prediction error by subtracting a prediction image from the current block. The transformation unit 1101 transforms the subtracted value (i.e., a prediction error) into a frequency coefficient, and outputs the obtained frequency coefficient to the quantization unit 1102. The quantization unit 1102 quantizes the input frequency coefficient, and outputs the obtained quantized value to the inverse quantization unit 1103 and the entropy coding unit 1109.

[0178] The inverse quantization unit 1103 inverse quantizes the sample values ​​(i.e., quantized values) output from the quantization unit 1102, and outputs frequency coefficients to the inverse transformation unit 1104. The inverse transformation unit 1104 performs inverse frequency transformation on the frequency coefficients to transform the frequency coefficients into sample values ​​of an image, i.e., pixel values, and outputs the resulting sample values ​​to the addition unit 1122.

[0179] The adder 1122 adds the sample values ​​output from the inverse transformer 1104 to pixel values ​​of the predicted image output from the intra predictor 1107 or the inter predictor 1108. The adder 1122 outputs the resulting sum to the block memory 1105 or the frame memory 1106 for further prediction.

[0180] As in the first embodiment, the parameter derivation unit 1110 derives parameters related to the image splicing process or parameters related to the camera from the images, and outputs the parameters to the image splicing unit 1111 and the entropy coding unit 1109. That is, the parameter derivation unit 1110 executes the processes of steps S2401 and S2402 shown in FIG. 27. For example, the input video may include these parameters, and in this case, the parameter derivation unit 1110 extracts and outputs the parameters included in the video. Alternatively, the input video may include parameters that are bases for deriving these parameters. In this case, the parameter derivation unit 1110 extracts base parameters included in the video, converts the extracted base parameters into the above-mentioned parameters, and outputs them.

[0181] The image splicing unit 1111 splices the reconstructed target image to another image using the parameters, as shown in step S303 of Fig. 26 and step S2403 of Fig. 27. Thereafter, the image splicing unit 1111 outputs the spliced ​​image to the frame memory 1106.

[0182] The intra prediction unit 1107 performs intra-screen prediction. That is, the intra prediction unit 1107 estimates the image of the processing target block using a reconstructed image included in the same picture as the picture of the processing target block stored in the block memory 1105. The inter prediction unit 1108 performs inter-screen prediction. That is, the inter prediction unit 1108 estimates the image of the processing target block using a reconstructed image included in a picture different from the picture of the processing target block stored in the frame memory 1106. At this time, the inter prediction unit 1108 may refer to a large image obtained by splicing a plurality of images by the image splicing unit 1111 stored in the frame memory 1106 as a reference image.

[0183] The entropy coding unit 1109 codes the quantized value output from the quantization unit 1102, obtains parameters from the parameter derivation unit 1110, and outputs a bitstream. That is, the entropy coding unit 1109 performs entropy coding on the quantized value and the parameters, and writes the parameters in the header of the bitstream.

[0184] [Decryption process] FIG. 29 is a flowchart showing an example of the video decoding process according to the present embodiment.

[0185] In step S401, the decoder parses a set of parameters from the header. Figure 12 shows possible locations of such a header in a compressed video bitstream. The parsed parameters include one or more parameters related to the position of the camera. The parsed parameters may further include one or more parameters related to the camera angle or instructions on how to stitch multiple images together. Other examples of parameters include camera and lens parameters (e.g. focal length, principal point, scale factor, type of image sensor used in the camera, etc.). Further examples of parameters include physical information related to the placement of the camera (e.g. camera position, camera angle, etc.).

[0186] Figure 15 shows one example of how images from two or more cameras can be stitched together. Figure 14 shows another example of how images from two or more cameras can be stitched together.

[0187] Next, in step S402, the decoding device decodes the image. The decoding process in step S402 may also be adapted based on the stitched image. For example, the decoding device may refer to the stitched larger image as a reference image in the motion compensation process, instead of an image of the same size as the decoded image (i.e., an unstitched image).

[0188] And finally, in step S403, the decoding device stitches the first image, which is the image reconstructed in step S402, with the second image based on the interpreted parameters to create a larger image. The stitched image may be used for prediction of future images (i.e., inter-frame prediction or motion compensation).

[0189] 27 is a flow chart showing a stitching process using the decoded parameters. In step S2401, the decoder determines the camera parameters or stitching parameters by decode the header for the current image. Similarly, in step S2402, the decoder determines the camera parameters or stitching parameters by decode the header for the other image. Finally, in step S2403, the decoder uses these decoded parameters to stitch the images together to create a larger image.

[0190] [Decryption device] FIG. 30 is a block diagram showing a configuration of a decoding device that decodes moving images in this embodiment.

[0191] The decoding device 1200 is a device that decodes an input coded video (i.e., an input bit stream) for each block and outputs a decoded video, and corresponds to the decoding device 200 of Embodiment 1. As shown in FIG. 30 , the decoding device 1200 includes an entropy decoding unit 1201, an inverse quantization unit 1202, an inverse transform unit 1203, a block memory 1204, a frame memory 1205, an adder 1222, an intra prediction unit 1206, an inter prediction unit 1207, and an image splicing unit 1208.

[0192] An input bitstream is input to an entropy decoding unit 1201. Thereafter, the entropy decoding unit 1201 performs entropy decoding on the input bitstream, and outputs a value obtained by the entropy decoding (i.e., a quantized value) to an inverse quantization unit 1202. The entropy decoding unit 1201 further decodes parameters from the input bitstream, and outputs the parameters to an image splicing unit 1208.

[0193] The image splicing unit 1208 uses the parameters to splice the reconstructed target image to another image, and then outputs the image obtained by splicing to the frame memory 1205.

[0194] The inverse quantization unit 1202 inverse quantizes the value obtained by the entropy decoding, and outputs the frequency coefficient to the inverse transform unit 1203. The inverse transform unit 1203 performs an inverse frequency transform on the frequency coefficient, converts the frequency coefficient into a sample value (i.e., a pixel value), and outputs the resultant pixel value to the adder unit 1222. The adder unit 1222 adds the resultant pixel value to the pixel value of the predicted image output from the intra prediction unit 1206 or the inter prediction unit 1207. The adder unit 1222 outputs the value obtained by the addition (i.e., a decoded image) to a display, and outputs the resultant value to the block memory 1204 or the frame memory 1205 for further prediction.

[0195] The intra prediction unit 1206 performs intra-screen prediction. That is, the intra prediction unit 1206 estimates an image of the block to be processed using a reconstructed image included in the same picture as the picture of the block to be processed stored in the block memory 1204. The inter prediction unit 1207 performs inter-screen prediction. That is, the inter prediction unit 1207 estimates an image of the block to be processed using a reconstructed image included in a picture different from the picture of the block to be processed stored in the frame memory 1205.

[0196] (Embodiment 4) [Encoding process] A method of performing a video coding process on an image captured using a non-rectilinear lens according to the fourth embodiment of this disclosure shown in FIG. 31 will be described.

[0197] FIG. 31 is a flowchart showing an example of the video encoding process according to the present embodiment.

[0198] In step S501, the encoding device writes a set of parameters into a header. Figure 12 shows a possible location of said header in a compressed video bitstream. The written parameters include one or more parameters relating to an identifier indicating whether the image is captured with a non-rectilinear lens. As shown in Figure 13, the captured image may be distorted due to the characteristics of the lens used during the capture of the image. One example of a written parameter is a parameter indicating the location of the center or principal axis of distortion.

[0199] Next, in step S502, the encoding device encodes the image by adaptive video encoding tools based on the written parameters. The adaptive video encoding tools include a motion vector prediction process. The set of adaptive video encoding tools may also include an intra-frame prediction process.

[0200] <In-screen prediction processing> Fig. 32 is a flowchart showing an intra-frame prediction process adapted based on the written parameters. As shown in Fig. 32, in step S2201, the encoding device determines a position in an image as a distortion center or a principal point based on the written parameters. Next, in step S2202, the encoding device predicts a sample group using spatially neighboring pixel values. The sample group is, for example, a group of pixels such as a block to be processed.

[0201] Finally, in step S2203, the encoding device performs a wrapping process on the predicted sample group using the determined distortion center or principal point to generate a block of predicted samples. For example, the encoding device may warp an image of the block of predicted samples and use the distorted image as the predicted image.

[0202] <Motion Vector Prediction> Fig. 33 is a flow chart showing a motion vector prediction process adapted based on the written parameters. As shown in Fig. 33, in step S2301, the encoding device determines a position in an image as a distortion center or principal point based on the written parameters. Next, in step S2302, the encoding device predicts a motion vector from spatially or temporally adjacent motion vectors.

[0203] Finally, in step S2303, the encoding apparatus corrects the direction of the predicted motion vector using the determined distortion center or principal point.

[0204] [Encoding device] FIG. 34 is a block diagram showing the configuration of a coding device that codes moving pictures in this embodiment.

[0205] The encoding device 1300 is a device for encoding an input video image for each block in order to generate an output bitstream, and corresponds to the encoding device 100 of Embodiment 1. As shown in Fig. 34, the encoding device 1300 includes a transform unit 1301, a quantization unit 1302, an inverse quantization unit 1303, an inverse transform unit 1304, a block memory 1305, a frame memory 1306, an intra prediction unit 1307, an inter prediction unit 1308, a subtraction unit 1321, an addition unit 1322, an entropy encoding unit 1309, and a parameter derivation unit 1310.

[0206] An image of the input video sequence (i.e., a current block) is input to the subtraction unit 1321, and the subtracted value is output to the transformation unit 1301. That is, the subtraction unit 1321 calculates a prediction error by subtracting a prediction image from the current block. The transformation unit 1301 transforms the subtracted value (i.e., a prediction error) into a frequency coefficient, and outputs the resulting frequency coefficient to the quantization unit 1302. The quantization unit 1302 quantizes the input frequency coefficient, and outputs the resulting quantized value to the inverse quantization unit 1303 and the entropy coding unit 1309.

[0207] The inverse quantization unit 1303 inverse quantizes the sample values ​​(i.e., quantized values) output from the quantization unit 1302, and outputs frequency coefficients to the inverse transform unit 1304. The inverse transform unit 1304 performs inverse frequency transform on the frequency coefficients to convert the frequency coefficients into sample values ​​of the image, i.e., pixel values, and outputs the resulting sample values ​​to the addition unit 1322.

[0208] The parameter derivation unit 1310 derives, from an image, one or more parameters (specifically, parameters indicating a distortion center or a principal point) related to an identifier indicating whether the image is captured by a non-rectilinear lens, as in the first embodiment. Then, the parameter derivation unit 1310 outputs the derived parameters to the intra prediction unit 1307, the inter prediction unit 1308, and the entropy coding unit 1309. For example, the input video may include these parameters, and in this case, the parameter derivation unit 1310 extracts and outputs the parameters included in the video. Alternatively, the input video may include base parameters for deriving these parameters. In this case, the parameter derivation unit 1310 extracts base parameters included in the video, converts the extracted base parameters into the above-mentioned parameters, and outputs them.

[0209] The adder 1322 adds the sample values ​​of the image output from the inverse transformer 1304 to pixel values ​​of the predicted image output from the intra predictor 1307 or the inter predictor 1308. The adder 1322 outputs the resulting sum to the block memory 1305 or the frame memory 1306 for further prediction.

[0210] The intra prediction unit 1307 performs intra prediction. That is, the intra prediction unit 1307 estimates an image of the block to be processed using a reconstructed image included in the same picture as the picture of the block to be processed stored in the block memory 1305. The inter prediction unit 1308 performs inter prediction. That is, the inter prediction unit 1308 estimates an image of the block to be processed using a reconstructed image included in a picture different from the picture of the block to be processed in the frame memory 1306.

[0211] Here, in this embodiment, the intra prediction unit 1307 and the inter prediction unit 1308 perform processing based on parameters derived by the parameter derivation unit 1310. That is, the intra prediction unit 1307 and the inter prediction unit 1308 perform processing according to the flowcharts shown in Figs. 32 and 33, respectively.

[0212] The entropy coding unit 1309 codes the quantized value output from the quantization unit 1302 and the parameters derived by the parameter derivation unit 1310, and outputs a bitstream. That is, the entropy coding unit 1309 writes the parameters into the header of the bitstream.

[0213] [Decryption process] FIG. 35 is a flowchart showing an example of the video decoding process in this embodiment.

[0214] In step S601, the decoding device parses a set of parameters from a header. Figure 12 shows possible locations of said header in a compressed video bitstream. The parsed parameters include one or more parameters relating to an identifier indicating whether the image has been captured with a non-rectilinear lens. As shown in Figure 13, the captured image may be distorted due to the characteristics of the lens used during capture of the image. One example of a parsed parameter is a parameter indicating the location of the center or principal axis of distortion.

[0215] Next, in step S602, the decoding device decodes the image based on these interpreted parameters by an adaptive video decoding tool. The adaptive video decoding tool includes a motion vector prediction process. The adaptive video decoding tool may also include an intra-frame prediction process. Note that the video decoding tool or the adaptive video decoding tool is the same as or corresponds to the above-mentioned video encoding tool or the adaptive video encoding tool.

[0216] <In-screen prediction processing> FIG. 32 is a flowchart showing an intra-frame prediction process adapted based on the interpreted parameters. As shown in FIG. 32, in step S2201, the decoding device determines a position in an image as a distortion center or principal point based on the interpreted parameters. Next, in step S2202, the decoding device predicts a group of samples using spatially neighboring pixel values. Finally, in step S2203, the decoding device performs a wrapping process on the predicted group of samples using the determined distortion center or principal point to generate a block of predicted samples. For example, the decoding device may distort an image of the block of predicted samples and use the distorted image as a predicted image.

[0217] <Motion Vector Prediction> FIG. 33 is a flow chart showing a motion vector prediction process adapted based on the interpreted parameters. As shown in FIG. 33, in step S2301, the decoding device determines a position in an image as a distortion center or principal point based on the interpreted parameters. Next, in step S2302, the decoding device predicts a motion vector from spatially or temporally adjacent motion vectors. Finally, in step S2303, the decoding device corrects the direction of the motion vector using the determined distortion center or principal point.

[0218] [Decryption device] FIG. 36 is a block diagram showing a configuration of a decoding device for decoding moving images in this embodiment.

[0219] The decoding device 1400 is a device for decoding an input coded video (i.e., an input bit stream) for each block and outputting a decoded video, and corresponds to the decoding device 200 in the first embodiment. As shown in FIG. 36 , the decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transform unit 1403, a block memory 1404, a frame memory 1405, an adder 1422, an intra prediction unit 1406, and an inter prediction unit 1407.

[0220] An input bitstream is input to an entropy decoding unit 1401. Thereafter, the entropy decoding unit 1401 performs entropy decoding on the input bitstream, and outputs a value obtained by the entropy decoding (i.e., a quantized value) to an inverse quantization unit 1402. The entropy decoding unit 1401 further decodes parameters from the input bitstream, and outputs the parameters to an inter prediction unit 1407 and an intra prediction unit 1406.

[0221] The inverse quantization unit 1402 inverse quantizes the value obtained by the entropy decoding, and outputs the frequency coefficient to the inverse transform unit 1403. The inverse transform unit 1403 performs an inverse frequency transform on the frequency coefficient to convert the frequency coefficient into a sample value (i.e., a pixel value), and outputs the resulting pixel value to the adder unit 1422. The adder unit 1422 adds the resulting pixel value to the pixel value of the predicted image output from the intra prediction unit 1406 or the inter prediction unit 1407. The adder unit 1422 outputs the value obtained by the addition (i.e., a decoded image) to a display, and outputs the obtained value to the block memory 1404 or the frame memory 1405 for further prediction.

[0222] The intra prediction unit 1406 performs intra-screen prediction. That is, the intra prediction unit 1406 predicts the image of the block to be processed using a reconstructed image included in the same picture as the picture of the block to be processed stored in the block memory 1404. The inter prediction unit 1407 performs inter-screen prediction. That is, the inter prediction unit 1407 estimates the image of the block to be processed using a reconstructed image included in a picture different from the picture of the block to be processed stored in the frame memory 1405.

[0223] Here, in this embodiment, the inter prediction unit 1407 and the intra prediction unit 1406 adapt the process based on the interpreted parameters. That is, the inter prediction unit 1407 and the intra prediction unit 1406 perform the process according to the flowcharts shown in Figs. 32 and 33 as the process by the adaptive video decoding tool.

[0224] (summary) Although examples of the encoding device and decoding device of the present disclosure have been described above using each embodiment, the encoding device and decoding device according to one aspect of the present disclosure are not limited to these embodiments.

[0225] For example, in each of the above embodiments, the encoding device encodes a video using a parameter related to image distortion or a parameter related to image splicing, and the decoding device decodes the encoded video using those parameters. However, the encoding device and the decoding device according to one aspect of the present disclosure do not need to perform encoding or decoding using those parameters. In other words, they do not need to perform processing using the adaptive video encoding tool and the adaptive video decoding tool in the above embodiments.

[0226] FIG. 37 is a block diagram of an encoding device according to one embodiment of the present disclosure.

[0227] An encoding device 1500 according to an aspect of the present disclosure is a device corresponding to the encoding device 100 of the first embodiment, and includes, as shown in Fig. 37, a transform unit 1501, a quantization unit 1502, an inverse quantization unit 1503, an inverse transform unit 1504, a block memory 1505, a frame memory 1506, an intra prediction unit 1507, an inter prediction unit 1508, a subtraction unit 1521, an addition unit 1522, and an entropy encoding unit 1509. Note that the encoding device 1500 does not include the parameter derivation units 910, 1110, and 1310.

[0228] The above components included in the encoding device 1500 execute the same processes as those in the above-mentioned first to fourth embodiments, but do not execute processes using an adaptive video encoding tool. That is, the adder 1522, the intra prediction unit 1507, and the inter prediction unit 1508 execute processes for encoding without using parameters derived by the parameter derivation units 910, 1110, and 1310 in the second to fourth embodiments, respectively.

[0229] Furthermore, the encoding device 1500 obtains a video and parameters related to the video, generates a bitstream by encoding the video without using the parameters, and writes the above-mentioned parameters into the bitstream. Specifically, the entropy encoding unit 1509 writes the parameters into the bitstream. Note that the parameters may be written into the bitstream at any position.

[0230] Moreover, each image (i.e., picture) included in the above-mentioned video input to the encoding device 1500 may be a distortion-corrected image or a stitched image obtained by stitching together images from multiple views. A distortion-corrected image is a rectangular image obtained by correcting the distortion of an image captured by a wide-angle lens such as a non-rectilinear lens. Such an encoding device 1500 encodes a video including the distortion-corrected image or stitched image.

[0231] Here, the quantization unit 1502, the inverse quantization unit 1503, the inverse transform unit 1504, the intra prediction unit 1507, the inter prediction unit 1508, the subtraction unit 1521, the addition unit 1522, and the entropy coding unit 1509 are configured as, for example, a processing circuit. Furthermore, the block memory 1505 and the frame memory 1506 are configured as memories.

[0232] That is, the encoding device 1500 includes a processing circuit and a memory connected to the processing circuit. The processing circuit uses the memory to obtain parameters related to at least one of a first process for correcting distortion of an image captured by a wide-angle lens and a second process for stitching together a plurality of images, generates an encoded image by encoding the image or an image to be processed based on the plurality of images, and writes the parameters into a bitstream including the encoded image.

[0233] As a result, since the above-mentioned parameters are written into the bit stream, the image being coded or decoded can be appropriately handled by using the parameters.

[0234] Here, the writing of the parameters may include writing the parameters into a header in a bitstream. Also, the coding of the image to be processed may include coding each block included in the image to be processed by applying a coding process based on the parameters to the block. Here, the coding process may include at least one of an inter-picture prediction process and an image reconstruction process.

[0235] Thereby, for example, as in the second embodiment, by using the inter prediction process and the image reconstruction process as an adaptive video coding tool, it is possible to appropriately code a processing target image, for example, a distorted image or a spliced ​​image, and as a result, it is possible to improve the coding efficiency for the processing target image.

[0236] In addition, when writing the parameters, parameters related to the above-mentioned second process may be written into a header in the bit stream, and when encoding the image to be processed, the encoding process for each block included in the image to be processed obtained by the second process may be omitted based on the parameters.

[0237] 21 and 22 in the second embodiment, it is possible to omit coding of blocks included in images that will not be looked at by the user in the near future among a plurality of images included in a stitched image, thereby reducing the processing load and the amount of code.

[0238] In addition, in writing the parameters, at least one of the position and the camera angle of each of the plurality of cameras may be written to a header in the bit stream as a parameter related to the second process. In encoding the image to be processed, the image to be processed, which is one of the plurality of images, may be encoded, and the image to be processed may be spliced ​​with another image of the plurality of images using the parameters written to the header.

[0239] This makes it possible to use a large image obtained by splicing, for example, as in the third embodiment, for inter-picture prediction or motion compensation, thereby improving coding efficiency.

[0240] In addition, in writing the parameters, at least one of a parameter indicating whether an image is captured by a wide-angle lens and a parameter related to distortion caused by a wide-angle lens may be written into a header in a bit stream as a parameter related to the above-mentioned first process. In addition, in encoding the image to be processed, for each block included in the image to be processed, which is an image captured by a wide-angle lens, the encoding process based on the parameters written into the header may be applied to the block, thereby encoding the block. Here, the encoding process may include at least one of a motion vector prediction process and an intra-screen prediction process.

[0241] Thus, for example, as in the fourth embodiment, by using the motion vector prediction process and the intra-frame prediction process as an adaptive video coding tool, it is possible to appropriately code a processing target image, for example, a distorted image, and as a result, it is possible to improve the coding efficiency of the distorted image.

[0242] The encoding process may also include one of an inter-screen prediction process and an intra-screen prediction process, and the prediction process may include a wrapping process that is a process of arranging or rearranging a plurality of pixels included in an image.

[0243] This allows for correcting the distortion of the image to be processed and performing inter-frame prediction processing appropriately based on the corrected image, as in, for example, embodiment 2. Also, as in, for example, embodiment 4, it is possible to perform intra-frame prediction processing on a distorted image and appropriately distort the predicted image obtained by the processing in accordance with the distorted image to be processed. As a result, it is possible to improve the coding efficiency of the distorted image.

[0244] The encoding process may also include an inter-prediction process for curved, diagonal or angular image boundaries, including padding of the image using parameters written in the above-mentioned header.

[0245] This allows inter-picture prediction processing to be performed appropriately, as in the second embodiment, for example, and coding efficiency to be improved.

[0246] In addition, the encoding process may include an inter-screen prediction process and an image reconstruction process, each of which may include a process for replacing pixel values ​​with predetermined values ​​based on parameters written in the above-mentioned header.

[0247] This makes it possible to appropriately perform inter prediction processing and image reconstruction processing, as in the second embodiment, for example, and improve coding efficiency.

[0248] In addition, in encoding the image to be processed, the encoded image to be processed may be reconstructed, and an image obtained by splicing the reconstructed image to be processed with the other image described above may be stored in memory as a reference frame to be used in the inter-screen prediction process.

[0249] This makes it possible to use a large image obtained by splicing, for example, as in the third embodiment, for inter-picture prediction or motion compensation, thereby improving coding efficiency.

[0250] The encoding devices of the above-mentioned second to fourth embodiments encode a video including distorted images, a video including spliced ​​images, or a video including images not spliced ​​from multiple views. However, the encoding device of the present disclosure may correct the distortion of images included in the video in order to encode the video, or may not correct the distortion. When the distortion is not corrected, the encoding device obtains a video including images whose distortion has been corrected in advance by another device, and encodes the video. Similarly, the encoding device of the present disclosure may splice images from multiple views included in the video in order to encode the video, or may not splice images. When the splicing is not performed, the encoding device obtains a video including images in which images from multiple views have been spliced ​​in advance by another device, and encodes the video. Moreover, the encoding device of the present disclosure may perform all or only a part of the distortion correction. Furthermore, the encoding device of the present disclosure may perform all or only a part of the splicing of images from multiple views.

[0251] FIG. 38 is a block diagram of a decoding device according to one embodiment of the present disclosure.

[0252] A decoding device 1600 according to one embodiment of the present disclosure is a device equivalent to the decoding device 200 of embodiment 1, and as shown in FIG. 38, includes an entropy decoding unit 1601, an inverse quantization unit 1602, an inverse transform unit 1603, a block memory 1604, a frame memory 1605, an intra prediction unit 1606, an inter prediction unit 1607, and an adder unit 1622.

[0253] The above components included in the decoding device 1600 execute the same processes as those in the above-mentioned Embodiments 1 to 4, but do not execute processes using an adaptive video decoding tool. That is, the adder 1622, the intra predictor 1606, and the inter predictor 1607 execute processes for decoding without using the above-mentioned parameters included in the bitstream.

[0254] Moreover, the decoding device 1600 obtains a bitstream, extracts coded video and parameters from the bitstream, and decodes the coded video without using the parameters. Specifically, the entropy decoding unit 1601 reads the parameters from the bitstream. Note that the parameters may be written in any position in the bitstream.

[0255] Moreover, each image (i.e., coded picture) included in the bitstream input to the decoding device 1600 may be a distortion-corrected image or a stitched image obtained by stitching together images from multiple views. A distortion-corrected image is a rectangular image obtained by correcting the distortion of an image captured by a wide-angle lens such as a non-rectilinear lens. Such a decoding device 1600 decodes a video including the distortion-corrected image or stitched image.

[0256] Here, the entropy decoding unit 1601, the inverse quantization unit 1602, the inverse transform unit 1603, the intra prediction unit 1606, the inter prediction unit 1607, and the addition unit 1622 are configured as, for example, a processing circuit. Furthermore, the block memory 1604 and the frame memory 1605 are configured as memories.

[0257] That is, the decoding device 1600 includes a processing circuit and a memory connected to the processing circuit. The processing circuit uses the memory to obtain a bit stream including an encoded image, decodes from the bit stream parameters related to at least one of a first process for correcting distortion of an image captured by a wide-angle lens and a second process for stitching together a plurality of images, and decodes the encoded image.

[0258] This allows the image being coded or decoded to be properly handled by using the above parameters interpreted from the bitstream.

[0259] Here, the parameter may be interpreted from a header in a bitstream. Also, the coded image may be decoded by applying a decoding process based on the parameters to each block included in the coded image. Here, the decoding process may include at least one of an inter-picture prediction process and an image reconstruction process.

[0260] As a result, for example, as in embodiment 2, by using inter-picture prediction processing and image reconstruction processing as adaptive video decoding tools, it is possible to appropriately decode encoded images that are, for example, distorted images or spliced ​​images.

[0261] In addition, when interpreting the parameters, the parameters related to the above-mentioned second process are read from the header in the bitstream, and when decoding the encoded image, for each block contained in the encoded image generated by encoding the image obtained by the second process, the decoding process for that block may be omitted based on the parameters.

[0262] 21 and 22 in the second embodiment, it is possible to omit decoding of each block included in an image that will not be looked at by the user in the near future among a plurality of images included in a spliced ​​image that is an encoded image, and as a result, it is possible to reduce the processing load.

[0263] In addition, in the parameter interpretation, at least one of the position and the camera angle of each of the plurality of cameras may be interpreted from a header in the bitstream as a parameter related to the second process. In addition, in the coded image decoding, a coded image generated by coding one of the plurality of images may be decoded, and the decoded coded image may be spliced ​​with another image of the plurality of images using the parameter interpreted from the header.

[0264] This makes it possible to use a large image obtained by splicing, as in the third embodiment, for inter-picture prediction or motion compensation, and to appropriately decode a bitstream with improved coding efficiency.

[0265] In addition, in the parameter interpretation, at least one of a parameter indicating whether an image is captured by a wide-angle lens and a parameter related to distortion caused by a wide-angle lens may be interpreted from a header in a bit stream as a parameter related to the above-mentioned first process.In addition, in the decoding of an encoded image, for each block included in an encoded image generated by encoding an image captured by a wide-angle lens, a decoding process based on the parameter interpreted from the header may be applied to the block, thereby decoding the block.Here, the decoding process may include at least one of a motion vector prediction process and an intra-screen prediction process.

[0266] As a result, for example, as in the fourth embodiment, by using the motion vector prediction process and the intra-frame prediction process as an adaptive video decoding tool, it is possible to appropriately decode a coded image that is, for example, a distorted image.

[0267] The decoding process may include one of an inter-prediction process and an intra-prediction process, and the prediction process may include a wrapping process that is a process of arranging or rearranging a plurality of pixels included in an image.

[0268] This allows for the distortion of an encoded image to be corrected and for inter-prediction processing to be performed appropriately based on the corrected image, as in, for example, embodiment 2. Also, as in, for example, embodiment 4, intra-prediction processing is performed on a distorted encoded image, and the resulting predicted image can be appropriately distorted to match the distorted encoded image. As a result, it is possible to appropriately predict the encoded image, which is a distorted image.

[0269] The decoding process may also include an inter-prediction process for curved, diagonal or angular image boundaries, and may include image padding using parameters read from the above-mentioned header.

[0270] This makes it possible to appropriately perform inter-picture prediction processing, for example, as in the second embodiment.

[0271] In addition, the decoding process may include an inter-frame prediction process and an image reconstruction process, each of which may include a process for replacing pixel values ​​with predetermined values ​​based on parameters interpreted from the above-mentioned header.

[0272] This makes it possible to appropriately perform inter prediction processing and image reconstruction processing, for example, as in the second embodiment.

[0273] In addition, when decoding a coded image, the coded image may be decoded, and an image obtained by splicing the decoded coded image with the other image described above may be stored in memory as a reference frame to be used in the inter-picture prediction process.

[0274] This makes it possible to use a large image obtained by splicing, for example, for inter-picture prediction or motion compensation, as in the third embodiment.

[0275] Note that the decoding devices of the above-mentioned second to fourth embodiments decode a bitstream including a distorted image, a bitstream including a spliced ​​image, or a bitstream including images from multiple views that are not spliced ​​together. However, the decoding device of the present disclosure may correct the distortion of the images included in the bitstream in order to decode the bitstream, or may not correct the distortion. When the distortion is not corrected, the decoding device obtains a bitstream including an image whose distortion has been corrected in advance by another device, and decodes the bitstream. Similarly, the decoding device of the present disclosure may splice images from multiple views included in the bitstream in order to decode the bitstream, or may not splice images. When the splicing is not performed, the decoding device obtains a bitstream including a large image generated in advance by splicing images from multiple views together by another device, and decodes the bitstream. Also, the decoding device of the present disclosure may perform all or only a part of the distortion correction. Furthermore, the decoding device of the present disclosure may perform all or only a part of the splicing of images from multiple views.

[0276] (Other embodiments) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, etc. Furthermore, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (programs) recorded on a recording medium such as a ROM. The software may be distributed by downloading, etc., or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, each functional block can also be realized by hardware (dedicated circuitry).

[0277] Furthermore, the processes described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. The processor that executes the above program may be either single or multiple. That is, centralized processing or distributed processing may be performed.

[0278] The present invention is not limited to the above-mentioned embodiments, and various modifications are possible, which are also included within the scope of the present invention.

[0279] Further, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, and an image coding / decoding device equipped with both. Other configurations in the system can be appropriately changed depending on the case.

[0280] [Usage example] 39 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.

[0281] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may be configured to connect any of the above elements in combination. Each device may be directly or indirectly connected to each other via a telephone network or short-distance wireless communication, without going through the base stations ex106 to ex110, which are fixed wireless stations. In addition, the streaming server ex103 is connected to each device such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101, etc. In addition, the streaming server ex103 is connected to a terminal in a hot spot in an airplane ex117, etc., via a satellite ex116.

[0282] Instead of the base stations ex106 to ex110, wireless access points or hot spots may be used. The streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0283] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that is compatible with a mobile communication system generally called 2G, 3G, 3.9G, 4G, and 5G in the future.

[0284] The home appliance ex118 is a refrigerator or an appliance included in a home fuel cell cogeneration system.

[0285] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live distribution and the like. In live distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) performs the encoding process described in each of the above embodiments on still image or video content photographed by a user using the terminal, multiplexes the video data obtained by encoding with sound data obtained by encoding sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.

[0286] On the other hand, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal in an airplane ex117, or the like, capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device functions as an image decoding device according to one aspect of the present invention.

[0287] [Distributed processing] The streaming server ex103 may be a plurality of servers or computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content distribution may be realized by a network that connects a large number of edge servers distributed around the world. In a CDN, an edge server that is physically close to the client is dynamically assigned according to the client. The content is cached and distributed to the edge server, thereby reducing delays. In addition, when an error occurs or the communication state changes due to an increase in traffic, the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the part of the network where a failure has occurred, thereby realizing high-speed and stable distribution.

[0288] In addition to the distributed processing of the distribution itself, the encoding processing of the captured data may be performed by each terminal, may be performed by the server side, or may be shared among the terminals. As an example, in the encoding processing, a processing loop is generally performed twice. In the first loop, the complexity of the image or the amount of code is detected for each frame or scene. In the second loop, processing is performed to maintain the image quality and improve the encoding efficiency. For example, the terminal performs the first encoding processing, and the server side that receives the content performs the second encoding processing, thereby improving the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode almost in real time, the data encoded once by the terminal can be received and played back by other terminals, making it possible to perform more flexible real-time distribution.

[0289] As another example, the camera ex113 etc. extracts features from an image, compresses data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning of the image, for example, by determining the importance of an object from the features and switching the quantization precision. The feature data is particularly effective in improving the precision and efficiency of motion vector prediction when the server performs recompression. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a large processing load such as CABAC (context-adaptive binary arithmetic coding).

[0290] As another example, in a stadium, a shopping mall, a factory, etc., there may be a plurality of video data in which almost the same scene has been shot by a plurality of terminals. In this case, using the plurality of terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, coding processing is assigned to each of them, for example, in units of GOPs (Group of Pictures), in units of pictures, or in units of tiles obtained by dividing a picture, for distributed processing. This reduces delays and realizes better real-time performance.

[0291] In addition, since the multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referenced. Alternatively, the server may receive the encoded data from each terminal and change the reference relationship between the multiple data, or correct or replace the pictures themselves and re-encode them. This makes it possible to generate a stream with improved quality and efficiency for each piece of data.

[0292] The server may also perform transcoding to change the encoding format of the video data before distributing it. For example, the server may convert an MPEG-based encoding format into a VP-based encoding format, or convert H.264 into H.265.

[0293] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, descriptions such as "server" or "terminal" are used to indicate the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0294] [3D, multi-angle] In recent years, it has become common to integrate and use images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are almost synchronized with each other. The videos taken by the devices are integrated based on the relative positional relationship between the devices that is obtained separately, or on areas where feature points included in the videos match.

[0295] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. If the server can obtain the relative positional relationship between the shooting terminals, the server may generate a 3D shape of the scene based on not only 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point clouds, etc., or may generate images to be transmitted to the receiving terminal by selecting or reconstructing images from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0296] In this way, the user can enjoy a scene by selecting any video corresponding to each shooting terminal, or can enjoy content in which a video from any viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, sound may be collected from multiple different angles, just like the video, and the server may multiplex the sound from a specific angle or space with the video and transmit it in accordance with the video.

[0297] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and the left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0298] In the case of an AR image, the server superimposes virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may obtain or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect them to create superimposed data. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to a request for virtual object information, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and deliver it to the decoding device. Note that the superimposed data has an α value indicating the transparency in addition to RGB, and the server may set the α value of the part other than the object created from the three-dimensional data to 0, etc., and encode the data in a state in which the part is transparent. Alternatively, the server may generate data in which a predetermined value of RGB value is set to the background like a chromakey, and the part other than the object is the background color.

[0299] Similarly, the decoding process of the distributed data may be performed by each client terminal, or may be performed by the server side, or may be shared among them. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request may be received by other terminals, decoded, and the decoded signal may be transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-enabled terminals themselves, data with good image quality can be reproduced. In another example, while large-sized image data is received by a TV or the like, a part of the area, such as tiles into which the picture is divided, may be decoded and displayed on the viewer's personal terminal. This allows the viewer to share the overall picture while checking his / her own area of ​​responsibility or the area he / she wants to check in more detail at hand.

[0300] In the future, it is expected that content will be seamlessly received by switching appropriate data for the currently connected communication using delivery system standards such as MPEG-DASH under circumstances where multiple short-distance, medium-distance, or long-distance wireless communication is available, regardless of whether indoors or outdoors. This allows users to freely select and switch in real time not only their own terminals but also decoding devices or display devices such as displays installed indoors and outdoors. In addition, decoding can be performed while switching the decoding device and the display device based on the user's location information, etc. This makes it possible to move while displaying map information on the wall or part of the ground of a neighboring building where a displayable device is embedded while moving to a destination. It is also possible to switch the bit rate of the received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be accessed from the receiving terminal in a short time, or when it is copied to an edge server in a content delivery service.

[0301] [Scalable Coding] The switching of contents will be described using a scalable stream compressed and coded by applying the video coding method shown in each of the above embodiments, as shown in FIG. 40. The server may have multiple streams with the same content but different qualities as individual streams, but may be configured to switch contents by taking advantage of the characteristics of a temporal / spatial scalable stream realized by coding in layers as shown in the figure. In other words, the decoding side can freely switch and decode low-resolution content and high-resolution content by determining which layer to decode according to an internal factor such as performance and an external factor such as the state of the communication band. For example, if you want to continue watching a video you were watching on your smartphone ex115 while on the move on a device such as an Internet TV after you get home, the device can decode the same stream up to a different layer, reducing the burden on the server side.

[0302] Furthermore, as described above, in addition to the configuration that realizes scalability in which pictures are coded for each layer and an enhancement layer exists above a base layer, the enhancement layer may include meta-information based on image statistics, etc., and the decoding side may generate high-quality content by super-resolving pictures of the base layer based on the meta-information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta-information includes information for specifying linear or nonlinear filter coefficients used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least squares calculation used in the super-resolution process.

[0303] Alternatively, a picture may be divided into tiles or the like according to the meaning of an object in an image, and the decoding side may decode only a part of the area by selecting a tile to be decoded. Also, by storing the attribute of an object (person, car, ball, etc.) and its position in a video (coordinate position in the same image, etc.) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 41, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of a main object.

[0304] Meta information may also be stored in units consisting of multiple pictures, such as streams, sequences, or random access units, etc. This allows the decoding side to obtain the time when a specific person appears in the video, and by combining this with picture-by-picture information, it is possible to identify the picture in which an object exists and the position of the object within the picture.

[0305] [Web page optimization] FIG. 42 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. FIG. 43 is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As shown in FIG. 42 and FIG. 43, a web page may include multiple link images that are links to image content, and the appearance of the web page differs depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) displays a still image or I-picture that each content has as a link image, displays a video such as a gif animation using multiple still images or I-pictures, or receives only the base layer to decode and display the video, until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image enters the screen.

[0306] When a link image is selected by a user, the display device gives top priority to decoding the base layer. If the HTML constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. In order to ensure real-time performance, before selection or when the communication bandwidth is very tight, the display device decodes and displays only forward-reference pictures (I-pictures, P-pictures, and B-pictures with forward reference only), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of decoding the content to the start of display). The display device may also ignore the reference relationship between pictures and roughly decode all B-pictures and P-pictures with forward reference, and perform normal decoding as the number of received pictures increases over time.

[0307] [Automatic driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0308] In this case, since a car, drone, or airplane including a receiving terminal moves, the receiving terminal can realize seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting the location information of the receiving terminal at the time of a reception request. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information according to the user's selection, the user's situation, or the state of the communication band.

[0309] In this manner, in the content supply system ex100, the client can receive, decode, and play back encoded information transmitted by the user in real time.

[0310] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distributors, but also low-quality, short-duration content from individuals via unicast or multicast distribution. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing processing before encoding processing. This can be achieved, for example, by the following configuration.

[0311] During shooting, in real time or after accumulating, the server performs recognition processing such as shooting errors, scene search, semantic analysis, and object detection from the original image or encoded data. Then, based on the recognition result, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes that are less bright than other pictures or are out of focus, emphasizes object edges, changes color, and performs other editing. The server encodes the edited data based on the editing result. It is also known that if the shooting time is too long, the viewer rating will decrease, and the server may automatically clip not only scenes of less importance as described above but also scenes with little movement based on the image processing result so that the content will be within a specific time range depending on the shooting time. Alternatively, the server may generate a digest based on the result of the semantic analysis of the scene and encode it.

[0312] In addition, there are cases where personal contents contain images that infringe copyrights, moral rights, portrait rights, etc., and the scope of sharing may exceed the intended scope, which may be inconvenient for individuals. Therefore, for example, the server may change the image to an unfocused image of a person's face on the periphery of the screen, or the inside of a house, and encode it. The server may also recognize whether the image to be encoded contains a face of a person other than a person registered in advance, and if so, may perform processing such as blurring the face. Alternatively, as pre-processing or post-processing of encoding, the user may specify a person or background area that he or she wishes to process in the image from the viewpoint of copyright, etc., and the server may replace the specified area with another image or blur the focus. If it is a person, the image of the face part can be replaced while tracking the person in the video.

[0313] In addition, since viewing of personal content with a small amount of data requires real-time performance, the decoding device first receives the base layer as a top priority, and performs decoding and playback, although this depends on the bandwidth. The decoding device may receive an enhancement layer during this time, and when playback is looped or otherwise played two or more times, play high-quality video including the enhancement layer. In this way, if the stream is scalably encoded, it is possible to provide an experience in which the video is rough when not selected or when viewing begins, but the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played the first time and a second stream encoded with reference to the first video are configured as one stream.

[0314] [Other use cases] Moreover, these encoding or decoding processes are generally processed in an LSIex500 possessed by each terminal. The LSIex500 may be a single chip or may be configured with multiple chips. Note that software for encoding or decoding moving images may be incorporated into some kind of recording medium (such as a CD-ROM, a flexible disk, or a hard disk) that can be read by the computer ex111 or the like, and the encoding or decoding process may be performed using the software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera may be transmitted. The video data in this case is data that has been encoded by the LSIex500 possessed by the smartphone ex115.

[0315] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0316] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is carried and transmitted over broadcasting radio waves using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to the encoding process and decoding process.

[0317] [Hardware configuration] FIG. 44 is a diagram showing a smartphone ex115. FIG. 45 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying the video captured by the camera unit ex465 and the decoded data of the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded data such as captured video or still images, recorded audio, received video or still images, and e-mail, or decoded data, and a slot unit ex464 which is an interface unit with a SIMex468 for identifying a user and authenticating access to various data including a network. In addition, an external memory may be used instead of the memory unit ex467.

[0318] In addition, a main control unit ex460, which comprehensively controls the display unit ex458 and the operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a bus ex470.

[0319] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state by supplying power to each unit from the battery pack.

[0320] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, a ROM, and a RAM. During a telephone call, a voice signal collected by a voice input unit ex456 is converted into a digital voice signal by a voice signal processing unit ex454, which is then subjected to spectrum spreading processing by a modulation / demodulation unit ex452, and the digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit ex451 is then transmitted via an antenna ex450. In addition, the received data is amplified, and subjected to frequency conversion processing and analog-to-digital conversion processing, and the spectrum inverse spreading processing by a modulation / demodulation unit ex452 is then performed, and the analog voice signal is converted into an analog voice signal by a voice signal processing unit ex454, which is then output from a voice output unit ex457. During a data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 by operating an operation unit ex466 or the like of the main unit, and transmission and reception processing is performed in the same manner. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / separation unit ex453. The audio signal processing unit ex454 also encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing the video or still images, and sends the encoded audio data to the multiplexing / separation unit ex453. The multiplexing / separation unit ex453 multiplexes the encoded video data and the encoded audio data by a predetermined method, and performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.

[0321] When receiving a video attached to an e-mail or chat, or a video linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / separation unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and displays the video or still image contained in the linked video file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 also decodes the audio signal, and audio is output from the audio output unit ex457. Note that since real-time streaming is widespread, there may be cases where audio playback is socially inappropriate depending on the user's situation. Therefore, as an initial value, a configuration in which only video data is played without playing audio signals is preferable. The audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0322] In addition, although the smartphone ex115 has been described as an example here, three types of implementation formats are possible for the terminal: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcasting system, multiplexed data in which music data and the like are multiplexed with video data is received or transmitted, but the multiplexed data may include text data related to the video in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0323] Although the main control unit ex460 including the CPU controls the encoding or decoding process, terminals often have a GPU. Therefore, a configuration may be used in which a wide area is processed collectively by utilizing the performance of the GPU using a memory shared by the CPU and GPU, or a memory whose addresses are managed so that they can be used in common. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processing of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transformation and quantization collectively in units such as pictures by the GPU, rather than by the CPU. [Industrial Applicability]

[0324] The present disclosure can be applied to devices such as televisions, digital video recorders, car navigation systems, mobile phones, digital cameras, or digital video cameras, including encoding devices that encode images or decoding devices that decode encoded images. [Explanation of symbols]

[0325] 1500 encoding device 1501 Conversion unit 1502 Quantization section 1503 Inverse quantization section 1504 Inverse conversion unit 1505 Block Memory 1506 Frame Memory 1507 Intra prediction unit 1508 Inter Prediction Unit 1509 Entropy coding unit 1521 Subtraction section 1522 Addition section 1600 Decoding Device 1601 Entropy Decoding Unit 1602 Inverse quantization section 1603 Inverse conversion unit 1604 Block Memory 1605 Frame Memory 1606 Intra prediction unit 1607 Inter Prediction Unit 1622 Addition section

Claims

1. A processing circuit; a memory coupled to the processing circuit; The processing circuitry uses the memory to: A stitched image is generated by stitching multiple images together. obtaining a parameter that identifies an empty area in the stitched image that is generated by the stitching process; performing an inter-picture prediction process on the stitched image; writing said parameters into a bitstream; the inter-prediction process includes a padding process of replacing pixel values ​​in the free space with values ​​in another area in the stitched image that is not the free space, the value of the other region is the value of the pixel closest to the free region, The inter-screen prediction process is performed on an image block basis. Encoding device.

2. A processing circuit; a memory coupled to the processing circuit; The processing circuitry uses the memory to: Obtaining parameters from the bitstream that identify free space that will be generated by a stitching process that stitches together a plurality of images; By performing the stitching process, a stitched image is generated; performing an inter-picture prediction process on the stitched image; the inter-prediction process includes a padding process of replacing pixel values ​​in the free space with values ​​in another area in the stitched image that is not the free space, the value of the other region is the value of the pixel closest to the free region, The inter-screen prediction process is performed on an image block basis. Decryption device.