Encoding method and decoding method
By obtaining parameters related to the image bonding process in the encoding device, generating the bonded image and performing entropy encoding, the problem of low encoding efficiency in the prior art is solved, and more efficient image coding and decoding are achieved.
Patent Information
- Application Number
- CN202210528629.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-05-27
- Filing Date
- 2017-05-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2037-05-23
AI Technical Summary
The existing encoding devices and decoding devices are difficult to properly handle the encoded or decoded images, resulting in low encoding efficiency.
By obtaining parameters related to multiple image bonding processes, these parameters are written to the bitstream, and using these parameters to bond the multiple images, the objects in the image are continued, the bonded image is generated, and the bitstream is generated by entropy encoding and the bonded image is saved as a reference frame.
The appropriate handling of encoded or decoded images is achieved, and encoding efficiency is improved, especially when processing 360 degree moving images and images captured by non-linear lenses.
Smart Images

Figure CN114979649B_ABST
Abstract
Description
[0001] This application is a divisional of an invention patent application filed on May 23, 2017, with application number 201780031614.8 and invention name “Encoding device, decoding device, encoding method and decoding method”. Technical Field
[0002] The present disclosure relates to a device and method for encoding an image, and a device and method for decoding an encoded image. Background Art
[0003] Currently, HEVC has been established as a standard for image coding (for example, refer to non-patent document 1). However, in the transmission and storage of next-generation videos (for example, 360-degree moving images), coding efficiency that exceeds the current coding performance is required. In addition, some research and experiments related to the compression of moving images captured by wide-angle lenses such as non-rectilinear lenses have been conducted so far. In these studies, distortion aberrations are eliminated by manipulating image samples to make the image of the processing object linear before encoding it. For this purpose, image processing technology is generally used.
[0004] Prior art literature
[0005] Non-patent literature
[0006] Non-patent document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the invention
[0007] Problems to be solved by the invention
[0008] However, conventional encoding devices and decoding devices have a problem of being unable to appropriately handle encoded or decoded images.
[0009] Therefore, the present disclosure provides an encoding device and the like that can appropriately handle an encoded or decoded image.
[0010] Means used to solve problems
[0011] A coding method for a technical solution related to the present disclosure obtains parameters related to a process of joining a plurality of images; writes the parameters into a bit stream; generates a joined image by joining the plurality of images using the parameters so that objects in the plurality of images are continuous; generates the bit stream by entropy coding; and saves the joined image as a reference frame to be used in an inter-frame prediction process, wherein the parameters are used to determine a plurality of positions of the plurality of images in the joined image.
[0012] In addition, these inclusive or specific technical solutions can also be implemented by systems, methods, integrated circuits, computer programs or non-temporary recording media such as computer-readable CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs and recording media.
[0013] Effects of the Invention
[0014] The encoding device of the present disclosure can appropriately handle encoded or decoded images. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a block diagram showing the functional structure of the encoding device related to embodiment 1.
[0016] Figure 2 This is a diagram showing an example of block division in implementation mode 1.
[0017] Figure 3 This is a table showing the transformation basis functions corresponding to each transformation type.
[0018] Figure 4A This is a diagram showing an example of the shape of a filter used in ALF.
[0019] Figure 4B This is a diagram showing another example of the shape of the filter used in ALF.
[0020] Figure 4C This is a diagram showing another example of the shape of the filter used in ALF.
[0021] Figure 5 This is a diagram showing 67 intra prediction modes of intra prediction.
[0022] Figure 6 This is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0023] Figure 7 This is a diagram used to explain pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0024] Figure 8 This is a diagram for explaining a model assuming uniform linear motion.
[0025] Fig. 9 This is a diagram for explaining derivation of a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks.
[0026] Fig.10 This is a block diagram showing the functional structure of a decoding device according to Embodiment 1.
[0027] Fig.11This is a flowchart showing an example of motion picture encoding processing according to the second embodiment.
[0028] Fig.12 This is a diagram showing possible positions of headers in which parameters can be written in the bit stream of the second embodiment.
[0029] Fig.13 It is a diagram showing a captured image according to the second embodiment and a processed image after image correction processing.
[0030] Fig.14 This is a diagram showing a joined image generated by joining a plurality of images through a joining process according to the second embodiment.
[0031] Fig.15 It is a diagram showing the arrangement of a plurality of cameras according to the second embodiment and a joined image including a blank area generated by joining images captured by these cameras.
[0032] Fig.16 This is a flowchart showing the inter-frame prediction processing or motion compensation according to the second embodiment.
[0033] Fig.17 This is a diagram showing an example of barrel distortion generated by a non-linear lens or a fish-eye lens in the second embodiment.
[0034] Fig.18 This is a flowchart showing a variation of the inter-frame prediction process or motion compensation according to the second embodiment.
[0035] Fig.19 This is a flowchart showing the image reconstruction processing according to the second embodiment.
[0036] Fig. 20 This is a flowchart showing a modified example of the image reconstruction process according to the second embodiment.
[0037] Fig.21 This is a diagram showing an example of partial encoding processing or partial decoding processing on a joined image according to the second embodiment.
[0038] Fig. 22 This is a diagram showing another example of partial encoding processing or partial decoding processing on a joined image according to the second embodiment.
[0039] Fig.23 This is a block diagram of an encoding device according to embodiment 2.
[0040] Fig.24 This is a flowchart showing an example of motion picture decoding processing according to the second embodiment.
[0041] Fig.25 This is a block diagram of a decoding device according to embodiment 2.
[0042] Fig.26 This is a flowchart showing an example of motion picture encoding processing according to the third embodiment.
[0043] Fig. 27 This is a flowchart showing an example of the joining process according to the third embodiment.
[0044] Fig.28 This is a block diagram of an encoding device according to embodiment 3.
[0045] Fig.29 This is a flowchart showing an example of a motion picture decoding process according to the third embodiment.
[0046] Fig.30 This is a block diagram of a decoding device according to Embodiment 3.
[0047] Fig.31 This is a flowchart showing an example of motion picture encoding processing according to the fourth embodiment.
[0048] Fig.32 This is a flowchart showing the intra-frame prediction processing of implementation mode 4.
[0049] Fig.33 This is a flowchart showing the motion vector prediction processing of implementation mode 4.
[0050] Fig.34 This is a block diagram of an encoding device according to embodiment 4.
[0051] Fig.35 This is a flowchart showing an example of motion picture decoding processing according to the fourth embodiment.
[0052] Fig.36 This is a block diagram of a decoding device according to embodiment 4.
[0053] Fig.37 This is a block diagram of an encoding device according to one aspect of the present disclosure.
[0054] Fig.38 This is a block diagram of a decoding device according to one aspect of the present disclosure.
[0055] Fig.39 It is an overall structural diagram of the content supply system that realizes content distribution services.
[0056] Fig.40 This is a diagram showing an example of a coding structure in scalable coding.
[0057] Fig.41 This is a diagram showing an example of a coding structure in hierarchical coding.
[0058] Fig.42 This is a diagram showing an example of a display screen of a web page.
[0059] Fig.43 This is a diagram showing an example of a display screen of a web page.
[0060] Fig.44 The diagram shows an example of a smart phone.
[0061] Fig.45 This is a block diagram showing a structural example of a smart phone. DETAILED DESCRIPTION
[0062] Hereinafter, the embodiments will be described in detail with reference to the drawings.
[0063] In addition, the embodiments described below are all inclusive or specific examples. The numerical values, shapes, materials, components, configuration positions and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and do not limit the meaning of the claims. In addition, among the components of the following embodiments, the components that are not recorded in the independent claims representing the highest concept are described as arbitrary components.
[0064] (Implementation Method 1)
[0065] [Overview of Encoding Device]
[0066] First, an overview of the encoding device according to Embodiment 1 will be described. Figure 1 1 is a block diagram showing a functional structure of a coding apparatus 100 according to Embodiment 1. The coding apparatus 100 is a moving picture / image coding apparatus that encodes a moving picture / image in units of blocks.
[0067] like Figure 1 As shown, the encoding device 100 is a device that encodes an image in block units, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126 and a prediction control unit 128.
[0068] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.
[0069] Hereinafter, each component included in the encoding device 100 will be described.
[0070] [Division Department]
[0071] The division unit 102 divides each picture included in the input motion image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides the picture into blocks of a fixed size (e.g., 128×128). The fixed-size block is sometimes referred to as a coding tree unit (CTU). Furthermore, the division unit 102 divides the fixed-size blocks into blocks of a variable size (e.g., less than 64×64) based on a recursive quadtree and / or binary tree block division. The variable-size block is sometimes referred to as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In addition, in the present embodiment, it is not necessary to distinguish between CU, PU, and TU, and part or all of the blocks in the picture may become the processing units of CU, PU, or TU.
[0072] Figure 2 FIG. 1 is a diagram showing an example of block division in Implementation 1. Figure 2 In the figure, the solid lines represent the block boundaries of the quadtree block partition, and the dotted lines represent the block boundaries of the binary tree block partition.
[0073] Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into four square 64×64 blocks (quadtree block division).
[0074] The upper left 64×64 block is further divided vertically into two rectangular 32×64 blocks, and the left 32×64 block is further divided vertically into two rectangular 16×64 blocks (binary tree block division). As a result, the upper left 64×64 block is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.
[0075] The upper right 64×64 block is horizontally divided into two rectangular 64×32 blocks 14 and 15 (binary tree block division).
[0076] The lower left 64×64 block is divided into 4 square 32×32 blocks (quadtree block division). The upper left block and the lower right block of the 4 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into 2 rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into 2 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into 2 32×16 blocks (binary tree block division). As a result, the lower left 64×64 block is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20 and two 32×16 blocks 21, 22.
[0077] The lower right 64×64 block 23 is not divided.
[0078] As above, in Figure 2 In the example, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block partitioning. Such a partitioning is sometimes referred to as QTBT (quad-tree plus binary tree) partitioning.
[0079] In addition, Figure 2 In the example, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary tree block division). Sometimes the division including such ternary tree block division is called MBT (multi type tree) division.
[0080] [Subtraction Department]
[0081] The subtracting unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in units of blocks divided by the dividing unit 102. That is, the subtracting unit 104 calculates the prediction error (also referred to as residual) of the encoding target block (hereinafter referred to as the current block). Then, the subtracting unit 104 outputs the calculated prediction error to the transforming unit 106.
[0082] The original signal is an input signal of the encoding device 100 and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, a signal representing an image may also be referred to as a sample.
[0083] [Conversion Department]
[0084] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0085] Alternatively, the transform unit 106 may adaptively select a transform type from a plurality of transform types and transform the prediction error into a transform coefficient using a transform basis function corresponding to the selected transform type. Such a transform is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0086] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. Figure 3 Where N represents the number of input pixels. The selection of a transform type from among these multiple transform types may be based on, for example, the type of prediction (intra prediction and inter prediction) or may depend on the intra prediction mode.
[0087] Information indicating whether such EMT or AMT is applied (e.g., called an AMT flag) and information indicating the selected transform type are signaled at the CU level. In addition, the signaling of such information does not need to be limited to the CU level, but can also be other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0088] In addition, the transformation unit 106 may also re-transform the transformation coefficient (transformation result). Such re-transformation is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 re-transforms each sub-block (for example, a 4×4 sub-block) contained in the block of transformation coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information about the transformation matrix used in NSST are signaled at the CU level. In addition, the signaling of this information does not necessarily need to be limited to the CU level, but can also be other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0089] [Quantitative Department]
[0090] The quantization unit 108 quantizes the transform coefficient output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficient of the current block in a predetermined scanning order, and quantizes the transform coefficient based on a quantization parameter (QP) corresponding to the scanned transform coefficient. Next, the quantization unit 108 outputs the quantized transform coefficient of the current block (hereinafter referred to as quantized coefficient) to the entropy coding unit 110 and the inverse quantization unit 112.
[0091] The prescribed order is the order for quantization / inverse quantization of transform coefficients. For example, the prescribed scanning order is defined in ascending order (in order from low frequency to high frequency) or descending order (in order from high frequency to low frequency) of frequency.
[0092] The so-called quantization parameter is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.
[0093] [Entropy coding unit]
[0094] The entropy coding unit 110 generates a coded signal (coded bit stream) by performing variable-length coding on the quantized coefficients that are input from the quantization unit 108. Specifically, the entropy coding unit 110 binarizes the quantized coefficients and arithmetically codes the binary signal, for example.
[0095] [Inverse Quantization Unit]
[0096] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. Furthermore, the inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114.
[0097] [Inverse transformation unit]
[0098] The inverse transform unit 114 restores the prediction error by inversely transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform performed by the transform unit 106. The prediction error restored by the inverse transform unit 114 is output to the addition unit 116.
[0099] In addition, since the restored prediction error loses information due to quantization, it does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error includes a quantization error.
[0100] [Addition Department]
[0101] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction signal input from the prediction control unit 128. The adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a local decoded block.
[0102] [Block Memory]
[0103] The block memory 118 is a storage unit for storing blocks in a picture to be coded (hereinafter referred to as a current picture) which are blocks to be referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116 .
[0104] [Loop filter section]
[0105] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. The so-called loop filter is a filter used in the encoding loop (intra-loop filter), and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0106] In ALF, a least square error filter is used to remove coding distortion. For example, one filter is selected from a plurality of filters based on the direction and activity of the local gradient for each 2×2 sub-block in the current block.
[0107] Specifically, first, the sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The sub-blocks are classified based on the direction and activity of the gradient. For example, the classification value C (e.g., C=5D+A) is calculated using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4). And, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).
[0108] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (for example, horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantizing the addition result.
[0109] Based on the result of such classification, a filter to be used for a sub-block is determined from a plurality of filters.
[0110] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is adopted. Figure 4A to Figure 4C The diagrams show a plurality of examples of filter shapes used in ALF. Figure 4A represents a 5×5 diamond filter, Figure 4B represents a 7×7 diamond filter, Figure 4C Represents a 9×9 diamond filter. The information representing the shape of the filter is signaled at the picture level. In addition, the signaling of the information representing the shape of the filter does not need to be limited to the picture level, and can also be other levels (for example, sequence level, slice level, tile level, CTU level or CU level).
[0111] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, regarding brightness, whether to apply ALF is determined at the CU level, and regarding color difference, whether to apply ALF is determined at the picture level. The information indicating the on / off of ALF is signaled at the picture level or the CU level. In addition, the signaling of the information indicating the on / off of ALF does not need to be limited to the picture level or the CU level, and can also be other levels (for example, the sequence level, the slice level, the tile level, or the CTU level).
[0112] The coefficient sets of a plurality of selectable filters (e.g., filters within 15 or 25) are signaled at the picture level. In addition, the signaling of the coefficient set does not need to be limited to the picture level, but may also be other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0113] [Frame Memory]
[0114] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120 .
[0115] [Intra-frame prediction unit]
[0116] The intra prediction unit 124 performs intra prediction (also referred to as intra-screen prediction) of the current block by referring to the blocks in the current picture stored in the block memory 118, and generates a prediction signal (intra prediction signal). Specifically, the intra prediction unit 124 performs intra prediction by referring to samples (e.g., luminance values, color difference values) of blocks adjacent to the current block, and generates the intra prediction signal, and outputs the intra prediction signal to the prediction control unit 128.
[0117] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0118] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0119] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 5 This diagram shows 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows indicate 33 directions specified by the H.265 / HEVC standard, and dotted arrows indicate 32 additional directions.
[0120] In addition, in the intra-frame prediction of the chrominance block, the luminance block can also be referenced. That is, the chrominance component of the current block can also be predicted based on the luminance component of the current block. Such intra-frame prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra-frame prediction mode of the chrominance block with reference to the luminance block (for example, called CCLM mode) can also be added as one of the intra-frame prediction modes of the chrominance block.
[0121] The intra prediction unit 124 may also correct the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction. Intra prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intraprediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is signaled, for example, at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but may also be other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0122] [Inter-frame prediction unit]
[0123] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-screen prediction) of the current block by referring to a reference picture stored in the frame memory 122 and different from the current picture, and generates a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 126 performs motion estimation (motion estimation) on the current block or sub-block in the reference picture. In addition, the inter-frame prediction unit 126 performs motion compensation using motion information (e.g., motion vector) obtained by motion estimation, and generates an inter-frame prediction signal for the current block or sub-block. In addition, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0124] The motion information used in motion compensation is signaled. In the signaling of the motion vector, a predicted motion vector (motion vector predictor) may be used. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0125] In addition, the inter-frame prediction signal may be generated using not only the motion information of the current block obtained by motion estimation but also the motion information of the adjacent blocks. Specifically, the inter-frame prediction signal may be generated in sub-block units within the current block by weighted addition of the prediction signal based on the motion information obtained by motion estimation and the prediction signal based on the motion information of the adjacent blocks. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0126] In such an OBMC mode, information indicating the size of a sub-block used for OBMC (e.g., OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is applied (e.g., OBMC flag) is signaled at the CU level. In addition, the signaling level of this information does not need to be limited to the sequence level and the CU level, and may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0127] Alternatively, the motion information may be derived on the decoding device side without being signaled. For example, a merge mode specified in the H.265 / HEVC standard may be used. Alternatively, the motion information may be derived, for example, by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.
[0128] Here, a mode for performing motion estimation on the decoding device side is described. The mode for performing motion estimation on the decoding device side is sometimes referred to as a PMMVD (pattern matched motion vector derivation) mode or a FRUC (frame rate up-conversion) mode.
[0129] First, one of the candidates included in the merge list is selected as the starting position for the pattern matching search. As the pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.
[0130] In the first pattern matching, pattern matching is performed between two blocks that are two blocks in two different reference pictures and that are along the motion trajectory of the current block.
[0131] Figure 6 This is a diagram used to illustrate pattern matching (bidirectional matching) between two blocks along a motion trajectory. Figure 6 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block).
[0132] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0133] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (eg, upper and / or left adjacent blocks)) and a block in the reference picture.
[0134] Figure 7 is a diagram used to illustrate pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 7 As shown, in the second pattern matching, the motion vector of the current block is derived by searching the reference picture (Ref0) for a block that best matches the block adjacent to the current block (Cur block) in the current picture (CurPic).
[0135] Information indicating whether such a FRUC mode is applied (e.g., called a FRUC flag) is signaled at the CU level. In addition, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the method of pattern matching (first pattern matching or second pattern matching) (e.g., called a FRUC mode flag) is signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and may be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0136] Alternatively, the motion information may be derived on the decoding device side using a method different from motion estimation. For example, the correction amount of the motion vector may be calculated pixel by pixel using surrounding pixel values based on a model assuming uniform linear motion.
[0137] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as a BIO (bi-directional optical flow) mode.
[0138] Figure 8 This is a diagram used to illustrate a model that assumes uniform linear motion. Figure 8 In (v x , v y) represents the velocity vector, τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.
[0139] At this time, the velocity vector (v x , v y ) is assumed to be a uniform linear motion, (MVx0, MVy0) and (MVx1, MVy1) are expressed as (v x τ0,v y τ0) and (-v x τ1, -v y τ1), the following optical flow equation (1) holds.
[0140] [Formula 1]
[0141]
[0142] Here, I (k) Represents the brightness value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation indicates that the sum of (i) the temporal differential of the brightness value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of the block unit obtained from the merge list, etc. is corrected in pixel units.
[0143] In addition, the motion vector may be derived on the decoding device side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0144] Here, a mode for deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes referred to as an affine motion compensation prediction mode.
[0145] Fig. 9 FIG. 1 is a diagram for explaining the derivation of a motion vector in a sub-block unit based on motion vectors of a plurality of adjacent blocks. Fig. 9In the example, the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vector of the adjacent block, and the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vector of the adjacent sub-block. Furthermore, using the two motion vectors v0 and v1, the motion vectors (v x , v y ).
[0146] [Formula 2]
[0147]
[0148] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a preset weight coefficient.
[0149] In such an affine motion compensation prediction mode, the method of deriving the motion vectors of the upper left and upper right control points may also include several different modes. Information indicating such an affine motion compensation prediction mode (e.g., called an affine flag) is signaled at the CU level. In addition, the signaling of the information indicating the affine motion compensation prediction mode does not need to be limited to the CU level, and may also be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0150] [Prediction Control Department]
[0151] The prediction control unit 128 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.
[0152] [Overview of Decoding Device]
[0153] Next, an overview of a decoding device that can decode the coded signal (coded bit stream) output from the coding device 100 described above will be described. Fig.10 2 is a block diagram showing a functional structure of a decoding device 200 according to Embodiment 1. The decoding device 200 is a moving picture / image decoding device that decodes a moving picture / image in units of blocks.
[0154] like Fig.10 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.
[0155] The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. In addition, the decoding device 200 may also be implemented by one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0156] Hereinafter, each component included in the decoding device 200 will be described.
[0157] [Entropy decoding unit]
[0158] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. Next, the entropy decoding unit 202 debinarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks.
[0159] [Inverse Quantization Unit]
[0160] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) which is the input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantization unit 204 outputs the inversely quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0161] [Inverse transformation unit]
[0162] The inverse transform unit 206 restores the prediction error by inversely transforming the transform coefficients that are input from the inverse quantization unit 204 .
[0163] For example, when the information interpreted from the coded bit stream indicates that EMT or AMT is applied (eg, the AMT flag is true), the inverse transform unit 206 inversely transforms the transform coefficients of the current block based on the information indicating the interpreted transform type.
[0164] In addition, for example, when the information interpreted from the encoded bit stream indicates that NSST is applied, the inverse transform unit 206 re-transforms the transformed transform coefficients (transformation results).
[0165] [Addition Department]
[0166] The adding unit 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction signal input from the prediction control unit 220 . The adding unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212 .
[0167] [Block Memory]
[0168] The block memory 210 is a storage unit for storing blocks in a decoding target picture (hereinafter referred to as a current picture) which are reference blocks in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208 .
[0169] [Loop filter section]
[0170] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208 , and outputs the reconstructed block after filtering to the frame memory 214 , a display device, and the like.
[0171] When the information indicating the on / off of ALF interpreted from the coded bit stream indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0172] [Frame Memory]
[0173] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212 .
[0174] [Intra-frame prediction unit]
[0175] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction based on the intra prediction mode interpreted from the coded bit stream and referring to the block in the current picture stored in the block memory 210. Specifically, the intra prediction unit 216 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, color difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0176] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0177] Furthermore, when the information decoded from the coded bit stream indicates application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.
[0178] [Inter-frame prediction unit]
[0179] The inter-frame prediction unit 218 predicts the current block with reference to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or a sub-block (e.g., a 4×4 block) in the current block. For example, the inter-frame prediction unit 126 generates an inter-frame prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., a motion vector) interpreted from the coded bit stream, and outputs the inter-frame prediction signal to the prediction control unit 128.
[0180] When the information interpreted from the coded bit stream indicates that the OBMC mode is applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks.
[0181] In addition, when the information interpreted from the coded bit stream indicates that the FRUC mode is applied, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) interpreted from the coded stream to derive motion information. Furthermore, the inter-frame prediction unit 218 performs motion compensation using the derived motion information.
[0182] In addition, when the BIO mode is applied, the inter-frame prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when the information interpreted from the coded bit stream indicates that the affine motion compensation prediction mode is applied, the inter-frame prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of multiple adjacent blocks.
[0183] [Prediction Control Department]
[0184] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the adder 208 as a prediction signal.
[0185] (Implementation Method 2)
[0186] Next, a part of the processing performed in the encoding device 100 and the decoding device 200 configured as above will be specifically described with reference to the drawings. In order to further expand the benefits of the present disclosure, it is obvious to those skilled in the art that the various embodiments described below may be combined.
[0187] The encoding device and decoding device of this embodiment can be used for encoding and decoding of arbitrary multimedia data, and more specifically, can be used for encoding and decoding of images captured by a non-linear (eg, fisheye) camera.
[0188] Here, in the above-mentioned prior art, the same moving picture coding tools as before are used to compress the processed images and the images directly photographed by the linear lens. In the prior art, there is no moving picture coding tool specially customized for compressing such processed images in a different way.
[0189] Usually, images are first taken with multiple cameras, and the images taken by the multiple cameras are joined together to create a larger image as a 360-degree image. In order to be able to display the image more appropriately on a flat display or to be able to more easily detect objects in the image using machine learning technology, image transformation processing including "defishing" or image correction to make it linear is sometimes performed before encoding the image. However, in this image transformation processing, image samples are usually interpolated, so duplicate parts occur in the information retained in the image. In addition, blank areas are sometimes formed in the image by the joining process and the image transformation process, which are usually filled with default pixel values (such as black pixels). Such problems caused by the joining process and the image transformation process become factors that reduce the encoding efficiency of the encoding process.
[0190] In order to solve these problems, in the present embodiment, an adaptive motion picture coding tool and an adaptive motion picture decoding tool are used as customized motion picture coding tools and motion picture decoding tools. In order to improve coding efficiency, the adaptive motion picture coding tool can adapt to image transformation processing or image splicing processing used to process images before the encoder. The present disclosure can reduce all duplications that occur in these processing by adapting the adaptive motion picture coding tool to the above-mentioned processing in the coding process. The adaptive motion picture decoding tool is also the same as the adaptive motion picture coding tool.
[0191] In this embodiment, the information of the image transformation process and / or the image splicing process is used to adapt the moving picture encoding tool and the moving picture decoding tool. Therefore, the moving picture encoding tool and the moving picture decoding tool can be applied to different types of processed images. Therefore, in this embodiment, the compression efficiency can be improved.
[0192] [Encoding Processing]
[0193] illustrate Fig.11 The second embodiment of the present disclosure is a method for encoding a moving image of an image captured using a non-linear lens. The non-linear lens is a wide-angle lens or an example thereof.
[0194] Fig.11 This is a flowchart showing an example of a moving picture encoding process according to the present embodiment.
[0195] In step S101, the encoding device writes parameters together into the header. Fig.12 Indicates the possible location of the above header in the compressed motion picture bitstream. The parameters to be written (i.e. Fig.12 The camera image parameters in the image correction process include one or more parameters related to the image correction process. For example, such parameters are Fig.12 As shown, the parameters are written into the video parameter set, sequence parameter set, picture parameter set, slice header or video system setting parameter set. That is, in this embodiment, the parameters written can be written into any header of the bitstream, or can be written into SEI (Supplemental Enhancement Information). In addition, the image correction processing is equivalent to the above-mentioned image transformation processing.
[0196] <Example of parameters for image correction processing>
[0197] like Fig.13 As shown, the captured image may also be distorted due to the characteristics of the lens used in capturing the image. In addition, image correction processing is used to correct the captured image linearly. In addition, a rectangular image is generated by linearly correcting the captured image. The written parameters include parameters used to determine the image correction processing used or used to describe it. As an example, the parameters used in the image correction processing include parameters that constitute a mapping table used to map the pixels of the input image to the desired output pixel values of the image correction processing. These parameters may also include one or more weight parameters for interpolation processing, and / or position parameters that determine the positions of the input pixels and output pixels of the image. As one of the embodiments capable of performing image correction processing, the mapping table used for image correction processing may also be used for all pixels in the corrected image.
[0198] Other examples of parameters used to describe image correction processing include a selection parameter for selecting one from a plurality of predefined correction algorithms, a direction parameter for selecting one from a plurality of prescribed directions of the correction algorithm, and / or a calibration parameter for correcting or fine-tuning the correction algorithm. For example, in the case where there are a plurality of predefined correction algorithms (for example, in the case where different algorithms are used in different types of lenses), the selection parameter is used to select one from these predefined algorithms. For example, in the case where there are more than two directions in which the correction algorithm can be applied (for example, in the case where the image correction processing can be performed in the horizontal direction, the vertical direction, or in either direction), the direction parameter selects one of these predefined directions. In the case where the image correction processing can be corrected, the image correction processing can be adjusted by the calibration parameter to make it suitable for different types of lenses.
[0199] <Example of parameters for joining process>
[0200] The written parameters may also include one or more parameters related to the joining process. Fig.14 and Fig.15 As shown, the image input into the encoding device may also be the result of combining and stitching multiple images from different cameras. The written parameters include, for example, parameters that provide information about the stitching process, such as the number of cameras, the center of distortion or the main axis of each camera, and the distortion level. In other examples of parameters described about the stitching process, parameters that determine the position of the stitched image are included, which are generated by repeated pixels from multiple images. These images each have areas that overlap with the angle of the camera, so they may also contain pixels that may appear in other images. In the stitching process, these repeated pixels are processed and reduced to generate a stitched image.
[0201] Other examples of parameters described in connection with the joining process include parameters for determining the layout of the joined image. For example, the configuration of the image in the joined image is different depending on the form of a 360-degree image such as an equirectangular projection method, a 3×2 layout of a cube, and a 4×3 layout of a cube. In addition, a 3×2 layout is a layout of 6 images configured in 3 columns and 2 rows, and a 4×3 layout is a layout of 12 images configured in 4 columns and 3 rows. Configuration parameters as the above parameters are used to determine the continuity of the image in a certain direction based on the configuration of the image. In motion compensation processing, pixels from other images or viewpoints (view) can be used in inter-frame prediction processing, and these images or viewpoints are determined by configuration parameters. Some images or pixels in images also need to be rotated to ensure continuity.
[0202] Other examples of parameters include parameters of the camera and lens (e.g., focal length, principal point, zoom factor, type of image sensor used in the camera, etc.) Other examples of parameters include physical information about the configuration of the camera (e.g., camera position, camera angle, etc.).
[0203] Next, in step S102, the encoding device encodes the image using an adaptive motion picture coding tool based on the written parameters. The adaptive motion picture coding tool includes an inter-frame prediction process. The complete set of adaptive motion picture coding tools may also include an image reconstruction process.
[0204] <Distortion Correction in Inter-frame Prediction>
[0205] Fig.16The flowchart is a flowchart showing the adaptive inter-frame prediction process when it is determined that the image is captured using a non-linear lens, when it is determined that the image is processed linearly, or when it is determined that the image is spliced from more than one image. Fig.16 As shown, in step S1901, the encoding device determines that a certain position in the image is the distortion center or the principal point based on the parameters written in the header. Fig.17 An example of distortion aberration generated by a fisheye lens is shown. In addition, a fisheye lens is an example of a wide-angle lens. As the distance from the center of distortion increases, the magnification decreases along the focal axis. Therefore, in step S1902, the encoding device may perform overlapping (lapping) processing on pixels in the image to correct the distortion or restore the correction that has been performed based on the center of distortion to make the image linear. That is, the encoding device performs image correction processing (i.e., overlapping processing) on the block of the distorted image that is the object of the encoding processing. Finally, the encoding device can perform block prediction of the block of the derived prediction sample in step S1903 based on the pixels of the overlapped image. In addition, the overlapping processing or overlapping of this embodiment is a process of configuring or reconfiguring pixels, blocks or images. In addition, the encoding device may also restore the prediction block that is the predicted block to the original distorted state before the image correction processing is performed, and use the distorted state of the expected block as the predicted image of the distorted processing object block. In addition, the predicted image and the processing object block are equivalent to the prediction signal and the current block of embodiment 1.
[0206] Other examples of adaptive inter-frame prediction processing include adaptive motion vector processing. Regarding the resolution of motion vectors, image blocks farther from the distortion center have lower resolutions than image blocks closer to the distortion center. For example, image blocks farther from the distortion center may have motion vector accuracy up to half-pixel accuracy. On the other hand, image blocks closer to the distortion center may have higher motion vector accuracy within 1 / 8 pixel accuracy. In adaptive motion vector accuracy, since there is a difference based on the image block position, the accuracy of the motion vector encoded in the bitstream may also be able to adapt according to the end position and / or start position of the motion vector. That is, the encoding device may also use parameters to make the accuracy of the motion vector different according to the position of the block.
[0207] Other examples of adaptive inter-frame prediction processing include adaptive motion compensation processing, in which pixels from different viewpoints may be used to predict image samples from a viewpoint of an object based on configuration parameters written in a header. For example, the configuration of images within a joined image is different in the form of a 360-degree image according to an equirectangular projection method, a 3×2 layout of a cube, a 4×3 layout of a cube, etc. Configuration parameters are used to determine the continuity of an image in a certain direction based on the configuration of the image. In motion compensation processing, pixels from other images or other viewpoints may be used for inter-frame prediction processing, and these images or viewpoints are determined by the configuration parameters. Some images or pixels in images may also need to be rotated to ensure continuity.
[0208] That is, the encoding device may also perform processing to ensure continuity. Fig.15 In the case of encoding the joined image shown in the figure, overlapping processing can also be performed based on its parameters. Specifically, among the five images (i.e., images A to D and the top viewpoint) included in the joined image, the top viewpoint is a 180-degree image, and images A to D are 90-degree images. Therefore, the space projected in the top viewpoint is continuous with the space projected in each of images A to D, and the space projected in image A is continuous with the space projected in image B. However, in the joined image, the top viewpoint is not continuous with images A, C, and D, and image A is not continuous with image B. Therefore, the encoding device performs the above-mentioned overlapping processing in order to improve the encoding efficiency. That is, the encoding device reconfigures each image included in the joined image. For example, the encoding device reconfigures each image so that image A is continuous with image B. As a result, the objects that are separated and projected in image A and image B are continuous, which can improve the encoding efficiency. In addition, such overlapping processing as a process of reconfiguring or configuring each image is also called frame packing.
[0209] <Padding in inter-frame prediction>
[0210] Fig.18 This is a flowchart showing a modified example of performing adaptive inter-frame prediction processing when it is determined that the image is captured using a non-linear lens, when it is determined that the image is processed linearly, or when it is determined that the image is spliced from two or more images. Fig.18 As shown, the encoding device determines the area of the image as a blank area in step S2001 based on the parameters written in the header. These blank areas are areas of the image that do not contain pixels of the captured image and are usually replaced with a specified pixel value (such as a black pixel). Fig.13 FIG. 1 is a diagram showing an example of these areas in an image. Fig.15is a diagram showing another example of these areas when a plurality of images are joined. Fig.18 In step S2002, the pixels in the determined areas are filled with values of other areas that are not blank areas of the image during the motion compensation process. The filled values may be values of the nearest pixels in the area that is not a blank area or values of the nearest pixels according to the physical three-dimensional space. Finally, in step S2003, the encoding device performs block prediction to generate a block of prediction samples based on the filled values.
[0211] <Distortion Correction in Image Reconstruction>
[0212] Fig.19 The flowchart is a flowchart showing the adaptive picture reconstruction process when it is determined that the image is photographed using a non-linear lens, or when it is determined that the image is processed linearly, or when it is determined that the image is spliced from two or more images. Fig.19 As shown, the encoding device determines the position in the image as the distortion center or the principal point in step S1801 based on the parameters written in the header. Fig.17 This shows an example of distortion aberration generated by a fisheye lens. As the axis of the focus moves away from the center of distortion, the magnification along the axis of the focus decreases. Therefore, in step S1802, the encoding device may also correct the distortion for the reconstructed pixels in the image based on the center of distortion, or perform overlapping processing to restore the correction performed to make the image linear. For example, the encoding device generates a reconstructed image by adding the image of the prediction error generated by the inverse transform to the prediction image. At this time, the encoding device performs overlapping processing to make the image of the prediction error and the prediction image linear.
[0213] Finally, in step S1803, the encoding device stores the blocks of the image reconstructed based on the pixels of the image after the superposition processing in the memory.
[0214] <Replacement of pixel values in image reconstruction>
[0215] Fig. 20 A modified example of performing adaptive picture reconstruction processing when it is determined that the image is captured using a non-linear lens, when it is determined that the image is processed linearly, or when it is determined that the image is spliced from two or more images is shown. Fig. 20 As shown, based on the parameters written in the header, in step S2101, the encoding device determines the area of the image as a blank area. These blank areas do not include the pixels of the image captured, and are usually areas of the image that are replaced with specified pixel values (such as black pixels). Fig.13 FIG. 1 is a diagram showing an example of these areas in an image. Fig.15 21 is a diagram showing another example of these regions when a plurality of images are joined. Next, in step S2102, the encoding device reconstructs the blocks of image samples.
[0216] Furthermore, in step S2103, the encoding device replaces the reconstructed pixels in these determined areas with prescribed pixel values.
[0217] <Omission of encoding processing>
[0218] exist Fig.11 In step S102, in other possible variations of the adaptive motion picture coding tool, the encoding process of the image may be omitted. That is, based on the parameters written about the layout configuration of the image and the information about the effective (active) viewpoint area based on the line of sight of the user's eyes or the direction of the head, the encoding device may also omit the encoding process of the image. That is, the encoding device performs partial encoding.
[0219] Fig.21 An example of the perspective of the user's line of sight or the direction of the head with respect to different viewpoints captured by different cameras is shown. As shown in the figure, the user's perspective is within the image captured by the camera only from viewpoint 1. In this example, images from other viewpoints do not need to be encoded because they are outside the user's perspective. Therefore, in order to reduce the complexity of encoding, or to reduce the transmission bit rate of compressed images, the encoding or transmission processing of these images can be omitted. In another possible example shown in the figure, since viewpoints 5 and 2 are physically close to the effective viewpoint 1, the image from viewpoint 5 and the image from viewpoint 2 are also encoded and transmitted. These images are not displayed to the observer or user at the current moment, but are displayed to the observer or user when the observer changes the direction of his or her head. These images are used to improve the user's audio-visual experience when the observer changes the direction of his or her head.
[0220] Fig. 22 Another example of representing the angle of the user's line of sight or the orientation of the head relative to different viewpoints captured by different cameras. Here, the effective line of sight area is within the image from viewpoint 2. Therefore, the image from viewpoint 2 is encoded and displayed to the user. Here, the encoding device predicts the range in which the observer's head is estimated to move recently, and defines a larger area as the range that is likely to become the line of sight area of future frames. The encoding device encodes and sends not only the image within the effective line of sight area of the object, but also the image from the viewpoint (other than viewpoint 2) in the larger future line of sight area, so that the viewpoint can be depicted more quickly at the observer. That is, not only the image from viewpoint 2, but also the image from viewpoint 2 is encoded. Fig. 22The images of the top viewpoint and viewpoint 1, which at least partly overlap with the possible viewing area, are also encoded and transmitted. The images from the remaining viewpoints (viewpoint 3, viewpoint 4, and bottom viewpoint) are not encoded, and the encoding process of these images is omitted.
[0221] [Encoding device]
[0222] Fig.23 This is a block diagram showing the structure of an encoding device for encoding a moving picture according to the present embodiment.
[0223] The encoding device 900 is a device for encoding an input moving image for each block in order to generate an output bit stream, and is equivalent to the encoding device 100 of the first embodiment. Fig.23 As shown, the encoding device 900 includes a transform unit 901, a quantization unit 902, an inverse quantization unit 903, an inverse transform unit 904, a block memory 905, a frame memory 906, an intra-frame prediction unit 907, an inter-frame prediction unit 908, a subtraction unit 921, an addition unit 922, an entropy encoding unit 909 and a parameter derivation unit 910.
[0224] The image of the input moving image (i.e., the processing target block) is input to the subtraction unit 921, and the value after subtraction is output to the transformation unit 901. That is, the subtraction unit 921 calculates the prediction error by subtracting the prediction image from the processing target block. The transformation unit 901 transforms the value after subtraction (i.e., the prediction error) into a frequency coefficient, and outputs the obtained frequency coefficient to the quantization unit 902. The quantization unit 902 quantizes the input frequency coefficient, and outputs the obtained quantized value to the inverse quantization unit 903 and the entropy coding unit 909.
[0225] The inverse quantization unit 903 inversely quantizes the sample values (i.e., quantized values) output from the quantization unit 902, and outputs frequency coefficients to the inverse transformation unit 904. The inverse transformation unit 904 performs inverse frequency transformation to transform the frequency coefficients into sample values of an image, i.e., pixel values, and outputs the obtained sample values to the addition unit 922.
[0226] The parameter derivation unit 910 derives parameters related to image correction processing, parameters related to the camera, or parameters related to the splicing processing based on the image, and outputs them to the inter-frame prediction unit 908, the addition unit 922, and the entropy coding unit 909. For example, these parameters may also be included in the input moving image. In this case, the parameter derivation unit 910 extracts the parameters included in the moving image and outputs them. Alternatively, the input moving image may also include parameters that serve as a basis for deriving these parameters. In this case, the parameter derivation unit 910 extracts the basic parameters included in the moving image, converts the extracted basic parameters into the above-mentioned parameters, and outputs them.
[0227] The adder 922 adds the sample value output from the inverse transform unit 904 to the pixel value of the predicted image output from the intra prediction unit 907 or the inter prediction unit 908. That is, the adder 922 performs image reconstruction processing to generate a reconstructed image. The adder 922 outputs the obtained addition value to the block memory 905 or the frame memory 906 for further prediction.
[0228] The intra prediction unit 907 performs intra prediction. That is, the intra prediction unit 907 estimates the image of the processing target block using the reconstructed image stored in the block memory 905 and included in the same picture as the picture of the processing target block. The inter prediction unit 908 performs inter prediction. That is, the inter prediction unit 908 estimates the image of the processing target block using the reconstructed image stored in the frame memory 906 and included in a picture different from the picture of the processing target block.
[0229] Here, in this embodiment, the inter-frame prediction unit 908 and the addition unit 922 make the processing adaptive based on the parameters derived by the parameter derivation unit 910. That is, the inter-frame prediction unit 908 and the addition unit 922 perform the processing according to the above-mentioned adaptive video coding tool. Fig.16 , Fig.18 , Fig.19 and Fig. 20 Processing of the flowchart shown.
[0230] The entropy coding unit 909 codes the quantized value output from the quantization unit 902 and the parameter derived by the parameter deriving unit 910, and outputs a bit stream. That is, the entropy coding unit 909 writes its parameters into the header of the bit stream.
[0231] [Decoding process]
[0232] Fig.24 This is a flowchart showing an example of a moving picture decoding process according to the present embodiment.
[0233] In step S201, the decoding device decodes a set of parameters from scratch. Fig.12 The decoded parameters include one or more parameters related to the image correction process.
[0234] <Example of parameters for image correction processing>
[0235] like Fig.13As shown, the captured image may also be distorted due to the characteristics of the lens used in capturing the image. In addition, image correction processing is used to correct the linearity of the captured image. The interpreted parameters include parameters used to determine the image correction processing used, or to record the image correction processing used. In the example of parameters used in the image correction processing, parameters are included that constitute a mapping table used to map pixels of an input image to desired output pixel values of the image correction processing. These parameters may also include one or more weight parameters for interpolation processing, or / and position parameters for determining the positions of input pixels and output pixels of the image. In one possible embodiment of the image correction processing, the mapping table used for the image correction processing may also be used for all pixels in the corrected image.
[0236] Other examples of parameters used to record image correction processing include a selection parameter for selecting one of a plurality of predefined correction algorithms, a direction parameter for selecting one of a plurality of specified directions of a correction algorithm, and / or a calibration parameter for correcting or fine-tuning the correction algorithm. For example, in the case where there are a plurality of predefined correction algorithms (for example, in the case where different algorithms are used in different types of lenses), the selection parameter is used to select one of these predefined algorithms. For example, in the case where there are more than two directions in which the correction algorithm can be applied (for example, in the case where image correction processing can be performed in the horizontal direction, the vertical direction, or in any direction), the direction parameter selects one of these predefined directions. For example, in the case where the image correction processing can be corrected, the image correction processing can be adjusted by the calibration parameter to be suitable for different types of lenses.
[0237] <Example of parameters for joining process>
[0238] The decoded parameters may also include one or more parameters related to the joining process. Fig.14 and Fig.15 As shown, the encoded image input to the decoding device may also be obtained as a result of a splicing process combining multiple images from different cameras. The decoded parameters include parameters that provide information related to the splicing process, such as the number of cameras, the center of distortion or the main axis of each camera, and the distortion level. As another example of parameters recorded about the splicing process, there are parameters that determine the position of the spliced image generated based on repeated pixels from multiple images. These images each have areas that overlap with the angle of the camera, so pixels that may also appear in other images may be included. In the splicing process, these repeated pixels are processed and reduced to generate a spliced image.
[0239] Other examples of parameters described in connection with the joining process include parameters for determining the layout of the joined image. For example, the configuration of the image within the joined image is different depending on the form of a 360-degree image such as an equirectangular projection method, a 3×2 layout of a cube, or a 4×3 layout of a cube. Configuration parameters, which are the above parameters, are used to determine the continuity of the image in a certain direction based on the configuration of the image. In the motion compensation process, pixels from other images or viewpoints can be used in the inter-frame prediction process, and these images or viewpoints are determined by the configuration parameters. There are also cases where some images or pixels in images need to be rotated to ensure continuity.
[0240] Other examples of parameters include parameters of the camera and lens (e.g., focal length, principal point, zoom factor, image sensor type, etc. used in the camera). Still other examples of parameters include physical information related to the configuration of the camera (e.g., camera position, camera angle, etc.).
[0241] Next, in step S202, the decoding device decodes the image by an adaptive motion picture decoding tool based on the decoded parameters. The adaptive motion picture decoding tool includes inter-frame prediction processing. The complete set of adaptive motion picture decoding tools may also include image reconstruction processing. In addition, the motion picture decoding tool or the adaptive motion picture decoding tool is the same as or corresponds to the above-mentioned motion picture encoding tool or the adaptive motion picture encoding tool.
[0242] <Distortion Correction in Inter-frame Prediction>
[0243] Fig.16 The flowchart is a flowchart showing the adaptive inter-frame prediction process when it is determined that the image is captured using a non-linear lens, or when it is determined that the image is processed linearly, or when it is determined that the image is spliced from more than one image. Fig.16 As shown, in step S1901, the decoding device determines that a certain position in the image is the distortion center or the principal point based on the parameters written in the header. Fig.17An example of distortion aberration generated by a fisheye lens is shown. As the focal axis moves away from the center of distortion, the magnification decreases along the focal axis. Therefore, in step S1902, the decoding device may also perform overlapping processing on the pixels in the image in order to correct the distortion based on the center of distortion or restore the correction performed to make the image linear. That is, the decoding device performs image correction processing (i.e., overlapping processing) on the block of the distorted image that is the object of decoding processing. Finally, the decoding device may perform block prediction of a block of prediction samples derived based on the pixels of the image that has been overlapped in step S1903. In addition, the decoding device may also restore the prediction block that is the predicted block to the original distorted state before the image correction processing is performed, and use the expected block in the distorted state as the predicted image of the distorted processing object block.
[0244] Other examples of adaptive inter-frame prediction processing include adaptive motion vector processing. Regarding the resolution of motion vectors, image blocks farther from the distortion center have lower resolution than image blocks closer to the distortion center. For example, image blocks farther from the distortion center may also have motion vector accuracy within half-pixel accuracy. On the other hand, image blocks closer to the distortion center may also have higher motion vector accuracy within 1 / 8 pixel accuracy. In adaptive motion vector accuracy, since there is a difference based on the image block position, the accuracy of the motion vector encoded in the bitstream may also be able to be adaptive based on the end position and / or start position of the motion vector. That is, the encoding device may also use parameters to make the accuracy of the motion vector different depending on the position of the block.
[0245] In other examples of adaptive inter-frame prediction processing, adaptive motion compensation processing may also be included, in which pixels from different viewpoints are used to predict image samples from the viewpoint of an object based on configuration parameters written in the header. For example, the configuration of images within the joined image is different in the form of 360-degree images according to equirectangular projection, 3×2 layout of a cube, 4×3 layout of a cube, etc. Configuration parameters are used to determine the continuity of the image in a certain direction based on the configuration of the image. In the motion compensation processing, pixels from other images or other viewpoints can be used for inter-frame prediction processing, and these images or viewpoints are determined by the configuration parameters. Some images or pixels in images may also need to be rotated to ensure continuity.
[0246] That is, the decoding device may also perform processing to ensure continuity. Fig.15In the case of the joint image encoding shown in the figure, the overlapping process can also be performed based on its parameters. Specifically, the decoding device, like the above encoding device, rearranges each image so that image A and image B are continuous. As a result, the objects separated and projected in image A and image B are continuous, which can improve the encoding efficiency.
[0247] <Padding in inter prediction>
[0248] Fig.18 This is a flowchart showing a modified example of performing adaptive inter-frame prediction processing when it is determined that the image is captured using a non-linear lens, when it is determined that the image is processed linearly, or when it is determined that the image is spliced from two or more images. Fig.18 As shown, the decoding device determines in step S2001 that the area of the image is a blank area based on the parameters decoded from the beginning. These blank areas do not contain pixels of the captured image and are usually areas of the image replaced with a specified pixel value (such as a black pixel). Fig.13 Show examples of these regions within the image. Fig.15 is a diagram showing another example of these areas when a plurality of images are joined. Fig.18 In step S2002, the pixels in the determined area are filled with values of other areas of the image that are not blank areas in the motion compensation process. The value of the filling process may be the value of the nearest pixel in the area that is not a blank area, or the value of the nearest pixel, depending on the physical three-dimensional space. Finally, in step S2003, the decoding device performs block prediction to generate a block of prediction samples based on the value after the filling process.
[0249] <Distortion Correction in Image Reconstruction>
[0250] Fig.19 The flowchart is a flowchart showing the adaptive image reconstruction process when it is determined that the image is captured using a non-linear lens, or when it is determined that the image is processed linearly, or when it is determined that the image is spliced from two or more images. Fig.19 As shown, the decoding device determines the position in the image as the distortion center or the principal point in step S1801 based on the parameters interpreted from the beginning. Fig.17This shows an example of distortion aberration generated by a fisheye lens. As the axis of the focus moves away from the center of distortion, the magnification along the axis of the focus decreases. Therefore, in step S1802, the decoding device performs an overlap process to correct the distortion of the reconstructed pixels in the image based on the center of distortion, or to restore the correction that has been performed to make the image linear. For example, the decoding device generates a reconstructed image by adding the image of the prediction error generated by the inverse transform to the prediction image. At this time, the decoding device performs an overlap process to make the image of the prediction error and the prediction image linear.
[0251] Finally, in step S1803, the decoding apparatus stores blocks of the reconstructed image in a memory based on the pixels of the image subjected to the overlap processing.
[0252] <Replacement of pixel values in image reconstruction>
[0253] Fig. 20 A modified example of performing adaptive image reconstruction processing when it is determined that the image is captured using a non-linear lens, when it is determined that the image is processed linearly, or when it is determined that the image is spliced from more than one image is shown. Fig. 20 As shown, based on the parameters decoded from the beginning, in step S2001, the decoding device determines the area of the image as a blank area. These blank areas do not contain pixels of the captured image, and are usually areas of the image replaced with a specified pixel value (such as a black pixel). Fig.13 Represent examples of these regions in the image. Fig.15 21 is a diagram showing another example of these regions when a plurality of images are joined. Next, in step S2102, the decoding apparatus reconstructs the blocks of image samples.
[0254] Furthermore, in step S2103, the decoding device replaces the reconstructed pixels within these determined areas with prescribed pixel values.
[0255] <Omission of decoding processing>
[0256] exist Fig.24 In step S202, in other possible variations of the adaptive motion picture decoding tool for images, the decoding process of the images may be omitted. That is, based on the parameters interpreted about the layout configuration of the images and the information about the effective viewpoint area based on the line of sight of the user's eyes or the direction of the head, the decoding device may also omit the decoding process of the images. That is, the decoding device performs partial decoding process.
[0257] Fig.21An example of the perspective of the user's line of sight or the direction of the head with respect to different viewpoints captured by different cameras is shown. As shown in the figure, the user's perspective is within the image captured by the camera from viewpoint 1 only. In this example, the images from other viewpoints do not need to be decoded because they are outside the user's perspective. Therefore, in order to reduce the complexity of decoding or to reduce the transmission bit rate of compressed images, the decoding process or display process of these images can be omitted. In another possible example shown in the figure, since viewpoints 5 and 2 are physically close to the effective viewpoint 1, the image from viewpoint 5 and the image from viewpoint 2 are also decoded. These images are not displayed to the observer or user at the current moment, but are displayed to the observer or user when the observer changes the direction of his or her head. By reducing the time to decode and display the viewpoint according to the movement of the user's head, when the user changes the direction of his or her head, these images are displayed as early as possible in order to improve the user's audio-visual experience.
[0258] Fig. 22 Another example of representing the angle of the user's line of sight or the orientation of the head relative to different viewpoints captured by different cameras. Here, the effective line of sight area is within the image from viewpoint 2. Therefore, the image from viewpoint 2 is decoded and displayed to the user. Here, the decoding device predicts the range in which the observer's head is estimated to move recently, and defines a larger area as the range that is likely to become the line of sight area of future frames. The decoding device decodes not the image within the effective line of sight area of the object, but also the image from the viewpoint (other than viewpoint 2) in the larger future line of sight area. That is, not only the image from viewpoint 2, but also the image from viewpoint 2 is decoded. Fig. 22 The images of the top viewpoint and viewpoint 1, which at least partly overlap with the possible sightline area, are also decoded. Thus, the images are displayed so that the viewpoints can be depicted more quickly at the observer's position. The images from the remaining viewpoints (viewpoint 3, viewpoint 4, and the viewpoint below) are not decoded, and the decoding process of these images is omitted.
[0259] [Decoding device]
[0260] Fig.25 This is a block diagram showing the structure of a decoding device for decoding a moving picture according to the present embodiment.
[0261] The decoding device 1000 is a device for decoding an input coded moving image (i.e., an input bit stream) for each block in order to generate a decoded moving image, and is equivalent to the decoding device 200 of the first embodiment. Fig.25 As shown, the decoding device 1000 includes an entropy decoding unit 1001, an inverse quantization unit 1002, an inverse transformation unit 1003, a block memory 1004, a frame memory 1005, an addition unit 1022, an intra-frame prediction unit 1006 and an inter-frame prediction unit 1007.
[0262] The input bit stream is input to the entropy decoding unit 1001. Then, the entropy decoding unit 1001 performs entropy decoding on the input bit stream, and outputs the value obtained by the entropy decoding (i.e., the quantized value) to the inverse quantization unit 1002. The entropy decoding unit 1001 also interprets the parameters from the input bit stream and outputs the parameters to the inter-frame prediction unit 1007 and the addition unit 1022.
[0263] The inverse quantization unit 1002 inversely quantizes the value obtained by entropy decoding, and outputs the frequency coefficient to the inverse transformation unit 1003. The inverse transformation unit 1003 performs inverse frequency transformation on the frequency coefficient, transforms the frequency coefficient into a sample value (i.e., a pixel value), and outputs the obtained pixel value to the addition unit 1022. The addition unit 1022 adds the obtained pixel value to the pixel value of the predicted image output from the intra-frame prediction unit 1006 or the inter-frame prediction unit 1007. In other words, the addition unit 1022 performs image reconstruction processing to generate a reconstructed image. The addition unit 1022 outputs the value obtained by the addition (i.e., the decoded image) to the display, and outputs the obtained value to the block memory 1004 or the frame memory 1005 for further prediction.
[0264] The intra prediction unit 1006 performs intra prediction. That is, the intra prediction unit 1006 estimates the image of the processing target block using the reconstructed image stored in the block memory 1004 and included in the same picture as the picture of the processing target block. The inter prediction unit 1007 performs inter prediction. That is, the inter prediction unit 1007 estimates the image of the processing target block using the reconstructed image stored in the frame memory 1005 and included in a picture different from the picture of the processing target block.
[0265] Here, in this embodiment, the inter-frame prediction unit 1007 and the addition unit 1022 make the processing based on the decoded parameters adaptive. That is, the inter-frame prediction unit 1007 and the addition unit 1022 perform the processing according to the adaptive video decoding tool as described above. Fig.16 , Fig.18 , Fig.19 and Fig. 20 Processing of the flowchart shown.
[0266] (Implementation method 3)
[0267] [Encoding Processing]
[0268] illustrate Fig.26 Embodiment 3 of the present disclosure shown here is a method of performing a moving picture encoding process on an image captured using a non-linear lens.
[0269] Fig.26 This is a flowchart showing an example of a moving picture encoding process according to the present embodiment.
[0270] In step S301, the encoding device writes the parameters together into the header. Fig.12 Indicates the possible position of the header in the compressed moving picture bit stream. The written parameters include one or more parameters related to the position of the camera. The written parameters may also include one or more parameters related to the camera angle or parameters related to instructions for a method of joining multiple images.
[0271] Other examples of parameters include parameters of the camera and lens (e.g., focal length, principal point, zoom factor, type of image sensor used in the camera, etc.) Still other examples of parameters include physical information related to the configuration of the camera (e.g., camera position, camera angle, etc.).
[0272] In the present embodiment, the above-mentioned parameters written to the head are also referred to as camera parameters or stitching parameters.
[0273] Fig.15 An example of a method of joining images from two or more cameras is shown. Fig.14 Another example of a method of joining images from two or more cameras is shown.
[0274] Next, in step S302, the encoding device encodes the image. In step S302, the encoding process may also be adapted based on the spliced image. For example, the encoding device may refer to a larger spliced image as a reference image instead of an image of the same size as the decoded image (i.e., an image that is not spliced) in the motion compensation process.
[0275] Finally, in step S303, the encoding device combines the first image, which is the image encoded and reconstructed in step S302, with the second image based on the written parameters to create a larger image. The combined image can also be used for prediction of future frames (i.e., inter-frame prediction or motion compensation).
[0276] Fig. 27 2401, the encoding device determines the camera parameters or the splicing parameters based on the parameters written to the target image. Similarly, in step S2402, the encoding device determines the camera parameters or the splicing parameters of other images based on the parameters written to other images. Finally, in step S2403, the encoding device uses these determined parameters to splice the images to create a larger image. These determined parameters are written to the header. In addition, the encoding device may also perform overlapping processing or frame encapsulation to further improve the encoding efficiency by configuring or reconfiguring multiple images.
[0277] [Encoding device]
[0278] Fig.28 This is a block diagram showing the structure of an encoding device for encoding a moving picture according to the present embodiment.
[0279] The encoding device 1100 is a device for encoding an input moving image for each block in order to generate an output bit stream, and is equivalent to the encoding device 100 of the first embodiment. Fig.28 As shown, the encoding device 1100 includes a transform unit 1101, a quantization unit 1102, an inverse quantization unit 1103, an inverse transform unit 1104, a block memory 1105, a frame memory 1106, an intra-frame prediction unit 1107, an inter-frame prediction unit 1108, a subtraction unit 1121, an addition unit 1122, an entropy encoding unit 1109, a parameter derivation unit 1110 and an image splicing unit 1111.
[0280] The image of the input moving image (i.e., the processing target block) is input to the subtraction unit 1121, and the value after subtraction is output to the transformation unit 1101. That is, the subtraction unit 1121 calculates the prediction error by subtracting the prediction image from the processing target block. The transformation unit 1101 transforms the value after subtraction (i.e., the prediction error) into a frequency coefficient, and outputs the obtained frequency coefficient to the quantization unit 1102. The quantization unit 1102 quantizes the input frequency coefficient, and outputs the obtained quantized value to the inverse quantization unit 1103 and the entropy coding unit 1109.
[0281] The inverse quantization unit 1103 inversely quantizes the sample values (i.e., the quantized original) output from the quantization unit 1102, and outputs the frequency coefficients to the inverse transformation unit 1104. The inverse transformation unit 1104 transforms the frequency coefficients into sample values of the image, i.e., pixel values, by performing inverse frequency transformation on the frequency coefficients, and outputs the resulting sample values to the addition unit 1122.
[0282] The adder 1122 adds the sample value output from the inverse transform unit 1104 to the pixel value of the predicted image output from the intra prediction unit 1107 or the inter prediction unit 1108. The adder 1122 outputs the obtained added value to the block memory 1105 or the frame memory 1106 for further prediction.
[0283] The parameter derivation unit 1110 derives parameters related to the image splicing process or camera-related parameters from the image, and outputs them to the image splicing unit 1111 and the entropy coding unit 1109, as in the first embodiment. Fig. 27The processing of steps S2401 and S2402 shown in the figure. For example, the input moving image may contain these parameters. In this case, the parameter derivation unit 1110 extracts the parameters contained in the moving image and outputs them. Alternatively, the input moving image may contain parameters that are the basis for deriving these parameters. In this case, the parameter derivation unit 1110 extracts the basic parameters contained in the moving image, converts the extracted basic parameters into the above-mentioned parameters, and outputs them.
[0284] The image joining unit 1111 is as follows Fig.26 Step S303 and Fig. 27 As shown in step S2403 of FIG. 1 , the reconstructed target image is joined with other images using the parameters. Then, the image joining unit 1111 outputs the joined image to the frame memory 1106 .
[0285] The intra prediction unit 1107 performs intra prediction. That is, the intra prediction unit 1107 estimates the image of the processing target block using the reconstructed image stored in the block memory 1105 and included in the same picture as the picture of the processing target block. The inter prediction unit 1108 performs inter prediction. That is, the inter prediction unit 1108 estimates the image of the processing target block using the reconstructed image stored in the frame memory 1106 and included in the picture different from the picture of the image of the processing target block. At this time, the inter prediction unit 1108 may refer to a larger image stored in the frame memory 1106 obtained by splicing a plurality of images by the image splicing unit 1111 as a reference image.
[0286] The entropy coding unit 1109 codes the quantized value output from the quantization unit 1102, obtains the parameter from the parameter derivation unit 1110, and outputs a bitstream. That is, the entropy coding unit 1109 entropy codes the quantized value and the parameter, and writes the parameter into the header of the bitstream.
[0287] [Decoding process]
[0288] Fig.29 This is a flowchart showing an example of a moving picture decoding process according to the present embodiment.
[0289] In step S401, the decoding device decodes a set of parameters from scratch. Fig.12Indicates the possible position of the above-mentioned header in the compressed video bitstream. The decoded parameters include one or more parameters related to the position of the camera. The decoded parameters may also include one or more parameters related to the camera angle, or parameters related to instructions for a method of joining multiple images. Other examples of parameters include parameters of the camera and the lens (for example, the focal length, principal point, zoom factor, type of image sensor used in the camera, etc.). Another example of parameters includes physical information related to the configuration of the camera (for example, the position of the camera, the angle of the camera, etc.).
[0290] Fig.15 An example of a method of joining images from two or more cameras is shown. Fig.14 Another example of a method of joining images from two or more cameras is shown.
[0291] Next, in step S402, the decoding device decodes the image. The decoding process in step S402 may also be made adaptive based on the spliced image. For example, in the motion compensation process, the decoding device may refer to the spliced larger image as a reference image instead of an image of the same size as the decoded image (i.e., an image that is not spliced).
[0292] Furthermore, in step S403, the decoding device finally joins the first image, which is the image reconstructed in step S402, with the second image based on the decoded parameters to create a larger image. The image obtained by joining can also be used for prediction of future images (i.e., inter-frame prediction or motion compensation).
[0293] Fig. 27 2401, the decoding device determines the camera parameters or the splicing parameters by decoding the header of the target image. Similarly, the decoding device determines the camera parameters or the splicing parameters by decoding the header of other images in step S2402. Finally, in step S2403, the decoding device splices the images using the decoded parameters to create a larger image.
[0294] [Decoding device]
[0295] Fig.30 This is a block diagram showing the structure of a decoding device for decoding a moving picture according to the present embodiment.
[0296] The decoding device 1200 is a device that decodes the input coded moving picture (i.e., the input bit stream) for each block and outputs the decoded moving picture, and is equivalent to the decoding device 200 of the first embodiment. Fig.30As shown, the decoding device 1200 includes an entropy decoding unit 1201, an inverse quantization unit 1202, an inverse transformation unit 1203, a block memory 1204, a frame memory 1205, an addition unit 1222, an intra-frame prediction unit 1206, an inter-frame prediction unit 1207 and an image splicing unit 1208.
[0297] The input bit stream is input to the entropy decoding unit 1201. Then, the entropy decoding unit 1201 performs entropy decoding on the input bit stream, and outputs the value obtained by the entropy decoding (i.e., the quantized value) to the inverse quantization unit 1202. The entropy decoding unit 1201 also decodes parameters from the input bit stream and outputs the parameters to the image splicing unit 1208.
[0298] The image joining unit 1208 joins the reconstructed target image with other images using the parameters, and then outputs the joined image to the frame memory 1205 .
[0299] The inverse quantization unit 1202 inversely quantizes the value obtained by entropy decoding, and outputs the frequency coefficient to the inverse transformation unit 1203. The inverse transformation unit 1203 performs inverse frequency transformation on the frequency coefficient, transforms the frequency coefficient into a sample value (i.e., a pixel value), and outputs the resulting pixel value to the addition unit 1222. The addition unit 1222 adds the resulting pixel value to the pixel value of the predicted image output from the intra-frame prediction unit 1206 or the inter-frame prediction unit 1207. The addition unit 1222 outputs the value obtained by the addition (i.e., the decoded image) to the display, and outputs the obtained value to the block memory 1204 or the frame memory 1205 for further prediction.
[0300] The intra prediction unit 1206 performs intra prediction. That is, the intra prediction unit 1206 estimates the image of the processing target block using the reconstructed image included in the same picture as the picture of the processing target block stored in the block memory 1204. The inter prediction unit 1207 performs inter prediction. That is, the inter prediction unit 1207 estimates the image of the processing target block using the reconstructed image included in the picture different from the picture of the processing target block stored in the frame memory 1205.
[0301] (Implementation 4)
[0302] [Encoding Processing]
[0303] illustrate Fig.31 Embodiment 4 of the present disclosure shown here is a method of performing a moving picture encoding process on an image captured using a non-linear lens.
[0304] Fig.31 This is a flowchart showing an example of a moving picture encoding process according to the present embodiment.
[0305] In step S501, the encoding device writes the parameters together into the header. Fig.12 The possible position of the header in the compressed moving image bit stream is indicated. The written parameters include one or more parameters related to an identification code indicating whether the image is taken by a non-linear lens. Fig.13 As shown, the captured image may be distorted due to the characteristics of the lens used for capturing the image. An example of the written parameter is a parameter indicating the position of the center or main axis of the distortion.
[0306] Next, in step S502, the encoding device encodes the image using an adaptive motion picture coding tool based on the written parameters. The adaptive motion picture coding tool includes a motion vector prediction process. The set of adaptive motion picture coding tools may also include an intra-frame prediction process.
[0307] <Intra-frame prediction processing>
[0308] Fig.32 FIG. 1 is a flowchart showing adaptive intra-frame prediction processing based on written parameters. Fig.32 As shown, in step S2201, the encoding device determines a position in the image as a distortion center or a principal point based on the written parameters. Next, in step S2202, the encoding device predicts a sample group using spatially adjacent pixel values. The sample group is, for example, a pixel group of a processing target block.
[0309] Finally, in step S2203, the encoding device performs overlapping processing on the predicted sample group using the determined distortion center or principal point to generate a block of predicted samples. For example, the encoding device may also distort the image of the block of predicted samples and use the distorted image as the predicted image.
[0310] <Motion Vector Prediction>
[0311] Fig.33 FIG. 1 is a flowchart showing adaptive motion vector prediction processing based on written parameters. Fig.33 As shown, in step S2301, the encoding device determines a certain position in the image as a distortion center or a principal point based on the written parameters. Next, in step S2302, the encoding device predicts a motion vector based on a motion vector adjacent in space or time.
[0312] Finally, in step S2303, the encoding device uses the determined distortion center or principal point to correct the direction of the predicted motion vector.
[0313] [Encoding device]
[0314] Fig.34This is a block diagram showing the structure of an encoding device that encodes a moving image in this embodiment.
[0315] The encoding device 1300 is a device for encoding an input moving image for each block in order to generate an output bit stream, and is equivalent to the encoding device 100 of the first embodiment. Fig.34 As shown, the encoding device 1300 includes a transformation unit 1301, a quantization unit 1302, an inverse quantization unit 1303, an inverse transformation unit 1304, a block memory 1305, a frame memory 1306, an intra-frame prediction unit 1307, an inter-frame prediction unit 1308, a subtraction unit 1321, an addition unit 1322, an entropy coding unit 1309 and a parameter derivation unit 1310.
[0316] The image of the input moving image (i.e., the processing target block) is input to the subtraction unit 1321, and the value after subtraction is output to the transformation unit 1301. That is, the subtraction unit 1321 calculates the prediction error by subtracting the prediction image from the processing target block. The transformation unit 1301 transforms the value after subtraction (i.e., the prediction error) into a frequency coefficient, and outputs the resulting frequency coefficient to the quantization unit 1302. The quantization unit 1302 quantizes the input frequency coefficient, and outputs the resulting quantized value to the inverse quantization unit 1303 and the entropy coding unit 1309.
[0317] The inverse quantization unit 1303 inversely quantizes the sample values (i.e., quantized values) output from the quantization unit 1302, and outputs the frequency coefficients to the inverse transformation unit 1304. The inverse transformation unit 1304 performs inverse frequency transformation on the frequency coefficients, transforms the frequency coefficients into sample values of the image, i.e., pixel values, and outputs the resulting sample values to the addition unit 1322.
[0318] The parameter derivation unit 1310 derives one or more parameters (specifically, parameters indicating the center of distortion or the principal point) related to the identification code indicating whether the image is captured by a non-linear lens, based on the image, similarly to the first embodiment. The parameter derivation unit 1310 outputs the derived parameters to the intra-frame prediction unit 1307, the inter-frame prediction unit 1308, and the entropy coding unit 1309. For example, these parameters may also be included in the input motion image. In this case, the parameter derivation unit 1310 extracts the parameters included in the motion image and outputs them. Alternatively, the input motion image may also include parameters that serve as a basis for deriving these parameters. In this case, the parameter derivation unit 1310 extracts the basic parameters included in the motion image, transforms the extracted basic parameters into the above-mentioned parameters, and outputs them.
[0319] The adder 1322 adds the sample value of the image output from the inverse transform unit 1304 to the pixel value of the predicted image output from the intra prediction unit 1307 or the inter prediction unit 1308. The adder 1322 outputs the obtained added value to the block memory 1305 or the frame memory 1306 for further prediction.
[0320] The intra prediction unit 1307 performs intra prediction. That is, the intra prediction unit 1307 estimates the image of the processing target block using the reconstructed image contained in the same picture as the picture of the processing target block stored in the block memory 1305. The inter prediction unit 1308 performs inter prediction. That is, the inter prediction unit 1308 estimates the image of the processing target block using the reconstructed image contained in the picture different from the picture of the processing target block stored in the frame memory 1306.
[0321] Here, in this embodiment, the intra prediction unit 1307 and the inter prediction unit 1308 perform processing based on the parameters derived by the parameter derivation unit 1310. That is, the intra prediction unit 1307 and the inter prediction unit 1308 perform processing based on Fig.32 and Fig.33 Processing of the flowchart shown.
[0322] The entropy coding unit 1309 codes the quantized value output from the quantization unit 1302 and the parameter derived by the parameter derivation unit 1310, and outputs the result to a bitstream. That is, the entropy coding unit 1309 writes the parameter to the header of the bitstream.
[0323] [Decoding process]
[0324] Fig.35 This is a flowchart showing an example of a moving picture decoding process according to the present embodiment.
[0325] In step S601, the decoding device decodes a set of parameters from the header. Fig.12 The decoded parameters include one or more parameters related to an identification code indicating whether the image is captured by a non-linear lens. Fig.13 As shown in FIG. 1 , the captured image may be distorted by the characteristics of the lens used for capturing the image. An example of the decoded parameter is a parameter indicating the position of the center or main axis of the distortion.
[0326] Next, in step S602, the decoding device decodes the image by an adaptive motion picture decoding tool based on the decoded parameters. The adaptive motion picture decoding tool includes a motion vector prediction process. The adaptive motion picture decoding tool may also include an intra-frame prediction process. In addition, the motion picture decoding tool or the adaptive motion picture decoding tool is the same as or corresponds to the above-mentioned motion picture encoding tool or the adaptive motion picture encoding tool.
[0327] <Intra-frame prediction processing>
[0328] Fig.32 is a flowchart showing adaptive intra-frame prediction processing based on the interpreted parameters. Fig.32 As shown, in step S2201, the decoding device determines a position in the image as a distortion center or a principal point based on the decoded parameters. Next, in step S2202, the decoding device predicts a sample group using spatially adjacent pixel values. Finally, in step S2203, the decoding device performs overlapping processing on the predicted sample group using the determined distortion center or principal point to generate a block of predicted samples. For example, the decoding device may also distort the image of the block of predicted samples and use the distorted image as the predicted image.
[0329] <Motion Vector Prediction>
[0330] Fig.33 is a flowchart showing adaptive motion vector prediction processing based on the interpreted parameters. Fig.33 As shown, in step S2301, the decoding device determines a position in the image as a distortion center or a principal point based on the decoded parameters. Next, in step S2302, the decoding device predicts a motion vector based on a motion vector adjacent in space or time. Finally, in step S2303, the decoding device uses the determined distortion center or principal point to correct the direction of the motion vector.
[0331] [Decoding device]
[0332] Fig.36 This is a block diagram showing the structure of a decoding device for decoding a moving picture according to the present embodiment.
[0333] The decoding device 1400 is a device for decoding an input coded moving picture (i.e., an input bit stream) for each block and outputting the decoded moving picture, and is equivalent to the decoding device 200 of the first embodiment. Fig.36 As shown, the decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, a block memory 1404, a frame memory 1405, an addition unit 1422, an intra-frame prediction unit 1406 and an inter-frame prediction unit 1407.
[0334] The input bit stream is input to the entropy decoding unit 1401. Then, the entropy decoding unit 1401 performs entropy decoding on the input bit stream and outputs the value obtained by the entropy decoding (i.e., the quantized value) to the inverse quantization unit 1402. The entropy decoding unit 1401 also interprets parameters from the input bit stream and outputs the parameters to the inter-frame prediction unit 1407 and the intra-frame prediction unit 1406.
[0335] The inverse quantization unit 1402 inversely quantizes the value obtained by entropy decoding, and outputs the frequency coefficient to the inverse transformation unit 1403. The inverse transformation unit 1403 performs inverse frequency transformation on the frequency coefficient, transforms the frequency coefficient into a sample value (i.e., a pixel value), and outputs the resulting pixel value to the addition unit 1422. The addition unit 1422 adds the resulting pixel value to the pixel value of the predicted image output from the intra-frame prediction unit 1406 or the inter-frame prediction unit 1407. The addition unit 1422 outputs the value obtained by the addition (i.e., the decoded image) to the display, and outputs the obtained value to the block memory 1404 or the frame memory 1405 for further prediction.
[0336] The intra prediction unit 1406 performs intra prediction. That is, the intra prediction unit 1406 predicts the image of the processing target block using the reconstructed image included in the same picture as the picture of the processing target block stored in the block memory 1404. The inter prediction unit 1407 performs inter prediction. That is, the inter prediction unit 1407 estimates the image of the processing target block using the reconstructed image included in the picture different from the picture of the processing target block stored in the frame memory 1405.
[0337] Here, in this embodiment, the inter-frame prediction unit 1407 and the intra-frame prediction unit 1406 adapt the processing based on the decoded parameters. That is, the inter-frame prediction unit 1407 and the intra-frame prediction unit 1406 perform the processing according to the adaptive video decoding tool. Fig.32 and Fig.33 Processing of the flowchart shown.
[0338] (Summarize)
[0339] As mentioned above, although an example of the encoding device and the decoding device of the present disclosure has been described using each embodiment, the encoding device and the decoding device according to one aspect of the present disclosure are not limited to these embodiments.
[0340] For example, in each of the above-mentioned embodiments, the encoding device encodes the moving image using parameters related to image distortion or parameters related to image splicing, and the decoding device decodes the encoded moving image using these parameters. However, the encoding device and decoding device of one form of the present disclosure may not perform encoding or decoding using these parameters. In other words, the processing using the adaptive moving image encoding tool and the adaptive moving image decoding tool of the above-mentioned embodiments may not be performed.
[0341] Fig.37 This is a block diagram of an encoding device according to one aspect of the present disclosure.
[0342] The encoding device 1500 according to one aspect of the present disclosure is a device equivalent to the encoding device 100 according to Embodiment 1. Fig.37 As shown, the encoding device 1500 includes a transform unit 1501, a quantization unit 1502, an inverse quantization unit 1503, an inverse transform unit 1504, a block memory 1505, a frame memory 1506, an intra-frame prediction unit 1507, an inter-frame prediction unit 1508, a subtraction unit 1521, an addition unit 1522, and an entropy coding unit 1509. In addition, the encoding device 1500 may not include the parameter derivation units 910, 1110, and 1310.
[0343] The above-mentioned components included in the encoding device 1500 perform the same processing as in the above-mentioned embodiments 1 to 4, but do not perform processing using the adaptive moving picture coding tool. That is, the adding unit 1522, the intra-frame prediction unit 1507, and the inter-frame prediction unit 1508 perform processing for encoding without using the parameters derived by the parameter derivation units 910, 1110, and 1310 of the embodiments 2 to 4, respectively.
[0344] In addition, the encoding device 1500 obtains a moving image and parameters related to the moving image, encodes the moving image without using the parameters, generates a bitstream, and writes the parameters to the bitstream. Specifically, the entropy encoding unit 1509 writes the parameters to the bitstream. In addition, the position of the parameters written to the bitstream can be any position.
[0345] In addition, each image (i.e., picture) included in the above-mentioned moving image input to the encoding device 1500 may be an image after distortion correction, or may be a joined image obtained by joining images from multiple viewpoints. The distortion-corrected image is a rectangular image obtained by correcting the distortion of an image photographed by a wide-angle lens such as a non-linear lens. Such an encoding device 1500 encodes a moving image including the distortion-corrected image or the joined image.
[0346] Here, the quantization unit 1502, the inverse quantization unit 1503, the inverse transformation unit 1504, the intra prediction unit 1507, the inter prediction unit 1508, the subtraction unit 1521, the addition unit 1522, and the entropy coding unit 1509 are configured as a processing circuit, for example. Furthermore, the block memory 1505 and the frame memory 1506 are configured as memories.
[0347] That is, the encoding device 1500 includes a processing circuit and a memory connected to the processing circuit. The processing circuit uses the memory to obtain parameters related to at least one of a first process of correcting distortion of an image captured by a wide-angle lens and a second process of joining a plurality of images, generates an encoded image by encoding an image to be processed based on the image or the plurality of images, and writes the parameters into a bit stream including the encoded image.
[0348] As a result, the above-mentioned parameters are written into the bit stream, so that the coded or decoded image can be appropriately handled by using the parameters.
[0349] Here, in writing the parameter, the parameter may be written into a header in a bitstream. In addition, in encoding the image to be processed, the block may be encoded by adapting the encoding process based on the parameter to each block included in the image to be processed. Here, the encoding process may include at least one of an inter-frame prediction process and an image reconstruction process.
[0350] Thus, for example, as in Embodiment 2, by using inter-frame prediction processing and image reconstruction processing as adaptive moving picture coding tools, it is possible to appropriately code a processing target image such as a distorted image or a spliced image. As a result, it is possible to improve the coding efficiency of the processing target image.
[0351] Alternatively, in the writing of parameters, the parameters related to the second processing are written into the header in the bit stream, and in the encoding of the image of the processing object, the encoding processing of each block contained in the image of the processing object obtained by the second processing is omitted based on its parameters.
[0352] Thus, for example, as in Embodiment 2 Fig.21 and Fig. 22 As shown, it is possible to omit the encoding of each block included in an image that the user will not focus on in the near future among the plurality of images included in the joined image. As a result, it is possible to reduce the processing load and the amount of code.
[0353] Furthermore, in writing the parameters, at least one of the positions and camera angles of the plurality of cameras may be written into the header of the bitstream as a parameter related to the second process. Furthermore, in encoding the image to be processed, the image to be processed, which is one of the plurality of images, may be encoded and the image to be processed may be joined with the other images of the plurality of images using the parameters written into the header.
[0354] As a result, for example, as in Embodiment 3, a larger image obtained by concatenation can be used in inter-frame prediction or motion compensation, thereby improving encoding efficiency.
[0355] Furthermore, in writing the parameters, as parameters related to the first process, at least one of a parameter indicating whether the image is captured by a wide-angle lens and a parameter related to distortion aberration generated by the wide-angle lens may be written into the header of the bitstream. Furthermore, in encoding the image to be processed, the block may be encoded by adapting the encoding process based on the parameters written in the header to each block included in the image to be processed as the image captured by the wide-angle lens. Here, the encoding process may include at least one of a motion vector prediction process and an intra-frame prediction process.
[0356] Thus, for example, by using motion vector prediction processing and intra-frame prediction processing as adaptive motion picture coding tools as in Embodiment 4, it is possible to appropriately code a processing target image such as a distorted image. As a result, it is possible to improve coding efficiency of distorted images.
[0357] Furthermore, the encoding process may include a prediction process of one of an inter-frame prediction process and an intra-frame prediction process, and the prediction process may include a superposition process as a process of arranging or rearranging a plurality of pixels included in an image.
[0358] Thus, for example, as in Embodiment 2, the distortion of the image to be processed can be corrected, and inter-frame prediction processing can be appropriately performed based on the corrected image. In addition, for example, as in Embodiment 4, intra-frame prediction processing can be performed on the distorted image, and the predicted image obtained by the processing can be appropriately distorted to match the distorted image to be processed. As a result, the coding efficiency of the distorted image can be improved.
[0359] Furthermore, the encoding process may also include an inter-frame prediction process, which is a process for curved, slanted, or angular image boundaries, including a padding process for the image using the parameters written to the above-mentioned header.
[0360] As a result, for example, as in Embodiment 2, inter-frame prediction processing can be appropriately performed, and encoding efficiency can be improved.
[0361] Furthermore, the encoding process may include an inter-frame prediction process and an image reconstruction process, each of which includes a process for replacing a pixel value with a predetermined value based on a parameter written in the header.
[0362] As a result, for example, as in the second embodiment, inter-frame prediction processing and image reconstruction processing can be appropriately performed, and encoding efficiency can be improved.
[0363] Furthermore, in encoding of an image to be processed, the encoded image to be processed may be reconstructed, and an image obtained by joining the reconstructed image to be processed and the other images may be stored in a memory as a reference frame used in an inter-frame prediction process.
[0364] As a result, for example, as in Embodiment 3, a larger image obtained by joining can be used for inter-frame prediction or motion compensation, thereby improving encoding efficiency.
[0365] In addition, the encoding device of the above-mentioned embodiments 2 to 4 encodes a moving image including a distorted image, a moving image including a joined image, or a moving image including images from multiple viewpoints that have not been joined. However, the encoding device of the present disclosure may correct the distortion of the image included in the moving image for the encoding of the moving image, or may not correct the distortion. In the case of not correcting the distortion, the encoding device pre-acquires a moving image including an image in which the distortion has been corrected by another device, and encodes the moving image. Similarly, the encoding device of the present disclosure may splice images from multiple viewpoints included in the moving image for the encoding of the moving image, or may not splice them. In the case of not splicing, the encoding device pre-acquires a moving image including an image in which images from multiple viewpoints have been spliced by another device, and encodes the moving image. In addition, the encoding device of the present disclosure may perform all or only a part of the correction of the distortion. Furthermore, the encoding device of the present disclosure may perform all or only a part of the splicing of images from multiple viewpoints.
[0366] Fig.38 This is a block diagram of a decoding device according to one aspect of the present disclosure.
[0367] A decoding device 1600 according to one aspect of the present disclosure is a device equivalent to the decoding device 200 according to Embodiment 1. Fig.38 As shown, it includes an entropy decoding unit 1601, an inverse quantization unit 1602, an inverse transformation unit 1603, a block memory 1604, a frame memory 1605, an intra-frame prediction unit 1606, an inter-frame prediction unit 1607 and an addition unit 1622.
[0368] The above components included in the decoding device 1600 perform the same processing as in the above embodiments 1 to 4, but do not perform processing using the adaptive video decoding tool. That is, the adding unit 1622, the intra prediction unit 1606, and the inter prediction unit 1607 perform processing for decoding without using the above parameters included in the bitstream.
[0369] In addition, the decoding device 1600 obtains a bit stream, extracts a coded moving image and parameters from the bit stream, and decodes the coded moving image without using its parameters. Specifically, the entropy decoding unit 1601 interprets the parameters from the bit stream. In addition, the position of the parameters written in the bit stream can be any position.
[0370] In addition, each image (i.e., encoded picture) included in the bit stream input to the decoding device 1600 may be an image after distortion correction or a joined image obtained by joining images from multiple viewpoints. The distortion-corrected image is a rectangular image obtained by correcting the distortion of an image photographed by a wide-angle lens such as a non-linear lens. Such a decoding device 1600 decodes a moving image including the distortion-corrected image or the joined image.
[0371] Here, the entropy decoding unit 1601, the inverse quantization unit 1602, the inverse transformation unit 1603, the intra prediction unit 1606, the inter prediction unit 1607, and the addition unit 1622 are configured as a processing circuit, for example. Furthermore, the block memory 1604 and the frame memory 1605 are configured as memories.
[0372] That is, the decoding device 1600 includes a processing circuit and a memory connected to the processing circuit. The processing circuit uses the memory to obtain a bit stream including a coded image, interprets parameters related to at least one of a first process of correcting distortion of an image captured by a wide-angle lens and a second process of joining a plurality of images from the bit stream, and decodes the coded image.
[0373] Thus, by using the above-mentioned parameters interpreted from the bit stream, it is possible to appropriately handle the encoded or decoded image.
[0374] Here, in the interpretation of the parameter, the parameter may be interpreted from the header in the bitstream. In addition, in the decoding of the coded image, the block may be decoded by adapting the decoding process based on the parameter to the block according to the block included in the coded image. Here, the decoding process may also include at least one of an inter-frame prediction process and an image reconstruction process.
[0375] Thus, for example, as in the second embodiment, by using the inter-frame prediction process and the image reconstruction process as adaptive moving picture decoding tools, for example, a distorted image or a coded image that is a joined image can be appropriately decoded.
[0376] Alternatively, in the interpretation of parameters, parameters related to the second process are interpreted from a header in the bit stream, and in the decoding of the encoded image, decoding processing for each block included in the encoded image generated by encoding the image obtained by the second process is omitted based on the parameters.
[0377] Thus, for example, as in Embodiment 2 Fig.21 and Fig. 22 As shown, it is possible to omit decoding of each block included in an image that the user will not focus on in the near future among a plurality of images included in a joined image as a coded image. As a result, it is possible to reduce the processing load.
[0378] Furthermore, in the interpretation of the parameters, at least one of the positions and camera angles of each of the plurality of cameras may be interpreted from a header in the bit stream as a parameter related to the second process. Furthermore, in the decoding of the coded image, a coded image generated by encoding one of the plurality of images may be decoded, and the decoded coded image may be joined with other images of the plurality of images using the parameters interpreted from the header.
[0379] As a result, for example, as in Embodiment 3, a larger image obtained by joining can be used for inter-frame prediction or motion compensation, and a bit stream with improved coding efficiency can be appropriately decoded.
[0380] Furthermore, in the interpretation of the parameters, as parameters related to the above-mentioned first process, at least one of a parameter indicating whether the image is photographed by a wide-angle lens and a parameter related to distortion aberration generated by the wide-angle lens may be interpreted from a header in the bitstream. Furthermore, in the decoding of the encoded image, for each block included in the encoded image generated by encoding the image photographed by a wide-angle lens, a decoding process based on the parameters interpreted from the header may be adapted to the block, thereby decoding the block. Here, the decoding process may also include at least one of the motion vector prediction process and the intra-frame prediction process.
[0381] Thus, for example, as in the fourth embodiment, by using the motion vector prediction process and the intra-frame prediction process as adaptive motion picture decoding tools, it is possible to appropriately decode a coded image that is, for example, a distorted image.
[0382] Furthermore, the decoding process may include a prediction process of one of an inter-frame prediction process and an intra-frame prediction process, and the prediction process may include a superposition process as a process of arranging or rearranging a plurality of pixels included in an image.
[0383] Thus, for example, as in Embodiment 2, the distortion of the coded image can be corrected, and inter-frame prediction processing can be appropriately performed based on the corrected image. Also, for example, as in Embodiment 4, intra-frame prediction processing can be performed on the distorted coded image, and the resulting predicted image can be appropriately distorted to match the distorted coded image. As a result, the coded image, which is a distorted image, can be appropriately predicted.
[0384] Furthermore, the decoding process may also include an inter-frame prediction process, which is a process for curved, inclined or angular image boundaries, including a padding process for the image using parameters read from the above-mentioned header.
[0385] This makes it possible to appropriately perform inter-frame prediction processing as in the second embodiment, for example.
[0386] Furthermore, the decoding process may include an inter-frame prediction process and an image reconstruction process, each of which includes a process for replacing a pixel value with a predetermined value based on a parameter read from the header.
[0387] Thereby, for example, as in the second embodiment, it is possible to appropriately perform inter-frame prediction processing and image reconstruction processing.
[0388] Furthermore, in decoding of a coded image, the coded image may be decoded, and an image obtained by joining the decoded coded image and the other image may be stored in a memory as a reference frame used in an inter-frame prediction process.
[0389] As a result, for example, as in Embodiment 3, a larger image obtained by joining can be used for inter-frame prediction or motion compensation.
[0390] In addition, the decoding devices of the above-mentioned embodiments 2 to 4 decode a bit stream containing a distorted image, a bit stream containing a joined image, or a bit stream containing unjoined images from multiple viewpoints. However, the decoding device of the present disclosure may correct the distortion of the image contained in the bit stream for decoding of the bit stream, or may not correct the distortion. In the case of not correcting the distortion, the decoding device pre-acquires a bit stream containing an image whose distortion is corrected by another device, and decodes the bit stream. Similarly, the decoding device of the present disclosure may splice images from multiple viewpoints contained in the bit stream for decoding of the bit stream, or may not splice them. In the case of not splicing, the decoding device acquires a bit stream containing a larger image generated by splicing images from multiple viewpoints in advance by another device, and decodes the bit stream. In addition, the decoding device of the present disclosure may perform all or only part of the correction of the distortion. Furthermore, the decoding device of the present disclosure may perform all or only part of the splicing of images from multiple viewpoints.
[0391] (Other embodiments)
[0392] In each of the above embodiments, the functional blocks can be generally implemented by an MPU and a memory. In addition, each of the functional blocks is generally implemented by a program execution unit of a processor or the like reading and executing software (program) recorded in a recording medium such as a ROM. The software can be distributed by downloading or the like, or can be distributed by recording in a recording medium such as a semiconductor memory. In addition, each functional block can of course be implemented by hardware (dedicated circuit).
[0393] In addition, the processing described in each embodiment can be implemented by centralized processing using a single device (system), or can be implemented by decentralized processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing can be performed or decentralized processing can be performed.
[0394] The present disclosure is not limited to the above-described embodiments, and various modifications can be made, which are also included in the scope of the present disclosure.
[0395] Furthermore, an application example of the moving picture encoding method (image encoding method) or the moving picture decoding method (image decoding method) shown in each of the above embodiments and a system using the same are described here. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding and decoding device having both. Other structures in the system can be appropriately changed according to the situation.
[0396] [Example of use]
[0397] Fig.39 1 is a diagram showing the overall configuration of a content providing system ex100 for realizing content distribution services. A communication service providing area is divided into cells of desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.
[0398] In the content providing system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content providing system ex100 may be connected by combining some of the above elements. Instead of the base stations ex106 to ex110, which are fixed wireless stations, each device may be directly connected to the communication network ex104 such as a telephone line, a cable TV, or an optical communication. Furthermore, each device may be directly or indirectly connected to each other via a telephone network or short-distance wireless communication. In addition, the streaming media server ex103 is connected to each device such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101. Furthermore, the streaming server ex103 is connected to a terminal or the like in a hotspot in an airplane ex117 via a satellite ex116.
[0399] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly connected to the aircraft ex117 without going through the satellite ex116.
[0400] The camera ex113 is a device capable of taking still images and moving images like a digital camera, etc. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that supports mobile communication systems generally called 2G, 3G, 3.9G, 4G, and 5G in the future.
[0401] The home appliance ex118 is a refrigerator or equipment included in a household fuel cell cogeneration system.
[0402] In the content supply system ex100, by connecting a terminal having a camera function to the streaming server ex103 via the base station ex106 or the like, on-site distribution can be performed. In on-site distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) performs the encoding process described in the above-mentioned embodiments on still images or moving image contents photographed by a user using the terminal, multiplexes the image data obtained by the encoding and the audio data obtained by encoding the audio corresponding to the image, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.
[0403] On the other hand, the streaming server ex103 streams the content data sent by the client that requested it. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc. that can decode the coded data. Each device that receives the distributed data decodes and reproduces the received data. That is, each device functions as an image decoding device according to one technical solution of the present invention.
[0404] [Distributed processing]
[0405] In addition, the streaming media server ex103 can also be a plurality of servers or a plurality of computers, and is a structure that distributes data by distributing processing or recording. For example, the streaming media server ex103 can also be implemented by CDN (Contents Delivery Network), and content distribution is achieved by connecting many edge servers scattered in the world to a network connected between edge servers. In CDN, physically closer edge servers are dynamically allocated according to the client. And, by caching and distributing the content to the edge server, delay can be reduced. In addition, in the case of a certain error or when the communication state changes due to an increase in the amount of communication, it is possible to distribute the processing with multiple edge servers, or switch the distribution subject to other edge servers, or bypass the part of the network where the failure occurred and continue to distribute, so that high-speed and stable distribution can be achieved.
[0406] In addition, the distributed processing of the distribution itself is not limited to the processing. The encoding processing of the photographic data can be performed by each terminal, on the server side, or shared. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or code amount of the image in the frame or scene unit is detected. In addition, in the second cycle, processing is performed to maintain the image quality and improve the encoding efficiency. For example, the first encoding process is performed by the terminal, and the second encoding process is performed by the server side that receives the content. It is possible to improve the quality and efficiency of the content while reducing the processing load in each terminal. In this case, if there is a requirement for receiving and decoding in approximately real time, the first encoded data performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.
[0407] As another example, the camera ex113 or the like extracts feature quantities from an image, compresses data on the feature quantities as metadata, and transmits the data to the server. The server determines the importance of an object based on the feature quantities, switches the quantization accuracy, and performs compression corresponding to the meaning of the image. Feature quantity data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression in the server. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding) and the server may perform encoding with a heavy processing load such as CABAC (context adaptive binary arithmetic coding).
[0408] As another example, in a stadium, shopping mall, or factory, there are multiple video data of roughly the same scene captured by multiple terminals. In this case, multiple terminals that captured the video and other terminals and servers that did not capture the video are used as needed, and the encoding process is distributed, for example, by GOP (Group of Picture) units, picture units, or tile units obtained by dividing the picture. This can reduce delays and achieve better real-time performance.
[0409] In addition, since the plurality of image data are substantially the same scene, the server may also manage and / or instruct the image data photographed by each terminal to refer to each other. Alternatively, the server may receive the encoded data from each terminal and change the reference relationship between the plurality of data, or may modify or replace the image itself and re-encode it. In this way, a stream with improved quality and efficiency of each data may be generated.
[0410] In addition, the server may perform transcoding to change the encoding method of the video data and distribute the video data. For example, the server may convert the encoding method of the MPEG system into VP profile, or convert H.264 into H.265.
[0411] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following uses "server" or "terminal" as the subject of the process, but part or all of the process performed by the server can also be performed by the terminal, and part or all of the process performed by the terminal can also be performed by the server. In addition, the same is true for the decoding process.
[0412] [3D, multi-angle]
[0413] In recent years, there has been an increasing number of cases where images or videos taken by terminals such as multiple cameras ex113 and / or smartphone ex115 that are substantially synchronized with each other, or images or videos taken by terminals at different angles, are combined and used. The images taken by the terminals are combined based on the relative positional relationship between the terminals obtained separately, or on regions where feature points in the images are consistent.
[0414] The server can not only encode two-dimensional moving images, but also encode still images automatically or at a user-specified time based on scene analysis of moving images and send them to the receiving terminal. When the server is able to obtain the relative positional relationship between the photographing terminals, it can generate not only two-dimensional moving images but also three-dimensional shapes of the scene based on images taken from the same scene but from different angles. In addition, the server can also encode three-dimensional data generated by point clouds, etc., and can also select or reconstruct images from images taken by multiple terminals based on the results of identifying or tracking people or objects using three-dimensional data to generate images to be sent to the receiving terminal.
[0415] In this way, the user can arbitrarily select each image corresponding to each camera terminal to enjoy the scene, and can also enjoy the content of the image cut out of any viewpoint from the three-dimensional data reconstructed using multiple images or images. Furthermore, like the image, the sound can also be collected from multiple different angles, and the server matches the image and multiplexes the sound from a specific angle or space and sends it.
[0416] In addition, in recent years, content that establishes correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server produces viewpoint images for the right eye and the left eye respectively, and can be encoded to allow reference between viewpoint images through Multi-View Coding (MVC) or encoded as different streams without reference to each other. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.
[0417] In the case of AR images, the server may also superimpose virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device obtains or maintains the virtual object information and three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and produces superimposed data by smoothly connecting them. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to the entrustment of the virtual object information. The server produces superimposed data according to the three-dimensional data maintained in the server, matching the received viewpoint movement, encoding the superimposed data and distributing it to the decoding device. In addition, the superimposed data has an alpha value that presents transmittance in addition to RGB. The server sets the alpha value of the part other than the object produced according to the three-dimensional data to 0, etc., and encodes the part in a transparent state. Alternatively, the server may also set the RGB value of a specified value as the background, as in the case of a chroma key, to generate data in which the part other than the object is the background color.
[0418] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, or it can be performed on the server side, or it can be shared and performed. As an example, a terminal may first send a request to the server, and other terminals may receive the content corresponding to the request and perform decoding processing, and send the decoded signal to a device with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communicative terminal itself, data with better image quality can be reproduced. In addition, as another example, larger-sized image data can also be received by a TV, etc., and a portion of the area such as the tiles after the image is divided can be decoded and displayed by the personal terminal of the viewer. In this way, while making the overall image shared, it is possible to confirm one's own area of responsibility or the area that wants to be confirmed in more detail at hand.
[0419] In addition, in the future, when multiple short-range, medium-range or long-range wireless communications can be used both indoors and outdoors, it is expected that the distribution system standards such as MPEG-DASH will be used to seamlessly receive content while switching appropriate data for the communication being connected. As a result, the user can not only use his own terminal, but also freely select a decoding device or display device such as a display set indoors or outdoors and switch in real time. In addition, it is possible to decode while switching the decoding terminal and the display terminal based on its own location information. As a result, it is also possible to display map information on a part of the wall or ground of a building next to a displayable device while moving to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as being able to cache the encoded data from the receiving terminal to an accessible server in a short time, or copying it to the edge server of the content distribution service.
[0420] [Classification Code]
[0421] To switch content, use Fig.40 1 and 2, which are compression-encoded hierarchical (scalable) streams applied with the moving picture coding method described in the above embodiments. The server may have multiple streams with the same content but different qualities as a single stream, but may also have a structure to switch the content by utilizing the characteristics of hierarchical streams in time and space realized by layered coding as shown in the figure. That is, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band, so that the decoding side can freely switch between decoding low-resolution content and high-resolution content. For example, if a video that is viewed on a smartphone ex115 while on the move is to be viewed on a device such as an Internet TV after returning home, the device only needs to decode the same stream into different layers, thereby reducing the burden on the server side.
[0422] Furthermore, in addition to the hierarchical structure in which the image is encoded for each layer and the extended layer exists above the base layer as described above, the extended layer may include meta-information such as statistical information based on the image, and the decoding side may generate high-definition content by super-resolving the image of the base layer based on the meta-information. The so-called super-resolution may be either improvement of the SN ratio or expansion of the resolution at the same resolution. The meta-information includes information for determining linear or nonlinear filter coefficients used in super-resolution processing, or information for determining parameter values in filter processing, machine learning, or least squares calculation used in super-resolution processing.
[0423] Alternatively, a picture may be divided into tiles according to the meaning of an object in the image, and the decoding side may decode only a part of the area by selecting the tile to be decoded. In addition, by storing the attributes of the object (a person, a car, a ball, etc.) and the position in the image (the coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and determine the tile that includes the object. For example, Fig.41 As shown, the meta information is stored using a data storage structure different from pixel data such as SEI messages in HEVC. The meta information indicates, for example, the position, size, or color of the main object.
[0424] In addition, the meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding side can obtain the time when a specific person appears in the image, and by matching the information of the picture unit, it can determine the picture where the object exists and the position of the object in the picture.
[0425] [Web page optimization]
[0426] Fig.42 This is a diagram showing an example of a display screen of a web page on the computer ex111 or the like. Fig.43 1 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. Fig.42 and Fig.43 As shown, there are cases where a web page includes a plurality of link images that are links to image contents, and the way in which they are visible differs depending on the device being browsed. When a plurality of link images are visible on the screen, before a user explicitly selects a link image, or before a link image approaches the center of the screen, or before the entire link image enters the screen, the display device (decoding device) displays a still image or I picture of each content as a link image, or displays an image such as a GIF animation using a plurality of still images or I pictures, or receives only the base layer and decodes and displays the image.
[0427] When a linked image is selected by the user, the display device decodes the base layer with the highest priority. In addition, if there is information indicating that the content is hierarchical in the HTML constituting the web page, the display device can also decode to the extended layer. In addition, in order to ensure real-time performance, before selection or when the communication band is very tight, the display device can reduce the delay between the decoding time and the display time of the head picture (the delay from the start of decoding of the content to the start of display) by decoding and displaying only the pictures for forward reference (I pictures, P pictures, and B pictures for forward reference only). In addition, the display device can also forcibly ignore the reference relationship of the pictures and set all B pictures and P pictures as forward references and decode them roughly, and perform normal decoding as the number of pictures received increases over time.
[0428] [Automatic driving]
[0429] Furthermore, when still images or video data such as two-dimensional or three-dimensional map information are transmitted and received for the purpose of automatic driving or driving support of a vehicle, the receiving terminal may receive weather or construction information as meta-information in addition to image data belonging to one or more layers, and decode them by establishing a correspondence between them. Furthermore, the meta-information may belong to a layer or may be multiplexed with image data alone.
[0430] In this case, since the car, drone, or airplane including the receiving terminal is moving, the receiving terminal can seamlessly receive and decode while switching base stations ex106 to ex110 by transmitting the location information of the receiving terminal when receiving a request. In addition, the receiving terminal can dynamically switch the degree to which meta-information is received or the degree to which map information is updated according to the user's selection, the user's condition, or the state of the communication band.
[0431] As described above, in the content providing system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0432] [Distribution of Personal Content]
[0433] Furthermore, in the content supply system ex100, not only high-quality, long-duration contents provided by video distributors but also low-quality, short-duration contents provided by individuals can be unicasted or multicasted. It is expected that such personal contents will increase in the future. In order to make personal contents better, the server may perform encoding after editing. This can be achieved, for example, by the following structure.
[0434] After the shooting is done in real time or stored, the server performs recognition processing such as shooting errors, scene search, meaning analysis and object detection based on the original image or encoded data. In addition, based on the recognition results, the server manually or automatically corrects focus deviation or hand shaking, deletes less important scenes such as scenes with lower brightness than other pictures or scenes that are not in focus, or emphasizes the edges of the object, or changes the color tone. The server encodes the edited data based on the editing results. In addition, it is known that if the shooting time is too long, the viewing rate will decrease. The server can also automatically crop not only the less important scenes as mentioned above, but also the scenes with less movement based on the image processing results according to the shooting time, so as to become the content within a specific time range. Alternatively, the server can also generate a summary based on the results of the scene's meaning analysis and encode it.
[0435] In addition, there are cases where content that infringes copyright, author's personality rights or portrait rights is written into personal content as it is, and there are also cases where the scope of sharing exceeds the desired scope, which is inconvenient for individuals. Thus, for example, the server can also forcibly change the faces of people in the peripheral part of the screen, or the home, etc. to an out-of-focus image for encoding. In addition, the server can also identify whether the face of a person different from the pre-registered person is captured in the encoded object image, and if so, apply mosaic or other processing to the face part. Alternatively, as a pre-processing or post-processing of the encoding, the user can also specify the person or background area that he wants to process the image from the perspective of copyright, etc., and the server replaces the specified area with another image, or blurs the focus, etc. If it is a person, it is possible to replace the image of part of the face while tracking the person in the moving image.
[0436] In addition, since the viewing and listening of personal content with a small amount of data has a strong demand for real-time performance, the decoding device first receives, decodes, and reproduces the base layer with the highest priority, although it also depends on the bandwidth. The decoding device may also receive the extended layer during this period, and when the playback is played back more than twice, such as when the playback is looped, the extended layer is also included to play back high-quality images. In this way, if the stream is hierarchically encoded, it is possible to provide an experience in which the stream gradually becomes smoother and the image becomes better, although the moving image is relatively rough when it is not selected or at the beginning of viewing. In addition to hierarchical encoding, if the relatively rough stream played back for the first time and the second stream encoded with reference to the moving image played back for the first time are composed of one stream, the same experience can be provided.
[0437] [Other usage examples]
[0438] In addition, these encoding or decoding processes are usually processed in the LSI ex500 included in each terminal. The LSI ex500 can be a single chip or a structure composed of multiple chips. In addition, the software for image encoding or decoding can be installed in a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by the computer ex111, etc., and the encoding process and decoding process can be performed using the software. Furthermore, when the smart phone ex115 has a camera, the moving image data obtained by the camera can be transmitted. In this case, the moving image data is the data encoded by the LSI ex500 included in the smart phone ex115.
[0439] In addition, LSIex500 may also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. If the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then obtains and reproduces the content.
[0440] Furthermore, the content supply system ex100 is not limited to the content supply system ex100 via the Internet ex101, and at least one of the video encoding devices (video encoding devices) or video decoding devices (video decoding devices) in the above-mentioned embodiments can be incorporated into a digital broadcasting system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by using broadcasting radio waves such as satellites, the content supply system ex100 is different from the structure that is easy for unicast in that it is suitable for multicast, but the same application can be made to the encoding process and the decoding process.
[0441] [Hardware Structure]
[0442] Fig.44 is a diagram showing a smartphone ex115. Fig.45 1 is a diagram showing a configuration example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of shooting video and still images, and a display unit ex458 such as a liquid crystal display for displaying decoded data such as the video shot by the camera unit ex465 and the video received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466 such as an operation panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing coded data or decoded data such as shot video, still images, recorded sound, received video or still images, mails, etc., and a slot unit ex464 as an interface unit with a SIM ex468 for identifying a user and authenticating access to various data such as a network.
[0443] In addition, the main control unit ex460 that comprehensively controls the display unit ex458 and the operation unit ex466, the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / separation unit ex453, the sound signal processing unit ex454, the slot unit ex464, and the memory unit ex467 are interconnected via a bus ex470.
[0444] When the power key is turned on by a user's operation, the power circuit unit ex461 supplies power from the battery pack to each unit, thereby activating the smartphone ex115 to be able to operate.
[0445] The smartphone ex115 performs processes such as phone calls and data communications under the control of a main control unit ex460 including a CPU, ROM, and RAM. During a phone call, a voice signal collected by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, which is then subjected to spectrum diffusion processing by the modulation / demodulation unit ex452, and then subjected to digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, and then transmitted via the antenna ex450. In addition, received data is amplified, subjected to frequency conversion processing and analog-to-digital conversion processing, subjected to spectrum inverse diffusion processing by the modulation / demodulation unit ex452, and then converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. During data communications, text, still image, or video data inputted through the operation unit ex466 of the main unit is sent to the main control unit ex460 through the operation input control unit ex462, and similarly subjected to transmission and reception processing. In the data communication mode, when transmitting video, still images, or video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of taking a video or still image by the camera unit ex465, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a predetermined manner, and performs modulation and conversion processing on the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.
[0446] When receiving an e-mail with attached images and / or sounds, or an image linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of image data and a bit stream of sound data, and supplies the encoded image data to the image signal processing unit ex455 and the encoded sound data to the sound signal processing unit ex454 via the synchronous bus ex470. The image signal processing unit ex455 decodes the image signal by a moving image decoding method corresponding to the moving image encoding method shown in the above-mentioned embodiments, and displays the image or still image included in the linked moving image file on the display unit ex458 via the display control unit ex459. In addition, the sound signal processing unit ex454 decodes the sound signal and outputs the sound from the sound output unit ex457. In addition, since real-time streaming is becoming more and more popular, the reproduction of sound may be socially inappropriate depending on the user's situation. Therefore, as an initial value, it is preferable to have a structure in which the sound signal is not reproduced and only the image data is reproduced. The sound may be reproduced synchronously only when the user performs an operation such as clicking on the video data.
[0447] In the above description, the smartphone ex115 is used as an example. However, as a terminal, three types of installation are conceivable, namely, a transmitting terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcasting system, the description is made on the assumption that multiplexed data in which audio data and the like are multiplexed in video data is received and transmitted. However, in addition to audio data, character data related to the video may be multiplexed in the multiplexed data, or video data itself may be received or transmitted instead of the multiplexed data.
[0448] In addition, although the main control unit ex460 including a CPU is assumed to control the encoding or decoding process, the terminal is often equipped with a GPU. Therefore, a structure can be made to process a larger area together by using a memory shared by the CPU and GPU, or a memory that manages addresses so that they can be used together, and utilizing the performance of the GPU. In this way, the encoding time can be shortened, and real-time performance can be ensured to achieve low latency. In particular, it is efficient if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset) and transformation / quantization processing are performed together by the GPU in units of pictures, etc. instead of the CPU.
[0449] Industrial Applicability
[0450] The present disclosure can be used in, for example, a television, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, or the like, an encoding device that encodes an image or a decoding device that decodes an encoded image.
[0451] Description of symbols
[0452] 1500 Encoding device
[0453] 1501 Transformation Department
[0454] 1502 Quantitative Department
[0455] 1503 Inverse Quantization Unit
[0456] 1504 Inverse Transformation Unit
[0457] 1505 Block Memory
[0458] 1506 Frame Memory
[0459] 1507 Intra-frame prediction unit
[0460] 1508 Inter-frame prediction unit
[0461] 1509 Entropy Coding Department
[0462] 1521 Subtraction Department
[0463] 1522 Addition Department
[0464] 1600 Decoding Device
[0465] 1601 Entropy Decoding Unit
[0466] 1602 Inverse Quantization Department
[0467] 1603 Inverse Transformation Unit
[0468] 1604 Block Memory
[0469] 1605 frame memory
[0470] 1606 Intra-frame prediction unit
[0471] 1607 Inter-frame prediction unit
[0472] 1622 Addition Department
Claims
1. A coding method, characterized in that: obtaining parameters related to a process of joining a plurality of images; Write the above parameters into the bitstream; generating a joined image by joining the plurality of images using the parameters so that objects in the plurality of images are continuous; Generate the above bit stream by entropy coding; as well as The above-mentioned joined image is saved as a reference frame to be used in the inter-frame prediction process, wherein The parameters are used to determine multiple positions of the multiple images in the joined image.
2. A decoding method, characterized in that: Entropy decode the current picture from the bitstream; parsing parameters related to a process of joining a plurality of images from the bit stream; generating a joined image by joining the plurality of images using the parameters so that objects in the plurality of images are continuous; storing the above-mentioned joined image as a reference frame to be used in inter-frame prediction processing; and The inter-frame prediction process is performed on the current picture after entropy decoding using the reference frame, wherein The parameters are used to determine multiple positions of the multiple images in the joined image.
Citation Information
Patent Citations
Moving image encoding apparatus and moving image encoding method
CN101960856A