Image decoding device, image encoding device, and bitstream generating device

The image encoding method enhances encoding efficiency by selectively using transformation components based on block size and image characteristics, addressing the inefficiencies of higher-order motion information in existing methods.

JP2025102849AActive Publication Date: 2025-07-08SUN PATENT TRUST
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025051129
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-09-30
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2034-07-01

AI Technical Summary

Technical Problem

Existing image encoding methods, such as those in the HEVC standard, face challenges in improving encoding efficiency due to the increased amount of encoded motion information and processing load when using higher-order motion information like affine transformation.

Method used

An image encoding method that selectively uses transformation components, including translation, rotation, scaling, and shear, by encoding selection information to reduce the amount of information and processing load, allowing flexible selection based on block size and image characteristics.

Benefits of technology

The method improves encoding efficiency by reducing the amount of information and processing load while maintaining prediction accuracy, enabling higher-quality image encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102849000001_ABST
    Figure 2025102849000001_ABST
Patent Text Reader

Abstract

To improve coding efficiency.SOLUTION: An image decoding device 200 includes a processing circuit and a storage device accessible from the processing circuit, and the processing circuit uses the storage device to decode from bitstream information that is applied to a plurality of blocks included in a sequence and is used to determine the number of parameters to be used in prediction based on an affine transformation, and generates a predicted image for each of the plurality of blocks included in the sequence by performing prediction based on an affine transformation using a number of parameters selected on the basis of the information, and candidates for the number of parameters to be used in prediction include 4.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image encoding method and an image decoding method.

Background Art

[0002] In the HEVC (High Efficiency Video Coding) standard, which is the latest moving image encoding standard, various studies have been conducted to improve the encoding efficiency (see, for example, Non-Patent Document 1). This method is a standard of the ITU-T (International Telecommunication Union Telecommunication Standardization Sector) shown by H.26x and the ISO / IEC standard shown by MPEG-x, and was studied as the next video encoding standard after the standard shown by H.264 / AVC or MPEG-4 AVC.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In such image encoding methods and image decoding methods, it is desired to improve the encoding efficiency.

[0006] Therefore, an object of the present invention is to provide an image encoding method or an image decoding method capable of improving the encoding efficiency.

Means for Solving the Problems

[0007] An image decoding apparatus according to an aspect of the present invention includes a processing circuit and a storage device accessible from the processing circuit. The processing circuit uses the storage device to decode, from a bitstream, information applied to a plurality of blocks included in a sequence and used for determining the number of parameters used in prediction in prediction based on an affine transformation, and generates a predicted image for each of the plurality of blocks included in the sequence by performing prediction based on an affine transformation using the number of parameters selected based on the information. The candidates for the number of parameters used in the prediction include 4.

[0008] Also, an image encoding apparatus according to an aspect of the present invention includes a processing circuit and a storage device accessible from the processing circuit. The processing circuit uses the storage device to determine the number of parameters used in prediction in prediction based on an affine transformation, generates a predicted image for each of the plurality of blocks included in a sequence by performing prediction based on an affine transformation using the determined number of parameters, encodes information applied to the plurality of blocks included in the sequence and used for determining the number of parameters, and the candidates for the number of parameters used in the prediction include 4.

[0009] Note that these general or specific aspects may be implemented in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM (Compact Disc Read Only Memory), or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

Advantages of the Invention

[0010] The present invention can provide an image encoding method or an image decoding method capable of improving encoding efficiency.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38A

Figure 38B

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51A

Figure 51B

BEST MODE FOR CARRYING OUT THE INVENTION

[0012] (Knowledge underlying the invention) In the conventional image encoding method, only the information regarding translation is used as the motion information.

[0013] However, in the captured moving image, when there are movements such as zooming in and out by the camera's zoom operation or rotation of the subject, the movement cannot be appropriately expressed only by translation. Therefore, at the time of encoding, a method of improving the prediction accuracy by reducing the size of the prediction block is used.

[0014] On the other hand, methods of using higher-order motion information such as affine transformation as motion information have also been studied. For example, affine transformation can represent three types of transformations, i.e., zooming in and out, rotation, and shear, in addition to translation, and can represent transformations such as rotation of the subject described above. As a result, the quality of the generated prediction image is improved. Also, since the prediction unit can be increased, the encoding efficiency can be improved.

[0015] However, while translation can be represented by two-dimensional information, affine transformation requires at least six-dimensional information to represent the three types of transformations in addition to translation. Thus, the increase in the number of dimensions required for motion information causes problems such as an increase in the amount of encoded motion information and an increase in the amount of calculation required for the estimation process of motion information.

[0016] Regarding such problems, the following method is disclosed in the prior art (Patent Document 1). At the time of encoding, in each prediction block, estimation of motion information related only to translation and estimation of higher-order motion information such as affine transformation are performed. Among these, the method determined to have higher encoding efficiency is selected. A flag indicating whether the motion information is translation or affine transformation and the motion information corresponding to the flag are encoded. As a result, while taking advantage of the benefits of using higher-order motion information, the amount of information can be reduced, so the encoding efficiency is improved. However, this method has a problem that the amount of code when representing higher-order motion information cannot be sufficiently reduced.

[0017] In response to this problem, an image encoding method according to one aspect of the present invention is an image encoding method for encoding an image, including: a selection step of selecting two or more transformation components out of a plurality of transformation components including a translation component and a plurality of non-translation components as reference information indicating a reference destination of a current block to be encoded; a prediction step of generating a predicted image using the reference information; an image encoding step of encoding the current block using the predicted image; a selection information encoding step of encoding selection information for specifying the two or more selected transformation components from among the plurality of transformation components; and a reference information encoding step of encoding the reference information of the current block using reference information of an encoded block different from the current block.

[0018] Accordingly, in this image encoding method, any transformation component can be selected from among a plurality of transformation components including translation and a plurality of non-translation. Therefore, in an encoding method using high-order motion information, the encoding efficiency of this image encoding method can be improved.

[0019] For example, the plurality of non-translation components may include a rotation component, a scaling component, and a shear component.

[0020] For example, the selection information may include a flag corresponding to each of the plurality of transformation components and indicating whether the corresponding transformation component is selected.

[0021] Accordingly, in this image encoding method, an affine matrix can be divided into a plurality of transformation components, and selection and non-selection can be specified for each of them, so that the amount of information can be reduced.

[0022] For example, in the selection step, one encoding level is selected from a plurality of encoding levels indicating different sets, the sets each including some or all of the plurality of transformation components, and the two or more transformation components included in the set indicated by the selected encoding level are selected, and the selection information may indicate the selected encoding level.

[0023] As a result, the image encoding method can further reduce the amount of information. Also, the image encoding method can reduce the processing load for selection.

[0024] For example, in the selection information encoding step, one piece of the selection information that is commonly used for the image including the current block may be encoded.

[0025] As a result, the image encoding method can further reduce the amount of information. Also, the image encoding method can reduce the processing load for selection.

[0026] For example, in the selection step, the two or more transform components may be selected according to the size of the current block, and the selection information may indicate the size.

[0027] As a result, the image encoding method can further reduce the amount of information.

[0028] For example, the plurality of non-translational components include a rotation component, a scaling component, and a shear component. In the selection step, when the size of the current block is less than a first threshold value, the shear component may not be selected.

[0029] As a result, the image encoding method can reduce the processing load more by restricting the transform components that are less likely to contribute to the improvement of the prediction accuracy.

[0030] For example, in the selection step, the two or more transform components may be preferentially selected in the order of the translational component, the rotation component, the scaling component, and the shear component.

[0031] As a result, the image encoding method can achieve more efficient processing by increasing the priority of the transform components that contribute to the improvement of the prediction accuracy.

[0032] Also, an image decoding method according to an aspect of the present invention is an image decoding method for decoding a bitstream obtained by encoding an image, including a selection information decoding step of decoding selection information for specifying two or more transform components from among a plurality of transform components including a translational component and a plurality of non-translational components from the bitstream, a selection step of selecting the two or more transform components specified by the decoded selection information as reference information indicating a reference destination of a current block to be decoded, a reference information decoding step of decoding the reference information of the current block from the bitstream using reference information of a decoded block different from the current block, a prediction step of generating a prediction image using the reference information, and an image decoding step of decoding the current block from the bitstream using the prediction image.

[0033] Thereby, the image decoding method can decode a bitstream with improved encoding efficiency.

[0034] For example, the plurality of non-translational components may include four transform components: a rotation component, a scaling component, and a shear component.

[0035] For example, the selection information may include a flag corresponding to each of the plurality of transform components and indicating whether the corresponding transform component is selected.

[0036] For example, the selection information is a set including some or all of the plurality of transform components, and indicates one of a plurality of encoding levels indicating different sets from each other. In the selection step, the two or more transform components included in the set indicated by the encoding level indicated by the selection information may be selected.

[0037] For example, in the selection information decoding step, one selection information commonly used for the image including the current block may be decoded.

[0038] For example, the selection information indicates the size of the current block, and in the selection step, the two or more conversion components may be selected according to the size of the current block.

[0039] For example, the plurality of non-translational components include a rotation component, a scaling component, and a shear component. In the selection step, when the size of the current block is equal to or less than a first threshold value, the shear component may not be selected.

[0040] For example, the plurality of non-translational components include a rotation component, a scaling component, and a shear component. In the selection step, the two or more conversion components may be preferentially selected in the order of the translational component, the rotation component, the scaling component, and the shear component.

[0041] Further, an image encoding apparatus according to an aspect of the present invention is an image encoding apparatus that encodes an image, and includes a processing circuit and a storage device accessible from the processing circuit. The processing circuit executes the image encoding method using the storage device.

[0042] Accordingly, the image encoding apparatus can select any conversion component from among a plurality of conversion components including translation and a plurality of non-translations. Therefore, the image encoding apparatus can improve the encoding efficiency in an encoding method using high-order motion information.

[0043] Further, an image decoding apparatus according to an aspect of the present invention is an image decoding apparatus that decodes a bit stream obtained by encoding an image, and includes a processing circuit and a storage device accessible from the processing circuit. The processing circuit executes the image decoding method using the storage device.

[0044] Accordingly, the image decoding apparatus can decode a bit stream with improved encoding efficiency.

[0045] Note that these general or specific aspects may be implemented in a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, method, integrated circuit, computer program, and recording medium.

[0046] Hereinafter, embodiments will be described in detail with reference to the drawings as appropriate. However, detailed descriptions of well-known matters and redundant descriptions of substantially the same configurations may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art.

[0047] Note that all of the embodiments described below show specific examples of the present invention. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present invention. In addition, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.

[0048] (Embodiment 1) An embodiment of an image encoding apparatus using the image encoding method according to this embodiment will be described. The image encoding apparatus according to this embodiment selects an arbitrary transform component from among a plurality of transform components represented by an affine transform, and generates a prediction image using the selected transform component and motion information. Further, a bitstream including information indicating the selected transform component is generated. Thereby, the image encoding apparatus can improve the encoding efficiency.

[0049] FIG. 1 is a block diagram of an image encoding apparatus 100 according to this embodiment. The image encoding apparatus 100 includes a block division unit 101, a subtraction unit 102, a frequency conversion unit 103, a quantization unit 104, an entropy encoding unit 105, an inverse quantization unit 106, an inverse frequency conversion unit 107, an addition unit 108, an intra prediction unit 109, a loop filter 110, a frame memory 111, an inter prediction unit 112, and a switching unit 113.

[0050] The image encoding device 100 generates a bitstream 126 by encoding the input image 121.

[0051] FIG. 2 is a flowchart of the encoding process by the image encoding device 100 in the present embodiment.

[0052] First, the block division unit 101 divides the input image 121, which is a still image or a moving image including one or more pictures, into encoding blocks 122 that are units of encoding processing (S101).

[0053] Next, the intra prediction unit 109 or the inter prediction unit 112 generates a prediction block 134 using the decoding block 129 or the decoded image 131 for each encoding block 122 (S102). The details of this process will be described later.

[0054] Next, the subtraction unit 102 generates a difference block 123 that is the difference between the encoding block 122 and the prediction block 134 (S103). The frequency conversion unit 103 generates a coefficient block 124 by frequency-converting the difference block 123. The quantization unit 104 generates a coefficient block 125 by quantizing the coefficient block 124 (S104).

[0055] Next, the entropy encoding unit 105 generates a bitstream 126 by entropy-encoding the coefficient block 125 (S105).

[0056] On the other hand, in order to generate the decoding block 129 and the decoded image 131 used for generating the prediction block 134 of the subsequent block or picture, the inverse quantization unit 106 generates a coefficient block 127 by inverse-quantizing the coefficient block 125. The inverse frequency conversion unit 107 restores the difference block 128 by inverse-frequency-converting the coefficient block 127 (S106).

[0057] Next, the addition unit 108 generates a decoded block 129 (reconstructed image) by adding the prediction block 134 used in step S102 and the difference block 128 (S107). This decoded block 129 is used for intra prediction processing by the intra prediction unit 109. Also, the loop filter 110 generates a decoded image 131 by performing loop filter processing on the decoded block 129. The frame memory 111 stores the decoded image 131. This decoded image 131 is used for inter prediction processing by the inter prediction unit 112.

[0058] These series of processes are repeatedly performed until the encoding process for the entire input image 121 is completed (S108).

[0059] Note that the frequency conversion and quantization in step S104 and the inverse quantization and inverse frequency conversion in step S106 may be sequentially performed as separate processes or may be performed collectively. Also, quantization is a process of digitizing values sampled at a predetermined interval by associating them with predetermined levels. Inverse quantization is a process of returning the values obtained by quantization to the values in the original interval. In the field of data compression, quantization means a process of dividing values into coarser intervals than the original, and inverse quantization means a process of re-dividing the coarser intervals into the original finer intervals. In the field of codec technology, quantization and inverse quantization are sometimes referred to as rounding, truncating, or scaling.

[0060] Next, regarding the prediction block generation process in step S102, an explanation will be given with reference to FIG. 3. FIG. 3 is a flowchart of the prediction block generation process (S102) according to the present embodiment.

[0061] First, the image encoding device 100 determines whether the prediction processing method to be performed on the prediction block to be processed is intra prediction or inter prediction (S121).

[0062] When the prediction processing method is intra prediction (intra prediction in S121), the intra prediction unit 109 generates a prediction block 130 by intra prediction processing (S122). Further, the switching unit 113 outputs the generated prediction block 130 as a prediction block 134.

[0063] On the other hand, when the prediction processing method is inter prediction (inter prediction in S121), the inter prediction unit 112 generates a prediction block 132 by inter prediction processing (S123). Further, the switching unit 113 outputs the generated prediction block 132 as a prediction block 134.

[0064] Note that the image encoding device 100 may not perform step S121, but perform both the processes of steps S122 and S123, perform cost calculation on each of the obtained prediction blocks using an R-D optimization model (the following (Equation 1)), etc., and select a method with a small cost, that is, a high encoding efficiency.

[0065] [Number]

[0066] In (Equation 1), D represents encoding distortion, for example, the sum of absolute differences between the original pixel values of the encoding target block and the generated prediction image. Further, R represents the generated code amount, for example, the code amount required to encode motion information, etc. for generating a prediction block. Further, λ is a Lagrange undetermined multiplier. Thereby, it becomes possible to select an appropriate prediction mode from intra prediction and inter prediction, and the encoding efficiency can be improved.

[0067] Subsequently, regarding the inter prediction processing in step S123, it will be described with reference to FIGS. 4 and 5.

[0068] FIG. 4 is a block diagram showing a configuration example of the inter prediction unit 112. The inter prediction unit 112 includes a coding information selection unit 141, a motion estimation unit 142, a motion compensation unit 143, and a motion information calculation unit 144.

[0069] FIG. 5 is a flowchart of the inter-prediction process (S123) according to the present embodiment.

[0070] First, the encoding information selection unit 141 selects a conversion component (deformation component) of the motion information to be used by using the input image (code block 122) or the like (S141). Here, the conversion component is information indicating various conversions (deformations) (for example, translation, rotation, scaling, and shear), and is a coefficient or the like in various conversions. As shown in FIG. 6, in the present embodiment, in addition to translation, a plurality of non-translation motions such as rotation, scaling, and shear are used for the motion information.

[0071] Next, the motion estimation unit 142 generates motion information 151 by performing a motion estimation process related to the conversion component of the motion information selected in step S141 using the input image (code block 122) and the decoded image 131 (S142).

[0072] Next, the motion compensation unit 143 generates a prediction block 132 by performing a motion compensation process using the motion information 151 obtained in step S142 and the decoded image 131 (S143).

[0073] Next, the motion information calculation unit 144 calculates differential motion information, which is the difference between the motion information 151 obtained in step S142 and the motion information of an encoded block that is spatially or temporally adjacent to the current block to be encoded (S144).

[0074] Thus, the inter-prediction process ends. Note that the differential motion information calculated in step S144 is output as motion information 133 to the entropy encoding unit 105. The entropy encoding unit 105 encodes the motion information 133 and outputs it included in the bit stream 126.

[0075] Next, the selection of the conversion component of the motion information (S141) will be described in detail with reference to FIG. 7. FIG. 7 is a flowchart of the selection of the conversion component of the motion information (S141) according to the present embodiment.

[0076] First, the encoding information selection unit 141 obtains the motion component between the input image (encoding block 122) and the decoded image 131 (S161). For example, the encoding information selection unit 141 extracts the feature amounts of SIFT (Scale-Invariant Feature Transform) for both the input image and the decoded image 131. The encoding information selection unit 141 estimates a homography matrix representing various transformations from the feature amount of the input image and the feature amount of the decoded image 131, and sets it as the motion component between the images. Here, an example using the feature amount of SIFT as the feature amount set as the motion component is shown, but it is not limited to this.

[0077] Next, the encoding information selection unit 141 selects the conversion component of the motion information used for encoding according to the motion component obtained in step S161 (S162). Further, the encoding information selection unit 141 selects the conversion component of the motion information used for encoding according to the size of the prediction block (S163). Through the above series of processes, the process of step S141 ends.

[0078] Note that step S162 and step S163 may be performed in this order, in the reverse order, in parallel, or only one of them may be performed.

[0079] In this embodiment, the "motion component" indicates a feature amount (motion between images) obtained by analyzing an image using a predetermined analysis method, and the "motion information" will be described as indicating a matrix, a vector, or coefficients constituting them, etc. used for encoding, which is determined based on the motion component.

[0080] Hereinafter, details of the selection of the conversion component using the motion component (step S162) and the selection of the conversion component using the size of the prediction block (step S163) will be described.

[0081] FIG. 8 is a flowchart of the selection of the conversion component using the motion component according to this embodiment (step S162).

[0082] First, the coding information selection unit 141 extracts translation, rotation, scaling, and shear transformation components from the motion components (S181).

[0083] Here, the affine matrix a can be expressed as shown in the following (Equation 2). As shown in the following (Equation 2), the affine matrix a is the sum of the product of a matrix representing scaling with a scaling factor k (k in the x - direction and k x in the y - direction), a matrix representing rotation with a rotation angle θ, and matrices representing shear with shear angles φ respectively, and a matrix representing translation by c in the x - direction and f in the y - direction. y ) and a matrix representing translation by c in the x - direction and f in the y - direction.

[0084]

Equation

[0085] From the affine matrix a expressed as in the above (Equation 2), it is possible to extract the scaling factor k, rotation angle θ, shear angle φ, and translation components. The motion information is described as information having the above four transformation components. For example, it is expressed as (c, f, θ, k x , k y , φ).

[0086] When both c and f, which are translation components, are 0, there is no translation component in the motion information. Otherwise, there is a translation component in the motion information. When the scaling factor k is 1, there is no scaling component in the motion information. Otherwise, there is a scaling component in the motion information. When the rotation angle θ is 0, there is no rotation component in the motion information. Otherwise, there is a rotation component in the motion information. When the shear angle φ is 0, there is no shear component in the motion information. Otherwise, there is a shear component in the motion information.

[0087] Note that, here, an example in which the thresholds of the respective conversion components of translation, rotation, scaling, and shear are (0, 0), 0, (1, 1), and 0 has been shown, but the present invention is not limited thereto. These thresholds may be values considering the features of the image, the accuracy of the imaging device, etc. Further, the scaling may not be decomposed into the x-direction and the y-direction, and the same magnification may be used in the x-direction and the y-direction. Thereby, the amount of code and the amount of processing can be reduced.

[0088] In this way, in step S181, the encoding information selection unit 141 extracts various conversion components from the motion information. Note that, here, the description has been made for the case where the motion information is represented by using an affine matrix, but the available motion information is not limited thereto. Hereinafter, it is assumed that the rotation angle represents the rotation component, the magnification represents the scaling component, the shear angle represents the shear component, and the translation vector represents the translation among the various conversion components.

[0089] Subsequently, the encoding information selection unit 141 performs a selection process of determining whether or not to use the translation. First, in step S181, the encoding information selection unit 141 determines whether or not the translation component indicated by the motion vector or the like has been extracted from the motion components (S182).

[0090] When the translation component exists in the motion components (Yes in S182), the encoding information selection unit 141 selects the translation (S183). On the other hand, when the translation component does not exist in the motion components (No in S182), the encoding information selection unit 141 does not select the translation (S184).

[0091] Next, the encoding information selection unit 141 determines whether or not to use the rotation. In step S181, the encoding information selection unit 141 determines whether or not the rotation component indicated by the rotation angle θ or the like has been extracted from the motion components (S185).

[0092] When the rotation component exists in the motion components (Yes in S185), the encoding information selection unit 141 selects the rotation (S186). When the rotation component does not exist in the motion components (No in S185), the encoding information selection unit 141 does not select the rotation (S187).

[0093] Next, the encoding information selection unit 141 determines whether to use scaling. In step S181, the encoding information selection unit 141 determines whether a scaling component indicated by a scaling ratio k or the like has been extracted from the motion component (S188). If a scaling component exists in the motion component (Yes in S188), the encoding information selection unit 141 selects scaling (S189). If no scaling component exists in the motion component (No in S188), the encoding information selection unit 141 does not select scaling (S190).

[0094] Finally, the encoding information selection unit 141 determines whether to use shear. In step S181, the encoding information selection unit 141 determines whether a shear component indicated by a shear angle φ or the like has been extracted from the motion component (S191).

[0095] If a shear component exists in the motion component (Yes in S191), the encoding information selection unit 141 selects shear (S192). If no shear component exists in the motion component (No in S191), the encoding information selection unit 141 does not select shear (S193). Through the above series of processes, the process of step S162 ends.

[0096] Next, with reference to FIG. 9, step S163 shown in FIG. 7 will be described. FIG. 9 is a flowchart of the selection of conversion components (step S163) using the block size according to the present embodiment.

[0097] First, the encoding information selection unit 141 acquires the block size of the prediction block (S201).

[0098] First, the encoding information selection unit 141 determines whether to use translation. In step S201, the encoding information selection unit 141 determines whether the block size obtained is 4×4 blocks or less (S202).

[0099] If the block size is larger than a 4×4 block (No in S202), the encoding information selection unit 141 selects translation (S203). On the other hand, if the block size is 4×4 blocks or less (Yes in S202), the encoding information selection unit 141 does not select translation (S204).

[0100] Subsequently, the encoding information selection unit 141 determines whether to use rotation. The encoding information selection unit 141 determines whether the block size obtained in step S201 is 8×8 blocks or less (S205).

[0101] If the block size is larger than an 8×8 block (No in S205), the encoding information selection unit 141 selects rotation (S206). On the other hand, if the block size is 8×8 blocks or less (Yes in S205), the encoding information selection unit 141 does not select rotation (S207).

[0102] Next, the encoding information selection unit 141 determines whether to use scaling. The encoding information selection unit 141 determines whether the block size obtained in step S201 is 16×16 blocks or less (S208). If the block size is larger than a 16×16 block (No in S208), the encoding information selection unit 141 selects scaling (S209). On the other hand, if the block size is 16×16 blocks or less (Yes in S208), the encoding information selection unit 141 does not select scaling (S210).

[0103] Finally, the encoding information selection unit 141 determines whether to use shear. The encoding information selection unit 141 determines whether the block size obtained in step S201 is 32×32 blocks or less (S211).

[0104] If the block size is larger than a 32×32 block (No in S211), the encoding information selection unit 141 selects shear (S212). On the other hand, if the block size is 32×32 blocks or less (Yes in S211), the encoding information selection unit 141 does not select shear (S213). Through the above series of processes, the process of step S163 ends.

[0105] Here, in steps S202, S205, S208, and S211, cases where 4×4, 8×8, 16×16, and 32×32 are used as examples of the threshold values of the block size have been described. However, the threshold value of the block size is not limited to this and may be any size. Also, the encoding information selection unit 141 may switch the threshold value according to the characteristics of the image. Thereby, the encoding efficiency can be improved.

[0106] Also, here, an example in which steps S162 and S163 are sequentially performed in this order has been described. However, the order may be reversed, or some or all of these processes may be performed simultaneously. Also, only one of steps S162 and S163 may be performed.

[0107] When both steps S162 and S163 are performed, for example, the encoding information selection unit 141 selects the conversion when the conversion component exists and the block size is larger than the threshold value, and does not select the conversion in other cases. In other words, the encoding information selection unit 141 selects the conversion determined to be selected in both steps S162 and S163, and does not select the conversion determined not to be selected in at least one of steps S162 and S163.

[0108] Also, the encoding information selection unit 141 may first perform the size comparison in step S163 and determine the presence or absence of a conversion component (S162) only for the conversion determined to be selected. Alternatively, the encoding information selection unit 141 may first determine the presence or absence of a conversion component in step S162 and perform the block size comparison (S163) only for the conversion in which the conversion component exists.

[0109] Furthermore, the order of steps S182 to S193 shown in FIG. 8 may be arbitrary, and part or all of them may be performed in parallel. Similarly, the order of S202 to S213 shown in FIG. 9 may be arbitrary, and part or all of them may be performed in parallel.

[0110] However, as verified by the inventors, since the shear component has fewer effective images compared to other conversion components, it was difficult to contribute to the improvement of the prediction accuracy. Therefore, the priority of the shear component may be set lower than that of other conversion components. For example, the priority can be lowered by performing the determination process of the shear component after the determination processes of other components, or by setting the threshold value of the block size larger than the threshold values of other components.

[0111] In step S142 following step S141 shown in FIG. 5, the motion estimation unit 142 performs motion estimation on the conversion component selected in step S141. For example, the motion estimation unit 142 calculates a residual signal (difference block) while changing each of a plurality of conversion components such as the magnitude of translation and the magnification rate by a fixed value. Then, the motion estimation unit 142 determines, as the motion information 151, the combination of conversion components corresponding to the smallest residual signal among the obtained plurality of residual signals.

[0112] Next, in step S143, the motion compensation unit 143 generates a prediction image (prediction block 132) from the decoded image 131 using the motion information 151 obtained in step S142.

[0113] Regarding the last step S144, the details will be described with reference to FIGS. 10 to 14.

[0114] FIG. 10 is a flowchart of the differential motion information calculation process (S144) according to the present embodiment.

[0115] First, the motion information calculation unit 144 derives predicted motion information from the motion information used in one or more encoded blocks that are spatially or temporally adjacent to the current block to be encoded (S221). Note that the predicted motion information, similar to the motion information, includes a translation component, a rotation component, a scaling component, and a shear component. Hereinafter, the various conversion components included in the predicted motion information are also referred to as predicted conversion components. Also, the translation component, rotation component, scaling component, and shear component included in the predicted motion information are also referred to as the predicted translation component, predicted rotation component, predicted scaling component, and predicted shear component, respectively.

[0116] For example, the motion information calculation unit 144 acquires a plurality of conversion components of the same type included in the motion information of a plurality of encoded blocks for each type of predicted conversion component, and derives the predicted conversion component of that type from the acquired plurality of conversion components. For example, the motion information calculation unit 144 derives the average value of the conversion components of each of the plurality of encoded blocks that refer to the same reference picture as the predicted conversion component. Note that, as a method for deriving the predicted conversion component, all calculation methods such as those used in HEVC and the like can be applied.

[0117] First, the motion information calculation unit 144 calculates a differential translation component, which is the differential motion information of the translation component (S222).

[0118] FIG. 11 is a flowchart of this calculation process (S222) of the differential translation component.

[0119] First, the motion information calculation unit 144 determines whether there is a translational component in the motion information 151 estimated in step S142 (S241). For example, when both the x-direction and y-direction components included in the translational component are zero, the motion information calculation unit 144 determines that there is no translational component, and when at least one of the x-direction and y-direction components is not zero, the motion information calculation unit 144 determines that there is a translational component. Note that the motion information calculation unit 144 may determine that there is no translational component when the magnitude of the translational component is smaller than a certain predetermined value. Also, the motion information calculation unit 144 may switch the predetermined value according to the characteristics of the image. Further, the motion information calculation unit 144 may determine that there is always a translational component (or there is never a translational component) when the characteristics of the image satisfy a predetermined condition.

[0120] Also, when it is determined in step S184 of FIG. 8 or step S204 of FIG. 9 that translation is not selected, the motion information calculation unit 144 may determine that there is no translational component.

[0121] When there is no translational component in the motion information 151 (No in S241), the motion information calculation unit 144 sets the translation flag indicating "translation exists" to OFF, and assigns the translation flag to the motion information 133 to be encoded (S242).

[0122] When there is a translational component in the motion information 151 (Yes in S241), the motion information calculation unit 144 sets the translation flag to ON, and assigns the translation flag to the motion information 133 (S243).

[0123] Next, the motion information calculation unit 144 determines whether there is a translational component in the predicted motion information derived in step S221 (S244). When there is no translational component in the predicted motion information (No in S244), the motion information calculation unit 144 sets the predicted translational component to 0 (S245).

[0124] On the other hand, when there is a translational component in the predicted motion information (Yes in S244), the motion information calculation unit 144 sets the translational component of the predicted motion information as the predicted translational component (S246).

[0125] Next, the motion information calculation unit 144 calculates a differential translational component by subtracting the predicted translational component obtained in step S245 or S246 from the translational component of the motion information 151 obtained in step S142, and adds the calculated differential translational component to the motion information 133 (S247).

[0126] Subsequently, the motion information calculation unit 144 calculates a differential translational component which is the differential motion information of the rotational component (S223).

[0127] FIG. 12 is a flowchart of this calculation process (S223) of the differential rotational component.

[0128] First, the motion information calculation unit 144 determines whether there is a rotational component in the motion information 151 estimated in step S142 (S261). For example, the motion information calculation unit 144 may determine whether the rotational component is zero, similar to the determination of translation (S241), or may compare the magnitude of the rotation angle with a predetermined value. Also, the motion information calculation unit 144 may switch the predetermined value according to the characteristics of the image. For example, if the image has characteristics that are likely to contribute to the improvement of the prediction accuracy of rotation, the motion information calculation unit 144 may always determine that there is a rotational component, and if not, may always determine that there is no rotational component. When it is determined that there is always a rotational component, for example, the rotation angle θ is determined in advance based on the time conversion of the angle from the rotation center, and the motion information calculation unit 144 may always use that value.

[0129] Also, when it is determined in step S187 of FIG. 8 or step S207 of FIG. 9 that rotation is not selected, the motion information calculation unit 144 may determine that there is no rotational component.

[0130] When there is no rotation component in the motion information 151 (No in S261), the motion information calculation unit 144 sets the rotation flag indicating "rotation exists" to OFF, and attaches the rotation flag to the motion information 133 to be encoded (S262).

[0131] When there is a rotation component in the motion information 151 (Yes in S261), the motion information calculation unit 144 sets the rotation flag to ON, and attaches the rotation flag to the motion information 133 (S263).

[0132] Next, the motion information calculation unit 144 determines whether there is a rotation component in the predicted motion information acquired in step S221 (S264).

[0133] When there is no rotation component in the predicted motion information (No in S264), the motion information calculation unit 144 sets the predicted rotation component to 0 (S265). When there is a rotation component in the predicted motion information (Yes in S264), the motion information calculation unit 144 sets the rotation component of the predicted motion information as the predicted rotation component (S266).

[0134] The motion information calculation unit 144 calculates the differential rotation component by subtracting the predicted rotation component obtained in step S265 or S266 from the rotation component of the motion information obtained in step S142, and adds the calculated differential rotation component to the motion information 133 (S267).

[0135] Next, the motion information calculation unit 144 calculates the differential expansion / contraction component, which is the differential motion information of the expansion / contraction component (S224).

[0136] FIG. 13 is a flowchart of the calculation process (S224) of this differential expansion / contraction component.

[0137] First, the motion information calculation unit 144 determines whether there is a zoom component in the motion information 151 estimated in step S142 (S281). For example, the motion information calculation unit 144 may determine whether the zoom component is zero (the value is 1), similar to step S241 or S261, or may compare the magnitude of the zoom component with a predetermined value. Also, the motion information calculation unit 144 may switch the predetermined value according to the characteristics of the image. For example, the motion information calculation unit 144 may set the predetermined value based on the temporal change of the angle of view in the case of a video when zooming in or out is performed.

[0138] Also, if the image has characteristics such that the zoom component is likely to greatly contribute to the prediction accuracy, the motion information calculation unit 144 may always determine that there is a zoom component, and if not, may always determine that there is no zoom component. For example, when the input image 121 is a video when zooming in or out is performed, the motion information calculation unit 144 may always determine that there is a zoom component. When determining that it always exists, for example, the value of the zoom component is determined in advance based on the temporal change of the angle of view, and the motion information calculation unit 144 may always use that value.

[0139] Also, in step S190 of FIG. 8 or step S210 of FIG. 9, if it is determined that zooming is not selected, the motion information calculation unit 144 may determine that there is no zoom component.

[0140] If there is no zoom component in the motion information 151 (No in S281), the motion information calculation unit 144 sets the zoom flag indicating "zooming exists" to OFF and attaches the zoom flag to the motion information 133 to be encoded (S282).

[0141] If there is a zoom component in the motion information 151 (Yes in S281), the motion information calculation unit 144 sets the zoom flag to ON and attaches the zoom flag to the motion information 133 (S283).

[0142] Next, the motion information calculation unit 144 determines whether there is a magnification / reduction component in the predicted motion information acquired in step S221 (S284).

[0143] If there is no magnification / reduction component in the predicted motion information (No in S284), the motion information calculation unit 144 sets the predicted magnification / reduction component to 1 (S285). If there is a magnification / reduction component in the predicted motion information (Yes in S284), the motion information calculation unit 144 sets the magnification / reduction component of the predicted motion information as the predicted magnification / reduction component (S286).

[0144] The motion information calculation unit 144 calculates a differential magnification / reduction component by subtracting the predicted magnification / reduction component obtained in step S284 or S285 from the magnification / reduction component of the motion information obtained in step S142, and adds the calculated differential magnification / reduction component to the motion information 133 (S287).

[0145] Finally, the motion information calculation unit 144 calculates a differential shear component which is the differential motion information of the shear component (S225).

[0146] FIG. 14 is a flowchart of this calculation process (S225) of the differential shear component.

[0147] First, the motion information calculation unit 144 determines whether there is a shear component in the motion information 151 estimated in step S142 (S301). For example, the motion information calculation unit 144 may determine whether the shear component is zero in the same manner as in step S241, S261, or S281, or may compare the magnitude of the shear component with a predetermined value. Also, the motion information calculation unit 144 may switch the predetermined value according to the characteristics of the image. Further, the motion information calculation unit 144 may determine that there is always (or never) a shear component when the characteristics of the image satisfy a predetermined condition.

[0148] Also, when it is determined in step S193 of FIG. 8 or step S213 of FIG. 9 that shear is not selected, the motion information calculation unit 144 may determine that there is no shear component.

[0149] When there is no shear component in the motion information 151 (No in S301), the motion information calculation unit 144 sets the shear flag indicating "shear exists" to OFF, and attaches the shear flag to the motion information 133 to be encoded (S302).

[0150] When there is a shear component in the motion information 151 (Yes in S301), the motion information calculation unit 144 sets the shear flag to ON, and attaches the shear flag to the motion information 133 (S303).

[0151] Next, the motion information calculation unit 144 determines whether there is a shear component in the predicted motion information acquired in step S221 (S304).

[0152] When there is no shear component in the predicted motion information (No in S304), the motion information calculation unit 144 sets the predicted shear component to 0 (S305). When there is a shear component in the predicted motion information (Yes in S304), the motion information calculation unit 144 sets the shear component of the predicted motion information as the predicted shear component (S306).

[0153] The motion information calculation unit 144 calculates the differential shear component by subtracting the predicted shear component obtained in step S305 or S306 from the shear component of the motion information obtained in step S142, and adds the calculated differential shear component to the motion information 133 (S307).

[0154] Through the above processing, the processing of step S144 ends.

[0155] Note that the processing order of steps S222 to S225 is not limited to this order and may be in any order. Also, some or all of these processes may be performed in parallel. Also, as described above, when the conversion components that can be used for encoding are limited due to the size of the current block, etc., the motion information calculation unit 144 may perform only the processing for the allowed conversion components.

[0156] FIG. 15 and FIG. 16 are diagrams showing an example of the motion information 133 to be encoded. As shown in FIG. 15, the motion information 133 includes a translation flag 161, a rotation flag 162, a scaling flag 163, a shear flag 164, a differential translation component 171, a differential rotation component 172, a differential scaling component 173, and a differential shear component 174. Note that the meaning of each piece of information is as described above.

[0157] Further, only when each flag is on, the motion information 133 includes a differential motion component corresponding to the flag. For example, as shown in FIG. 16, when the translation flag 161 and the rotation flag 162 are ON and the scaling flag 163 and the shear flag 164 are OFF, the motion information 133 includes the differential translation component 171 and the differential rotation component 172, and does not include the differential scaling component 173 and the differential shear component 174.

[0158] (Effect) As described above, the image encoding apparatus 100 according to the present embodiment selects only necessary conversion components using features of an image or the like, and encodes information (translation flag, rotation flag, scaling flag, and shear flag) indicating the selected conversion components and the selected conversion components. Thereby, the image encoding apparatus 100 can selectively use only effective conversion components, so that the encoding efficiency can be improved. As described above, the image encoding apparatus 100 can flexibly select various conversions and can reduce the motion information necessary for generating a predicted image, so that the encoding efficiency can be improved.

[0159] Further, the image encoding apparatus 100 selects conversion components using an estimation result of the type of conversion existing in the image being encoded, the size of a prediction block, or the like. Thereby, the image encoding apparatus 100 can limit the types of motion information to be searched during the prediction process, so that the processing speed can be increased.

[0160] Note that in the present embodiment, an affine transformation is used as the motion information, and translation, rotation, scaling, and shear are used as the types of conversion, but the types of conversion to be used are not limited to this.

[0161] For example, instead of the affinity matrix, a projective transformation matrix capable of representing a projective transformation may be used. In this case, trapezoidal transformation can be used as motion information. Thereby, the quality of the prediction block can be further improved, and an improvement in coding efficiency can be expected.

[0162] That is, the image encoding apparatus 100 according to the present embodiment only needs to divide the motion component into a plurality of transformation components and be able to determine whether each transformation component is used. That is, the image encoding apparatus 100 may use a matrix or transformation other than the above, or may divide the motion information into transformation components different from the above.

[0163] Also, in the above description, for each transformation component, an example of assigning a flag in block units during inter prediction is shown in steps S242, S243, S261, S263, S282, S283, S302, and S303. However, these flags may be specified in units of an image, a sequence, or a region into which the image is divided. For example, this region is a region in which the image is divided into four parts. Thereby, all the blocks included in a previously specified image or sequence before performing inter prediction are encoded using the same type of transformation component. Therefore, the determination process in block units can be omitted, so that the coding efficiency can be further improved. Also, since the number of flags in the bitstream can be reduced, the coding efficiency can be further improved.

[0164] Also, when there are no various conversion components in the predicted motion information, in steps S245, S265, S285, and S305, 0, 0, 1, and 0 are respectively substituted as the predicted conversion components, but the substituted values are not limited to this. For example, when the entire screen is rotating, the angle determined according to the characteristics of the image may be substituted as the predicted rotation component. Further, the process of substituting a predetermined angle such as 0 into the predicted conversion component and the process of substituting an angle or the like determined according to the characteristics of the image into the predicted conversion component may be switched according to a predetermined condition. Thus, when there are conversion components in the motion information but not in the predicted motion information, by performing prediction while switching the substitution values, it is possible to reduce the value of the differential motion information. Thereby, the coding efficiency can be further improved.

[0165] (Embodiment 2) In the above-described Embodiment 1, an example in which flags corresponding to each of a plurality of conversions are used as information indicating whether each of various conversions is used has been described. In this embodiment, a rank (coding level) is associated with a combination of various conversion components. And, as the above information, information specifying this coding level is used.

[0166] In this embodiment, instead of the processes of steps S162 and S163 described above, for example, the following processes are performed.

[0167] FIG. 17 is a diagram showing an example of the relationship between the coding level and the conversion components. As shown in FIG. 17, for example, coding level 1 includes translation, coding level 2 includes rotation and scaling, and coding level 3 includes shear.

[0168] The coding information selection unit 141 selects a coding level according to conditions such as the characteristics of the image or the allowable bandwidth.

[0169] Next, the encoding information selection unit 141 selects conversion components included in encoding levels equal to or lower than the selected encoding level. For example, in the case of encoding level 2, three types of conversion components, namely translation, rotation, and scaling, included in encoding level 2 and encoding level 1 are selected.

[0170] Also, encoding level information 181 indicating the selected encoding level is added to the motion information 133 to be encoded.

[0171] FIGs. 18 and 19 are diagrams showing an example of the motion information 133 to be encoded. As shown in FIG. 18, the motion information 133 includes encoding level information 181, a differential translation component 171, a differential rotation component 172, a differential scaling component 173, and a differential shear component 174.

[0172] Also, the motion information 133 includes differential motion components corresponding to the conversion components included in encoding levels equal to or lower than the encoding level indicated by the encoding level information 181. For example, as shown in FIG. 19, when the encoding level information 181 indicates encoding level 2, the motion information 133 includes the differential translation component 171, the differential rotation component 172, and the differential scaling component 173 at encoding levels 1 and 2, and does not include the differential shear component 174 at encoding level 3.

[0173] As described above, in this embodiment, the amount of information required for encoding can be further reduced, so that the encoding efficiency can be improved. Here, an example of using the levels of various conversions has been described, but the present invention is not limited thereto, and information indicating a combination of conversion components may be used. For example, in the above description, an example in which all of the conversion components included below the selected level are specified has been described, but the following information may also be used. For example, if the level is 0, not all conversion components are used; if the level is 1, only translation is used; if the level is 2, translation and rotation are used; if the level is 3, translation and scaling are used; if the level is 4, rotation and scaling are used; if the level is 5, translation, rotation, and scaling are used; if the level is 6, translation and shear are used; if the level is 7, translation, rotation, scaling, and shear may be used. In this way, a combination of conversion components may be specified for each level, and the conversion components associated with each level may be selected.

[0174] (Embodiment 3) As a result of verification by the inventors, it has been found that among translation, rotation, scaling, and shear, shear is less likely to contribute to an improvement in prediction accuracy compared to other conversions. Therefore, in this embodiment, an example of assigning priorities to a plurality of conversion components will be described.

[0175] As a first method of assigning priorities to a plurality of conversion components, the threshold of the block size is changed as in the process of step S163 shown in FIG. 9 described above. Specifically, the threshold of shear is set larger than that of other conversions such as rotation that can greatly improve the prediction accuracy. Thereby, when the block size is small, the use of conversion components whose prediction accuracy is difficult to improve can be restricted, so that a decrease in encoding efficiency can be suppressed. Similarly, since scaling does not contribute much to an improvement in prediction accuracy with a small block size, a larger threshold is set compared to translation and rotation.

[0176] When the inventors verified with general moving images, it was found that depending on the features of the images, although the rankings were different, it was easier to improve the prediction accuracy in the order of translation, rotation, scaling, and shear. Therefore, the priorities are set in this order. This makes it possible to limit the selection of transformation components with low priorities according to the state of the image features, size, or bandwidth, etc.

[0177] Next, a second method of assigning priorities to a plurality of transformation components will be described. Here, an example of a method for restricting shear will be described.

[0178] First, how to set the coefficients of the affine transformation will be briefly explained. As described in Non-Patent Document 2, an evaluation function E(a) with the affine matrix a shown in the following (Equation 3) as a variable is set, and the coefficients of the affine matrix are set by solving a minimization problem. That is, the coefficients (motion information) that minimize the evaluation function E(a) are derived.

[0179]

Equation

[0180] When shear is restricted, the orthogonality of the x-axis and y-axis is maintained before and after the affine transformation. That is, the coefficients of the affine transformation shown in the above (Equation 2) need to satisfy the following constraint conditions (Equation 4).

[0181]

Equation

[0182] Therefore, constraint conditions are added to the evaluation function in the above (Equation 3) as shown in the following (Equation 5). Due to this constraint condition, the larger the shear component, the larger the value of the evaluation function E(a), so it becomes difficult to select the coefficients including the shear component. This method is called the penalty method. Also, by increasing the positive constant μ, the influence of the constraint condition becomes larger. By obtaining the coefficients of the affine matrix that minimize the evaluation function E(a) in the following (Equation 5) in this way, it becomes possible to calculate the coefficients with the shear component restricted.

[0183]

Number

[0184] The inventors set a small μ, solved the minimization problem, used the optimal solution at that time in the result of repeating the search a certain number of times as the new initial value, and repeated the process of solving the minimization problem again with the value of μ increased until μ reached a predetermined magnitude, thereby calculating the coefficient. Note that the method for obtaining the coefficient is not limited to this method.

[0185] Also, although a method of restricting (lowering the priority) shear has been described here, the same technique can be applied to other conversions.

[0186] (Other Variants) In the above-described Embodiments 1 to 3, a method of encoding differential motion information, which is the difference between motion information and predicted motion information, has been described, but the method is not limited thereto.

[0187] For example, in HEVC, there is a motion prediction method called the merge mode. In this merge mode, differential information is not encoded, and a predicted motion vector selected from a plurality of predicted motion vector candidates is used as the motion vector of the current block. Then, selection information indicating the selected predicted motion vector candidate is encoded. Also in such a case, similar to the above-described Embodiments 1 to 3, the image encoding apparatus may select one or more transform components from a plurality of transform components and generate a predicted image using the selected transform components. Further, the image encoding apparatus may encode information (the above flag or encoding level information) indicating the selected (used) transform components.

[0188] For example, the image encoding device encodes the selected transform component based on whether there are various transform components in the prediction information predicted from the motion information of blocks spatially or temporally adjacent to the current block. In this case, when all of the motion information of a plurality of adjacent blocks includes a rotation component, the image encoding device may use the average value of the rotation components as the rotation component of the current block. Also, when only one of the motion information of a plurality of adjacent blocks includes a rotation component and the rotation component is less than a predetermined value, the image encoding device may determine not to use the rotation component for the current block. Or, the image encoding device may set the rotation components of other adjacent blocks to zero, calculate the average value of the plurality of rotation components, and use the calculated average value as the rotation component.

[0189] Methods such as using such an average value and methods using temporal scaling processing are known, and these modifications may be applied to the present embodiment.

[0190] (Embodiment 4) In the present embodiment, one embodiment of an image decoding device that decodes the bitstream generated by the image encoding device according to the above-described Embodiment 1 will be described.

[0191] FIG. 20 is a block diagram of an image decoding device 200 according to the present embodiment. The image decoding device 200 includes an entropy decoding unit 201, an inverse quantization unit 202, an inverse frequency conversion unit 203, an addition unit 204, an intra prediction unit 205, a loop filter 206, a frame memory 207, an inter prediction unit 208, and a switching unit 209.

[0192] The image decoding device 200 generates a decoded image 227 by decoding the bitstream 221. For example, the bitstream 221 is generated by the above-described image encoding device 100.

[0193] FIG. 21 is a flowchart of the image decoding process according to the present embodiment.

[0194] First, the entropy decoding unit 201 decodes the motion information 228 from the bitstream 221 obtained by encoding a still image or a moving image including one or more pictures (S401). Also, the entropy decoding unit 201 decodes the coefficient block 222 from the bitstream 221 (S402).

[0195] The inverse quantization unit 202 generates a coefficient block 223 by inverse quantizing the coefficient block 222. The inverse frequency conversion unit 203 restores the difference block 224 by performing inverse frequency conversion on the coefficient block 223 (S403).

[0196] Next, the intra prediction unit 205 or the inter prediction unit 208 generates a prediction block 230 using the motion information 228 decoded in step S401 and the decoded block (S404). Specifically, the intra prediction unit 205 generates a prediction block 226 by intra prediction processing. The inter prediction unit 208 generates a prediction block 229 by inter prediction processing. The switching unit 209 outputs one of the prediction blocks 226 and 229 as the prediction block 230.

[0197] Next, the addition unit 204 generates a decoded block 225 by adding the difference block 224 and the prediction block 230 (S405). This decoded block 225 is used for the intra prediction processing by the intra prediction unit 205.

[0198] Next, the image decoding device 200 determines whether all the blocks included in the bitstream 221 have been decoded (S406). For example, the image decoding device 200 makes this determination according to whether the input bitstream 221 has ended. If the decoding of all blocks has not been completed (No in S406), the processes from step S401 onward are performed on the next block. If the decoding of all the blocks included in the bitstream 221 has ended (Yes in S406), the loop filter 206 generates a decoded image 227 (reconstructed image) by combining all the decoded blocks and performing loop filter processing (S407). The frame memory 207 stores the decoded image 227. This decoded image 227 is used for the inter prediction processing by the inter prediction unit 208.

[0199] Note that the inverse quantization and inverse frequency conversion in step S403 may be sequentially performed as separate processes, or may be performed in a batch. In current mainstream coding standards such as HEVC, the inverse quantization and inverse frequency conversion are performed in a batch. Also, on the decoding side, similar to the encoding side (Embodiment 1), expressions such as scaling may be used in some cases.

[0200] Next, the motion information decoding process (S401) will be described with reference to FIG. 22. FIG. 22 is a flowchart of the motion information decoding process (S401) according to the present embodiment.

[0201] First, the entropy decoding unit 201 decodes differential motion information from the bitstream 221 (S421). Also, the entropy decoding unit 201 decodes information for deriving predicted motion information (S422). Note that the information for deriving predicted motion information and the differential motion information are included in the motion information 228. Next, the inter prediction unit 208 derives predicted motion information from the information obtained in step S422, and generates motion information from the derived predicted motion information and the differential motion information obtained in step S421 (S423).

[0202] Details of the differential motion information decoding (S421) and motion information generation (S423) will be described below.

[0203] Regarding the decoding process (S421) of differential motion information, an explanation will be given with reference to FIG. 23. FIG. 23 is a flowchart of the decoding process (S421) of differential motion information according to the present embodiment.

[0204] First, the entropy decoding unit 201 performs a decoding process for the differential translation component (S441 to S443). First, the entropy decoding unit 201 decodes a translation flag indicating whether there is a translation (indicating whether the differential translation component is included in the motion information 228) from the bit stream 221 (S441). Note that the configuration of the motion information 228 and the meanings of various types of information included in the motion information 228 are the same as the configuration of the motion information 133 and the meanings of various types of information included in the motion information 133 in Embodiment 1. Next, the entropy decoding unit 201 determines whether there is a translation using the translation flag (S442).

[0205] When it is shown by the translation flag that there is a translation (Yes in S442), the entropy decoding unit 201 decodes the differential translation component from the bit stream 221 (S443). When it is shown by the translation flag that there is no translation (No in S442), the entropy decoding unit 201 does not decode the differential translation component and proceeds to the next processing step (S444).

[0206] Subsequently, the entropy decoding unit 201 performs a decoding process for the differential rotation component (S444 to S446). The entropy decoding unit 201 decodes a rotation flag indicating whether there is a rotation (indicating whether the differential rotation component is included in the motion information 228) from the bit stream 221 (S444). The entropy decoding unit 201 determines whether there is a rotation using the rotation flag (S445).

[0207] When it is shown that rotation exists based on the rotation flag (Yes in S445), the entropy decoding unit 201 decodes the differential rotation component from the bit stream 221 (S446). When it is shown that rotation does not exist based on the rotation flag (No in S445), the entropy decoding unit 201 performs the next processing step (S447).

[0208] Next, the entropy decoding unit 201 performs the decoding process of the differential expansion / contraction component (S447 - S449). The entropy decoding unit 201 decodes an expansion / contraction flag from the bit stream 221 that indicates whether expansion / contraction exists (whether the differential expansion / contraction component is included in the motion information 228) (S447). The entropy decoding unit 201 determines whether expansion / contraction exists using the expansion / contraction flag (S448).

[0209] When it is shown that expansion / contraction exists based on the expansion / contraction flag (Yes in S448), the entropy decoding unit 201 decodes the differential expansion / contraction component from the bit stream 221 (S449). When it is shown that expansion / contraction does not exist based on the expansion / contraction flag (No in S448), the entropy decoding unit 201 performs the next processing step (S450).

[0210] Finally, the entropy decoding unit 201 performs the decoding process of the differential shear component (S450 - S452). The entropy decoding unit 201 decodes a shear flag from the bit stream 221 that indicates whether the shear component exists (whether the differential shear component is included in the motion information 228) (S450). The entropy decoding unit 201 determines whether the shear component exists using the shear flag (S451).

[0211] When it is indicated that shear exists according to the shear flag (Yes in S451), the entropy decoding unit 201 decodes the differential shear component from the bit stream 221 (S452). Thereafter, the entropy decoding unit 201 ends the differential motion information decoding process (S421). When it is indicated that no rotation exists according to the shear flag (No in S451), the entropy decoding unit 201 ends the differential motion information decoding process (S421).

[0212] Subsequently, the motion information generation process (S423) will be described with reference to FIGS. 24 to 28.

[0213] FIG. 24 is a flowchart of the motion information generation process (S423) according to the present embodiment.

[0214] First, the inter prediction unit 208 derives predicted motion information using the information acquired in step S422. Specifically, the inter prediction unit 208 acquires the motion information of decoded blocks that are spatially or temporally adjacent to the current block to be decoded, and derives predicted motion information using the acquired motion information (S461).

[0215] First, the inter prediction unit 208 generates a translation component from the differential motion information and the predicted motion information (S462). FIG. 25 is a flowchart of this translation component generation process (S462).

[0216] First, the inter prediction unit 208 determines whether a differential translation component was acquired in step S421 (S481).

[0217] When a differential translation component has been acquired (Yes in S481), the inter prediction unit 208 determines whether a translation component exists in the predicted motion information acquired in step S461 (S482).

[0218] When there is no translational movement in the predicted movement information (No in S482), the inter-prediction unit 208 sets the predicted translational movement component to 0 (S483). When there is a translational movement component in the predicted movement information (Yes in S482), the inter-prediction unit 208 sets the translational movement component of the predicted movement information as the predicted translational movement component (S484).

[0219] The inter-prediction unit 208 generates a translational movement component by adding the predicted translational movement component obtained in step S483 or S484 to the differential translational movement component (S485).

[0220] When there is no differential translational movement component (No in S482), the inter-prediction unit 208 ends the process regarding the translational movement component (S462).

[0221] Next, the inter-prediction unit 208 generates a rotational component from the differential movement information and the predicted movement information (S463). FIG. 26 is a flowchart of this rotational component generation process (S463).

[0222] First, the inter-prediction unit 208 determines whether the differential rotational component was obtained in step S421 (S501).

[0223] When the differential rotational component has been obtained (Yes in S501), the inter-prediction unit 208 determines whether there is a rotational component in the predicted movement information obtained in step S461 (S502).

[0224] When there is no rotation in the predicted movement information (No in S502), the inter-prediction unit 208 sets the predicted rotational component to 0 (S503). When there is a rotational component in the predicted movement information (Yes in S502), the inter-prediction unit 208 sets the rotational component of the predicted movement information as the predicted rotational component (S504).

[0225] The inter-prediction unit 208 generates a rotational component by adding the predicted rotational component obtained in step S503 or S504 to the differential rotational component (S505).

[0226] When there is no differential rotation component (No in S501), the inter prediction unit 208 ends the process related to the rotation component (S463).

[0227] Subsequently, the inter prediction unit 208 generates a scaling component from the differential motion information and the predicted motion information (S464). FIG. 27 is a flowchart of this generation process (S464) of the scaling component.

[0228] First, the inter prediction unit 208 determines whether the differential scaling component was acquired in step S421 (S521).

[0229] When the differential scaling component has been acquired (Yes in S521), the inter prediction unit 208 determines whether there is a scaling component in the predicted motion information acquired in step S461 (S522).

[0230] When there is no scaling in the predicted motion information (No in S522), the inter prediction unit 208 sets the predicted scaling component to 1 (S523). When there is a scaling component in the predicted motion information (Yes in S523), the inter prediction unit 208 sets the scaling component of the predicted motion information as the predicted scaling component (S524).

[0231] The inter prediction unit 208 generates the scaling component by adding the predicted scaling component acquired in steps S523 and S524 to the differential scaling component (S525).

[0232] When the differential scaling component does not exist (No in S521), the inter prediction unit 208 ends the process related to the scaling component (S464).

[0233] Finally, the inter prediction unit 208 generates a shear component from the differential motion information and the predicted motion information (S465). FIG. 28 is a flowchart of this generation process (S465) of the shear component.

[0234] First, the inter prediction unit 208 determines whether the differential shear component has been acquired in step S421 (S541).

[0235] If the differential shear component has been acquired (Yes in S541), the inter prediction unit 208 determines whether there is a shear component in the predicted motion information acquired in step S461 (S542).

[0236] If there is no shear in the predicted motion information (No in S542), the inter prediction unit 208 sets the predicted shear component to 0 (S543). If there is a shear component in the predicted motion information (Yes in S542), the inter prediction unit 208 sets the shear component of the predicted motion information as the predicted shear component (S544).

[0237] The inter prediction unit 208 generates a shear component by adding the predicted shear component acquired in steps S543 and S544 to the differential shear component (S545).

[0238] If the differential shear component does not exist (No in S541), the inter prediction unit 208 ends the process related to the shear component (S465).

[0239] After the inter prediction unit 208 performs these series of processes for each block, it ends the motion information generation process (S423).

[0240] As described in Embodiment 1, each flag may be provided for each block, or may be provided in units of an image, a sequence, or a region into which an image is divided. When a flag is provided in a unit other than the block unit and a conversion component is specified by the flag, the image decoding apparatus performs decoding processing using the same type of conversion component for all blocks included in the specified unit. Thereby, the image decoding apparatus can omit determination steps and the like, and execute decoding processing with high prediction accuracy with a small amount of information and a small processing load.

[0241] Also, similar to Embodiment 1, the order of the generation processes (S462 to S465) of various conversion components is not limited to the order shown in FIG. 24. Also, part or all of the processes in steps S462 to S465 may be performed in parallel. Or, superiority and inferiority may be set for these processes, and the processes may be performed in order from the process with the highest priority.

[0242] Also, when encoding level information indicating the encoding level is used instead of various flags as in Embodiment 2, the image decoding apparatus 200 decodes the encoding level information from the bit stream 221 instead of the various flags described above. Further, the image decoding apparatus determines the presence or absence of conversion components using the information indicated by the encoding level information instead of the various flags.

[0243] Also, as described in Embodiment 1, when the presence or absence of various conversion components is specified by the block size, the image decoding apparatus 200 decodes the block size from the bit stream 221 and determines the presence or absence of various conversion components using the block size.

[0244] (Effect) As described above, the image decoding apparatus 200 according to the present embodiment can decode the bit stream 221 in which only the conversion components necessary for generating the prediction block are encoded. Further, since the image decoding apparatus 200 can decode the bit stream 221 generated so that the amount of code is reduced using high-dimensional motion prediction, a higher-quality image can be reproduced.

[0245] In the present embodiment, an example in which translation, rotation, enlargement / reduction, and shear are included in the motion information as conversion components has been described, but the available conversions are not limited to this. For example, trapezoidal conversion may be used. Thereby, since a complex conversion can be expressed, the image decoding apparatus can generate a more accurate prediction image.

[0246] That is, the image decoding apparatus 200 according to the present embodiment only needs to be able to divide the motion information into a plurality of transform components and determine whether to decode each transform component or not. That is, the image decoding apparatus 200 may use transform components other than the above.

[0247] Further, the method of the present embodiment can be applied not only to the case of decoding the differential motion information which is the difference between the motion information and the predicted motion information, but also to, for example, the merge mode.

[0248] (Embodiment 5) In the above-described Embodiments 1 to 4, an example in which the affine transformation is used for inter prediction has been described. In the present embodiment, an example in which the affine transformation is used for intra prediction will be described.

[0249] As shown in FIG. 29, the image encoding apparatus (or image decoding apparatus) according to the present embodiment uses the pixel values of the already processed (encoded or decoded) blocks in the same picture as the current block as the pixel values of the current block. That is, the pixel values of the processed block are copied to the current block. Further, in the present embodiment, as shown in FIG. 30, various transformations such as rotation, scaling, and shear are used in addition to translation during this copy.

[0250] That is, in the present embodiment, the intra prediction unit 109 (or 205) generates a prediction block using the reference information indicating the processed block. This reference information includes various transform components such as a translation component, a rotation component, a scaling component, and a shear component.

[0251] In such a case, similar to the above-described Embodiments 1 to 4, by encoding (or decoding) information (various flags or encoding levels) indicating whether each of the various transform components is used or not, an arbitrary transform component can be selectively used.

[0252] In intra prediction, among rotation, scaling, and shearing, in the order of scaling, rotation, and shearing, it is easier to contribute to the improvement of prediction accuracy. Therefore, it is preferable to assign priorities in this order. Note that as the method for assigning priorities to a plurality of transformations, the method described in Embodiment 3 can be used.

[0253] Also, this priority may be changed according to the type of the image or the like. For example, in the case of a natural image or the like, priorities may be assigned in the order of the above-described scaling, rotation, and shearing, and in the case of screen content such as map information, priorities may be assigned in the order of rotation, scaling, and shearing.

[0254] As described above in Embodiments 1 to 3 and 5, the image encoding device according to the embodiment is an image encoding device that encodes an image, and performs the processing shown in FIG. 31.

[0255] First, the image encoding device selects two or more transformation components from among a plurality of transformation components including a translation component and a plurality of non-translation components as reference information indicating a reference destination of a current block to be encoded (S601). Here, the reference information is motion information in inter prediction or information indicating a processed block of a reference destination in intra prediction. The plurality of non-translation components include, for example, a rotation component, a scaling component, and a shearing component.

[0256] Next, the image encoding device generates a prediction image using the reference information (S602). Next, the image encoding device encodes the current block using the prediction image (S603).

[0257] Further, the image encoding device encodes selection information for specifying two or more selected transform components from among a plurality of transform components (S604). For example, the selection information corresponds to each of the plurality of transform components and includes flags (translation flag, rotation flag, scale flag, and shear flag) indicating whether or not the corresponding transform component is selected. Alternatively, the selection information is a set including some or all of the plurality of transform components, and indicates one encoding level from among a plurality of encoding levels indicating different sets. In this case, in step S601, two or more transform components included in the set indicated by the encoding level indicated by the selection information are selected.

[0258] Also, the image encoding device may encode the selection information on a per-block basis, on a per-picture basis, on a per-sequence basis, or on a per-region basis in which a picture is divided. That is, the image encoding device may encode one selection information commonly used for an image including the current block, may encode one selection information commonly used for a sequence including the current block, or may encode one selection information commonly used for a region including the current block.

[0259] Next, the image encoding device encodes the reference information of the current block using the reference information of an encoded block different from the current block (S605). For example, the image encoding device encodes differential reference information that is the difference between the reference information of the encoded block and the reference information of the current block.

[0260] Note that, in step S601, the image encoding device may select two or more transform components according to the size of the current block. In this case, the selection information indicates the size of the current block. For example, the selection information includes information indicating the size of the maximum coding unit (CU) and information indicating whether or not to further divide each coding unit.

[0261] Also, for example, when the size of the current block is less than the first threshold, the image encoding device does not select the shear component.

[0262] Also, in step S601, the image encoding device may select two or more conversion components in order of priority, such as a translation component, a rotation component, a scaling component, and a shear component. For example, the priority can be set using the method described in Embodiment 3.

[0263] Also, as described in Embodiments 4 and 5, the image decoding device according to the embodiment is an image decoding device that decodes a bitstream obtained by encoding an image, and performs the processing shown in FIG. 32.

[0264] First, the image decoding device decodes selection information for specifying two or more conversion components from among a plurality of conversion components including a translation component and a plurality of non-translation components from the bitstream (S611). Here, the plurality of non-translation components include, for example, a rotation component, a scaling component, and a shear component. Also, the selection information corresponds to each of the plurality of conversion components, and includes flags (translation flag, rotation flag, scaling flag, and shear flag) indicating whether or not the corresponding conversion component is selected. Alternatively, the selection information is a set including some or all of the plurality of conversion components, and indicates one coding level from a plurality of coding levels indicating different sets.

[0265] Also, the image decoding device may decode the selection information for each block, for each picture, for each sequence, or for each region into which the picture is divided. That is, the image decoding device may decode one selection information commonly used for an image including the current block, may decode one selection information commonly used for a sequence including the current block, or may decode one selection information commonly used for a region including the current block.

[0266] Also, the image decoding device selects two or more transform components specified by the decoded selection information as reference information indicating the reference destination of the current block to be decoded (S612). Here, the reference information is motion information in inter prediction or information indicating a processed block of the reference destination in intra prediction. When the selection information indicates an encoding level, the image decoding device selects two or more transform components included in the set indicated by the encoding level indicated by the selection information.

[0267] Next, the image decoding device decodes the reference information of the current block from the bitstream using the reference information of the decoded block different from the current block (S613). Specifically, the image decoding device decodes the differential reference information of the selected transform component from the bitstream. Next, the image decoding device generates reference information by adding the obtained differential decoded information and the reference information of the decoded block for each selected transform component.

[0268] Next, the image decoding device generates a predicted image using the reference information (S614). Next, the image decoding device decodes the current block from the bitstream using the predicted image (S615).

[0269] Note that the selection information indicates the size of the current block, and in step S612, the image decoding device may select two or more transform components according to the size of the current block. For example, the selection information includes information indicating the size of the maximum coding unit (CU) and information indicating whether to further divide each coding unit. The image decoding device determines the current block size using this information.

[0270] Also, for example, when the size of the current block is less than or equal to the first threshold, the image decoding device does not select the shear component.

[0271] Also, in step S612, the image encoding apparatus may select two or more conversion components in order of priority of the translation component, rotation component, scaling component, and shear component. For example, the priority can be set using the method described in Embodiment 3.

[0272] As described above, the image encoding method and the image decoding method according to the embodiment have been described. However, the present invention is not limited to this embodiment.

[0273] Also, each processing unit included in the image encoding apparatus and the image decoding apparatus according to the above embodiment is typically realized as an LSI which is an integrated circuit. These may be individually formed into one chip, or may be formed into one chip so as to include part or all of them.

[0274] Also, the integration into an integrated circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.

[0275] In each of the above embodiments, each component may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0276] In other words, the image encoding device and the image decoding device include a processing circuitry and a storage electrically connected to (accessible from) the processing circuitry. The processing circuitry includes at least one of dedicated hardware and a program execution unit. Further, when the processing circuitry includes the program execution unit, the storage stores a software program executed by the program execution unit. The processing circuitry uses the storage to execute the encoding method or the decoding method according to the above embodiment.

[0277] Furthermore, the present invention may be the above software program or a non-transitory computer-readable recording medium on which the above program is recorded. Needless to say, the above program can be distributed via a transmission medium such as the Internet.

[0278] Also, all the numbers used above are for illustrative purposes to specifically describe the present invention, and the present invention is not limited to the illustrated numbers.

[0279] Also, the division of the functional blocks in the block diagram is an example, and a plurality of functional blocks may be realized as one functional block, one functional block may be divided into a plurality, or some functions may be transferred to other functional blocks. Also, the functions of a plurality of functional blocks having similar functions may be processed by a single piece of hardware or software in parallel or in time division.

[0280] Also, the order in which the steps included in the above encoding method or decoding method are executed is for illustrative purposes to specifically describe the present invention, and may be an order other than the above. Also, some of the above steps may be executed simultaneously (in parallel) with other steps.

[0281] As described above, the encoding device and the decoding device according to one or more aspects of the present invention have been described based on the embodiments. However, the present invention is not limited to these embodiments. As long as the gist of the present invention is not deviated from, various modifications conceived by those skilled in the art applied to these embodiments, or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects of the present invention.

[0282] (Embodiment 6) By recording a program for realizing the configuration of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments on a storage medium, the processing shown in each of the above embodiments can be easily implemented in an independent computer system. The storage medium may be any medium capable of recording a program, such as a magnetic disk, an optical disk, a magneto-optical disk, an IC card, a semiconductor memory, or the like.

[0283] Furthermore, here, application examples of the moving image encoding method (image encoding method) and the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding / decoding device including an image encoding device using the image encoding method and an image decoding device using the image decoding method. Other configurations in the system can be appropriately changed as the case may be.

[0284] FIG. 33 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed radio stations, are installed in each cell.

[0285] This content supply system ex100 has various devices such as a computer ex111, a PDA (Personal Digital Assistant) ex112, a camera ex113, a mobile phone ex114, and a game console ex115 connected thereto via an Internet service provider ex102, a telephone network ex104, and a base station ex106 to ex110 on the Internet ex101.

[0286] However, the content supply system ex100 is not limited to the configuration as shown in FIG. 33, and any elements may be combined and connected. Also, each device may be directly connected to the telephone network ex104 without going through ex110 from the base station ex106 which is a fixed radio station. Further, each device may be directly connected to each other via short-range wireless or the like.

[0287] The camera ex113 is a device capable of shooting moving images such as a digital video camera, and the camera ex116 is a device capable of shooting still images and moving images such as a digital camera. Also, the mobile phone ex114 may be a mobile phone of the GSM (registered trademark) (Global System for Mobile Communications) system, CDMA (Code Division Multiple Access) system, W-CDMA (Wideband-Code Division Multiple Access) system, or LTE (Long Term Evolution) system, HSPA (High Speed Packet Access), or PHS (Personal Handyphone System), etc., and any of them is acceptable.

[0288] In the content supply system ex100, cameras ex113 etc. are connected to the streaming server ex103 through the base station ex109 and the telephone network ex104, enabling live distribution and the like. In live distribution, for the content captured by the user using the camera ex113 (for example, the video of a music live etc.), encoding processing is performed as described in each of the above embodiments (that is, it functions as an image encoding device according to an aspect of the present invention), and is transmitted to the streaming server ex103. On the other hand, the streaming server ex103 stream-distributes the content data transmitted to the requested client. As clients, there are a computer ex111, a PDA ex112, a camera ex113, a mobile phone ex114, a game machine ex115, etc. that can decode the encoded data. Each device that has received the distributed data decodes and plays back the received data (that is, it functions as an image decoding device according to an aspect of the present invention).

[0289] Note that the encoding processing of the captured data may be performed by the camera ex113, may be performed by the streaming server ex103 that performs the data transmission processing, or may be shared between them. Similarly, the decoding processing of the distributed data may be performed by the client, may be performed by the streaming server ex103, or may be shared between them. Also, not limited to the camera ex113, still image and / or moving image data captured by the camera ex116 may be transmitted to the streaming server ex103 via the computer ex111. In this case, the encoding processing may be performed by any of the camera ex116, the computer ex111, and the streaming server ex103, or may be shared between them.

[0290] Also, these encoding and decoding processes are generally processed in a computer ex111 or an LSI ex500 included in each device. The LSI ex500 may be a single-chip or a multi-chip configuration. Note that software for moving image encoding and decoding may be incorporated into some recording medium (such as a CD-ROM, a flexible disk, a hard disk, etc.) readable by a computer ex111 or the like, and encoding and decoding processes may be performed using the software. Further, when the mobile phone ex114 has a camera, moving image data acquired by the camera may be transmitted. The moving image data at this time is data encoded by the LSI ex500 included in the mobile phone ex114.

[0291] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers, which may distribute, record, and deliver data in a distributed manner.

[0292] As described above, in the content supply system ex100, the client can receive and reproduce the encoded data. Thus, in the content supply system ex100, the client can receive, decode, and reproduce the information transmitted by the user in real time, and a user without special rights or facilities can also realize personal broadcasting.

[0293] Note that, not limited to the example of the content supply system ex100, as shown in FIG. 34, at least any one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can be incorporated into the digital broadcast system ex200. Specifically, in the broadcasting station ex201, multiplexed data in which music data and the like are multiplexed with video data is transmitted to the communication or satellite ex202 via radio waves. This video data is data encoded by the moving image encoding method described in the above embodiments (that is, data encoded by the image encoding device according to one aspect of the present invention). The receiving broadcast satellite ex202 transmits radio waves for broadcasting, and the antenna ex204 of a home capable of receiving satellite broadcasts receives these radio waves. The received multiplexed data is decoded and reproduced by a device such as a television (receiver) ex300 or a set-top box (STB) ex217 (that is, it functions as an image decoding device according to one aspect of the present invention).

[0294] Also, it is possible to implement the moving image decoding device or the moving image encoding device shown in the above embodiments in a reader / recorder ex218 that reads and decodes multiplexed data recorded on a recording medium ex215 such as a DVD or a BD, or encodes a video signal onto the recording medium ex215 and, in some cases, multiplexes and writes it with a music signal. In this case, the reproduced video signal is displayed on the monitor ex219, and the video signal can be reproduced by other devices or systems using the recording medium ex215 on which the multiplexed data is recorded. Further, a moving image decoding device may be implemented in the set-top box ex217 connected to the cable ex203 for cable television or the antenna ex204 for satellite / terrestrial wave broadcast, and this may be displayed on the monitor ex219 of the television. At this time, instead of the set-top box, the moving image decoding device may be incorporated into the television.

[0295] FIG. 35 is a diagram showing a television (receiver) ex300 using the moving image decoding method and the moving image encoding method described in each of the above embodiments. The television ex300 acquires or outputs multiplexed data in which audio data is multiplexed with video data via an antenna ex204 or a cable ex203 or the like that receives the above broadcast, a tuner ex301, a modulation / demodulation unit ex302 that demodulates the received multiplexed data or modulates the multiplexed data to be transmitted externally, and a multiplexing / demultiplexing unit ex303 that separates the demodulated multiplexed data into video data and audio data, or multiplexes the video data and audio data encoded by a signal processing unit ex306.

[0296] Further, the television ex300 has a signal processing unit ex306 having an audio signal processing unit ex304 and a video signal processing unit ex305 (functioning as an image encoding device or an image decoding device according to an aspect of the present invention) that decode audio data and video data respectively, or encode respective information, and an output unit ex309 having a speaker ex307 that outputs the decoded audio signal and a display unit ex308 such as a display that displays the decoded video signal. Further, the television ex300 has an interface unit ex317 having an operation input unit ex312 or the like that receives an input of a user operation. Further, the television ex300 has a control unit ex310 that comprehensively controls each unit, and a power supply circuit unit ex311 that supplies power to each unit. The interface unit ex317 may have, in addition to the operation input unit ex312, a bridge ex313 connected to an external device such as a reader / recorder ex218, a slot unit ex314 that enables a recording medium ex216 such as an SD card to be mounted, a driver ex315 connected to an external recording medium such as a hard disk, a modem ex316 connected to a telephone network, and the like. Note that the recording medium ex216 enables electrical recording by a nonvolatile / volatile semiconductor memory element for storing. Each unit of the television ex300 is connected to each other via a synchronization bus.

[0297] First, a configuration in which the TV ex300 decodes and plays back the multiplexed data acquired from the outside by the antenna ex204 or the like will be described. The TV ex300 receives a user operation from a remote controller ex220 or the like, and based on the control of a control unit ex310 having a CPU or the like, separates the multiplexed data demodulated by a modulation / demodulation unit ex302 by a multiplexing / demultiplexing unit ex303. Further, the TV ex300 decodes the separated audio data by an audio signal processing unit ex304, and decodes the separated video data by using the decoding method described in each of the above embodiments by a video signal processing unit ex305. The decoded audio signal and video signal are output from an output unit ex309 toward the outside. When outputting, these signals may be temporarily stored in buffers ex318, ex319, etc. so that the audio signal and the video signal are played back synchronously. Also, the TV ex300 may read the multiplexed data from recording media ex215, ex216 such as a magnetic / optical disk or an SD card, instead of from a broadcast or the like. Next, a configuration in which the TV ex300 encodes an audio signal or a video signal and transmits it to the outside or writes it to a recording medium or the like will be described. The TV ex300 receives a user operation from a remote controller ex220 or the like, and based on the control of the control unit ex310, encodes the audio signal by an audio signal processing unit ex304, and encodes the video signal by using the encoding method described in each of the above embodiments by a video signal processing unit ex305. The encoded audio signal and video signal are multiplexed by the multiplexing / demultiplexing unit ex303 and output to the outside. When multiplexing, these signals may be temporarily stored in buffers ex320, ex321, etc. so that the audio signal and the video signal are synchronized. Note that a plurality of buffers ex318, ex319, ex320, ex321 may be provided as shown in the figure, or a configuration in which one or more buffers are shared may be employed. Further, in addition to what is shown in the figure, for example, data may be stored in a buffer as a buffer material for avoiding system overflow and underflow between the modulation / demodulation unit ex302 and the multiplexing / demultiplexing unit ex303 or the like.

[0298] In addition to acquiring audio data and video data from broadcasts, recording media, etc., the TV ex300 is configured to accept AV inputs from microphones and cameras, and may perform encoding processing on the data acquired therefrom. Here, the TV ex300 has been described as being configured to perform the above encoding processing, multiplexing, and external output, but it may be configured such that these processes cannot be performed and only the above reception, decoding processing, and external output are possible.

[0299] Also, when reading or writing multiplexed data from a recording medium using the reader / writer ex218, the above decoding processing or encoding processing may be performed by either the TV ex300 or the reader / writer ex218, or the TV ex300 and the reader / writer ex218 may share the processing with each other.

[0300] As an example, FIG. 36 shows the configuration of the information reproduction / recording unit ex400 when reading or writing data from / to an optical disk. The information reproduction / recording unit ex400 includes elements ex401, ex402, ex403, ex404, ex405, ex406, and ex407 described below. The optical head ex401 irradiates a laser spot on the recording surface of the recording medium ex215, which is an optical disk, to write information, and detects the reflected light from the recording surface of the recording medium ex215 to read information. The modulation recording unit ex402 electrically drives a semiconductor laser built in the optical head ex401 to modulate the laser light according to the recording data. The reproduction demodulation unit ex403 amplifies the reproduction signal obtained by electrically detecting the reflected light from the recording surface by a photodetector built in the optical head ex401, separates and demodulates the signal components recorded on the recording medium ex215, and reproduces the necessary information. The buffer ex404 temporarily holds the information to be recorded on the recording medium ex215 and the information reproduced from the recording medium ex215. The disk motor ex405 rotates the recording medium ex215. The servo control unit ex406 moves the optical head ex401 to a predetermined information track while controlling the rotational drive of the disk motor ex405, and performs a tracking process of the laser spot. The system control unit ex407 controls the entire information reproduction / recording unit ex400. The above-described reading and writing processes are realized by the system control unit ex407 using various information held in the buffer ex404, generating and adding new information as necessary, and causing the modulation recording unit ex402, the reproduction demodulation unit ex403, and the servo control unit ex406 to cooperate, and performing information recording and reproduction through the optical head ex401. The system control unit ex407 is composed of, for example, a microprocessor, and executes those processes by executing a reading and writing program.

[0301] In the above, the optical head ex401 has been described as irradiating a laser spot, but a configuration in which higher-density recording is performed using near-field light may also be used.

[0302] Fig. 37 shows a schematic diagram of a recording medium ex215 which is an optical disc. On the recording surface of the recording medium ex215, guide grooves are formed in a spiral shape. On the information track ex230, address information indicating the absolute position on the disc is recorded in advance by a change in the shape of the grooves. This address information includes information for specifying the position of a recording block ex231 which is a unit for recording data. In a device for recording or playing back, the recording block can be specified by playing back the information track ex230 and reading the address information. Also, the recording medium ex215 includes a data recording area ex233, an inner peripheral area ex232, and an outer peripheral area ex234. The area used for recording user data is the data recording area ex233. The inner peripheral area ex232 and the outer peripheral area ex234 arranged inside or outside the data recording area ex233 are used for specific purposes other than recording user data. The information playback / recording unit ex400 reads and writes encoded audio data, video data, or multiplexed data obtained by multiplexing these data to the data recording area ex233 of such a recording medium ex215.

[0303] In the above, an optical disc such as a single-layer DVD or BD has been described as an example, but it is not limited to these, and an optical disc having a multi-layer structure and capable of recording not only on the surface may be used. Also, an optical disc having a structure for performing multi-dimensional recording / playback such as recording information using lights of different colors with different wavelengths at the same location on the disc or recording layers of different information from different angles may be used.

[0304] Also, in the digital broadcast system ex200, it is also possible for a vehicle ex210 having an antenna ex205 to receive data from a satellite ex202 or the like and play back a video on a display device such as a car navigation ex211 of the vehicle ex210. Note that as for the configuration of the car navigation ex211, for example, a configuration in which a GPS receiving unit is added to the configuration shown in Fig. 35 can be considered, and the same can be considered for a computer ex111, a mobile phone ex114, or the like.

[0305] FIG. 38A is a diagram showing a mobile phone ex114 using the moving image decoding method and the moving image encoding method described in the above embodiment. The mobile phone ex114 includes an antenna ex350 for transmitting and receiving radio waves to and from a base station ex110, a camera unit ex365 capable of taking video and still images, and a display unit ex358 such as a liquid crystal display for displaying data obtained by decoding video images captured by the camera unit ex365, video images received by the antenna ex350, and the like. The mobile phone ex114 further includes a main body unit having an operation key unit ex366, an audio output unit ex357 such as a speaker for outputting audio, an audio input unit ex356 such as a microphone for inputting audio, a memory unit ex367 for storing encoded or decoded data such as captured video images, still images, recorded audio, or received video images, still images, and mails, or a slot unit ex364 which is an interface unit with a recording medium for storing data in the same manner.

[0306] Furthermore, a configuration example of the mobile phone ex114 will be described with reference to FIG. 38B. In the mobile phone ex114, a power supply circuit unit ex361, an operation input control unit ex362, a video signal processing unit ex355, a camera interface unit ex363, an LCD (Liquid Crystal Display) control unit ex359, a modulation / demodulation unit ex352, a multiplexing / demultiplexing unit ex353, an audio signal processing unit ex354, a slot unit ex364, and a memory unit ex367 are connected to each other via a bus ex370 with respect to a main control unit ex360 that comprehensively controls each unit of the main body unit including the display unit ex358 and the operation key unit ex366.

[0307] When the end call and the power key are turned on by a user's operation, the power supply circuit unit ex361 supplies power to each unit from a battery pack to activate the mobile phone ex114 to an operable state.

[0308] When the mobile phone ex114 is in the voice call mode, based on the control of the main control unit ex360 having a CPU, ROM, RAM, etc., the voice signal picked up by the voice input unit ex356 is converted into a digital voice signal by the voice signal processing unit ex354, and this is subjected to spread spectrum processing by the modulation / demodulation unit ex352. After performing digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex351, it is transmitted via the antenna ex350. Also, when the mobile phone ex114 is in the voice call mode, the received data received via the antenna ex350 is amplified, subjected to frequency conversion processing and analog-to-digital conversion processing, subjected to spread spectrum reverse processing by the modulation / demodulation unit ex352, converted into an analog voice signal by the voice signal processing unit ex354, and then output from the voice output unit ex357.

[0309] Furthermore, when sending an email in the data communication mode, the text data of the email input by operating the operation key unit ex366 of the main body unit, etc. is sent to the main control unit ex360 via the operation input control unit ex362. The main control unit ex360 performs spread spectrum processing on the text data by the modulation / demodulation unit ex352, and after performing digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex351, it is transmitted to the base station ex110 via the antenna ex350. When receiving an email, almost the reverse process is performed on the received data, and it is output to the display unit ex358.

[0310] When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex355 compresses and encodes the video signal supplied from the camera unit ex365 by the moving image encoding method shown in each of the above embodiments (i.e., functions as an image encoding device according to an aspect of the present invention), and sends the encoded video data to the multiplexing / demultiplexing unit ex353. Also, the voice signal processing unit ex354 encodes the voice signal picked up by the voice input unit ex356 while the camera unit ex365 is imaging video, still images, etc., and sends the encoded voice data to the multiplexing / demultiplexing unit ex353.

[0311] The multiplexing / demultiplexing unit ex353 multiplexes the encoded video data supplied from the video signal processing unit ex355 and the encoded audio data supplied from the audio signal processing unit ex354 in a predetermined manner, and the resulting multiplexed data is subjected to spread spectrum processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex352, and after digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex351, it is transmitted via the antenna ex350.

[0312] When receiving data of a moving image file linked to a homepage or the like in the data communication mode, or when receiving an e-mail with video and / or audio attached, in order to decode the multiplexed data received via the antenna ex350, the multiplexing / demultiplexing unit ex353 separates the multiplexed data into a bit stream of video data and a bit stream of audio data, supplies the encoded video data to the video signal processing unit ex355 via the synchronization bus ex370, and supplies the encoded audio data to the audio signal processing unit ex354. The video signal processing unit ex355 decodes the video signal by decoding it by a moving image decoding method corresponding to the moving image encoding method shown in each of the above embodiments (i.e., functions as an image decoding device according to an aspect of the present invention), and from the display unit ex358 via the LCD control unit ex359, for example, video and still images included in a moving image file linked to a homepage are displayed. Also, the audio signal processing unit ex354 decodes the audio signal, and audio is output from the audio output unit ex357.

[0313] Also, terminals such as the mobile phone ex114 can have three implementation forms: a transmission / reception type terminal having both an encoder and a decoder, a transmission terminal having only an encoder, and a reception terminal having only a decoder, similar to the TV ex300. Further, in the digital broadcast system ex200, although it has been described as receiving and transmitting multiplexed data in which music data and the like are multiplexed with video data, data in which character data related to video is multiplexed in addition to audio data may be used, or the video data itself instead of the multiplexed data may be used.

[0314] Thus, it is possible to use the moving image encoding method or the moving image decoding method shown in each of the above embodiments in any of the devices and systems described above, and by doing so, the effects described in each of the above embodiments can be obtained.

[0315] Furthermore, the present invention is not limited to the above-described embodiments, and various modifications or corrections can be made without departing from the scope of the present invention.

[0316] (Embodiment 7) It is also possible to generate video data by appropriately switching, as necessary, between the moving image encoding method or apparatus shown in each of the above embodiments and a moving image encoding method or apparatus compliant with different standards such as MPEG-2, MPEG4-AVC, and VC-1.

[0317] Here, when generating a plurality of video data compliant with different standards respectively, it is necessary to select a decoding method corresponding to each standard when decoding. However, since it is impossible to identify which standard the video data to be decoded complies with, there arises a problem that an appropriate decoding method cannot be selected.

[0318] To solve this problem, the multiplexed data obtained by multiplexing audio data or the like with the video data has a configuration including identification information indicating which standard the video data complies with. A specific configuration of the multiplexed data including the video data generated by the moving image encoding method or apparatus shown in each of the above embodiments will be described below. The multiplexed data is a digital stream in the form of an MPEG-2 transport stream.

[0319] FIG. 39 is a diagram showing the configuration of multiplexed data. As shown in FIG. 39, the multiplexed data is obtained by multiplexing one or more of a video stream, an audio stream, a presentation graphics stream (PG), and an interactive graphics stream. The video stream shows the main video and sub-video of a movie, the audio stream (IG) shows the main audio part of the movie and the sub-audio mixed with the main audio, and the presentation graphics stream shows the subtitles of the movie. Here, the main video refers to the normal video displayed on the screen, and the sub-video refers to the video displayed in a small screen within the main video. Also, the interactive graphics stream shows an interactive screen created by arranging GUI components on the screen. The video stream is encoded by the moving image encoding method or apparatus shown in each of the above embodiments, or a moving image encoding method or apparatus compliant with conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. The audio stream is encoded in a format such as Dolby AC-3, Dolby Digital Plus, MLP, DTS, DTS-HD, or linear PCM.

[0320] Each stream included in the multiplexed data is identified by a PID. For example, 0x1011 is assigned to the video stream used for the video of the movie, 0x1100 to 0x111F are assigned to the audio stream, 0x1200 to 0x121F are assigned to the presentation graphics, 0x1400 to 0x141F are assigned to the interactive graphics stream, 0x1B00 to 0x1B1F are assigned to the video stream used for the sub-video of the movie, and 0x1A00 to 0x1A1F are assigned to the audio stream used for the sub-audio mixed with the main audio, respectively.

[0321] FIG. 40 is a diagram schematically showing how multiplexed data is multiplexed. First, a video stream ex235 composed of a plurality of video frames and an audio stream ex238 composed of a plurality of audio frames are respectively converted into PES packet sequences ex236 and ex239, and then converted into TS packets ex237 and ex240. Similarly, the data of the presentation graphics stream ex241 and the interactive graphics ex244 are respectively converted into PES packet sequences ex242 and ex245, and further converted into TS packets ex243 and ex246. The multiplexed data ex247 is constituted by multiplexing these TS packets into one stream.

[0322] FIG. 41 shows in more detail how a video stream is stored in a PES packet sequence. The first stage in FIG. 41 shows a video frame sequence of the video stream. The second stage shows the PES packet sequence. As shown by the arrows yy1, yy2, yy3, yy4 in FIG. 41, the I picture, B picture, and P picture, which are a plurality of Video Presentation Units in the video stream, are divided for each picture and stored in the payload of the PES packet. Each PES packet has a PES header, and the PES header stores a PTS (Presentation Time-Stamp), which is the display time of the picture, and a DTS (Decoding Time-Stamp), which is the decoding time of the picture.

[0323] Figure 42 shows the format of the TS packet that is finally written to the multiplexed data. The TS packet is a fixed-length packet of 188 bytes composed of a 4-byte TS header that holds information such as the PID for identifying the stream and a 184-byte TS payload for storing data. The above PES packet is split and stored in the TS payload. In the case of a BD-ROM, a 4-byte TP_Extra_Header is added to the TS packet to form a 192-byte source packet, which is written to the multiplexed data. Information such as ATS (Arrival_Time_Stamp) is described in the TP_Extra_Header. ATS indicates the transfer start time to the PID filter of the decoder for the TS packet. As shown in the lower part of Figure 42, source packets are arranged in the multiplexed data, and the number incremented from the beginning of the multiplexed data is called SPN (Source Packet Number).

[0324] In addition, the TS packets included in the multiplexed data include, in addition to each stream such as video, audio, and subtitles, PAT (Program Association Table), PMT (Program Map Table), PCR (Program Clock Reference), etc. PAT indicates what the PID of the PMT used in the multiplexed data is, and the PID of PAT itself is registered as 0. PMT has the PID of each stream such as video, audio, and subtitles included in the multiplexed data and the attribute information of the stream corresponding to each PID, and also has various descriptors regarding the multiplexed data. The descriptor includes copy control information for instructing permission / non-permission of copying the multiplexed data. PCR has the information of the STC time corresponding to the ATS when the PCR packet is transferred to the decoder in order to synchronize the ATC (Arrival Time Clock), which is the time axis of ATS, with the STC (System Time Clock), which is the time axis of PTS and DTS.

[0325] FIG. 43 is a diagram for explaining in detail the data structure of a PMT. At the head of the PMT, a PMT header describing the length of the data included in the PMT and the like is arranged. After that, a plurality of descriptors regarding the multiplexed data are arranged. The above copy control information and the like are described as descriptors. After the descriptors, a plurality of stream information regarding each stream included in the multiplexed data are arranged. The stream information is composed of a stream descriptor that describes a stream type, a PID of the stream, and attribute information of the stream (such as a frame rate and an aspect ratio) in order to identify a compression codec of the stream and the like. The stream descriptors exist in the same number as the number of streams existing in the multiplexed data.

[0326] When recording on a recording medium or the like, the above multiplexed data is recorded together with a multiplexed data information file.

[0327] As shown in FIG. 44, the multiplexed data information file is management information of the multiplexed data, corresponds one-to-one with the multiplexed data, and is composed of multiplexed data information, stream attribute information, and an entry map.

[0328] As shown in FIG. 44, the multiplexed data information is composed of a system rate, a reproduction start time, and a reproduction end time. The system rate indicates the maximum transfer rate of the multiplexed data to the PID filter of a system target decoder described later. The interval of the ATSs included in the multiplexed data is set to be equal to or less than the system rate. The reproduction start time is the PTS of the first video frame of the multiplexed data, and the reproduction end time is set to be the PTS of the last video frame of the multiplexed data plus the reproduction interval for one frame.

[0329] As shown in Fig. 45, for each stream included in the multiplexed data, stream attribute information is registered for each PID. The attribute information has different information for each of the video stream, audio stream, presentation graphics stream, and interactive graphics stream. The video stream attribute information includes information such as what compression codec the video stream is compressed with, what the resolution of the individual picture data constituting the video stream is, what the aspect ratio is, and what the frame rate is. The audio stream attribute information includes information such as what compression codec the audio stream is compressed with, how many channels are included in the audio stream, what language it corresponds to, and what the sampling frequency is. These information are used for initialization of the decoder before playback by the player and the like.

[0330] In the present embodiment, among the above multiplexed data, the stream type included in the PMT is used. Also, when the multiplexed data is recorded on the recording medium, the video stream attribute information included in the multiplexed data information is used. Specifically, in the moving image encoding method or apparatus shown in each of the above embodiments, a step or means for setting unique information indicating that it is video data generated by the moving image encoding method or apparatus shown in each of the above embodiments for the stream type included in the PMT or the video stream attribute information is provided. With this configuration, it becomes possible to distinguish the video data generated by the moving image encoding method or apparatus shown in each of the above embodiments from the video data conforming to other standards.

[0331] Also, the steps of the moving image decoding method in this embodiment are shown in FIG. 46. In step exS100, the stream type included in the PMT from the multiplexed data or the video stream attribute information included in the multiplexed data information is acquired. Next, in step exS101, it is determined whether the stream type or the video stream attribute information indicates that the multiplexed data is generated by the moving image encoding method or apparatus shown in each of the above embodiments. And when it is determined that the stream type or the video stream attribute information is generated by the moving image encoding method or apparatus shown in each of the above embodiments, in step exS102, decoding is performed by the moving image decoding method shown in each of the above embodiments. Also, when the stream type or the video stream attribute information indicates that it conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, in step exS103, decoding is performed by a moving image decoding method conforming to the conventional standard.

[0332] In this way, by setting a new unique value for the stream type or the video stream attribute information, it is possible to determine whether decoding can be performed by the moving image decoding method or apparatus shown in each of the above embodiments when decoding. Therefore, even when multiplexed data conforming to different standards is input, an appropriate decoding method or apparatus can be selected, so that decoding can be performed without causing an error. Also, the moving image encoding method or apparatus, or the moving image decoding method or apparatus shown in this embodiment can be used in any of the devices and systems described above.

[0333] (Embodiment 8) The moving image encoding method and apparatus, moving image decoding method and apparatus shown in the above embodiments are typically realized by an LSI which is an integrated circuit. As an example, FIG. 47 shows the configuration of an LSIex500 integrated into one chip. LSIex500 includes elements ex501, ex502, ex503, ex504, ex505, ex506, ex507, ex508, ex509 described below, and each element is connected via a bus ex510. The power supply circuit section ex505 starts up to an operable state by supplying power to each section when the power supply is on.

[0334] For example, when performing encoding processing, based on the control of the control section ex501 which has a CPUex502, a memory controller ex503, a stream controller ex504, a drive frequency control section ex512, etc., LSIex500 inputs an AV signal from a microphone ex117, a camera ex113, etc. through the AV I / Oex509. The input AV signal is temporarily stored in an external memory ex511 such as an SDRAM. Based on the control of the control section ex501, the stored data is appropriately divided into multiple times according to the processing amount and processing speed and sent to the signal processing section ex507, where encoding of the audio signal and / or encoding of the video signal is performed. Here, the encoding process of the video signal is the encoding process described in the above embodiments. The signal processing section ex507 further performs processes such as multiplexing the encoded audio data and the encoded video data as the case may be, and outputs it to the outside through the stream I / Oex506. This output multiplexed data is transmitted toward the base station ex107 or written to the recording medium ex215. Note that when multiplexing, it is advisable to temporarily store the data in the buffer ex508 so as to be synchronized.

[0335] Note that in the above, the memory ex511 was described as an external configuration of the LSIex500, but it may also be a configuration included inside the LSIex500. The buffer ex508 is not limited to one, and a plurality of buffers may be provided. Also, the LSIex500 may be integrated into one chip or may be made up of multiple chips.

[0336] Also, in the above description, the control unit ex501 is assumed to include the CPU ex502, the memory controller ex503, the stream controller ex504, the drive frequency control unit ex512, etc., but the configuration of the control unit ex501 is not limited to this configuration. For example, the signal processing unit ex507 may further include a CPU. By providing a CPU inside the signal processing unit ex507 as well, it becomes possible to further improve the processing speed. Also, as another example, the CPU ex502 may include the signal processing unit ex507, or a part of the signal processing unit ex507, for example, a voice signal processing unit. In such a case, the control unit ex501 has a configuration including the signal processing unit ex507 or the CPU ex502 having a part thereof.

[0337] Here, it is described as an LSI, but depending on the integration level, it may also be referred to as an IC, a system LSI, a super LSI, or an ultra LSI.

[0338] Also, the method of integrating into an integrated circuit is not limited to LSI, and it may be realized by a dedicated circuit or a general-purpose processor. After manufacturing the LSI, an FPGA (Field Programmable Gate Array) that can be programmed, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used. Such a programmable logic device can typically execute the moving image encoding method or the moving image decoding method shown in each of the above embodiments by loading a program constituting software or firmware or reading it from a memory or the like.

[0339] Furthermore, if an integrated circuit technology that replaces the LSI appears due to the progress of semiconductor technology or a derived other technology, naturally, the integration of functional blocks may be performed using that technology. The application of biotechnology or the like is possible as an example.

[0340] (Embodiment 9) When decoding video data generated by the moving image encoding method or apparatus shown in each of the above embodiments, the processing amount is considered to increase compared to the case of decoding video data compliant with conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. Therefore, in the LSIex500, it is necessary to set the driving frequency higher than the driving frequency of the CPUex502 when decoding video data compliant with the conventional standards. However, when the driving frequency is increased, there arises a problem that the power consumption increases.

[0341] To solve this problem, a moving image decoding apparatus such as the TVex300 and the LSIex500 is configured to identify which standard the video data complies with and switch the driving frequency according to the standard. FIG. 48 shows the configuration ex800 in this embodiment. The driving frequency switching unit ex803 sets the driving frequency high when the video data is generated by the moving image encoding method or apparatus shown in each of the above embodiments. Then, it instructs the decoding processing unit ex801 that executes the moving image decoding method shown in each of the above embodiments to decode the video data. On the other hand, when the video data is video data compliant with the conventional standards, the driving frequency is set lower than when the video data is generated by the moving image encoding method or apparatus shown in each of the above embodiments. Then, it instructs the decoding processing unit ex802 compliant with the conventional standards to decode the video data.

[0342] More specifically, the drive frequency switching unit ex803 is composed of the CPU ex502 and the drive frequency control unit ex512 in FIG. 47. Also, the decoding processing unit ex801 that executes the moving image decoding method shown in each of the above embodiments, and the decoding processing unit ex802 that conforms to the conventional standard correspond to the signal processing unit ex507 in FIG. 47. The CPU ex502 identifies which standard the video data conforms to. Then, based on the signal from the CPU ex502, the drive frequency control unit ex512 sets the drive frequency. Also, based on the signal from the CPU ex502, the signal processing unit ex507 decodes the video data. Here, for the identification of the video data, for example, it is conceivable to use the identification information described in Embodiment 7. The identification information is not limited to that described in Embodiment 7, and any information that can identify which standard the video data conforms to may be used. For example, when it is possible to identify which standard the video data conforms to based on an external signal that identifies whether the video data is used for a television or for a disk, etc., such an external signal may be used for identification. Also, the selection of the drive frequency in the CPU ex502 can be considered to be performed based on, for example, a look-up table that associates the standard of the video data as shown in FIG. 50 with the drive frequency. By storing the look-up table in the buffer ex508 or the internal memory of the LSI and having the CPU ex502 refer to this look-up table, it is possible to select the drive frequency.

[0343] FIG. 49 shows the steps of implementing the method of the present embodiment. First, in step exS200, the signal processing unit ex507 acquires identification information from the multiplexed data. Next, in step exS201, the CPU ex502 identifies whether the video data is generated by the encoding method or apparatus shown in each of the above embodiments based on the identification information. If the video data is generated by the encoding method or apparatus shown in each of the above embodiments, in step exS202, the CPU ex502 sends a signal for setting a high driving frequency to the driving frequency control unit ex512. Then, the driving frequency control unit ex512 sets a high driving frequency. On the other hand, if it is shown that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, in step exS203, the CPU ex502 sends a signal for setting a low driving frequency to the driving frequency control unit ex512. Then, the driving frequency control unit ex512 sets a lower driving frequency than when the video data is generated by the encoding method or apparatus shown in each of the above embodiments.

[0344] Furthermore, in conjunction with the switching of the driving frequency, it is possible to further enhance the power saving effect by changing the voltage applied to the LSI ex500 or the apparatus including the LSI ex500. For example, when setting a low driving frequency, it is conceivable to set a lower voltage applied to the LSI ex500 or the apparatus including the LSI ex500 compared to when setting a high driving frequency.

[0345] Also, the method of setting the driving frequency may be to set a high driving frequency when the processing amount during decoding is large, and to set a low driving frequency when the processing amount during decoding is small, and is not limited to the above-described setting method. For example, when the processing amount for decoding video data conforming to the MPEG4-AVC standard is larger than the processing amount for decoding video data generated by the moving image encoding method or apparatus shown in each of the above embodiments, it is conceivable to reverse the setting of the driving frequency compared to the above-described case.

[0346] Furthermore, the method of setting the driving frequency is not limited to a configuration that lowers the driving frequency. For example, when the identification information indicates that the video data is generated by the moving image encoding method or apparatus shown in each of the above embodiments, the voltage applied to the LSIex500 or an apparatus including the LSIex500 is set high. When it indicates that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, it is also conceivable to set the voltage applied to the LSIex500 or an apparatus including the LSIex500 low. As another example, when the identification information indicates that the video data is generated by the moving image encoding method or apparatus shown in each of the above embodiments, the driving of the CPUex502 is not stopped. When it indicates that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, since there is a margin in processing, it is also conceivable to temporarily stop the driving of the CPUex502. Even when the identification information indicates that the video data is generated by the moving image encoding method or apparatus shown in each of the above embodiments, if there is a margin in processing, it is also conceivable to temporarily stop the driving of the CPUex502. In this case, it is conceivable to set the stop time shorter than when it indicates that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1.

[0347] In this way, by switching the driving frequency according to the standard to which the video data conforms, it becomes possible to achieve power saving. Also, when driving an apparatus including the LSIex500 or the LSIex500 using a battery, it is possible to extend the life of the battery with the power saving.

[0348] (Embodiment 10) In devices and systems such as televisions and mobile phones described above, there may be cases where multiple video data conforming to different standards are input. In order to enable decoding even when multiple video data conforming to different standards are input in this way, the signal processing unit ex507 of LSIex500 needs to support multiple standards. However, if the signal processing unit ex507 corresponding to each standard is used individually, there will be problems such as an increase in the circuit scale of LSIex500 and an increase in cost.

[0349] To solve this problem, a configuration is adopted in which a decoding processing unit for executing the moving image decoding method shown in each of the above embodiments and a decoding processing unit conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1 are partially shared. This configuration example is shown in ex900 of FIG. 51A. For example, the moving image decoding method shown in each of the above embodiments and the moving image decoding method conforming to the MPEG4-AVC standard have some common processing contents in processes such as entropy encoding, inverse quantization, deblocking filter, and motion compensation. For the common processing contents, a decoding processing unit ex902 corresponding to the MPEG4-AVC standard is shared, and for other processing contents specific to an aspect of the present invention that do not conform to the MPEG4-AVC standard, a dedicated decoding processing unit ex901 is used. In particular, since an aspect of the present invention is characterized by motion compensation, for example, a dedicated decoding processing unit ex901 is used for motion compensation, and for any one or all of the other processes of inverse quantization, entropy decoding, and deblocking filter, it is conceivable to share the decoding processing unit. Regarding the sharing of the decoding processing unit, for the common processing contents, the decoding processing unit for executing the moving image decoding method shown in each of the above embodiments is shared, and for the processing contents specific to the MPEG4-AVC standard, a configuration using a dedicated decoding processing unit may be adopted.

[0350] Another example of sharing part of the processing is shown in ex1000 of FIG. 51B. In this example, a dedicated decoding processing unit ex1001 corresponding to the processing content specific to one aspect of the present invention, a dedicated decoding processing unit ex1002 corresponding to the processing content specific to other conventional standards, and a shared decoding processing unit ex1003 corresponding to the processing content common to the moving image decoding method according to one aspect of the present invention and the moving image decoding methods of other conventional standards are used. Here, the dedicated decoding processing units ex1001 and ex1002 are not necessarily specialized for the processing content specific to one aspect of the present invention or other conventional standards, and may be capable of executing other general-purpose processing. Also, the configuration of the present embodiment can be implemented by LSIex500.

[0351] In this way, by sharing the decoding processing unit for the processing content common to the moving image decoding method according to one aspect of the present invention and the moving image decoding methods of conventional standards, it is possible to reduce the circuit scale of the LSI and reduce the cost.

Industrial Applicability

[0352] The present invention is applicable to an image processing apparatus, an image capturing apparatus, and an image reproducing apparatus. Specifically, the present invention is applicable to a digital still camera, a movie camera, a mobile phone with a camera function, and a smartphone.

Explanation of Signs

[0353] 100 Image encoding apparatus 101 Block division unit 102 Subtraction unit 103 Frequency conversion unit 104 Quantization unit 105 Entropy encoding unit 106, 202 Inverse quantization unit 107, 203 Inverse frequency conversion unit 108, 204 Addition unit 109, 205 Intra prediction unit 110, 206 Loop filter 111, 207 Frame memory 112, 208 Inter-prediction unit 113, 209 Switching unit 121 Input image 122 Symbol block 123, 128, 224 Difference block 124, 125, 127, 222, 223 Coefficient block 126, 221 Bit stream 129, 225 Decoding block 130, 132, 134, 226, 229, 230 Prediction block 131, 227 Decoded image 133, 151, 228 Motion information 141 Encoding information selection unit 142 Motion estimation unit 143 Motion compensation unit 144 Motion information calculation unit 161 Translation flag 162 Rotation flag 163 Scaling flag 164 Shearing flag 171 Differential translation component 172 Differential rotation component 173 Differential scaling component 174 Differential shearing component 181 Encoding level information 200 Image decoding device 201 Entropy decoding unit

Claims

1. A processing circuit, and a storage device accessible from the processing circuit, wherein the processing circuit uses the storage device to decode, from a bitstream, information applied to a plurality of blocks included in a sequence, the information being used to determine the number of parameters used in prediction in an affine transformation-based prediction, generate a predicted image for each of the plurality of blocks included in the sequence by performing an affine transformation-based prediction using the number of parameters selected based on the information, wherein candidates for the number of parameters used in the prediction include 4, An image decoding apparatus.

2. A processing circuit, and a storage device accessible from the processing circuit, wherein the processing circuit uses the storage device to determine the number of parameters used in prediction in an affine transformation-based prediction, generate a predicted image for each of the plurality of blocks included in a sequence by performing an affine transformation-based prediction using the determined number of parameters, encode information applied to the plurality of blocks included in the sequence, the information being used to determine the number of parameters, wherein candidates for the number of parameters used in the prediction include 4, An image encoding apparatus.

3. A processing circuit, and a storage device accessible from the processing circuit, wherein the processing circuit uses the storage device to generate information applied to a plurality of blocks included in a sequence, the information being used to determine the number of parameters used in prediction processing in an affine transformation-based prediction, include the information in a bitstream, wherein the prediction processing includes processing for generating a predicted image for each of the plurality of blocks included in the sequence by performing an affine transformation-based prediction using the number of parameters selected based on the information, wherein candidates for the number of parameters used in the prediction include 4, A bitstream generation apparatus.

Citation Information

Patent Citations

  • Image encoding device and image decoding device

    JP1996065680A

  • Motion compression predict coding method for dynamic image, decoding method, coder and decoder

    JP1997224252A

  • Method for encoding and decoding digital image and device for encoding and decoding digital image using the same

    JP2004007804A

  • Image encoding method, image decoding method, image encoding device, image decoding device, image encoding program, image decoding program and recording medium recorded with the programs

    JP2004159132A

  • Image decoding device, image encoding device, and bitstream generating device

    JP7664546B2