Video encoding device and video decoding device

The moving image decoding and encoding apparatuses address synchronization issues by using estimated phase display information to adaptively manage resolution changes, ensuring synchronized image decoding even without the phase indication SEI message, thus preventing temporal shifts.

JP2025101791APending Publication Date: 2025-07-08SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023218806
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In video coding methods like ITU-T Rec. H.266, the phase indication SEI message for image resolution changes is not mandatory, leading to potential misalignment and temporal shifts when encoding and decoding images with different reduction ratios, causing synchronization issues during resolution conversion.

Method used

A moving image decoding apparatus that adaptively decodes data with changed resolution, incorporating estimated phase display information when the SEI message is absent, and a moving image encoding apparatus that generates encoded data with adaptive resolution and estimated phase information, ensuring synchronization between transmission and reception sides.

Benefits of technology

Prevents temporal shifts in images during resolution changes by providing estimated phase display information, maintaining synchronization and improving encoding efficiency even without the phase indication SEI message.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025101791000001_ABST
    Figure 2025101791000001_ABST
Patent Text Reader

Abstract

To provide information for performing image resolution conversion processing to the same picture size on the decoding side when encoded and decoded at different picture sizes with different reduction ratios in the same sequence.SOLUTION: In a video transmission system 1, a video decoding device 30 includes: an image decoding device 31 that decodes encoded data whose resolution has been adaptively changed to generate a decoded image; an image resolution conversion processing device 61 that changes the resolution of the decoded image; and an auxiliary extension information decoding device 91 that decodes phase display information for operating the image resolution conversion processing device 61. The video decoding device has an estimated value of the phase display information when the auxiliary extension information is not present.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a moving image encoding device and a decoding device.

Background Art

[0002] In order to efficiently transmit or record a moving image, a moving image encoding device that generates encoded data by encoding the moving image, and a moving image decoding device that generates a decoded image by decoding the encoded data are used.

[0003] Specific moving image encoding methods include, for example, the H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding) methods.

[0004] In such a moving image encoding method, an image (picture) constituting a moving image is managed by a hierarchical structure composed of a slice obtained by dividing the image, a coding tree unit (CTU) obtained by dividing the slice, a coding unit (sometimes called a coding unit (CU)) obtained by dividing the coding tree unit, and a transform unit (TU) obtained by dividing the coding unit, and is encoded / decoded for each CU.

[0005] Also, in such a moving image encoding method, usually, a predicted image is generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is encoded. Examples of the method for generating a predicted image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).

[0006] In the latest video coding method of Non-Patent Document 1, a method called RPR (Reference Picture Re-sampling) that can dynamically change the resolution of the coded image (picture) is adopted.

[0007] Also, in Non-Patent Document 2, as a technique for video coding and decoding, an auxiliary enhancement information SEI (Supplemental Enhancement Information) message for transmitting the properties of an image, display method, timing, etc. simultaneously with the coded data is defined. There is a phase indication SEI message (Phase Indication SEI message) that indicates the phase information of pixels when the resolution of the image is changed.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0009] In the video coding method disclosed in Non-Patent Document 1, a method called RPR (Reference Picture Re-sampling) that changes the picture size of the coded picture and can dynamically change the resolution is adopted. Also, since the phase indication information in Non-Patent Document 2 is SEI (auxiliary enhancement information), it is not information that must always be sent. When the phase indication SEI does not exist, reduction and If the phase display information for enlargement is not correct on the transmission side and the reception side, when encoding and decoding pictures with different reduction ratios in the same sequence and then performing image resolution conversion processing to the same picture size on the decoding side, there is a problem that the images are shifted in time at the timing when the resolution changes.

Means for Solving the Problem

[0010] A moving image decoding apparatus according to an aspect of the present invention includes an image decoding apparatus that adaptively decodes encoded data with a changed resolution to generate a decoded image, and an image resolution conversion processing apparatus that changes the resolution of the decoded image, and includes an auxiliary extension information decoding apparatus that decodes phase display information for operating the image resolution conversion processing apparatus, characterized by having an estimated value of the phase display information when the previous auxiliary extension information does not exist.

[0011] Also, a moving image encoding apparatus according to an aspect of the present invention includes an image reduction processing apparatus that adaptively reduces an image to the same size or reduces it, an image encoding apparatus that generates encoded data of an image with an adaptively changed resolution, an auxiliary extension information generation apparatus that creates phase display information for operating the image reduction processing apparatus, and includes an auxiliary extension information encoding apparatus that encodes the phase display information, characterized by having an estimated value of the phase display information when the previous auxiliary extension information does not exist.

Advantages of the Invention

[0012] With such a configuration, even when the phase display SEI message does not exist, the phase display information for reduction and enlargement can match the transmission side and the reception side. When encoding and decoding pictures with different reduction ratios in the same sequence and then performing image resolution conversion processing to the same picture size on the decoding side, the problem that the images are shifted in time at the timing when the resolution changes can be solved.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0014] (First Embodiment) Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0015] FIG. 1 is a schematic diagram showing the configuration of another moving image transmission system according to the present embodiment.

[0016] The moving image transmission system 1 is a system that transmits encoded data obtained by encoding an image and decodes and displays the transmitted encoded data. The moving image transmission system 1 includes a moving image encoding device 10, a network 21, a moving image decoding device 30, and an image display device 41.

[0017] The moving image encoding device 10 is composed of an image encoding device (image encoding unit) 11, an auxiliary extension information creation device (auxiliary extension information creation unit) 71, an auxiliary extension information encoding device (auxiliary extension information encoding unit) 81, and an image reduction processing device (image reduction processing unit) 51.

[0018] The moving image encoding device 10 creates a non-reduced or reduced image T2 from the input moving image T1 using the image reduction processing device 51, compresses and encodes the image using the image encoding device 11, analyzes the input moving image T1 and the partial decoded image T3 of the image encoding device 11, generates auxiliary extension information for input to the auxiliary extension processing device 61 using the auxiliary extension information creation device 51, encodes it using the auxiliary extension information encoding device 81 to generate encoded data Te, and sends it to the network 21. Note that the reduced image T2 includes the case of non-reduction.

[0019] The moving image decoding device 30 is composed of an image decoding device (image decoding unit) 31, an auxiliary extension information decoding device (auxiliary extension information decoding unit) 91, and an image resolution conversion processing device (image resolution conversion processing unit) 61.

[0020] The moving image decoding device 30 decodes the encoded data Te received from the network 21 using the image decoding device 31 and the auxiliary extension information decoding device 91, performs auxiliary extension processing on the decoded image Td1 using the auxiliary extension information in the auxiliary extension processing device 61, and outputs the enlarged image Td2 to the image display device 41. Note that the enlarged image Td2 includes the case of non-reduction.

[0021] The image display device 41 displays all or part of the auxiliary extended processing image Td2 output from the auxiliary extended processing device 61. The image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Also, when the image decoding device 31 has high processing capabilities, an image with high image quality is displayed, and when it has only low processing capabilities, an image that does not require high processing capabilities and display capabilities is displayed.

[0022] The network 21 transmits the encoded auxiliary extended information and the encoded data Te to the image decoding device 31. Part or all of the encoded auxiliary extended information may be included in the encoded data Te as the auxiliary extended information SEI. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Also, the network 21 may be replaced by a storage medium that records the encoded data Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).

[0023] In such a configuration, a framework is provided that enables efficient encoding and decoding of auxiliary extended information.

[0024] FIG. 4 is a conceptual diagram of an image (picture) to be processed in the moving image transmission system shown in FIG. 1, and shows the change in the resolution of the image over time. However, in FIG. 4, it is not distinguished whether the image is encoded or not. FIG. 4 shows that in the processing process of the moving image transmission system, the resolution is decreased to reduce the image (picture) size and the image An example of transmitting an image to the decoding device 31 is shown. As shown in FIG. 4, normally, the image processing device 51 performs a conversion to reduce the resolution of the image to the same as or lower than that of the input image in order to reduce the amount of information of the transmitted information.

[0025] <Operator> The operators used in this specification are described below.

[0026] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || indicates a logical OR.

[0027] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).

[0028] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). If c < a, it returns a; if c > b, it returns b; otherwise, it returns c (where a <= b).

[0029] abs(a) is a function that returns the absolute value of a.

[0030] Int(a) is a function that returns the integer value of a.

[0031] Floor(a) is a function that returns the largest integer less than or equal to a.

[0032] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0033] a / d represents the division of a by d (truncating the decimal part).

[0034] a ÷ d, and a over d represent the division of a by d (without rounding).

[0035] (Structure of the encoded data Te) Prior to the detailed description of the image encoding device 11 and the image decoding device 31 according to this embodiment, the data structure of the encoded data Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be described with reference to FIGS. 2 and 3.

[0036] The encoded data Te is a bitstream composed of a plurality of CVSs (Coded Video Sequences) and EoB (End of Bitstream) NAL units shown in FIG. 2. A CVS is composed of a plurality of AUs (Access Units) and an EoS (End of Sequence) NAL unit. The AU at the head of the CVS is called a CVSS (Coded Video Sequence Start) AU. A unit obtained by dividing a CVS for each layer is called a CLVS (Coded Layer Video Sequence). An AU consists of one or a plurality of PUs (Picture Units) of the same output time in one or more layers. If the multi-layer encoding method is not adopted, an AU consists of one PU. A PU is a unit of the encoded data of one decoded picture composed of a plurality of NAL units. A CLVS is composed of PUs of the same layer, and the PU at the head of the CLVS is called a CLVSS (Coded Layer Video Sequence Start) PU. The CLVSS PU is limited to PUs composed of IRAP (Intra Random Access Pictures) and GDR (Gradual Decoder Refresh Picture) that are randomly accessible. An NAL unit is composed of an NAL unit header and RBSP (Raw Byte Sequence Payload) data. The NAL unit header is composed of 6-bit nuh_layer_id indicating the layer value, 5-bit nuh_unit_type indicating the NAL unit type, and 3-bit nuh_temporal_id_plus1 which is a value obtained by adding 1 to the Temporal ID value, following 2-bit 0 data.

[0037] FIG. 3 is a diagram showing the hierarchical structure of data in the encoded data Te in PU units. The encoded data Te includes, by way of example, a sequence and a plurality of pictures constituting the sequence. FIG. 3 shows a diagram showing an encoded video sequence that defines a sequence SEQ, an encoded picture that defines a picture PICT, an encoded slice that defines a slice S, encoded slice data that defines slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.

[0038] In the encoded video sequence, a set of data that the image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in FIG. 3, the sequence SEQ includes a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information (SEI).

[0039] In the video parameter set VPS, in a moving image composed of a plurality of layers, a set of encoding parameters common to the plurality of moving images and a set of encoding parameters related to the plurality of layers and individual layers included in the moving image are defined.

[0040] In the sequence parameter set SPS, a set of encoding parameters that the image decoding device 31 refers to in order to decode the target sequence is defined. For example, the width and height of the picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS.

[0041] Here, the sequence parameter set SPS includes the following syntax elements. ·pic_width_max_in_luma_samples: A syntax element that specifies the width of the image with the maximum width among the images in a single sequence in luminance block units. Also, the value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY). Here, MinCbSizeY is a value determined by the minimum size of the luminance block. ·pic_height_max_in_luma_samples: A syntax element that specifies the height of the image with the maximum height among the images in a single sequence in luminance block units. Also, the value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY).

[0042] In the picture parameter set PPS, a set of encoding parameters that the image decoding device 31 refers to for decoding each picture in the target sequence is defined. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence.

[0043] Here, the picture parameter set PPS includes the following syntax elements.

[0044] pps_pic_width_in_luma_samples is a syntax element that specifies the width of the target picture. The value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and a value less than or equal to sps_pic_width_max_in_luma_samples.

[0045] pps_pic_height_in_luma_samples is a syntax element that specifies the height of the target picture. The value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and a value less than or equal to sps_pic_height_max_in_luma_samples.

[0046] The pps_conformance_window_flag is a flag indicating whether the conformance (cropping) window offset parameter will be subsequently notified, and is a flag indicating the location where the conformance window is to be displayed. When this flag is 1, the said parameter is notified; when it is 0, it indicates that there is no conformance window offset parameter.

[0047] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are offset values for specifying the left, right, top, and bottom positions of the picture output in the decoding process with respect to the rectangular area specified by the output picture coordinates. Also, when the value of conformance_window_flag is 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are presumed to be 0.

[0048] For the output picture, the horizontal direction ranges from SubWidthC * pps_conf_win_left_offset to pps_pic_width_in_luma_samples - (SubWidthC * pps_conf_win_right_offset + 1), and the vertical direction ranges from SubHeightC * pps_conf_win_top_offset to pps_pic_height_in_luma_samples - (SubHeightC * pps_conf_win_bottom_offset + 1) in terms of pixels.

[0049] Here, the variable ChromaFormatIdc of the chroma format is the value of sps_chroma_format_id, and the variables SubWidthC and SubHightC are values determined by this ChromaFormatIdc. In the case of the monochrome format, both SubWidthC and SubHightC are 1. In the case of the 4:2:0 format, both SubWidthC and SubHightC are 2. In the case of the 4:2:2 format, SubWidthC is 2 and SubHightC is 1. In the case of the 4:4:4 format, both SubWidthC and SubHightC are 1.

[0050] (Encoded Picture) In the encoded picture, a set of data that the image decoder 31 refers to for decoding the picture PICT to be processed is defined. As shown in FIG. 3, the picture PICT includes a picture header PH and slices 0 to NS - 1 (NS is the total number of slices included in the picture PICT).

[0051] (Encoded Slice) In the encoded slice, a set of data that the image decoder 31 refers to for decoding the slice S to be processed is defined. As shown in FIG. 3, the slice includes a slice header and slice data.

[0052] The slice header includes a set of encoding parameters that the image decoder 31 refers to for determining the decoding method of the target slice. The slice type specification information (slice_type) that specifies the slice type is an example of the encoding parameters included in the slice header.

[0053] The slice types that can be specified by the slice type specification information include: (1) I slice that uses only intra prediction during encoding, (2) P slice that uses single prediction (L0 prediction) or intra prediction during encoding, (3) B slice that uses single prediction (L0 prediction or L1 prediction), bi-prediction, or intra prediction during encoding, and so on. Note that inter prediction is not limited to single prediction and bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it refers to a slice including a block that can use inter prediction.

[0054] (Encoded slice data) In the encoded slice data, a set of data that the image decoding device 31 refers to for decoding the slice data to be processed is defined. The slice data includes CTUs as shown in the encoded slice header of FIG. 3. A CTU is a block of a fixed size (for example, 64x64) that constitutes a slice, and is sometimes referred to as the largest coding unit (LCU).

[0055] (Coding tree unit) FIG. 3 defines a set of data that the image decoding device 31 refers to for decoding the CTU to be processed. The CTU is divided into coding units (CUs), which are the basic units of the encoding process, by recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). A node of the tree structure obtained by recursive quadtree partitioning is called a coding node. Intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the topmost coding node.

[0056] (Coding unit) FIG. 3 defines a set of data that the image decoding device 31 refers to in order to decode the encoding unit to be processed. Specifically, the CU is composed of a CU header CUH, prediction parameters, transform parameters, quantization transform coefficients, etc. The prediction mode and the like are defined in the CU header.

[0057] The prediction process may be performed in units of CUs or in units of sub-CUs obtained by further dividing the CUs.

[0058] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to a prediction process performed between different pictures (for example, between display times, between layer images).

[0059] The transform and quantization process is performed in units of CUs, but the quantization transform coefficients may be entropy encoded in units of sub-blocks such as 4x4.

[0060] In addition, when it is described as "a flag indicating whether or not XX" in this specification, when the flag is other than 0 (for example, 1), it is considered that XX is true, and when the flag is 0, XX is false. In logical negation, logical product, etc., 1 is regarded as true and 0 is regarded as false (the same applies hereinafter). However, in actual devices and methods, other values can also be used as true values and false values.

[0061] (Configuration of the image decoding device) The configuration of the image decoding device 31 (FIG. 5) according to the present embodiment will be described.

[0062] The image decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (predicted image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation device) 308, an inverse quantization and inverse transform unit 311, an addition unit 312, and a prediction parameter derivation unit 320. Note that, in accordance with the image encoding device 11 described later, there is also a configuration in which the image decoding device 31 does not include the loop filter 305.

[0063] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice headers (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. The TU decoding unit 3024 decodes QP update information (quantization correction value) and quantization prediction error (residual coding) from the encoded data.

[0064] The predicted image generation unit 308 is configured to include an inter-predicted image generation unit 309 and an intra-predicted image generation unit 310.

[0065] The entropy decoding unit 301 performs entropy decoding on the encoded data Te input from the outside to decode individual codes (syntax elements).

[0066] The entropy decoding unit 301 outputs the decoded codes to the parameter decoding unit 302. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.

[0067] (Basic flow) FIG. 6 is a flowchart for explaining the schematic operation of the image decoding apparatus 31.

[0068] (S1100: Parameter set information decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.

[0069] (S1200: Slice information decoding) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.

[0070] Hereinafter, the image decoding device 31 derives a decoded image of each CTU by repeating the processes from S1300 to S5000 for each CTU included in the target picture.

[0071] (S1300: CTU Information Decoding) The CT information decoding unit 3021 decodes the CTU from the encoded data.

[0072] (S1400: CT Information Decoding) The CT information decoding unit 3021 decodes the CT from the encoded data.

[0073] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data.

[0074] (S1510: CU Information Decoding) The CU decoding unit 3022 decodes CU information, prediction information, etc. from the encoded data.

[0075] (S1520: TU Information Decoding) The TU decoding unit 3024 decodes QP update information, quantization prediction error, etc. from the encoded data. Note that the QP update information is a difference value from the quantization parameter prediction value qPpred which is a predicted value of the quantization parameter QP.

[0076] (S2000: Predicted Image Generation) The predicted image generation unit 308 generates a predicted image for each block included in the target CU based on the prediction information.

[0077] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation processing for each TU included in the target CU.

[0078] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the target CU by adding the predicted image supplied from the predicted image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transformation unit 311.

[0079] (S5000: Loop Filter) The loop filter 305 applies loop filters such as a deblocking filter, SAO, and ALF to the decoded image to generate a decoded image.

[0080] (Phase Indication SEI Message) Figure 8 shows the syntax of the Phase Indication SEI message. The Phase Indication SEI message provides the decoder with information regarding the pixel positions of luminance within the decoded picture with respect to the display screen. This information can be used by the decoder, for example, to ensure correct spatial alignment of the picture with the changed resolution when switching the resolution of the image.

[0081] The syntax elements pi_hor_phase_num and pi_hor_phase_den_minus1 specify the horizontal position of the luminance sampling position with respect to the display screen. The horizontal position pi_hor_phase_num ÷ (pi_hor_phase_den_minus1 + 1) is represented in units of the horizontal distance between the pixel positions of two adjacent luminances in the horizontal direction. The value of pi_hor_phase_num must be 0 or greater and pi_hor_phase_den_minus1 + 1 or less.

[0082] The syntax elements pi_ver_phase_num and pi_ver_phase_den_minus1 specify the vertical position of the luminance sampling position with respect to the display screen. The vertical position pi_ver_phase_num ÷ (pi_ver_phase_den_minus1 + 1) is represented in units of the vertical distance between the pixel positions of two adjacent luminance pixels in the vertical direction. The value of pi_ver_phase_num must be 0 or greater and pi_ver_phase_den_minus1 + 1 or less.

[0083] Note that when the number of output pixels of the horizontal luminance is equal to the width of the display screen and pi_hor_phase_num ÷ (pi_hor_phase_den_minus1 + 1) is equal to 1 / 2, the image is intended to have its resolution converted based on the center of the picture without applying a horizontal phase shift. Similarly, when the number of vertical luminance output pixels is equal to the height of the display screen and pi_ver_phase_num ÷ (pi_ver_phase_den_minus1 + 1) is equal to 1 / 2, the picture will have its resolution converted based on the center of the picture without applying a vertical phase shift.

[0084] The ratios a / b and c / d in FIG. 9 represent the horizontal and vertical positions of the luminance samples (marked with x) with respect to the display screen. a / b is equal to pi_hor_phase_num ÷ (pi_hor_phase_den_minus1 + 1), and c / d is equal to pi_ver_phase_num ÷ (pi_ver_phase_den_minus1 + 1).

[0085] To use this SEI message, the following variables need to be defined.

[0086] The width and height in pixel units of the luminance signal of the decoded picture for display are represented by the variables CroppedWidth and CroppedHeight, respectively. This value is defined in the SEI payload described later.

[0087] The phase display SEI message is applied to the current decoded picture and persists for all subsequent pictures in the current layer in output order, using the same value of CroppedWidth as the current picture and the same value of CroppedHeight as the current picture, until any of the following conditions are met. · When a new CLVS of the current layer is started. · When the encoded data ends. · When another phase display SEI message exists.

[0088] The phase representation information can be used for the process of converting the resolution. For example, when the resolution is changed, the positions of the pixels are offset by an amount proportional to the horizontal and vertical phase information at which the signal was transmitted. Specifically, in the present embodiment, it is used in the image resolution conversion processing apparatus 61.

[0089] Note that the phase representation information is applied to the luminance pixels of the decoded picture. The offset of the phase of the pixels of the chrominance signal can be derived from the encoded phase information in consideration of the positions of the pixels of the chrominance signal relative to the positions of the pixels of the luminance indicated by ChromaFormatIdc, vui_chroma_sample_loc_type_frame of VUI (Video Usability Information), or vui_chroma_sample_loc_type_top_field and vui_chroma_sample_loc_type_bottom_field in the case of a field picture.

[0090] When the variable ChromaFormatIdc is equal to 1 (4:2:0 chrominance format), and the value of vui_chroma_sample_loc_ type_frame, or vui_chroma_sample_loc_type_top_field and vui_chroma_sample_loc_type_bottom_field (if applicable) in the case of a field picture is equal to 6, or is presumed to be equal to 6, that is, when the value is unknown, the values of vui_chroma_sample_loc_type_frame, or vui_chroma_sample_loc_type_top_field and vui_chroma_sample_loc_type_bottom_field in the case of a field picture are presumed to be 0.

[0091] (SEI payload) FIG. 10 is a diagram showing a part of the syntax of an SEI payload which is a container of an SEI message. Specifically, it is encoded by the auxiliary extension information encoding apparatus 81 and decoded by the auxiliary extension information decoding apparatus 91.

[0092] PREFIX_SEI_NUT, which is called when nal_unit_type is PREFIX_SEI_NUT, indicates that the SEI is located before the slice data.

[0093] When payloadType is 212, the Phase Indication SEI message is called.

[0094] SUFFIX_SEI_NUT, which is called when nal_unit_type is SUFFIX_SEI_NUT, indicates that the SEI is located after the slice data.

[0095] The following variables are specified to decode the Phase Indication SEI message. · The variable CroppedWidth is set equal to pps_pic_width_in_luma_samples - SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset). · The variable CroppedHeight is set equal to pps_pic_height_in_luma_samples - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset).

[0096] In the moving image encoding method disclosed in Non-Patent Document 1, a method called RPR (Reference Picture Re-sampling) is adopted, which can change the picture size of the encoded picture and dynamically change the resolution. Also, since the phase display information in Non-Patent Document 2 is SEI (Supplemental Enhancement Information), it is not information that must always be sent. When the phase display SEI does not exist, if the phase display information for reduction and enlargement does not match between the transmission side and the reception side, when encoding and decoding pictures with different reduction rates in the same sequence and then performing image resolution conversion processing to the same picture size on the decoding side, there is a problem that the image is shifted in time at the timing when the resolution changes.

[0097] Therefore, in the present embodiment, when the phase display SEI message does not exist, estimated values of the syntax elements pi_hor_phase_num, pi_hor_phase_den_minus1, pi_ver_phase_num, and pi_ver_phase_den_minus1 of the phase display SEI message are defined, and the pixel positions of the luminance in the decoded picture with respect to the display screen when the encoded picture size is reduced are defined.

[0098] Specifically, it is estimated that the syntax elements pi_hor_phase_num, pi_hor_phase_den_minus1, pi_ver_phase_num, and pi_ver_phase_den_minus1 are the minimum values that satisfy the following equations.

[0099] pi_hor_phase_num / (pi_hor_phase_den_minus1 + 1) = CroppedWidth / (2 * OrgCroppedWidth) pi_ver_phase_num / (pi_ver_phase_den_minus1 + 1) = CroppedHeight / (2 * OrgCroppedHeight) Note that the variables OrgCroppedWidth and OrgCroppedHeight are defined as follows.

[0100] OrgCroppedWidth = sps_pic_width_max_in_luma_samples - SubWidthC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) OrgCroppedHeight = sps_pic_height_max_in_luma_samples - SubHeightC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) These equations mean that the value obtained by doubling the width of the display picture size defined in the SPS is used as the denominator, the width of the display picture size defined in the PPS is used as the numerator, the value of the denominator minus 1 of the reduced fraction is taken as pi_hor_phase_den_minus1, and the value of the reduced numerator is taken as pi_hor_phase_num. Also, the value obtained by doubling the height of the display picture size defined in the SPS is used as the denominator, the height of the display picture size defined in the PPS is used as the numerator, the value of the denominator minus 1 of the reduced fraction is taken as pi_ver_phase_den_minus1, and the value of the reduced numerator is taken as pi_ver_phase_num.

[0101] These estimated values indicate that the phase display information is at the upper left position of the picture in the SPS. Aligning the pixel positions to the upper left position of the picture means that even if the picture size is changed, the pixel positions of the corresponding luminance signals among pictures with different picture sizes are the same, which can be said to be a desirable pixel position in terms of coding efficiency.

[0102] In the present embodiment, as another estimated value of the phase display information when the phase display SEI message does not exist, the values of the syntax elements pi_hor_phase_num, pi_hor_phase_den_minus1, pi_ver_phase_num, and pi_ver_phase_den_minus1 may all be estimated as 1.

[0103] These estimated values indicate that the phase display information is at the center position of the picture. Aligning the pixel position with the center position of the picture means that the pixel position is evenly enlarged or reduced from the display surface, so it can be said to be a desirable pixel position.

[0104] FIG. 11 shows a flowchart of the processing of the moving image decoding apparatus 30 according to the present embodiment. The image decoding apparatus 31 performs image decoding processing (S2000). Next, it is confirmed whether the phase display SEI message exists (S2100). If the phase display SEI message exists, the auxiliary extension information decoding apparatus 91 performs decoding processing of the phase display SEI message (S2200). If the phase display SEI message does not exist, an estimated value of the phase display information is set (S2300). Using this phase display information, image resolution conversion processing (S2400) is performed.

[0105] The moving image encoding apparatus 10 generates phase display information in the auxiliary extension information creating apparatus 71 based on the reduction method used in the image reduction processing apparatus 51, and encodes the phase display SEI message in the auxiliary extension information encoding apparatus 81. When the moving image decoding apparatus 30 cannot handle the phase display SEI message, or when image reduction processing is performed using the estimated value of the phase display information, it is not always necessary to encode the phase display SEI message.

[0106] In this way, when the phase display SEI message does not exist, the problem can be solved by defining the estimated value of the syntax element of the phase display SEI message and defining the luminance pixel position in the decoded picture for the default display screen when the encoded picture size is reduced.

[0107] As another solution, for the syntax elements pi_hor_phase_num, pi_hor_phase_den_minus1, pi_ver_phase_num, and pi_ver_phase_den_minus1 of the phase representation SEI message, if they do not exist, they are estimated as values of all 0s, and in that case, the phase representation information is determined to be indeterminate. Then, the phase representation SEI message may be always sent in the CLVS where the RPR exists.

[0108] If the phase representation information for reduction and enlargement is not correct between the transmitting side and the receiving side, when encoding and decoding pictures with different picture sizes at different reduction rates in the same sequence, at the decoding side, for the same picture size As another solution to the problem that when performing image resolution conversion processing, the image will be temporally shifted at the timing when the resolution changes, in this embodiment, a method of adding a new syntax element to the SPS is shown.

[0109] In the SPS of Non-Patent Document 1, there are syntax elements sps_ref_pic_resampling_enabled_flag and sps_res_change_in_clvs_allowed_flag.

[0110] When sps_res_change_in_clvs_allowed_flag is equal to 1, it indicates that the picture size may be changed within the CLVS that refers to the SPS. When sps_res_change_in_clvs_allowed_flag is equal to 0, it indicates that the picture size will not be changed within the CLVS that refers to the SPS. If it does not exist, the value of sps_res_change_in_clvs_allowed_flag is estimated to be equal to 0.

[0111] If sps_res_change_in_clvs_allowed_flag is equal to 1, it indicates that the picture size may be changed within the CLVS referring to the SPS. If sps_res_change_in_clvs_allowed_flag is equal to 0, it indicates that the picture size will not be changed within the CLVS referring to the SPS. If it does not exist, the value of sps_res_change_in_clvs_allowed_flag is presumed to be equal to 0. In Non-Patent Document 1, a method called RPR (Reference Picture Re-sampling) is adopted, where sps_ref_pic_resampling_enabled_flag and sps_res_change_in_clvs_allowed_flag are set to 1, and the coded picture sizes pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples can be made smaller than the input picture sizes sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples, enabling dynamic change of the resolution of the coded picture.

[0112] In the present embodiment, as shown in FIG. 12, when both sps_ref_pic_resampling_enabled_flag and sps_res_change_in_clvs_allowed_flag are 1, sps_phase_center_position_flag is added as a new syntax element.

[0113] When sps_phase_center_position_flag is equal to 1, it indicates that the phase display information when the coded picture size is smaller than the input picture size is the center position of the picture. When sps_phase_center_position_flag is equal to 0, it indicates that the phase display information when the coded picture size is smaller than the input picture size is the upper left position of the picture. If it does not exist, it is presumed that the value of sps_phase_center_position_flag is equal to 0.

[0114] This flag can be used to determine whether the center position of the picture, which is a representative pixel position, or the upper left position of the picture is used during RPR. In the SPS, since the actually coded picture size is unknown, arbitrary phase display information cannot be defined, but it is possible to avoid the phase display information becoming indefinite.

[0115] Note that the syntax element sps_phase_center_position_flag may be used to extend and describe the SPS using sps_extension_flag.

[0116] By adopting such a configuration, when coding and decoding pictures with different reduction ratios in the same sequence, when the image resolution conversion process is performed to the same picture size on the decoding side, the problem that the image is shifted in time at the timing when the resolution changes can be solved.

[0117] Also, as another solution, in this embodiment, a method of adding a new syntax element to the PPS is shown.

[0118] As described above, the PPS in Non-Patent Document 1 has syntax elements such as pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, pps_conformance_window_flag, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset.

[0119] In this embodiment, as shown in FIG. 13, new syntax elements pps_phase_indication_flag, pps_hor_phase_num, pps_hor_phase_den_minus1, pps_ver_phase_num, and pps_ver_phase_den_minus1 are defined.

[0120] When pps_phase_indication_flag is equal to 1, it indicates that phase display information exists in the PPS. When pps_phase_indication_flag is equal to 0, it indicates that phase display information does not exist in the PPS.

[0121] Note that pps_phase_indication_flag can only take the value of 1 only when the values of both sps_ref_pic_resampling_enabled_flag and sps_res_change_in_clvs_allowed_flag are 1. A constraint may be imposed that when the values of both sps_ref_pic_resampling_enabled_flag and sps_res_change_in_clvs_allowed_flag are not both 1, pps_phase_indication_flag must take the value of 0.

[0122] When the value of pps_phase_indication_flag is 1, the syntax elements pps_hor_phase_num, pps_hor_phase_den_minus1, pps_ver_phase_num, and pps_ver_phase_den_minus1 exist. It is assumed that these syntax elements have the same semantics as the syntax elements pi_hor_phase_num, pi_hor_phase_den_minus1, pi_ver_phase_num, and pi_ver_phase_den_minus1 of the phase indication SEI message.

[0123] If the syntax elements pps_hor_phase_num, pps_hor_phase_den_minus1, pps_ver_phase_num, and pps_ver_phase_den_minus1 do not exist, it is estimated that there is phase indication information at the pixel position at the upper left position of the picture in the SPS. Specifically, the syntax elements pps_hor_phase_num, pps_hor_phase_den_minus1, pps_ver_phase_num, and pps_ver_phase_den_minus1 are estimated to be the minimum values that satisfy the following equations.

[0124] pps_hor_phase_num / (pps_hor_phase_den_minus1 + 1) = CroppedWidth / (2 * OrgCroppedWidth) pps_ver_phase_num / (pps_ver_phase_den_minus1 + 1) = CroppedHeight / (2 * OrgCroppedHeight) Note that the variables OrgCroppedWidth and OrgCroppedHeight are defined as follows.

[0125] OrgCroppedWidth = sps_pic_width_max_in_luma_samples - SubWidthC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) OrgCroppedHeight = sps_pic_height_max_in_luma_samples - SubHeightC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) These formulas mean that the denominator is twice the width of the display picture size defined in the SPS, the numerator is the width of the display picture size defined in the PPS, the value of the denominator of the reduced fraction minus 1 is pps_hor_phase_den_minus1, and the value of the reduced numerator is pps_hor_phase_num. Also, the denominator is twice the height of the display picture size defined in the SPS, the numerator is the height of the display picture size defined in the PPS, the value of the denominator of the reduced fraction minus 1 is pps_ver_phase_den_minus1, and the value of the reduced numerator is pps_ver_phase_num.

[0126] In this embodiment, the values of the syntax elements pps_hor_phase_num, pps_hor_phase_den_min us1, pps_ver_phase_num, and pps_ver_phase_den_minus1 may all be estimated to be 1. In this case, the estimated value of the phase display information is the center position of the picture.

[0127] Similar to the above-mentioned syntax element sps_phase_center_position_flag, the phase display information may be limited to the upper left position and the center position of the picture, and described using a flag called pps_phase_ecnter_position_flag.

[0128] The above syntax elements on the PPS may be described by extending the PPS using pps_extension_flag.

[0129] By adopting such a configuration, when encoding and decoding picture sizes with different reduction ratios in the same sequence, when performing image resolution conversion processing to the same picture size on the decoding side, the problem that the image is temporally shifted at the timing when the resolution changes can be solved.

[0130] As another problem, in the phase display SEI message disclosed in Non-Patent Document 2, the phase display method may not become a constant value when matching the pixel position at the upper left position of a desirable picture in terms of encoding efficiency.

[0131] In the phase display SEI message disclosed in Non-Patent Document 2, a method is adopted in which the syntax element becomes a constant when the pixel position is matched to the center position of the picture. However, in terms of encoding efficiency, matching the pixel position to the upper left position of the picture is desirable because the pixel positions of corresponding luminance signals match between pictures of different picture sizes.

[0132] In the present embodiment, a method for defining phase display information is shown based on the upper left position of the display picture.

[0133] FIG. 14 explains a method for defining phase display information based on the upper left position of the display picture before reduction. In this example, the reduced picture is reduced to half the size both in the horizontal and vertical directions. The circled marks are the pixel positions of the reduced picture. On the left side (a), it is the case where the pixel position of the reduced picture is at the upper left position, which coincides with the pixel position of the display picture before reduction. On the right side (b), it is the case where the pixel position of the reduced picture is at the center position, which exists at the center position of the 2x2 pixels of the display picture before reduction.

[0134] In this method, phase display information is defined based on the horizontal and vertical distances of the pixel positions of the reduced picture from the upper left position of the display picture before reduction. Assuming that the pixel interval before reduction is 1 and is expressed as a rational number, let the numerator of the phase display information in the horizontal direction be hor_phase_num, the denominator be hor_phase_den, the numerator of the phase display information in the vertical direction be ver_phase_num, and the denominator be hor_phase_den.

[0135] In the case of the left side (a) of FIG. 14, the values of hor_phase_num, hor_phase_den, ver_phase_num, and ver_phase_den are all set to 0. In the case of the right side (b), the value of hor_phase_num is 1, the value of hor_phase_den is 2, the value of ver_phase_num is 1, and ver_phase_den is 2.

[0136] In this way, by defining the phase display information, it is possible to solve the problem that the value does not become a constant value when matching the pixel position at the upper left position of the desired picture in terms of coding efficiency.

[0137] (Configuration of Image Encoding Device) Next, the configuration of the image encoding device 11 according to the present embodiment will be described. FIG. 7 is a block diagram showing the configuration of the image encoding device 11 according to the present embodiment. The image encoding device 11 includes a prediction image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, a prediction parameter derivation unit 120, and an entropy coding unit 104.

[0138] The prediction image generation unit 101 generates a prediction image for each CU.

[0139] The subtraction unit 102 subtracts the pixel value of the predicted image of the block input from the predicted image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform and quantization unit 103.

[0140] The transform and quantization unit 103 calculates transform coefficients for the prediction error input from the subtraction unit 102 by frequency conversion, and derives quantized transform coefficients by quantization. The transform and quantization unit 103 outputs the quantized transform coefficients to the parameter encoding unit 111 and the inverse quantization and inverse transform unit 105.

[0141] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 (Fig. 5) in the image decoding device 31, and the description thereof is omitted. The calculated prediction error is output to the addition unit 106.

[0142] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU encoding unit 1112 (prediction mode encoding unit). The CU encoding unit 1112 further includes a TU encoding unit 1114. The schematic operations of each module will be described below.

[0143] The header encoding unit 1110 performs encoding processing on parameters such as header information, segmentation information, prediction information, and quantized transform coefficients.

[0144] The CT information encoding unit 1111 encodes QT, MT (BT, TT) segmentation information, etc.

[0145] The CU encoding unit 1112 encodes CU information, prediction information, segmentation information, etc.

[0146] When the TU contains a prediction error, the TU encoding unit 1114 encodes QP update information and the quantized prediction error.

[0147] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter-prediction parameters and quantized transform coefficients to the parameter encoding unit 111.

[0148] The entropy encoding unit 104 receives the quantized transform coefficients and encoding parameters from the parameter encoding unit 111. The entropy encoding unit 104 entropy-encodes these to generate and output encoded data Te.

[0149] The prediction parameter derivation unit 120 derives an inter-prediction parameter and an intra-prediction parameter from the parameters input from the encoding parameter determination unit 110. The derived inter-prediction parameter and intra-prediction parameter are output to the parameter encoding unit 111.

[0150] The addition unit 106 adds the pixel values of the prediction block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transformation unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0151] The loop filter 107 performs a deblocking filter, SAO, and ALF on the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters, and may be configured with only the deblocking filter, for example.

[0152] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.

[0153] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at predetermined positions for each target picture and CU.

[0154] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters are the above-mentioned QT, BT, or TT segmentation information, prediction parameters, or parameters to be encoded generated in relation to these. The prediction image generation unit 101 generates a prediction image using these encoding parameters.

[0155] Incidentally, some parts of the image encoding apparatus 11 and the image decoding apparatus 31 in the above-described embodiments, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transformation unit 311, the addition unit 312, the predicted parameter derivation unit 320, the predicted image generation unit 101, the subtraction unit 102, the transformation / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transformation unit 105, the loop filter 107, the encoding parameter determination unit 110, the parameter encoding unit 111, and the predicted parameter derivation unit 120 may be realized by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" refers to a computer system built in either the image encoding apparatus 11 or the image decoding apparatus 31 and including hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or the like, or a storage device such as a hard disk built in a computer system. Furthermore, the "computer-readable recording medium" also includes something that holds a program dynamically for a short time, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and something that holds a program for a certain time, like a volatile memory inside a computer system serving as a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be realized in combination with a program already recorded in the computer system for realizing the aforementioned functions.

[0156] Further, part or all of the image encoding device 11 and the image decoding device 31 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be individually processorized, or part or all of them may be integrated and processorized. Also, the method of integrating into a circuit is not limited to LSI, and it may be realized by a dedicated circuit or a general-purpose processor. Further, when a technology for integrating into a circuit that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit using such technology may be used.

[0157] As described above, one embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0158] The embodiments of the present invention are not limited to the above-described embodiments, and various changes are possible within the scope shown in the claims. That is, embodiments obtained by combining technical means appropriately changed within the scope shown in the claims are also included in the technical scope of the present invention.

Industrial Applicability

[0159] The embodiments of the present invention can be suitably applied to a moving image decoding device that decodes encoded data in which image data is encoded, and a moving image encoding device that generates encoded data in which image data is encoded. Also, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding device and referred to by the moving image decoding device.

Explanation of Signs

[0160] 1 Moving image transmission system 30 Moving image decoding device 31 Image decoding device 301 Entropy decoding unit 302 Parameter decoding unit 305, 107 Loop filter​ 306 and 109 Reference Picture Memory 307 and 108 Prediction Parameter Memory 308 and 101 Prediction Image Generation Unit 311 and 105 Inverse Quantization and Inverse Transformation Unit 312 and 106 Addition Unit 320 Prediction Parameter Derivation Unit 10 Moving Picture Encoding Device 11 Image Encoding Device 102 Subtraction Unit 103 Transformation and Quantization Unit 104 Entropy Encoding Unit 110 Encoding Parameter Determination Unit 111 Parameter Encoding Unit 120 Prediction Parameter Derivation Unit 41 Image Display Device 51 Image Reduction Processing Device 61 Image Resolution Conversion Processing Device 71 Auxiliary Extension Information Creation Device 81 Auxiliary Extension Information Encoding Device 91 Auxiliary Extension Information Decoding Device

Claims

1. An image decoding device that decodes encoded data with an adaptively changed resolution to generate a decoded image, and an image resolution conversion processing device that changes the resolution of the decoded image, comprising an auxiliary extension information decoding device that decodes phase display information for operating the image resolution conversion processing device, A moving image decoding device characterized by having an estimated value of phase display information when the previous auxiliary extension information does not exist.

2. An image reduction processing device that adaptively enlarges or reduces an image, An image encoding device that generates encoded data of an image with an adaptively changed resolution, An auxiliary extension information generation device that creates phase display information for operating the image reduction processing device, comprising an auxiliary extension information encoding device that encodes the phase display information, A moving image encoding device characterized by having an estimated value of phase display information when the previous auxiliary extension information does not exist.