Receiving method and receiving device

By hierarchically encoding images into layers with controlled B-picture usage based on frame rate, the method addresses inefficiencies in conventional image coding, enhancing scalability and compression efficiency while managing delays effectively.

JP2026053738APending Publication Date: 2026-03-25PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional image coding and decoding methods are inefficient and may result in increased display latency and transmission delay due to limitations in hierarchical encoding structures, particularly when dealing with varying frame rates.

Method used

The method involves hierarchically encoding images into multiple layers, with the lowest layer containing I-pictures and P-pictures, and higher layers containing B-pictures, where the number of consecutive B-pictures is limited based on the frame rate, and the number of layers is determined to maintain efficient encoding and decoding without excessive display or transmission delay.

Benefits of technology

This approach allows for efficient encoding and decoding of images by increasing the number of layers while keeping display and transmission delays within acceptable limits, improving scalability and compression performance across different frame rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053738000001_ABST
    Figure 2026053738000001_ABST
Patent Text Reader

Abstract

Encode images efficiently. [Solution] The receiving method includes a receiving step of receiving a bitstream containing a video and control information encoded by hierarchically encoding multiple images contained in the video into a hierarchical structure with one or more layers, and a decoding step (S401) of decoding multiple images from the bitstream, wherein the lowest layer of the hierarchical structure includes I-pictures and P-pictures, and layers other than the lowest layer of the hierarchical structure include B-pictures, the control information includes information on the frame rate of the video, the number of layers is predetermined based on the frame rate of the video, and the number of consecutive B-pictures, which is the number of consecutive B-pictures in the display order, is less than or equal to a predetermined maximum number of consecutive B-pictures.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image encoding method for encoding an image, or an image decoding method for decoding an image. [Background technology]

[0002] Non-Patent Document 1 describes an image encoding method for encoding images (including moving images) and an image decoding method for decoding images. Furthermore, Non-Patent Document 2 describes operational regulations concerning encoding and decoding. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 12th Meeting: Geneva, CH, 14-23 Jan. 2013 JCTVC-L1003_v34.doc, High Efficiency Video Coding (HEVC) text specification draft 10 (for FDIS & Last Call) http: / / phenix.it-sudparis.eu / jct / doc_end_user / documents / 12_Geneva / wg11 / JCTVC-L1003-v34.zip [Non-Patent Document 2] The Association of Radio Industries and Businesses (ARIB) standard, ARIB STD-B32 version 2.8, 2-STD-B32v2_8.pdf, Video coding, audio coding, and multiplexing methods in digital broadcasting. http: / / www.arib.or.jp / english / html / overview / doc / 2-STD-B32v2_8.pdf [Overview of the project] [Problems that the invention aims to solve]

[0004] However, inefficient processing may be used in conventional image coding or decoding methods.

[0005] Therefore, the present invention aims to provide an image encoding method for efficiently encoding an image, or an image decoding method for efficiently decoding an image. [Means for solving the problem]

[0006] To achieve the above objective, a transmission method according to one aspect of the present invention includes a generation step of generating a bitstream containing a video and control information, which are encoded by hierarchically encoding a plurality of images contained in the video into a hierarchical structure having one or more layers, and a transmission step of transmitting the bitstream and the control information, wherein the lowest layer of the hierarchical structure includes I-pictures and P-pictures, the layers of the hierarchical structure other than the lowest layer include B-pictures, the control information includes information on the frame rate of the video, the number of layers is predetermined based on the frame rate of the video, and the number of consecutive B-pictures, which is the number of consecutive B-pictures in the display order, is less than or equal to a predetermined maximum number of consecutive B-pictures.

[0007] Furthermore, a receiving method according to one aspect of the present invention includes a receiving step of receiving a bitstream containing a moving image and control information, which are encoded by hierarchically encoding a plurality of images contained in the moving image into a hierarchical structure having one or more layers, and a decoding step of decoding the plurality of images from the bitstream, wherein the lowest layer of the hierarchical structure includes I-pictures and P-pictures, the layers of the hierarchical structure other than the lowest layer include B-pictures, the control information includes information on the frame rate of the moving image, the number of layers is predetermined based on the frame rate of the moving image, and the number of consecutive B-pictures, which is the number of consecutive B-pictures in the display order, is less than or equal to a predetermined maximum number of consecutive B-pictures.

[0008] Note that these general or specific aspects may be implemented in a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, method, integrated circuit, computer program, and recording medium.

Advantages of the Invention

[0009] The present invention can provide an image encoding method capable of efficiently encoding an image or an image decoding method capable of efficiently decoding an image.

Brief Description of the Drawings

[0010] [Figure 1] FIG. 1 is a diagram showing an example of an encoding structure. [Figure 2] FIG. 2 is a diagram showing the number of display delay pictures. [Figure 3] FIG. 3 is a block diagram of an image encoding apparatus according to Embodiment 1. [Figure 4] FIG. 4 is a flowchart of an image encoding process according to Embodiment 1. [Figure 5] FIG. 5 is a block diagram of a limit value setting unit according to Embodiment 1. [Figure 6] FIG. 6 is a flowchart of a limit value setting process according to Embodiment 1. [Figure 7] FIG. 7 is a block diagram of an encoding unit according to Embodiment 1. [Figure 8] FIG. 8 is a flowchart of an encoding process according to Embodiment 1. [Figure 9A] FIG. 9A is a diagram showing the number of transmission delay pictures according to Embodiment 1. [Figure 9B] FIG. 9B is a diagram showing the number of transmission delay pictures according to Embodiment 1. [Figure 9C] FIG. 9C is a diagram showing the number of transmission delay pictures according to Embodiment 1. [Figure 9D] FIG. 9D is a diagram showing the number of transmission delay pictures according to Embodiment 1. [Figure 10]FIG. 10 is a diagram showing an example of an encoding structure limit value according to Embodiment 1. [Figure 11A] FIG. 11A is a diagram showing an encoding structure according to Embodiment 1. [Figure 11B] FIG. 11B is a diagram showing an encoding structure according to Embodiment 1. [Figure 11C] FIG. 11C is a diagram showing an encoding structure according to Embodiment 1. [Figure 11D] FIG. 11D is a diagram showing an encoding structure according to Embodiment 1. [Figure 12A] ' FIG. 12A is a diagram showing the number of display delay pictures according to Embodiment 1. [Figure 12B] FIG. 12B is a diagram showing the number of display delay pictures according to Embodiment 1. [Figure 12C] FIG. 12C is a diagram showing the number of display delay pictures according to Embodiment 1. [Figure 12D] FIG. 12D is a diagram showing the number of display delay pictures according to Embodiment 1. [Figure 13] FIG. 13 is a block diagram of an image decoding apparatus according to Embodiment 2. [Figure 14] FIG. 14 is a flowchart of an image decoding process according to Embodiment 2. [Figure 15] FIG. 15 is a flowchart of an image encoding method according to Embodiment 1. [Figure 16] FIG. 16 is a flowchart of an image decoding method according to Embodiment 2. [Figure 17] FIG. 17 is an overall configuration diagram of a content supply system that realizes a content delivery service. [Figure 18] FIG. 18 is an overall configuration diagram of a digital broadcast system. [Figure 19] FIG. 19 is a block diagram showing an example of the configuration of a television. [Figure 20] FIG. 20 is a block diagram showing an example of the configuration of an information reproduction / recording unit that reads and writes information on a recording medium that is an optical disk. [Figure 21]Figure 21 shows an example of the structure of a recording medium, which is an optical disc. [Figure 22A] Figure 22A shows an example of a mobile phone. [Figure 22B] Figure 22B is a block diagram showing an example of a mobile phone configuration. [Figure 23] Figure 23 is a diagram showing the structure of the multiplexed data. [Figure 24] Figure 24 schematically shows how each stream is multiplexed in the multiplexed data. [Figure 25] Figure 25 shows in more detail how the video stream is stored in the PES packet sequence. [Figure 26] Figure 26 shows the structure of TS packets and source packets in multiplexed data. [Figure 27] Figure 27 shows the data structure of PMT. [Figure 28] Figure 28 shows the internal structure of the multiplexed data information. [Figure 29] Figure 29 shows the internal structure of stream attribute information. [Figure 30] Figure 30 shows the steps for identifying video data. [Figure 31] Figure 31 is a block diagram showing an example of the configuration of an integrated circuit that implements the video encoding method and video decoding method of each embodiment. [Figure 32] Figure 32 shows a configuration for switching the drive frequency. [Figure 33] Figure 33 shows the steps for identifying video data and switching the drive frequency. [Figure 34] Figure 34 shows an example of a lookup table that associates video data specifications with drive frequencies. [Figure 35A] Figure 35A shows an example of a configuration in which the signal processing module is shared. [Figure 35B] Figure 35B shows another example of a configuration in which the signal processing module is shared. [Modes for carrying out the invention]

[0011] (Knowledge that formed the basis of this invention) The inventors have found that the following problems arise with respect to the image encoding device for encoding images or the image decoding device for decoding images, as described in the "Background Art" section.

[0012] In recent years, advancements in digital video equipment technology have been remarkable, and there is an increasing trend for video signals (multiple pictures arranged in chronological order) output from video cameras or television tuners to be compressed and encoded, and for the resulting encoded signals to be recorded on recording media such as DVDs or hard disks.

[0013] One image encoding standard is H.264 / AVC (MPEG-4 AVC). Furthermore, the HEVC (High Efficiency Video Coding) standard (Non-Patent Document 1) is being considered as a next-generation standard. Regulations on how to implement image encoding standards are also being considered (Non-Patent Document 2).

[0014] Under the current operating regulations (Non-Patent Literature 2), the encoding structure is limited to three layers, as shown in Figure 1, which limits the maximum number of display delay pictures to two, as shown in Figure 2. The TemporalId shown in Figure 1 is an identifier for the layer of the encoding structure. A larger TemporalId indicates a deeper layer.

[0015] One square block represents a picture, and within the block is I x I is an I picture (predictive picture on screen), P x P-picture (forward-referencing predictive picture), B x This indicates a B-picture (two-way reference prediction picture). x / P x / B x of x This indicates the display order, showing the sequence in which the pictures will be displayed.

[0016] The arrows between pictures indicate reference relationships. For example, picture B1 generates a predicted image using pictures I0, B2, and P4 as reference images. It is also prohibited to use a picture with a TemporalId greater than its own as a reference image. Therefore, the picture decoding order is in ascending order of TemporalId, as shown in Figure 2: picture I0, picture P4, picture B2, picture B1, and picture B3.

[0017] By defining a hierarchy, it is possible to give the code sequence temporal scalability.

[0018] For example, if you want to obtain a 30fps video from a 60fps (frames per second) code sequence, the image decoder will decode only the pictures labeled TemporalId0 and TemporalId1 in Figure 1. This allows the image decoder to obtain a 30fps image. Since the decoded images must be output in order without any gaps, the image decoder outputs the pictures sequentially starting from picture I0 after decoding picture B2. Therefore, the number of display delay pictures is 2. Converting this to time, the display delay time is 2 / 30 seconds when the original frame rate is 30fps, and 2 / 60 seconds when the frame rate is 60fps.

[0019] By using a structure with high temporal scalability, when bandwidth is congested or when an image decoding device with low processing power is performing the decoding process, the image decoding device can decode only the pictures in the lower TemporalId hierarchy and display the resulting image. In this way, versatility is improved. However, allowing a deep hierarchical structure presents the challenge of increased display latency.

[0020] However, even if the number of display delay pictures is predetermined as described above, the display delay time will differ depending on the frame rate. At a frame rate lower than the standard frame rate (e.g., 24fps), the display delay time is 2 / 24 seconds, which is longer than the 2 / 30 seconds at 30fps.

[0021] An image encoding method according to one aspect of the present invention is an image encoding method for hierarchical encoding of an image, comprising: a step of determining the number of layers such that the number of layers in the hierarchical encoding is less than or equal to a maximum number of layers determined according to the frame rate; and an encoding step of generating a bitstream by hierarchically encoding the image with the determined number of layers.

[0022] According to this, the image encoding method can increase the number of layers while suppressing an increase in display delay time. Therefore, the image encoding method can encode images efficiently.

[0023] For example, if the frame rate is 60fps or less, the maximum number of layers may be 4 or less.

[0024] For example, if the frame rate is 120fps, the maximum number of layers may be 5.

[0025] For example, the image encoding method may further include a picture type determination step in which an image decoding device determines the picture type of the image such that the number of display delay pictures, which is the number of pictures from decoding to outputting the image, is less than or equal to the maximum number of pictures determined according to the frame rate, and the encoding step may encode the image with the determined picture type.

[0026] For example, in the picture type determination step, the picture type of the image may be determined such that the number of consecutive B pictures, which is the number of consecutive B pictures, is less than or equal to the maximum number of consecutive pictures determined according to the frame rate.

[0027] For example, the maximum number of pictures, the encoder transmission delay which is the time from when the image is input to the image encoding device until the bitstream is output, and the frame rate satisfy the following relationship: Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps])) The maximum number of consecutive pictures, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive pictures = int(encoder transmission delay [s] × frame rate [fps] - 1) The maximum number of hierarchical levels, the encoder transmission delay, and the frame rate may satisfy the following relationship. Maximum number of hierarchical levels = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1

[0028] For example, the maximum number of pictures [i] in each hierarchical level, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2 (n-i) )) The maximum number of consecutive pictures [i] in each hierarchical level, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive pictures [i] = int(encoder transmission delay [s] × frame rate [fps] / 2 (n-i) - 1) i is an integer less than or equal to the maximum number of hierarchical levels, indicating the hierarchical level, and n may indicate (the maximum number of hierarchical levels - 1).

[0029] Furthermore, an image decoding method according to one aspect of the present invention is an image decoding method for decoding a bitstream obtained by hierarchical coding of an image, comprising: an image decoding step of decoding the image from the bitstream; an information decoding step of decoding first information indicating the number of layers in the hierarchical coding from the bitstream; and a sorting step of sorting and outputting the decoded image using the number of layers indicated by the first information, wherein the number of layers is less than or equal to a maximum number of layers predetermined according to the frame rate of the bitstream.

[0030] According to this, the image decoding method can decode a bitstream obtained through efficient encoding.

[0031] For example, if the frame rate is 60fps or less, the maximum number of layers may be 4 or less.

[0032] For example, if the frame rate is 120fps, the maximum number of layers may be 5.

[0033] For example, in the information decoding step, the image decoding device may further decode from the bitstream a second piece of information indicating the number of display delay pictures, which is the number of pictures from the time the image is decoded until it is output. In the sorting step, the decoded image may be sorted and output using the number of layers indicated by the first piece of information and the number of display delay pictures indicated by the second piece of information.

[0034] For example, in the information decoding step, a third piece of information indicating the number of consecutive B-pictures, which is the number of consecutive B-pictures, may be decoded from the bitstream, and in the sorting step, the decoded image may be sorted and output using the number of layers indicated in the first piece of information, the number of display delay pictures indicated in the second piece of information, and the number of consecutive B-pictures indicated in the third piece of information.

[0035] For example, the following relationship is satisfied between the maximum number of pictures, the encoder transmission delay which is the time from when the image is input to the image encoding device until the bitstream is output, and the frame rate. Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps])) The following relationship is satisfied between the maximum number of consecutive frames, the encoder transmission delay, and the frame rate: Maximum number of consecutive frames = int(encoder transmission delay [s] × frame rate [fps] - 1) The following relationship may be satisfied between the maximum number of layers, the encoder transmission delay, and the frame rate. Maximum number of layers = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1

[0036] For example, the maximum number of pictures in each layer [i], the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2) (n-i) )) The maximum number of consecutive steps in each layer [i], the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive frames [i] = int(encoder transmission delay [s] × frame rate [fps] / 2) (n-i) -1) i is an integer less than or equal to the maximum number of levels, and represents a level, while n may represent (the maximum number of levels - 1).

[0037] Furthermore, an image encoding device according to one aspect of the present invention is an image encoding device for encoding an image, comprising a processing circuit and a storage device accessible from the processing circuit, wherein the processing circuit uses the storage device to execute the image encoding method.

[0038] According to this, the image encoding device can increase the number of layers while suppressing an increase in display delay time. Therefore, the image encoding device can encode images efficiently.

[0039] Furthermore, an image decoding device according to one aspect of the present invention is an image decoding device that decodes a bitstream obtained by encoding an image, and comprises a processing circuit and a storage device accessible from the processing circuit, wherein the processing circuit executes the image decoding method using the storage device.

[0040] According to this, the image decoding device can decode a bitstream obtained through efficient encoding.

[0041] These general or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0042] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of the present invention. The numerical values, shapes, materials, components, arrangement positions and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.

[0043] (Embodiment 1) The image encoding device according to this embodiment increases the number of layers when the frame rate is high. This makes it possible to increase the number of layers while suppressing an increase in display delay time.

[0044] <Overall Structure> Figure 3 is a block diagram showing the configuration of the image encoding device 100 in this embodiment.

[0045] The image encoding device 100 shown in Figure 3 generates a code sequence 155 (bitstream) by encoding the input image 153. This image encoding device 100 includes a limit value setting unit 101 and an encoding unit 102.

[0046] <Overall Operation> Next, the overall flow of the encoding process will be explained with reference to Figure 4. Figure 4 is a flowchart of the image encoding method according to this embodiment.

[0047] First, the limit value setting unit 101 sets the coding structure limit value 154 related to the coding structure in hierarchical coding (S101). Specifically, the limit value setting unit 101 sets the coding structure limit value 154 using the frame rate 151 and the transmission delay time limit value 152.

[0048] Next, the encoding unit 102 encodes the encoding structure restriction value 154 and uses the encoding structure restriction value 154 to encode the input image 153, thereby generating a code sequence 155 (S102).

[0049] <Configuration of the limit value setting unit 101> Figure 5 is a block diagram showing an example of the internal configuration of the limit value setting unit 101.

[0050] As shown in Figure 5, the limit value setting unit 101 includes a hierarchical number setting unit 111, a hierarchical number parameter setting unit 112, a display delay picture number setting unit 113, a B picture consecutive number setting unit 114, and a consecutive number parameter setting unit 115.

[0051] <Operation (Setting encoding structure limit values)> Next, an example of the limit value setting process (S101 in Figure 4) will be described with reference to Figure 6. Figure 6 is a flowchart of the limit value setting process according to this embodiment.

[0052] First, the layer number setting unit 111 sets the layer number 161 of the encoding structure using the frame rate 151 and the transmission delay time limit value 152 input from outside the image encoding device 100. For example, the layer number 161 is calculated by the following (Equation 1) (S111).

[0053] Number of layers = int(log2(transmission delay time limit [s] × frame rate [fps])) + 1 ... (Equation 1)

[0054] In the above (Equation 1), int(x) means a function that returns an integer with the fractional part of x truncated, and log2(x) means a function that returns the base 2 logarithm of x. The transmission delay time limit value 152 indicates the maximum time from when the input image 153 is input to the image encoding device 100 until the code sequence 155 of the input image 153 is output.

[0055] Next, the layer number parameter setting unit 112 sets the layer number parameter 163, sps_max_sub_layers_minus1, using the layer number 161 according to the following equation (Equation 2) (S112).

[0056] sps_max_sub_layers_minus1=Number of layers -1 (Formula 2)

[0057] Next, the limit value setting unit 101 sets TId to 0 (S113). TId is a variable used to identify the hierarchy and is used to identify the hierarchy to be processed in subsequent processing for each hierarchy.

[0058] Next, the display delay picture count setting unit 113 sets the display delay picture count 164 for the tier where TemporalId is TId using the frame rate 151, the transmission delay time limit value 152, and the tier number parameter 163 (S114). The display delay picture count 164 is the number of pictures from when picture decoding starts until the display of the picture starts during picture decoding. The display delay picture count 164 is calculated by the following (Equation 3).

[0059] Number of display delay pictures in the hierarchy where TemporalId is TId = int(log2(transmission delay time limit [s] × frame rate [fps] ÷ 2) (n-TId) )) ...(Formula 3)

[0060] In the above (Equation 3), n represents the maximum TemporalId and is the value of sps_max_sub_layers_minus1 calculated in step S112. The display delay picture count setting unit 113 sets the calculated number of display delay pictures in the layer where TemporalId is TId to sps_max_num_reorder_pics[TId].

[0061] Next, the B-picture consecutive number setting unit 114 sets the number of consecutive B-pictures 162 for the tier where TemporalId is TId, using the frame rate 151, the transmission delay time limit value 152, and the tier number parameter 163 (S115). The number of consecutive B-pictures 162 is the number of consecutive B-pictures and is calculated by the following (Equation 4).

[0062] Number of consecutive B-pictures in the hierarchy where TemporalId is TId = int(transmission delay time limit [s] × frame rate [fps] ÷ 2) (n-TId) -1)...(Formula 4)

[0063] Next, the continuity parameter setting unit 115 sets the continuity parameter 165 for the hierarchy where TemporalId is TId, using the B picture continuity 162 and the display delay picture count 164 (sps_max_num_reorder_pics[TId]) for the hierarchy where TemporalId is TId (S116). The continuity parameter 165 for the hierarchy where TemporalId is TId is set by the following (Equation 5).

[0064] The parameter for the number of consecutive TemporalIds in the hierarchy where TId = number of consecutive B-pictures in the hierarchy where TemporalId is TId - sps_max_num_reorder_pics[TId] + 1 ... (Equation 5)

[0065] The calculated continuity parameter 165 is set to sps_max_latency_increase_plus1[TId].

[0066] Next, the limit value setting unit 101 moves to the next processing level by adding 1 to TId (S117). Steps S114 to S117 are repeated until TId reaches 161 levels, that is, until processing of all levels is completed (S118).

[0067] In this example, the limit value setting unit 101 first sets the number of layers 161, and then sets the number of display delay pictures 164 and the number of consecutive B pictures 162 in each layer, but the setting order is not limited to this.

[0068] <Configuration of encoding unit 102> Figure 7 is a block diagram showing the internal configuration of the encoding unit 102. As shown in Figure 7, the encoding unit 102 includes an image rearrangement unit 121, a code block division unit 122, a subtraction unit 123, a transformation quantization unit 124, a variable-length encoding unit 125, an inverse transformation quantization unit 126, an addition unit 127, a frame memory 128, an intra prediction unit 129, an inter prediction unit 130, and a selection unit 131.

[0069] <Operation (encoding)> Next, the encoding process according to this embodiment (S102 in Figure 4) will be described with reference to Figure 8. Figure 8 is a flowchart of the encoding process according to this embodiment.

[0070] First, the variable-length coding unit 125 performs variable-length coding on sps_max_sub_layers_minus1, sps_max_num_reorder_pics[], and sps_max_latency_increase_plus1[], which are set by the limit value setting unit 101 (S121). Although sps_max_num_reorder_pics[] and sps_max_latency_increase_plus1[] exist for each layer, the variable-length coding unit 125 codes all of them.

[0071] Next, the image rearrangement unit 121 rearranges the input images 153 according to sps_max_sub_layers_minus1, sps_max_num_reorder_pics[], and sps_max_latency_increase_plus1[], and also determines the picture type of the input images 153 (S122).

[0072] The image rearrangement unit 121 performs this rearrangement using sps_max_num_reorder_pics[sps_max_sub_layers_minus1] and SpsMaxLatencyPictures. SpsMaxLatencyPictures is calculated by the following (Equation 6).

[0073] SpsMaxLatencyPictures=sps_max_num_reorder_pics[sps_max_sub_layers_minus1]+sps_max_latency_increase_plus1[sps_max_sub_layers_minus1]-1 (Formula 6)

[0074] Figures 9A to 9D illustrate this rearrangement. Because the rearrangement shown in Figures 9A to 9D occurs, the encoding unit 102 cannot begin encoding the input images 153 until multiple input images 153 have been input. In other words, a delay occurs between the input of the first input image 153 and the start of outputting the code sequence 155. This delay is the transmission delay time, and the transmission delay time limit value 152 mentioned above is the limit value of this transmission delay time.

[0075] Furthermore, Figures 9A to 9D show the number of transmission delay pictures corresponding to the encoding structure limit value of 154. Figure 9A shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 1 and SpsMaxLatencyPictures is 2. Figure 9B shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 2 and SpsMaxLatencyPictures is 3. Figure 9C shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 3 and SpsMaxLatencyPictures is 7. Figure 9B shows the number of delayed-transmission pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 4 and SpsMaxLatencyPictures is 7.

[0076] For example, in Figure 9A, images 0, 1, 2, and 3 are input to the image encoding device 100 in this order, and the image encoding device 100 encodes these images in the order of image 0, image 3, image 1, and image 2. Because the image encoding device 100 needs to send out the code sequence without any gaps, it does not start sending out the code sequence until image 3 is input. Therefore, a transmission delay of 3 pictures occurs between the input of image 0 and the start of code sequence transmission. The image rearrangement unit 121 also determines the picture type of each image and outputs information to the interpretation unit 130 indicating which image each image will use as a reference picture. Here, the picture types are I-picture, P-picture, and B-picture.

[0077] Next, the code block division unit 122 divides the input image 153 into code blocks 171 (S123).

[0078] Next, the intra-prediction unit 129 generates a prediction block for intra-prediction and calculates the cost of that prediction block (S124). The inter-prediction unit 130 generates a prediction block for inter-prediction and calculates the cost of that prediction block (S125). The selection unit 131 uses the calculated cost and other factors to determine the prediction mode and prediction block 177 to be used (S126).

[0079] Next, the subtraction unit 123 generates a difference block 172 by calculating the difference between the prediction block 177 and the code block 171 (S127). Next, the conversion quantization unit 124 generates conversion coefficients 173 by performing frequency conversion and quantization on the difference block 172 (S128). Next, the inverse conversion quantization unit 126 restores the difference block 174 by performing inverse quantization and inverse frequency conversion on the conversion coefficients 173 (S129). Next, the addition unit 127 generates a decoded block 175 by adding the prediction block 177 and the difference block 174 (S130). This decoded block 175 is stored in the frame memory 128 and used for prediction processing by the intra prediction unit 129 and the inter prediction unit 130.

[0080] Next, the variable-length coding unit 125 encodes prediction information 178 indicating the prediction mode used, etc. (S131), and encodes the conversion coefficient 173 (S132).

[0081] Then, processing moves to the next code block (S133), and the encoding unit 102 repeats steps S124 to S133 until processing of all code blocks in the picture is completed (S134).

[0082] Then, the encoding unit 102 repeats steps S122 to S134 until processing of all pictures is complete (S135).

[0083] <Effects> As described above, the image encoding device 100 according to this embodiment determines the encoding structure based on the frame rate 151 and the transmission delay time limit value 152. This allows the image encoding device 100 to deepen the layering without increasing the display delay time of the decoder or the transmission delay time of the encoder when the frame rate 151 is high, thereby improving time scalability. Furthermore, increasing the number of B-pictures improves compression performance.

[0084] Furthermore, even at various frame rates, the display delay time of the specified decoder and the transmission delay time of the encoder can be kept within limits.

[0085] This will be explained in more detail. Figure 10 shows the number of layers (161), the number of display delay pictures (164), and the number of consecutive B pictures (162) calculated using the frame rate (151) and the transmission delay time limit (152). Figure 10 also shows an example where the transmission delay time limit is 4 / 30 seconds.

[0086] Figures 11A to 11D show the encoded structures based on the conditions in Figure 10. Figure 11A shows the structure when the frame rate is 24 fps. Figure 11B shows the structure when the frame rate is 30 fps. Figure 11C shows the structure when the frame rate is 60 fps. Figure 11D shows the structure when the frame rate is 120 fps.

[0087] Furthermore, the number of delayed picture frames for sending the code sequence is shown in Figures 9A to 9D for frame rates of 24fps, 30fps, 60fps, and 120fps, respectively. Figures 12A to 12D show the number of delayed picture frames for frame rates of 24fps, 30fps, 60fps, and 120fps.

[0088] As shown in Figure 10, the transmission delay time does not exceed the limit of 4 / 30 seconds at all frame rates. Furthermore, under the current operational regulations (Non-Patent Literature 2), the encoding structure is limited to three layers, as shown in Figure 11B. In other words, the number of display delay pictures is limited to two, as shown in Figure 12B, and the number of transmission delay pictures in the code sequence is limited to four, as shown in Figure 9B. Also, at 30fps, the display delay time is 2 / 30 seconds and the transmission delay time is 4 / 30 seconds. In this embodiment, when the transmission delay time limit is set to 4 / 30 seconds, the transmission delay time does not exceed 4 / 30 seconds and the display delay time does not exceed 2 / 30 seconds, even if the number of layers is increased or decreased according to the frame rate.

[0089] Furthermore, in this embodiment, the image encoding device 100 determines the encoding structure using a limit value for the transmission delay time, rather than the display delay time. By limiting the encoding structure by the transmission delay time in this way, the encoding structure can be determined so as not to exceed both the display delay time of 2 / 30 seconds and the transmission delay time of 4 / 30 seconds in the current operating regulations (Non-Patent Literature 2). More specifically, if the number of display delay pictures is determined so as not to exceed the display delay time of 2 / 30 seconds in the current operating regulations (Non-Patent Literature 2), the number of display delay pictures at 120fps is 8 (8 / 120 seconds), and an encoding structure of up to 9 layers is permitted. However, if a 9-layer encoding structure is used, the number of transmission delay pictures becomes 256 (256 / 120 seconds), which greatly exceeds the transmission delay time of 4 / 30 seconds in the current operating regulations (Non-Patent Literature 2). On the other hand, focusing on the transmission delay time, if the number of transmission delay pictures is determined so as not to exceed a transmission delay time of 4 / 30 seconds, then at 120fps, the number of transmission delay pictures is limited to 16 (16 / 120 seconds), and the encoding structure is limited to 5 layers. In this case, the display delay time will not exceed 2 / 30 seconds. In this way, by limiting the transmission delay time, both the transmission delay time and the display delay time can be appropriately limited.

[0090] Furthermore, the image encoding device 100 sets a limit value for the encoding structure at each layer. This ensures that even in an image decoding device that only decodes pictures at layers with small TemporalId values, the display delay time does not exceed the specified time.

[0091] In the above description, the image encoding device 100 calculates encoding structure limit values ​​such as the number of layers using a formula. However, the table shown in Figure 10 may be stored in memory in advance, and the encoding structure limit values ​​corresponding to the frame rate 151 and the transmission delay time limit value 152 may be set by referring to this table. Alternatively, the image encoding device 100 may use both the table and the formula. For example, the image encoding device 100 may set the encoding structure limit values ​​using the table when the frame rate is 24 fps or less, and set the encoding structure limit values ​​using the formula when the frame rate exceeds 24 fps.

[0092] Furthermore, although the above description assumes that the image encoding device 100 uses an externally input transmission delay time limit value 152 and frame rate 151, it is not limited to this. For example, the image encoding device 100 may use a predetermined fixed value as at least one of the transmission delay time limit value 152 and frame rate 151. Alternatively, the image encoding device 100 may determine at least one of the transmission delay time limit value 152 and frame rate 151 according to the internal state of the buffer memory or the like.

[0093] Furthermore, the encoding structures shown in Figures 11A to 11D are examples only and are not limited to them. For example, the arrows indicating reference images are not limited to these; each picture simply needs to avoid using a picture with a TemporalId greater than its own as a reference image. For example, image B1 shown in Figure 11B may use image P4 as a reference image.

[0094] Furthermore, the above-mentioned limitations on the encoding structure (number of layers, number of consecutive B-pictures, and number of display delay pictures) are merely maximum values, and smaller values ​​may be used depending on the situation. For example, in Figure 10, when the frame rate is 30fps, the number of layers is 3, the number of consecutive B-pictures is 3[2], and the number of display delay pictures is 2[2], and the encoding structure shown in Figure 11B is displayed. However, the number of layers only needs to be 3 or less, and the number of consecutive B-pictures and the number of display delay pictures only need to be values ​​corresponding to a number of layers of 3 or less. For example, the number of layers may be 2, the number of consecutive B-pictures may be 2[1], and the number of display delay pictures may be 2[2]. In this case, for example, the encoding structure shown in Figure 11A is used. In that case, in the encoding structure encoding of step S121 shown in Figure 8, information indicating the encoding structure used is encoded.

[0095] Furthermore, in the above explanation, sps_max_num_reorder_pics is set to the number of display-delayed pictures, but sps_max_num_reorder_pics may also be a variable representing the number of pictures whose order is changed. For example, in the example shown in Figure 9C, input images 8, 4, and 2 are rearranged and encoded so that they appear before their input order (display order) positions. In this case, the number of pictures whose order is changed is 3, and this value of 3 may be set to sps_max_num_reorder_pics.

[0096] Furthermore, in the above explanation, the number of consecutive images parameter is set to sps_max_latency_increase_plus1, and the value sps_max_num_reorder_pics + sps_max_latency_increase_plus1 - 1 (SpsMaxLatencyPictures) is treated as the number of consecutive B-pictures. However, SpsMaxLatencyPictures may also represent the maximum number of picture decodes, which is the number of pictures decoded between the time a picture is stored in the buffer after decoded and the time it becomes available for display. For example, in the case of image P4 in Figure 12B, after the decoding of image P4 is complete, three images, B2, B1, and B3, are decoded before image P4 becomes available for display. Also, images B2, B3, and P4 are displayed in order. This maximum number of picture decodes, 3, may be set as SpsMaxLatencyPictures.

[0097] Furthermore, in this embodiment, sps_max_num_reorder_pics and sps_max_latency_increase_plus1 are set for each hierarchy and encoded, but this is not limited to this. For example, in a system that does not use temporal scalability, only the values ​​of sps_max_num_reorder_pics and sps_max_latency_increase_plus1 for the deepest hierarchy (the hierarchy with the largest TemporalId) may be set and encoded.

[0098] Furthermore, while the above explanation uses four frame rates—24fps, 30fps, 60fps, and 120fps—other frame rates may also be used. Additionally, frame rates may include decimal values, such as 29.97fps.

[0099] Furthermore, the processing in this embodiment may be implemented in software. This software may be distributed by download or other means. Alternatively, this software may be recorded on a recording medium such as a CD-ROM and distributed. This also applies to other embodiments described herein.

[0100] (Embodiment 2) This embodiment describes an image decoding device corresponding to the image encoding device described in Embodiment 1.

[0101] <Overall Structure> Figure 13 is a block diagram showing the configuration of the image decoding device 200 in this embodiment.

[0102] The image decoding device 200 shown in Figure 13 generates an output image 263 by decoding a code sequence 251. The code sequence 251 is, for example, a code sequence 155 generated by the image encoding device 100 in Embodiment 1. This image decoding device 200 includes a variable-length decoding unit 201, an inverse transformation quantization unit 202, an addition unit 203, a frame memory 204, an intra-prediction block generation unit 205, an inter-prediction block generation unit 206, a limit value decoding unit 208, an image rearrangement unit 209, and an encoding structure verification unit 210.

[0103] <Overall Operation> Next, the image decoding process according to this embodiment will be described with reference to Figure 14.

[0104] First, the variable-length decoding unit 201 decodes the encoding structure limit value 257 from the code sequence 251. This encoding structure limit value 257 includes sps_max_sub_layers_minus1, sps_max_num_reorder_pics, and sps_max_latency_increase_plus1. The meaning of this information is the same as in Embodiment 1. Next, the limit value decoding unit 208 obtains the number of layers by adding 1 to sps_max_sub_layers_minus1, obtains the number of consecutive B-pictures by the formula sps_max_num_reorder_pics + sps_max_latency_increase_plus1 - 1, and obtains the number of display delay pictures by sps_max_num_reorder_pics (S201). Furthermore, the limit value decoding unit 208 obtains the encoded structure 262 (number of layers, number of display delay pictures, and number of consecutive B pictures) from sps_max_num_reorder_pics and sps_max_latency_increase_plus1 of the TemporalId layers corresponding to the value of HighestTId252 input from the outside, and outputs the obtained encoded structure 262 to the image rearrangement unit 209 and the encoded structure verification unit 210. Here, HighestTId252 indicates the TemporalId of the highest layer to be decoded.

[0105] Next, the coding structure verification unit 210 checks whether each value of the coding structure 262 conforms to the operational specifications (S202). Specifically, the coding structure verification unit 210 uses the transmission delay time limit value 253 input from an external source and the frame rate 256 obtained by variable-length decoding of the code sequence 251 to calculate each limit value using the following equations (7) to (9), and determines whether the coding structure is less than or equal to the calculated limit value.

[0106] Number of layers = int(log2(transmission delay time limit [s] × frame rate [fps])) + 1 ... (Equation 7)

[0107] Display delay picture count [TId] = int(log2(transmission delay time limit [s] × frame rate [fps] ÷ 2) (n-TId))) ...(Formula 8)

[0108] B-Picture Consecutive Sequences [TId] = int(Transmission Delay Time Limit [s] × Frame Rate [fps] ÷ 2) (n-TId) -1)...(Formula 9)

[0109] If the encoded structure is greater than the limit value (Yes in S203), the encoded structure verification unit 210 displays an error message to that effect (S204) and terminates the decoding process.

[0110] Next, the variable-length decoding unit 201 decodes prediction information 255 indicating the prediction mode from the code sequence 251 (S205). If the prediction mode is intra-prediction (Yes in S206), the intra-prediction block generation unit 205 generates a prediction block 261 using intra-prediction (S207). On the other hand, if the prediction mode is inter-prediction (No in S206), the inter-prediction block generation unit 206 generates a prediction block 261 using inter-prediction (S208).

[0111] Next, the variable-length decoding unit 201 decodes the conversion coefficients 254 from the code sequence 251 (S209). Next, the inverse conversion quantization unit 202 reconstructs the difference block 258 by performing inverse quantization and inverse frequency conversion on the conversion coefficients 254 (S210). Next, the addition unit 203 generates a decoded block 259 by adding the difference block 258 and the prediction block 261 (S211). This decoded block 259 is stored in the frame memory 204 and used in the prediction block generation process by the intra-prediction block generation unit 205 and the inter-prediction block generation unit 206.

[0112] Then, the image decoding device 200 moves on to processing the next code block (S212) and repeats steps S205 to S212 until processing of all code blocks in the picture is completed (S213).

[0113] Note that the processing in steps S205 to S212 is performed only on pictures with a TemporalId of HighestTId 252 or less that are input from an external source.

[0114] Next, the image rearrangement unit 209 rearranges the decoded pictures according to the hierarchical encoding structure 262 of the HighestTId 252 input from the outside, and outputs the rearranged decoded pictures as the output image 263 (S214).

[0115] The image decoding device 200 then repeats steps S205 to S214 until processing of all pictures is complete (S215).

[0116] <Effects> As described above, the image decoding device 200 according to this embodiment can decode code sequences generated by efficient encoding. Furthermore, the image decoding device 200 can check whether the encoding structure conforms to the operational specifications, and if it does not, it can stop the decoding process and display an error.

[0117] In the above explanation, the image decoding device 200 is configured to decode only pictures at or below the HighestTId252 level, according to the HighestTId252 input from an external source. However, this is not always the case. The image decoding device 200 may always decode pictures at all levels. Alternatively, the image decoding device 200 may use a predetermined fixed value for HighestTId252 and always decode only pictures at or below the predetermined level indicated by HighestTId252.

[0118] Furthermore, although the above description states that the image decoding device 200 checks whether the encoding structure 262 conforms to the operational regulations, this function is not mandatory, and it is not necessary to check the encoding structure 262.

[0119] Furthermore, although the image decoding device 200 uses an externally input transmission delay time limit value 253, a predetermined fixed value may also be used as the transmission delay time limit value 253.

[0120] Other details are the same as in Embodiment 1 and will therefore be omitted.

[0121] Note that the order of each flow is not limited to the above, just as with the encoding side.

[0122] As described above in Embodiment 1 and Embodiment 2, the image encoding device 100 according to Embodiment 1 is an image encoding device that generates a code sequence 155 (bitstream) by hierarchical encoding of the input image 153, and performs the processing shown in Figure 15.

[0123] First, the image encoding device 100 determines the number of layers 161 in layer coding so that the number of layers 161 in layer coding is less than or equal to the maximum number of layers predetermined according to the frame rate (S301). Here, the maximum number of layers is the number of layers shown in Figure 10, for example, it is 2 when the frame rate is 24fps, 3 when the frame rate is 30fps, 4 when the frame rate is 60fps, and 5 when the frame rate is 120fps. In other words, when the frame rate is 60fps or higher, the maximum number of layers is 4 or higher. Also, when the frame rate is 60fps or lower, the maximum number of layers is 4 or lower. Also, when the frame rate is greater than 30fps, the maximum number of layers is greater than 3.

[0124] Furthermore, the image encoding device 100 determines the picture type of the input image 153 so that the number of display delay pictures 164 is less than or equal to the maximum number of pictures predetermined according to the frame rate. Here, the number of display delay pictures 164 is the number of pictures from the time the image decoding device starts decoding the code sequence 155 generated by the image encoding device 100 until it outputs (displays) the image. The picture type is either an I picture, a P picture, or a B picture. Here, the maximum number of pictures is the number of display delay pictures shown in Figure 10, which is, for example, 1 when the frame rate is 24fps, 2 when the frame rate is 30fps, 3 when the frame rate is 60fps, and 4 when the frame rate is 120fps. In other words, when the frame rate is 60fps or higher, the maximum number of pictures is 3 or more. When the frame rate is 60fps or lower, the maximum number of pictures is 3 or less. Additionally, if the frame rate is greater than 30fps, the maximum number of pictures is greater than 2.

[0125] Furthermore, the image encoding device 100 determines the picture type of the input image 153 such that the number of consecutive B-pictures, or the number of consecutive B-pictures 162, is less than or equal to a predetermined maximum number of consecutive B-pictures depending on the frame rate. Here, the maximum number of consecutive B-pictures is the number of display delay pictures shown in Figure 10, and for example, it is 2 when the frame rate is 24fps, 3 when the frame rate is 30fps, 7 when the frame rate is 60fps, and 15 when the frame rate is 120fps. In other words, when the frame rate is 60fps or higher, the maximum number of consecutive B-pictures is 7 or higher. Also, when the frame rate is 60fps or lower, the maximum number of consecutive B-pictures is 7 or lower. Also, when the frame rate is greater than 30fps, the maximum number of consecutive B-pictures is greater than 3.

[0126] Furthermore, as shown in Figure 10, the image encoding device 100 may determine the maximum number of layers, the maximum number of pictures, and the number of consecutive B-pictures according to the frame rate. In other words, the image encoding device 100 may set the maximum number of layers, the maximum number of pictures, and the number of consecutive B-pictures to be higher as the frame rate increases.

[0127] Furthermore, as described above, the number of layers 161, the number of display delay pictures 164, and the number of consecutive B pictures 162 are calculated using the frame rate 151 and the transmission delay time limit value 152, according to (Equation 1), (Equation 3), and (Equation 4) above. In other words, the maximum number of pictures, the encoder transmission delay (transmission delay time), which is the time from when the input image 153 is input to the image encoding device 100 until the code sequence 155 is output, and the frame rate satisfy the following relationship.

[0128] Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps]))

[0129] Furthermore, the maximum number of consecutive frames, encoder transmission delay, and frame rate satisfy the following relationship.

[0130] Maximum number of consecutive frames = int(encoder transmission delay [s] × frame rate [fps] - 1)

[0131] The maximum number of layers, encoder transmission delay, and frame rate satisfy the following relationship.

[0132] Maximum number of layers = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1

[0133] Furthermore, the following relationship is satisfied between the maximum number of pictures [i] in each layer, the encoder transmission delay, and the frame rate.

[0134] Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2) (n-i) ))

[0135] The maximum number of consecutive layers [i], encoder transmission delay, and frame rate satisfy the following relationship.

[0136] Maximum number of consecutive frames [i] = int(encoder transmission delay [s] × frame rate [fps] / 2) (n-i) -1)

[0137] Here, i is an integer less than or equal to the maximum number of levels, and represents a level. n represents (maximum number of levels - 1).

[0138] Next, the image encoding device 100 generates a code sequence 155 by hierarchically encoding the input image 153 with the determined number of layers 161 and picture type (S302). The image encoding device 100 also encodes first information (sps_max_sub_layers_minus1), second information (sps_max_num_reorder_pics), and third information (sps_max_latency_increase_plus1) indicating the determined number of layers 161, the number of display delay pictures 164, and the number of consecutive B pictures 162.

[0139] Furthermore, the image decoding device 200 according to Embodiment 2 is an image decoding device that generates an output image 263 by decoding a code sequence 251 (bitstream) obtained by hierarchical encoding of an image, and performs the processing shown in Figure 16.

[0140] First, the image decoding device 200 decodes the image from the code sequence 251 (S401).

[0141] Next, the image decoding device 200 decodes from the code sequence 251 first information (sps_max_sub_layers_minus1) indicating the number of layers in hierarchical coding (S402). For example, this number of layers is less than or equal to the maximum number of layers predetermined according to the frame rate of the code sequence 251.

[0142] Furthermore, the image decoding device 200 decodes a second piece of information (sps_max_num_reorder_pics) indicating the number of delayed display pictures from the code sequence 251. The image decoding device 200 also decodes a third piece of information (sps_max_latency_increase_plus1) indicating the number of consecutive B pictures from the code sequence 251.

[0143] Next, the image decoding device 200 rearranges and outputs the decoded image using the number of layers indicated by the first information, the number of display delay pictures indicated by the second information, and the number of consecutive B pictures indicated by the third information (S403).

[0144] The specific examples and limitations for the maximum number of layers, the maximum number of display delay pictures, and the maximum number of consecutive B pictures are the same as in the case of the image encoding device 100. Furthermore, the relationship between the maximum number of layers, the maximum number of pictures, and the maximum number of consecutive images, and the frame rate and encoder transmission delay are also the same as in the case of the image encoding device 100.

[0145] Although the image decoding device and image encoding device according to the embodiment have been described above, the present invention is not limited to this embodiment.

[0146] Furthermore, each processing unit included in the image decoding device or image encoding device according to the above embodiment is typically implemented as an LSI (Large-Scale Integrated Circuit). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0147] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.

[0148] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0149] In other words, the image decoding device and the image encoding device comprise a processing circuitry and a storage device electrically connected to (accessible from) the processing circuitry. The processing circuitry includes at least one of dedicated hardware and a program execution unit. The storage device also stores a software program executed by the program execution unit if the processing circuitry includes the program execution unit. The processing circuitry uses the storage device to execute the image decoding method or image encoding method according to the above embodiment.

[0150] Furthermore, the present invention may be the software program described above, or a non-temporary computer-readable recording medium on which the program is recorded. It goes without saying that the program can also be distributed via a transmission medium such as the Internet.

[0151] Furthermore, the figures used above are all illustrative to specifically explain the present invention, and the present invention is not limited to the illustrated figures.

[0152] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0153] Furthermore, the order in which the steps included in the above-described image decoding method or image encoding method are performed is illustrative for the purpose of specifically illustrating the present invention, and may be performed in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0154] Furthermore, the processing described in the above embodiment may be implemented by centralized processing using a single device (system), or by distributed processing using multiple devices. Also, the computer executing the above program may be one or multiple. In other words, centralized processing may be performed, or distributed processing may be performed.

[0155] Furthermore, the present invention is particularly effective in cases such as broadcasting to a large number of end users, where the functions of the receiving terminals are diverse. For example, a signal with the data structure described above is broadcast. A terminal such as a 4K2K television can process the full layer of data. On the other hand, a smartphone can process up to two layers. Also, depending on the bandwidth congestion, the transmitting device can send only the upper layers instead of the full layer. This enables flexible broadcasting and communication.

[0156] Although an image decoding device and an image encoding device according to one or more embodiments of the present invention have been described above based on embodiments, the present invention is not limited to these embodiments. Without departing from the spirit of the present invention, various modifications that a person skilled in the art can conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments of the present invention.

[0157] (Embodiment 3) By recording a program for realizing the configuration of the video encoding method (image encoding method) or video decoding method (image decoding method) shown in each of the above embodiments onto a storage medium, the processes shown in each of the above embodiments can be easily performed on an independent computer system. The storage medium can be anything that can record a program, such as a magnetic disk, optical disk, magneto-optical disk, IC card, or semiconductor memory.

[0158] Furthermore, here we will describe application examples of the video encoding method (image encoding method) and video decoding method (image decoding method) shown in each of the embodiments described above, and a system using them. The system is characterized by having an image encoding and decoding device consisting of an image encoding device using the image encoding method and an image decoding device using the image decoding method. Other configurations in the system can be appropriately modified as needed.

[0159] Figure 17 shows the overall configuration of the content supply system ex100 that realizes the content distribution service. The service area for the communication service is divided into cells of a desired size, and fixed radio stations, base stations ex106, ex107, ex108, ex109, and ex110, are installed in each cell.

[0160] This content supply system ex100 connects various devices such as computers ex111, PDAs (Personal Digital Assistants) ex112, cameras ex113, mobile phones ex114, and game consoles ex115 to the internet ex101, via the internet service provider ex102 and telephone network ex104, and base stations ex106 via ex110.

[0161] However, the content supply system ex100 is not limited to the configuration shown in Figure 17, and any combination of elements may be used for connection. Furthermore, each device may be directly connected to the telephone network ex104 from the base station ex106, which is a fixed radio station, without going through ex110. Also, each device may be directly connected to one another via short-range radio or the like.

[0162] Camera ex113 is a device capable of recording video, such as a digital video camera, and Camera ex116 is a device capable of taking still images and recording video, such as a digital camera. Furthermore, Mobile phone ex114 can be a mobile phone using GSM (Registered Trademark) (Global System for Mobile Communications), CDMA (Code Division Multiple Access), W-CDMA (Wideband-Code Division Multiple Access), LTE (Long Term Evolution), HSPA (High Speed ​​Packet Access), or PHS (Personal Handyphone System), and any of these is acceptable.

[0163] In the content supply system ex100, cameras ex113 and other devices are connected to the streaming server ex103 via the base station ex109 and the telephone network ex104, enabling live streaming and other functions. In live streaming, content captured by the user using the camera ex113 (for example, video of a music concert) is encoded as described in each of the embodiments above (i.e., it functions as an image encoding device according to one aspect of the present invention) and transmitted to the streaming server ex103. Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. Clients include computers ex111, PDAs ex112, cameras ex113, mobile phones ex114, game consoles ex115, etc., that are capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data (i.e., it functions as an image decoding device according to one aspect of the present invention).

[0164] The encoding process for the captured data may be performed by the camera ex113, the streaming server ex103 which handles data transmission, or the tasks may be shared between them. Similarly, the decoding process for the transmitted data may be performed by the client, the streaming server ex103, or the tasks may be shared between them. In addition, still images and / or video data captured by camera ex116 may be transmitted to the streaming server ex103 via computer ex111, not limited to camera ex113. In this case, the encoding process may be performed by camera ex116, computer ex111, or streaming server ex103, or the tasks may be shared among them.

[0165] Furthermore, these encoding and decoding processes are generally performed by the computer ex111 or the LSIex500 in each device. The LSIex500 may be a single chip or a configuration consisting of multiple chips. Alternatively, the software for video encoding and decoding may be embedded in some recording medium (CD-ROM, flexible disk, hard disk, etc.) that can be read by the computer ex111, and the encoding and decoding processes may be performed using that software. In addition, if the mobile phone ex114 has a camera, video data acquired by that camera may be transmitted. In this case, the video data is data that has been encoded by the LSIex500 in the mobile phone ex114.

[0166] Furthermore, the streaming server ex103 may consist of multiple servers or multiple computers that distribute, process, record, and distribute data.

[0167] As described above, the content supply system ex100 allows clients to receive and play back encoded data. In this way, the content supply system ex100 allows clients to receive, decode, and play back information transmitted by users in real time, enabling personal broadcasting even for users who do not possess special rights or equipment.

[0168] Furthermore, as shown in Figure 18, the digital broadcasting system ex200 can also incorporate at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) described in each of the above embodiments, not just the content supply system ex100. Specifically, at the broadcasting station ex201, multiplexed data, in which video data and music data are multiplexed, is transmitted via radio waves to the communication station or satellite ex202. This video data is data encoded by the video encoding method described in each of the above embodiments (i.e., data encoded by an image encoding device according to one aspect of the present invention). The broadcasting satellite ex202 receives this data and transmits radio waves for broadcasting, which are received by a home antenna ex204 capable of receiving satellite broadcasts. The received multiplexed data is decoded and reproduced by a device such as a television (receiver) ex300 or a set-top box (STB) ex217 (i.e., it functions as an image decoding device according to one aspect of the present invention).

[0169] Furthermore, the video decoding device or video encoding device described in each of the above embodiments can also be implemented in a reader / recorder ex218 that reads and decodes multiplexed data recorded on recording media ex215 such as DVDs and BDs, or encodes video signals onto recording media ex215, and, in some cases, multiplexes them with music signals before writing. In this case, the reproduced video signal is displayed on a monitor ex219, and the video signal can be reproduced on other devices or systems using the recording media ex215 on which the multiplexed data is recorded. Alternatively, the video decoding device may be implemented in a set-top box ex217 connected to a cable television cable ex203 or a satellite / terrestrial broadcast antenna ex204, and displayed on the television monitor ex219. In this case, the video decoding device may be built into the television itself instead of in a set-top box.

[0170] Figure 19 shows a television (receiver) ex300 using the video decoding method and video encoding method described in each of the above embodiments. The television ex300 includes a tuner ex301 that acquires or outputs multiplexed data in which audio data is multiplexed with video data via an antenna ex204 or cable ex203, etc. that receives the above broadcast, a modulation / demodulation unit ex302 that demodulates the received multiplexed data or modulates it into multiplexed data to be transmitted externally, and a multiplexing / separation unit ex303 that separates the demodulated multiplexed data into video data and audio data, or multiplexes the video data and audio data encoded by the signal processing unit ex306.

[0171] Furthermore, the TV ex300 includes a signal processing unit ex306 having an audio signal processing unit ex304 that decodes audio data and video data, respectively, or encodes the information of each, and a video signal processing unit ex305 (which functions as an image encoding device or image decoding device according to one aspect of the present invention), and an output unit ex309 having a speaker ex307 that outputs the decoded audio signal and a display unit ex308 such as a display that shows the decoded video signal. Furthermore, the TV ex300 has an interface unit ex317 having an operation input unit ex312 that receives user input, etc. Furthermore, the TV ex300 has a control unit ex310 that comprehensively controls each unit and a power supply circuit unit ex311 that supplies power to each unit. The interface unit ex317 may include, in addition to the operation input unit ex312, a bridge ex313 for connecting to external devices such as a reader / recorder ex218, a slot unit ex314 for inserting recording media such as an SD card ex216, a driver ex315 for connecting to external recording media such as a hard disk, and a modem ex316 for connecting to a telephone network. The recording media ex216 enables the electrical recording of information using non-volatile / volatile semiconductor memory elements that it stores. The various parts of the television ex300 are connected to each other via a synchronization bus.

[0172] First, we will describe a configuration in which the TV ex300 decodes and plays back multiplexed data acquired from an external source, such as the antenna ex204. The TV ex300 receives user input from a remote controller ex220 or the like, and based on the control of the control unit ex310, which has a CPU, the multiplexed data demodulated by the modulation / demodulation unit ex302 is separated by the multiplexing / separation unit ex303. Furthermore, the TV ex300 decodes the separated audio data with the audio signal processing unit ex304, and decodes the separated video data with the video signal processing unit ex305 using the decoding method described in each of the above embodiments. The decoded audio signal and video signal are output to the outside from the output unit ex309, respectively. When outputting, it is advisable to temporarily store these signals in buffers ex318, ex319, etc., so that the audio signal and video signal are played back in sync. In addition, the TV ex300 may read multiplexed data from recording media such as magnetic / optical disks or SD cards ex215, ex216, rather than from broadcasts, etc. Next, a configuration in which the TV ex300 encodes audio and video signals and transmits them externally or writes them to a recording medium will be described. The TV ex300 receives user operations from a remote controller ex220 or the like, and based on the control of the control unit ex310, encodes audio signals in the audio signal processing unit ex304 and encodes video signals in the video signal processing unit ex305 using the encoding method described in each of the embodiments above. The encoded audio and video signals are multiplexed in the multiplexing / decompression unit ex303 and output externally. When multiplexing, it is advisable to temporarily store these signals in buffers ex320, ex321, etc., so that the audio and video signals are synchronized. Note that there may be multiple buffers ex318, ex319, ex320, and ex321 as shown in the figure, or one or more buffers may be shared. Furthermore, in addition to what is shown in the figure, data may also be stored in buffers as a buffer to avoid system overflow and underflow, for example, between the modulation / demodulation unit ex302 and the multiplexing / decompression unit ex303.

[0173] Furthermore, in addition to acquiring audio and video data from broadcasts and recording media, the TV ex300 may also be configured to accept AV inputs from microphones and cameras, and may perform encoding processing on the data acquired from them. While the TV ex300 is described here as having the above-mentioned encoding processing, multiplexing, and external output capabilities, it may also be configured to only perform the above-mentioned reception, decoding, and external output functions, without being able to perform these processes.

[0174] Furthermore, when the reader / recorder ex218 reads or writes multiplexed data from the recording medium, the above decoding or encoding process may be performed by either the television ex300 or the reader / recorder ex218, or the television ex300 and the reader / recorder ex218 may share the task.

[0175] As an example, Figure 20 shows the configuration of the information playback / recording unit ex400 when reading or writing data from an optical disc. The information playback / recording unit ex400 comprises the elements ex401, ex402, ex403, ex404, ex405, ex406, and ex407, which are described below. The optical head ex401 writes information by irradiating the recording surface of the recording medium ex215, which is an optical disc, with a laser spot, and reads the information by detecting the reflected light from the recording surface of the recording medium ex215. The modulation recording unit ex402 electrically drives the semiconductor laser built into the optical head ex401 and modulates the laser light according to the recorded data. The playback / demodulation unit ex403 amplifies the playback signal electrically detected by a photodetector built into the optical head ex401 that detects the reflected light from the recording surface, separates and demodulates the signal components recorded on the recording medium ex215, and plays back the necessary information. The buffer ex404 temporarily holds information to be recorded on the recording medium ex215 and information played back from the recording medium ex215. The disk motor ex405 rotates the recording medium ex215. The servo control unit ex406 controls the rotational drive of the disk motor ex405, moves the optical head ex401 to a predetermined information track, and performs laser spot tracking. The system control unit ex407 controls the entire information playback / recording unit ex400. The above reading and writing processes are realized by the system control unit ex407 using various information held in the buffer ex404, generating and adding new information as needed, and performing information recording and playback through the optical head ex401 while coordinating the operation of the modulation recording unit ex402, the playback / demodulation unit ex403, and the servo control unit ex406. The system control unit ex407 is composed of, for example, a microprocessor and executes these processes by running read / write programs.

[0176] In the above explanation, the optical head ex401 was described as emitting a laser spot, but a configuration using near-field light for higher-density recording is also possible.

[0177] Figure 21 shows a schematic diagram of the recording medium ex215, which is an optical disc. Guide grooves are formed in a spiral shape on the recording surface of the recording medium ex215, and address information indicating the absolute position on the disc is recorded in advance on the information track ex230 by changes in the shape of the grooves. This address information includes information for identifying the position of the recording block ex231, which is the unit in which data is recorded, and the recording block can be identified by playing back the information track ex230 and reading the address information in a recording or playback device. The recording medium ex215 also includes a data recording area ex233, an inner circumference area ex232, and an outer circumference area ex234. The data recording area ex233 is the area used to record user data, and the inner circumference area ex232 and outer circumference area ex234, which are located inward or outward from the data recording area ex233, are used for specific purposes other than recording user data. The information playback / recording unit ex400 reads and writes encoded audio data, video data, or multiplexed data obtained by multiplexing these data to the data recording area ex233 of the recording medium ex215.

[0178] The above explanation used single-layer optical discs such as DVDs and Blu-ray discs as examples, but it is not limited to these; it may also be a multi-layer optical disc that can record on surfaces other than the surface. Furthermore, it may be an optical disc with a multi-dimensional recording / playback structure, such as recording information using light of various different wavelengths of color in the same location on the disc, or recording different layers of information from various angles.

[0179] Furthermore, in the digital broadcasting system ex200, it is possible to receive data from satellite ex202, etc., using a vehicle ex210 equipped with antenna ex205, and to play video on a display device such as the car navigation system ex211 located in the vehicle ex210. For example, the configuration of the car navigation system ex211 could be one of the configurations shown in Figure 19 with the addition of a GPS receiver, and similar configurations could be considered for the computer ex111, mobile phone ex114, etc.

[0180] Figure 22A shows a mobile phone ex114 using the video decoding method and video encoding method described in the above embodiment. The mobile phone ex114 includes an antenna ex350 for transmitting and receiving radio waves with a base station ex110, a camera unit ex365 capable of taking video and still images, and a display unit ex358 such as a liquid crystal display that displays decoded data such as video captured by the camera unit ex365 and video received by the antenna ex350. The mobile phone ex114 further includes a main unit having an operation key unit ex366, an audio output unit ex357 such as a speaker for outputting sound, an audio input unit ex356 such as a microphone for inputting sound, a memory unit ex367 for storing encoded or decoded data such as captured video, still images, recorded audio, or received video, still images, and emails, or a slot unit ex364 which is an interface unit with a recording medium for storing data.

[0181] Furthermore, an example of the configuration of the mobile phone ex114 will be explained using Figure 22B. In the mobile phone ex114, the main control unit ex360 comprehensively controls each part of the main body, which includes the display unit ex358 and the operation key unit ex366. The power supply circuit unit ex361, operation input control unit ex362, video signal processing unit ex355, camera interface unit ex363, LCD (Liquid Crystal Display) control unit ex359, modulation / demodulation unit ex352, multiplexing / decompression unit ex353, audio signal processing unit ex354, slot unit ex364, and memory unit ex367 are all connected to each other via the bus ex370.

[0182] When the user performs an action such as ending a call or turning on the power key, the power supply circuit unit ex361 supplies power from the battery pack to each component, thereby starting up the mobile phone ex114 into an operational state.

[0183] The mobile phone ex114, based on the control of the main control unit ex360 which has a CPU, ROM, RAM, etc., converts the audio signal picked up by the audio input unit ex356 into a digital audio signal in the audio signal processing unit ex354 during voice call mode, performs spread spectrum processing on this in the modulation / demodulation unit ex352, performs digital-to-analog conversion processing and frequency conversion processing in the transmission / reception unit ex351, and then transmits it via the antenna ex350. Also, during voice call mode, the mobile phone ex114 amplifies the received data received via the antenna ex350, performs frequency conversion processing and analog-to-digital conversion processing on it, performs despread spectrum processing in the modulation / demodulation unit ex352, converts it into an analog audio signal in the audio signal processing unit ex354, and then outputs it from the audio output unit ex357.

[0184] Furthermore, when sending an email in data communication mode, the text data of the email entered by operating the operation keys ex366 on the main unit is sent to the main control unit ex360 via the operation input control unit ex362. The main control unit ex360 performs spread spectrum processing on the text data in the modulation / demodulation unit ex352, and after digital-to-analog conversion and frequency conversion processing in the transmission / reception unit ex351, it is transmitted to the base station ex110 via the antenna ex350. When receiving an email, the received data is processed in almost the reverse order and output to the display unit ex358.

[0185] When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit ex355 compresses and encodes the video signal supplied from the camera unit ex365 using the video encoding method shown in each of the above embodiments (i.e., it functions as an image encoding device according to one aspect of the present invention), and sends the encoded video data to the multiplexing / separation unit ex353. The audio signal processing unit ex354 encodes the audio signal picked up by the audio input unit ex356 while the camera unit ex365 is capturing video, still images, etc., and sends the encoded audio data to the multiplexing / separation unit ex353.

[0186] The multiplexing / decompression unit ex353 multiplexes encoded video data supplied from the video signal processing unit ex355 and encoded audio data supplied from the audio signal processing unit ex354 in a predetermined manner. The resulting multiplexed data is then subjected to spread spectrum processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex352, and after digital-to-analog conversion and frequency conversion processing in the transmission / reception unit ex351, it is transmitted via the antenna ex350.

[0187] When receiving video data linked to a homepage or the like in data communication mode, or when receiving an email with attached video and / or audio, the multiplexing / decomposition unit ex353 separates the multiplexed data received via antenna ex350 into a video data bitstream and an audio data bitstream, and supplies the encoded video data to the video signal processing unit ex355 and the encoded audio data to the audio signal processing unit ex354 via the synchronization bus ex370. The video signal processing unit ex355 decodes the video signal by decoding it using a video decoding method corresponding to the video encoding method shown in each of the above embodiments (i.e., it functions as an image decoding device according to one aspect of the present invention), and the video and still images contained in the video file linked to a homepage are displayed on the display unit ex358 via the LCD control unit ex359. The audio signal processing unit ex354 decodes the audio signal, and the audio is output from the audio output unit ex357.

[0188] Furthermore, terminals such as the mobile phone ex114 mentioned above, like the television ex300, can be implemented in three ways: a transceiver-type terminal with both an encoder and a decoder, a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. In addition, while the digital broadcasting system ex200 was described as receiving and transmitting multiplexed data in which music data etc. is multiplexed with video data, it may also be data in which text data related to the video is multiplexed in addition to audio data, or it may be video data itself instead of multiplexed data.

[0189] Thus, the video encoding method or video decoding method shown in each of the above embodiments can be used in any of the above-described devices or systems, and by doing so, the effects described in each of the above embodiments can be obtained.

[0190] Furthermore, the present invention is not limited to the above-described embodiments, and various modifications or alterations are possible without departing from the scope of the present invention.

[0191] (Embodiment 4) It is also possible to generate video data by appropriately switching between the video encoding methods or devices shown in each of the above embodiments and video encoding methods or devices compliant with different standards such as MPEG-2, MPEG4-AVC, and VC-1, as needed.

[0192] When multiple video data sets, each conforming to a different standard, are generated, it becomes necessary to select a decoding method corresponding to each standard during the decoding process. However, because it is not possible to identify which standard the video data to be decoded conforms to, a problem arises in that the appropriate decoding method cannot be selected.

[0193] To solve this problem, the multiplexed data, which is created by multiplexing audio data and other data onto video data, is configured to include identification information indicating which standard the video data conforms to. The specific configuration of the multiplexed data, which includes video data generated by the video encoding method or apparatus shown in each of the above embodiments, is described below. The multiplexed data is a digital stream in MPEG-2 transport stream format.

[0194] Figure 23 is a diagram showing the structure of multiplexed data. As shown in Figure 23, multiplexed data is obtained by multiplexing one or more of the following: video stream, audio stream, presentation graphics stream (PG), and interactive graphics stream. The video stream represents the main and secondary video of a film, the audio stream (IG) represents the main audio portion of a film and the secondary audio mixed with the main audio, and the presentation graphics stream represents the subtitles of a film. Here, the main video refers to the normal video displayed on the screen, and the secondary video refers to the video displayed on a smaller screen within the main video. The interactive graphics stream represents the interactive screen created by placing GUI components on the screen. The video stream is encoded by the video encoding method or apparatus shown in each embodiment above, or by a video encoding method or apparatus conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. The audio stream is encoded using methods such as Dolby AC-3, Dolby Digital Plus, MLP, DTS, DTS-HD, or Linear PCM.

[0195] Each stream included in the multiplexed data is identified by a PID. For example, the video stream used for movie footage is assigned 0x1011, audio streams are assigned 0x1100 to 0x111F, presentation graphics are assigned 0x1200 to 0x121F, interactive graphics streams are assigned 0x1400 to 0x141F, video streams used for secondary footage in movies are assigned 0x1B00 to 0x1B1F, and audio streams used for secondary audio mixed with the main audio are assigned 0x1A00 to 0x1A1F.

[0196] Figure 24 schematically illustrates how multiplexed data is multiplexed. First, the video stream ex235, consisting of multiple video frames, and the audio stream ex238, consisting of multiple audio frames, are converted into PES packet sequences ex236 and ex239, respectively, and then into TS packets ex237 and ex240. Similarly, the data from the presentation graphics stream ex241 and the interactive graphics stream ex244 are converted into PES packet sequences ex242 and ex245, respectively, and then further into TS packets ex243 and ex246. The multiplexed data ex247 is constructed by multiplexing these TS packets into a single stream.

[0197] Figure 25 shows in more detail how a video stream is stored in a sequence of PES packets. The first row in Figure 25 shows a sequence of video frames in the video stream. The second row shows a sequence of PES packets. As indicated by the arrows yy1, yy2, yy3, and yy4 in Figure 25, the multiple Video Presentation Units in the video stream, namely I-pictures, B-pictures, and P-pictures, are separated picture by picture and stored in the payload of a PES packet. Each PES packet has a PES header, which contains the Presentation Time-Stamp (PTS), which is the time the picture was displayed, and the Decoding Time-Stamp (DTS), which is the time the picture was decoded.

[0198] Figure 26 shows the format of the TS packet that is ultimately written to the multiplexed data. The TS packet is a fixed-length 188-byte packet consisting of a 4-byte TS header containing information such as the PID that identifies the stream, and a 184-byte TS payload that stores the data. The PES packet is split and stored in the TS payload. In the case of BD-ROM, a 4-byte TP_Extra_Header is attached to the TS packet, forming a 192-byte source packet that is written to the multiplexed data. The TP_Extra_Header contains information such as the ATS (Arrival_Time_Stamp). The ATS indicates the start time of forwarding the TS packet to the decoder's PID filter. As shown in the lower part of Figure 26, the source packets are arranged in the multiplexed data, and the number that increments from the beginning of the multiplexed data is called the SPN (Source Packet Number).

[0199] Furthermore, the TS packets included in the multiplexed data contain not only individual streams such as video, audio, and subtitles, but also PAT (Program Association Table), PMT (Program Map Table), and PCR (Program Clock Reference). The PAT indicates the PID of the PMT used in the multiplexed data, and the PAT itself is registered with a PID of 0. The PMT contains the PIDs of each stream such as video, audio, and subtitles included in the multiplexed data, as well as attribute information of the stream corresponding to each PID, and also contains various descriptors related to the multiplexed data. These descriptors include copy control information that instructs whether to allow or deny copying of the multiplexed data. The PCR contains information about the STC time corresponding to the ATS to which the PCR packet is forwarded to the decoder, in order to synchronize the ATC (Arrival Time Clock), which is the time axis of the ATS, with the STC (System Time Clock), which is the time axis of the PTS / DTS.

[0200] Figure 27 is a diagram illustrating the data structure of a PMT in detail. At the beginning of a PMT is a PMT header that indicates the length of the data contained in the PMT. Following this are multiple descriptors related to the multiplexed data. The copy control information mentioned above is written as a descriptor. After the descriptors are multiple stream information entries for each stream contained in the multiplexed data. The stream information consists of stream descriptors that describe the stream type, stream PID, and stream attribute information (frame rate, aspect ratio, etc.) to identify the compression codec of the stream. There are as many stream descriptors as there are streams in the multiplexed data.

[0201] When recording to a recording medium, the above-mentioned multiplexed data is recorded together with the multiplexed data information file.

[0202] As shown in Figure 28, the multiplexed data information file is management information for the multiplexed data, has a one-to-one correspondence with the multiplexed data, and consists of multiplexed data information, stream attribute information, and an entry map.

[0203] As shown in Figure 28, the multiplexed data information consists of the system rate, playback start time, and playback end time. The system rate indicates the maximum transfer rate of the multiplexed data to the PID filter of the system target decoder, which will be described later. The interval of the ATS included in the multiplexed data is set to be less than or equal to the system rate. The playback start time is the PTS of the first video frame of the multiplexed data, and the playback end time is set by adding the playback interval of one frame to the PTS of the last video frame of the multiplexed data.

[0204] As shown in Figure 29, attribute information for each stream included in the multiplexed data is registered for each PID. Attribute information differs for video streams, audio streams, presentation graphics streams, and interactive graphics streams. Video stream attribute information includes details such as the compression codec used to compress the video stream, the resolution of each individual picture data component, the aspect ratio, and the frame rate. Audio stream attribute information includes details such as the compression codec used to compress the audio stream, the number of channels included, the languages ​​supported, and the sampling frequency. This information is used for initializing the decoder before playback by the player.

[0205] In this embodiment, the stream type included in the PMT is used from the multiplexed data. Furthermore, if the multiplexed data is recorded on the recording medium, the video stream attribute information included in the multiplexed data information is used. Specifically, in the video encoding method or apparatus shown in each embodiment, a step or means is provided to set unique information indicating that the video data was generated by the video encoding method or apparatus shown in each embodiment, for the stream type included in the PMT or the video stream attribute information. This configuration makes it possible to distinguish between video data generated by the video encoding method or apparatus shown in each embodiment and video data conforming to other standards.

[0206] Furthermore, Figure 30 shows the steps of the video decoding method in this embodiment. In step exS100, the stream type included in the PMT or the video stream attribute information included in the multiplexed data information is obtained from the multiplexed data. Next, in step exS101, it is determined whether the stream type or video stream attribute information indicates that the multiplexed data was generated by the video encoding method or device shown in each of the embodiments described above. If it is determined that the stream type or video stream attribute information was generated by the video encoding method or device shown in each of the embodiments described above, decoding is performed in step exS102 using the video decoding method shown in each of the embodiments described above. If the stream type or video stream attribute information indicates that it conforms to conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, decoding is performed in step exS103 using a video decoding method conforming to the conventional standard.

[0207] In this way, by setting new unique values ​​for the stream type or video stream attribute information, it is possible to determine whether the video can be decoded using the video decoding method or apparatus described in each of the embodiments above. Therefore, even when multiplexed data conforming to different standards is input, an appropriate decoding method or apparatus can be selected, making it possible to decode without errors. Furthermore, the video encoding method or apparatus, or video decoding method or apparatus, described in this embodiment can be used with any of the devices or systems described above.

[0208] (Embodiment 5) The video encoding method and apparatus, and video decoding method and apparatus described in each of the above embodiments are typically implemented using an integrated circuit (LSI). As an example, Figure 31 shows the configuration of a single-chip LSIex500. The LSIex500 comprises elements ex501, ex502, ex503, ex504, ex505, ex506, ex507, ex508, and ex509, which are described below, and each element is connected via a bus ex510. The power supply circuit ex505 starts up to an operational state by supplying power to each part when the power supply is turned on.

[0209] For example, when performing encoding processing, the LSIex500 receives AV signals from a microphone ex117, camera ex113, etc., via AV I / O ex509, based on the control of the control unit ex501, which has a CPU ex502, memory controller ex503, stream controller ex504, drive frequency control unit ex512, etc. The input AV signals are temporarily stored in an external memory ex511 such as SDRAM. Based on the control of the control unit ex501, the stored data is divided into multiple parts as appropriate depending on the amount of processing and processing speed, and sent to the signal processing unit ex507, where the audio signal encoding and / or video signal encoding are performed. Here, the video signal encoding process is the encoding process described in each of the embodiments above. The signal processing unit ex507 further performs processing such as multiplexing the encoded audio data and encoded video data, and outputs it externally from the stream I / O ex506. This output multiplexed data is transmitted to the base station ex107 or written to the recording medium ex215. When multiplexing, it is recommended to temporarily store the data in buffer ex508 to ensure synchronization.

[0210] Although the memory ex511 was described above as an external component of the LSIex500, it may also be an internal component of the LSIex500. Similarly, the buffer ex508 is not limited to one, but may have multiple buffers. Furthermore, the LSIex500 may be implemented on a single chip or across multiple chips.

[0211] Furthermore, while the above assumes that the control unit ex501 includes a CPU ex502, a memory controller ex503, a stream controller ex504, a drive frequency control unit ex512, etc., the configuration of the control unit ex501 is not limited to this configuration. For example, the signal processing unit ex507 may also have a CPU. By providing a CPU inside the signal processing unit ex507, it becomes possible to further improve the processing speed. Another example is that the CPU ex502 may include the signal processing unit ex507, or a part of the signal processing unit ex507, such as an audio signal processing unit. In such a case, the control unit ex501 will have a configuration that includes the signal processing unit ex507, or a CPU ex502 having a part of it.

[0212] Although we have used the term LSI here, depending on the degree of integration, they may also be called IC, system LSI, super LSI, or ultra LSI.

[0213] Furthermore, the method of integrated circuit implementation is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. After LSI manufacturing, FPGAs (Field Programmable Gate Arrays) that can be programmed, or reconfigurable processors that allow for the reconfiguration of the connections and settings of circuit cells inside the LSI, may also be used. Such programmable logic devices can typically execute the video encoding method or video decoding method described in each of the above embodiments by loading or reading a program that constitutes software or firmware from memory or the like.

[0214] Furthermore, if advancements in semiconductor technology or derivative technologies lead to the emergence of integrated circuit technologies that replace LSIs, then naturally, these technologies can be used to integrate functional blocks. The application of biotechnology, for example, is a possible possibility.

[0215] (Embodiment 6) When decoding video data generated by the video encoding methods or devices described in the above embodiments, the processing load is likely to increase compared to decoding video data conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. Therefore, in the LSIex500, it is necessary to set the drive frequency of the CPUex502 to a higher frequency than when decoding video data conforming to conventional standards. However, increasing the drive frequency leads to the problem of increased power consumption.

[0216] To solve this problem, video decoding devices such as the TV ex300 and LSI ex500 are configured to identify which standard the video data conforms to and to switch the drive frequency according to the standard. Figure 32 shows the configuration ex800 in this embodiment. The drive frequency switching unit ex803 sets a high drive frequency when the video data is generated by the video encoding method or device shown in each of the above embodiments. It then instructs the decoding processing unit ex801, which executes the video decoding method shown in each of the above embodiments, to decode the video data. On the other hand, when the video data is video data conforming to a conventional standard, the drive frequency is set lower than when the video data is generated by the video encoding method or device shown in each of the above embodiments. It then instructs the decoding processing unit ex802, which conforms to a conventional standard, to decode the video data.

[0217] More specifically, the drive frequency switching unit ex803 consists of the CPU ex502 and the drive frequency control unit ex512 shown in Figure 31. The decoding processing unit ex801, which executes the video decoding method shown in each of the above embodiments, and the decoding processing unit ex802, which conforms to conventional standards, correspond to the signal processing unit ex507 shown in Figure 31. The CPU ex502 identifies which standard the video data conforms to. Based on the signal from the CPU ex502, the drive frequency control unit ex512 sets the drive frequency. Based on the signal from the CPU ex502, the signal processing unit ex507 decodes the video data. Here, for example, the identification information described in Embodiment 4 can be used to identify the video data. The identification information is not limited to that described in Embodiment 4; any information that can identify which standard the video data conforms to is acceptable. For example, if it is possible to identify which standard the video data conforms to based on an external signal that identifies whether the video data is for use on a television or on a disc, then such an external signal may be used for identification. Furthermore, the selection of the drive frequency in CPUex502 can be performed based on a lookup table that associates the video data standard with the drive frequency, as shown in Figure 34. The lookup table can be stored in buffer ex508 or the internal memory of the LSI, and CPUex502 can select the drive frequency by referring to this lookup table.

[0218] Figure 33 shows the steps for implementing the method of this embodiment. First, in step exS200, the signal processing unit ex507 obtains identification information from the multiplexed data. Next, in step exS201, the CPU ex502 identifies, based on the identification information, whether or not the video data was generated by the encoding method or device shown in each of the embodiments described above. If the video data was generated by the encoding method or device shown in each of the embodiments described above, in step exS202, the CPU ex502 sends a signal to the drive frequency control unit ex512 to set a higher drive frequency. The drive frequency control unit ex512 then sets the drive frequency to a higher level. On the other hand, if the video data is compliant with conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, in step exS203, the CPU ex502 sends a signal to the drive frequency control unit ex512 to set a lower drive frequency. The drive frequency control unit ex512 then sets the drive frequency to a lower level than when the video data was generated by the encoding method or device shown in each of the embodiments described above.

[0219] Furthermore, by changing the voltage supplied to the LSIex500 or the device containing the LSIex500 in conjunction with the switching of the drive frequency, it is possible to further enhance the power saving effect. For example, when setting a lower drive frequency, it is conceivable to set the voltage supplied to the LSIex500 or the device containing the LSIex500 to a lower value compared to when the drive frequency is set to a higher value.

[0220] Furthermore, the method for setting the drive frequency is not limited to the above-described method; a higher drive frequency is set when the processing load during decoding is large, and a lower drive frequency is set when the processing load during decoding is small. For example, if the processing load for decoding video data compliant with the MPEG4-AVC standard is greater than the processing load for decoding video data generated by the video encoding method or device shown in each of the above embodiments, the drive frequency can be set in the opposite way to the above-described method.

[0221] Furthermore, the method for setting the drive frequency is not limited to a configuration that lowers the drive frequency. For example, if the identification information indicates that the video data was generated by the video encoding method or device shown in each of the above embodiments, the voltage supplied to the LSIex500 or the device including the LSIex500 can be set high. If the identification information indicates that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, the voltage supplied to the LSIex500 or the device including the LSIex500 can be set low. Another example is that if the identification information indicates that the video data was generated by the video encoding method or device shown in each of the above embodiments, the CPUex502 can be driven without stopping. If the identification information indicates that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, there is processing capacity, so the CPUex502 can be temporarily driven. Even if the identification information indicates that the video data was generated by the video encoding method or device shown in each of the above embodiments, if there is processing capacity, the CPUex502 can be temporarily driven. In this case, it is possible to set a shorter stop time compared to when indicating that the video data conforms to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1.

[0222] In this way, power saving can be achieved by switching the drive frequency according to the standard to which the video data conforms. Furthermore, if the LSIex500 or a device containing the LSIex500 is powered by batteries, the battery life can be extended as a result of the power saving.

[0223] (Embodiment 7) Televisions, mobile phones, and other devices and systems mentioned above may receive multiple video data streams conforming to different standards. To enable decoding of these streams, the LSIex500's signal processing unit, ex507, needs to support multiple standards. However, using separate ex507 signal processing units for each standard would increase the circuit size and cost of the LSIex500.

[0224] To solve this problem, the decoding processing unit for executing the video decoding method shown in each of the above embodiments is partially shared with the decoding processing unit conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. An example of this configuration is shown in ex900 of Figure 35A. For example, the video decoding method shown in each of the above embodiments and the video decoding method conforming to the MPEG4-AVC standard have some common processing content in processes such as entropy coding, inverse quantization, deblocking filter, and motion compensation. For common processing content, a decoding processing unit ex902 corresponding to the MPEG4-AVC standard is shared, and for other processing content specific to one aspect of the present invention that does not correspond to the MPEG4-AVC standard, a dedicated decoding processing unit ex901 is used. In particular, since one aspect of the present invention is characterized by hierarchical coding, for example, a dedicated decoding processing unit ex901 is used for inverse quantization, and the decoding processing unit is shared for any or all of the other processes such as entropy coding, inverse quantization, deblocking filter, and motion compensation. Regarding the sharing of the decoding processing unit, a configuration may be used in which the decoding processing unit for executing the video decoding method shown in each of the above embodiments is shared for common processing content, and a dedicated decoding processing unit is used for processing content specific to the MPEG4-AVC standard.

[0225] Furthermore, another example of partially sharing processing is shown in ex1000 of Figure 35B. In this example, a dedicated decoding processing unit ex1001 corresponding to processing content specific to one aspect of the present invention, a dedicated decoding processing unit ex1002 corresponding to processing content specific to other conventional standards, and a shared decoding processing unit ex1003 corresponding to processing content common to the video decoding method according to one aspect of the present invention and the video decoding method of other conventional standards are used. Here, the dedicated decoding processing units ex1001 and ex1002 are not necessarily specialized for processing content specific to one aspect of the present invention or other conventional standards, but may be capable of executing other general-purpose processing. Furthermore, the configuration of this embodiment can also be implemented on LSIex500.

[0226] Thus, by sharing the decoding processing unit for common processing content between the video decoding method according to one aspect of the present invention and the conventional video decoding method, it is possible to reduce the circuit size of the LSI and lower costs. [Industrial applicability]

[0227] The present invention can be applied to an image decoding method and apparatus, or an image encoding method and apparatus. Furthermore, the present invention can be used in high-resolution information display devices or imaging devices such as televisions, digital video recorders, car navigation systems, mobile phones, digital cameras, and digital video cameras that are equipped with an image decoding device. [Explanation of symbols]

[0228] 100 Image encoding device 101 Limit value setting section 102 Encoding section 111 Layer number setting section 112. Number of levels parameter setting section 113 Display delay picture count setting section 114 B Picture Consecutive Count Setting Section 115. Continuous Parameter Setting Section 121, 209 Image sorting section 122 Code block division section 123 Subtraction Unit 124 Transformation Quantization Unit 125 Variable-length coding unit 126, 202 Inverse Transform Quantization Unit 127, 203 Addition section 128,204 frame memory 129 Intra Prediction Unit 130 Interpretation Unit 131 Selection Section 151,256 frame rates 152, 253 Transmission delay time limit values 153 Input Images 154, 257 Encoding structure limit values 155, 251 code string 161 levels 162 consecutive B-pictures 163 Number of levels parameter 164 Number of pictures with display delay 165 consecutive parameters 171 Code Block 172, 174, 258 difference blocks 173, 254 conversion coefficients 175,259 decryption blocks 177, 261 prediction blocks 178, 255 Prediction Information 200 Image Decoders 201 Variable-length decoding unit 205 Intra Prediction Block Generation Unit 206 Interpretation Block Generation Unit 208 Limit Decoding Unit 210 Encoding structure confirmation section 252HighestTid 262 encoding structure 263 Output Image

Claims

1. A receiving step of receiving a bitstream containing a video and control information, which is encoded by hierarchically encoding multiple images contained in the video into a hierarchical structure with one or more layers, The decoding step includes decoding the plurality of images from the bitstream, The lowest level of the aforementioned hierarchical structure includes I-pictures and P-pictures, The layers of the aforementioned hierarchical structure, other than the lowest layer, include B-pictures. The control information includes information on the frame rate of the video, The number of layers is predetermined based on the frame rate of the moving image, The number of consecutive B-pictures in the display order, which is the number of consecutive B-pictures, is less than or equal to a predetermined maximum number of consecutive B-pictures. Reception method.

2. Processing circuit and The system includes a storage device accessible from the processing circuit, The processing circuit uses the memory device, The receiving method described in claim 1 is performed. Receiving device.