Transmission method and transmission device
The hierarchical encoding and decoding methods address inefficiencies in image processing by limiting layers and B pictures based on frame rate, ensuring efficient and scalable image coding and decoding.
Patent Information
- Application Number
- JP2025081724
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-06-05
- Filing Date
- 2025-05-15
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2034-06-04
AI Technical Summary
Existing image encoding and decoding methods are inefficient, particularly in handling high frame rates, leading to increased display delay times and limited scalability.
An image coding method that hierarchically encodes images into multiple layers, with the lowest layer containing I and P pictures and higher layers containing B pictures, and an image decoding method that decodes these layers efficiently based on frame rate, limiting the number of layers and B pictures to maintain efficient processing.
The methods allow for efficient image coding and decoding while maintaining low display delay times and improving scalability across various frame rates.
Smart Images

Figure 0007802994000001 
Figure 0007802994000002 
Figure 0007802994000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image coding method for coding an image or an image decoding method for decoding an image. [Background technology]
[0002] Non-Patent Document 1 describes a technique relating to an image coding method for coding an image (including a moving image) and an image decoding method for decoding an image. Non-Patent Document 2 also describes operational rules for coding and decoding. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 12th Meeting: Geneva, CH, 14-23 Jan. 2013 JCTVC-L1003_v34.doc, High Efficiency Video Coding (HEVC) text specification draft 10 (for FDIS & Last Call) http: / / phenix.it-sudparis.eu / jct / doc_end_user / documents / 12_Geneva / wg11 / JCTVC-L1003-v34.zip [Non-patent document 2] Association of Radio Industries and Businesses (ARIB) Standard ARIB STD-B32 Version 2.8 2-STD-B32v2_8.pdf, Video Coding, Audio Coding and Multiplexing Methods for Digital Broadcasting http: / / www.arib.or.jp / english / html / overview / doc / 2-STD-B32v2_8.pdf Summary of the Invention [Problem to be solved by the invention]
[0004] However, in some cases, inefficient processing is used in image encoding or decoding methods according to the prior art.
[0005] Therefore, an object of the present invention is to provide an image coding method for efficiently coding an image or an image decoding method for efficiently decoding an image. [Means for solving the problem]
[0006] In order to achieve the above object, a transmission method according to one embodiment of the present invention includes a generation step of generating a bitstream including a video encoded by hierarchically encoding a plurality of images included in the video into a hierarchical structure having one or more layers, and control information; and a transmission step of transmitting the bitstream and the control information, wherein the lowest layer of the hierarchical structure includes I pictures and P pictures, and layers other than the lowest layer of the hierarchical structure include B pictures, the control information includes information on the frame rate of the video, the number of layers is predetermined based on the frame rate of the video, and the number of consecutive B pictures, which is the number of consecutive B pictures in display order, is equal to or less than a predetermined maximum consecutive number.
[0007] Furthermore, a receiving method according to one aspect of the present invention includes a receiving step of receiving a bitstream including a moving image encoded by hierarchically encoding the plurality of images included in the moving image into a hierarchical structure having one or more layers, and control information; and a decoding step of decoding the plurality of images from the bitstream, wherein the lowest layer of the hierarchical structure includes I pictures and P pictures, and layers other than the lowest layer of the hierarchical structure include B pictures, the control information includes information on the frame rate of the moving image, and the number of layers is predetermined based on the frame rate of the moving image.
[0008] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]
[0009] The present invention can provide an image coding method that can efficiently code images or an image decoding method that can efficiently decode images. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing an example of an encoding structure. [Figure 2] FIG. 2 is a diagram showing the number of display delay pictures. [Figure 3] FIG. 3 is a block diagram of an image coding device according to the first embodiment. [Figure 4] FIG. 4 is a flowchart of the image coding process according to the first embodiment. [Figure 5] FIG. 5 is a block diagram of the limit value setting unit according to the first embodiment. [Figure 6] FIG. 6 is a flowchart of the limit value setting process according to the first embodiment. [Figure 7] FIG. 7 is a block diagram of the encoding unit according to the first embodiment. [Figure 8] FIG. 8 is a flowchart of the encoding process according to the first embodiment. [Figure 9A] FIG. 9A is a diagram showing the number of delivery delay pictures according to the first embodiment. [Figure 9B] FIG. 9B is a diagram showing the number of delivery delay pictures according to the first embodiment. [Figure 9C] FIG. 9C is a diagram showing the number of delivery delay pictures according to the first embodiment. [Figure 9D] FIG. 9D is a diagram showing the number of delivery delay pictures according to the first embodiment. [Figure 10]FIG. 10 is a diagram showing an example of coding structure limiting values according to the first embodiment. [Figure 11A] FIG. 11A is a diagram showing a coding structure according to the first embodiment. [Figure 11B] FIG. 11B is a diagram showing a coding structure according to the first embodiment. [Figure 11C] FIG. 11C is a diagram showing a coding structure according to the first embodiment. [Figure 11D] FIG. 11D is a diagram showing a coding structure according to the first embodiment. [Figure 12A] FIG. 12A is a diagram showing the number of display delay pictures according to the first embodiment. [Figure 12B] FIG. 12B is a diagram showing the number of display delay pictures according to the first embodiment. [Figure 12C] FIG. 12C is a diagram showing the number of display delay pictures according to the first embodiment. [Figure 12D] FIG. 12D is a diagram showing the number of display delay pictures according to Embodiment 1. [Figure 13] FIG. 13 is a block diagram of an image decoding device according to the second embodiment. [Figure 14] FIG. 14 is a flowchart of the image decoding process according to the second embodiment. [Figure 15] FIG. 15 is a flowchart of an image coding method according to the first embodiment. [Figure 16] FIG. 16 is a flowchart of an image decoding method according to the second embodiment. [Figure 17] FIG. 17 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 18] FIG. 18 is a diagram showing the overall configuration of a digital broadcasting system. [Figure 19] FIG. 19 is a block diagram showing an example of the configuration of a television. [Figure 20] FIG. 20 is a block diagram showing an example of the configuration of an information reproducing / recording unit that reads and writes information from and to a recording medium that is an optical disc. [Figure 21]FIG. 21 is a diagram showing an example of the structure of a recording medium that is an optical disc. [Figure 22A] FIG. 22A is a diagram showing an example of a mobile phone. [Figure 22B] FIG. 22B is a block diagram showing an example of the configuration of a mobile phone. [Figure 23] FIG. 23 is a diagram showing the structure of multiplexed data. [Figure 24] FIG. 24 is a diagram showing a schematic diagram of how each stream is multiplexed in multiplexed data. [Figure 25] FIG. 25 shows in more detail how a video stream is stored in a sequence of PES packets. [Figure 26] FIG. 26 is a diagram showing the structure of TS packets and source packets in multiplexed data. [Figure 27] FIG. 27 shows the data structure of a PMT. [Figure 28] FIG. 28 is a diagram showing the internal structure of the multiplexed data information. [Figure 29] FIG. 29 shows the internal structure of the stream attribute information. [Figure 30] FIG. 30 shows the steps for identifying video data. [Figure 31] FIG. 31 is a block diagram showing an example of the configuration of an integrated circuit that realizes the video encoding method and video decoding method according to each embodiment. [Figure 32] FIG. 32 is a diagram showing a configuration for switching the drive frequency. [Figure 33] FIG. 33 is a diagram showing steps for identifying video data and switching the drive frequency. [Figure 34] FIG. 34 is a diagram showing an example of a lookup table in which video data standards and drive frequencies are associated with each other. [Figure 35A] FIG. 35A is a diagram showing an example of a configuration in which modules of a signal processing unit are shared. [Figure 35B] FIG. 35B is a diagram showing another example of a configuration in which modules of a signal processing unit are shared. DETAILED DESCRIPTION OF THE INVENTION
[0011] (Findings that form the basis of the present invention) The present inventor has found that the following problems occur with the image encoding device that encodes an image or the image decoding device that decodes an image, as described in the "Background Art" section.
[0012] In recent years, technological advances in digital video equipment have been remarkable, and there are increasing opportunities for video signals (multiple pictures arranged in chronological order) output from video cameras or television tuners to be compressed and encoded, and for the resulting encoded signals to be recorded on recording media such as DVDs or hard disks.
[0013] H.264 / AVC (MPEG-4 AVC) is an image coding standard. The High Efficiency Video Coding (HEVC) standard (Non-Patent Document 1) is being considered as a next-generation standard. Furthermore, regulations on how to operate the image coding standard are also being considered (Non-Patent Document 2).
[0014] In the current operational specifications (Non-Patent Document 2), the coding structure is limited to three layers as shown in Figure 1, and therefore the maximum number of display latency pictures is limited to two as shown in Figure 2. TemporalId shown in Figure 1 is an identifier for the layer of the coding structure. The larger the TemporalId, the deeper the layer.
[0015] Each square block represents a picture, and the I x is an I picture (intra-predicted picture), P x is a P picture (forward reference predicted picture), B x indicates a B picture (bidirectional reference predicted picture). x / P x / B x of x indicates the display order, which indicates the order in which the pictures are displayed.
[0016] Arrows between pictures indicate reference relationships. For example, picture B1 generates a predicted image using pictures I0, B2, and P4 as reference images. It is prohibited to use a picture with a Temporal Id greater than its own as a reference image. Therefore, the picture decoding order is in ascending order of Temporal Id, as shown in Figure 2, which is picture I0, picture P4, picture B2, picture B1, and picture B3.
[0017] By defining layers, it is possible to provide time scalability to a code stream.
[0018] For example, when it is desired to obtain a 30 fps (frames per second) video from a 60 fps code stream, the image decoding device decodes only the pictures of TemporalId0 and TemporalId1 in FIG. 1. This allows the image decoding device to obtain a 30 fps image. Since the decoded images must be output in order without gaps, the image decoding device outputs pictures in order starting with picture I0 after decoding picture B2. Therefore, the number of display delay pictures is two. Converting this to time, the display delay is 2 / 30 seconds when the original frame rate is 30 fps, and 2 / 60 seconds when the frame rate is 60 fps.
[0019] By using a structure with high temporal scalability, when the bandwidth is congested or when an image decoding device with low processing power performs decoding processing, the image decoding device can decode only pictures in layers with small TemporalIds and display the resulting images. In this way, versatility is improved. However, allowing a deep layer structure poses a problem of increased display delay.
[0020] However, even if the number of display delay pictures is specified in advance as described above, the display delay time varies depending on the frame rate. For a frame rate (e.g., 24 fps) lower than a standard frame rate (e.g., 30 fps), the display delay time is 2 / 24 seconds, which is longer than the 2 / 30 seconds for 30 fps.
[0021] An image coding method according to one aspect of the present invention is an image coding method for hierarchically coding an image, and includes a number of layers determination step of determining the number of layers in the hierarchical coding so that the number of layers is less than or equal to a maximum number of layers determined according to a frame rate, and an encoding step of hierarchically coding the image with the determined number of layers to generate a bitstream.
[0022] According to this, the image coding method can increase the number of layers while suppressing an increase in display delay time, and therefore can code images efficiently.
[0023] For example, when the frame rate is 60 fps or less, the maximum number of layers may be four or less.
[0024] For example, if the frame rate is 120 fps, the maximum number of layers may be five.
[0025] For example, the image encoding method may further include a picture type determination step in an image decoding device, which determines a picture type of the image so that the number of display delay pictures, which is the number of pictures from decoding an image to outputting it, is equal to or less than a maximum number of pictures determined according to the frame rate, and in the encoding step, the image may be encoded using the determined picture type.
[0026] For example, in the picture type determination step, the picture type of the image may be determined so that the number of consecutive B pictures, that is, the number of consecutive B pictures, is equal to or less than the maximum consecutive number determined according to the frame rate.
[0027] For example, the maximum number of pictures, an encoder output delay, which is the time from when the image is input to the image encoding device until the bitstream is output, and the frame rate satisfy the following relationship: Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps])) the maximum consecutive number, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive frames = int (encoder transmission delay [s] x frame rate [fps] - 1) The maximum number of layers, the encoder output delay, and the frame rate may satisfy the following relationship:
[0028] Maximum number of layers = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1 For example, the maximum number of pictures [i] in each layer, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2 (n-i) )) The maximum consecutive number [i] of each layer, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive frames [i] = int (encoder transmission delay [s] × frame rate [fps] / 2 (n-i) -1) i may be an integer equal to or less than the maximum number of layers and indicate a layer, and n may indicate (the maximum number of layers - 1).
[0029] Furthermore, an image decoding method according to one embodiment of the present invention is an image decoding method for decoding a bitstream obtained by hierarchical encoding of an image, and includes an image decoding step for decoding the image from the bitstream, an information decoding step for decoding first information indicating the number of layers in the hierarchical encoding from the bitstream, and a sorting step for sorting and outputting the decoded image using the number of layers indicated by the first information, wherein the number of layers is equal to or less than a maximum number of layers predetermined according to the frame rate of the bitstream.
[0030] According to this, the image decoding method can decode a bitstream obtained by efficient encoding.
[0031] For example, when the frame rate is 60 fps or less, the maximum number of layers may be four or less.
[0032] For example, if the frame rate is 120 fps, the maximum number of layers may be five.
[0033] For example, in the information decoding step, the image decoding device may further decode from the bitstream second information indicating the number of display delay pictures, which is the number of pictures from when an image is decoded to when it is output, and in the reordering step, the decoded images may be reordered and output using the number of layers indicated by the first information and the number of display delay pictures indicated by the second information.
[0034] For example, in the information decoding step, third information indicating the number of consecutive B pictures, which is the number of consecutive B pictures, may be decoded from the bitstream, and in the rearrangement step, the decoded images may be rearranged and output using the number of layers indicated by the first information, the number of display delay pictures indicated by the second information, and the number of consecutive B pictures indicated by the third information.
[0035] For example, the maximum number of pictures, an encoder output delay, which is the time from when the image is input to the image encoding device until the bitstream is output, and the frame rate satisfy the following relationship: Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps])) the maximum consecutive number, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive frames = int (encoder transmission delay [s] x frame rate [fps] - 1) The maximum number of layers, the encoder output delay, and the frame rate may satisfy the following relationship:
[0036] Maximum number of layers = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1 For example, the maximum number of pictures [i] in each layer, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2 (n-i) )) The maximum consecutive number [i] of each layer, the encoder transmission delay, and the frame rate satisfy the following relationship: Maximum number of consecutive frames [i] = int (encoder transmission delay [s] × frame rate [fps] / 2 (n-i) -1) i may be an integer equal to or less than the maximum number of layers and indicate a layer, and n may indicate (the maximum number of layers - 1).
[0037] Furthermore, an image coding device according to one aspect of the present invention is an image coding device that codes an image, and includes a processing circuit and a storage device accessible from the processing circuit, and the processing circuit executes the image coding method using the storage device.
[0038] This allows the image encoding device to increase the number of layers while suppressing an increase in display delay time, thereby enabling the image encoding device to encode images efficiently.
[0039] In addition, an image decoding device according to one aspect of the present invention is an image decoding device that decodes a bitstream obtained by encoding an image, and is equipped with a processing circuit and a storage device accessible from the processing circuit, and the processing circuit executes the image decoding method using the storage device.
[0040] This allows the image decoding device to decode a bitstream obtained by efficient encoding.
[0041] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0042] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below represents a specific example of the present invention. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept will be described as optional components.
[0043] (Embodiment 1) The image coding device according to this embodiment increases the number of layers when the frame rate is high, thereby making it possible to increase the number of layers while suppressing an increase in display delay time.
[0044] <Overall structure> FIG. 3 is a block diagram showing the configuration of the image coding device 100 according to this embodiment.
[0045] 3 generates a code string 155 (bit stream) by encoding an input image 153. The image encoding device 100 includes a limit value setting unit 101 and an encoding unit .
[0046] <Movement (overall)> Next, the overall flow of the encoding process will be described with reference to Fig. 4. Fig. 4 is a flowchart of the image encoding method according to this embodiment.
[0047] First, the limit value setting unit 101 sets a coding structure limit value 154 related to a coding structure in hierarchical coding (S101). Specifically, the limit value setting unit 101 sets the coding structure limit value 154 using a frame rate 151 and a transmission delay time limit value 152.
[0048] Next, the encoding unit 102 encodes the encoding structure limiting value 154 and encodes the input image 153 using the encoding structure limiting value 154 to generate a code string 155 (S102).
[0049] <Configuration of limit value setting unit 101> FIG. 5 is a block diagram showing an example of the internal configuration of the limit value setting unit 101. As shown in FIG.
[0050] As shown in FIG. 5, the limit value setting unit 101 includes a layer count setting unit 111, a layer count parameter setting unit 112, a display delay picture count setting unit 113, a B-picture consecutive number setting unit 114, and a consecutive number parameter setting unit 115.
[0051] <Operation (Encoding structure limit value setting)> Next, an example of the limit value setting process (S101 in FIG. 4) will be described with reference to Fig. 6. Fig. 6 is a flowchart of the limit value setting process according to this embodiment.
[0052] First, the number-of-layers setting unit 111 sets the number of layers 161 of the coding structure using a frame rate 151 and a transmission delay time limit value 152 input from outside the image coding device 100. For example, the number of layers 161 is calculated by the following (Equation 1) (S111).
[0053] Number of layers = int(log2(transmission delay time limit [s] × frame rate [fps])) + 1 (Formula 1)
[0054] In the above (Equation 1), int(x) is a function that returns an integer obtained by truncating the decimal part of x, and log2(x) is a function that returns the logarithm of x with base 2. The output delay time limit value 152 indicates the maximum time from when an input image 153 is input to the image encoding device 100 until a code string 155 of the input image 153 is output.
[0055] Next, the layer count parameter setting unit 112 uses the layer count 161 to set sps_max_sub_layers_minus1, which is the layer count parameter 163, by the following (Equation 2) (S112).
[0056] sps_max_sub_layers_minus1=Number of layers -1 (Formula 2)
[0057] Next, the limit value setting unit 101 sets TId to 0 (S113). TId is a variable for identifying a tier, and is used to identify the tier to be processed in the subsequent processing for each tier.
[0058] Next, the display latency picture number setting unit 113 sets the display latency picture number 164 of the layer whose TemporalId is TId, using the frame rate 151, the transmission latency limit value 152, and the layer number parameter 163 (S114). The display latency picture number 164 is the number of pictures from the start of picture decoding to the start of picture display during picture decoding. The display latency picture number 164 is calculated by the following (Equation 3).
[0059] Number of display delay pictures for the layer where TemporalId is TId = int(log2(transmission delay time limit value [s] × frame rate [fps] ÷ 2 (n-TId) )) ...(Formula 3)
[0060] In the above (Equation 3), n indicates the maximum TemporalId and is the value of sps_max_sub_layers_minus1 calculated in step S112. The display latency picture number setting unit 113 sets the calculated number of display latency pictures of the layer whose TemporalId is TId to sps_max_num_reorder_pics[TId].
[0061] Next, the consecutive B-picture count setting unit 114 sets the number of consecutive B-pictures 162 for the layer whose TemporalId is TId (S115) using the frame rate 151, the transmission delay time limit value 152, and the layer count parameter 163. The consecutive B-picture count 162 is the number of consecutive B-pictures, and is calculated by the following (Equation 4).
[0062] Number of consecutive B pictures in the layer where TemporalId is TId = int(transmission delay time limit [s] × frame rate [fps] ÷ 2 (n-TId) -1)...(Formula 4)
[0063] Next, the continuation number parameter setting unit 115 sets the continuation number parameter 165 for the layer whose TemporalId is TId, using the number of consecutive B pictures 162 and the number of display delay pictures 164 (sps_max_num_reorder_pics[TId]) for the layer whose TemporalId is TId (S116). The continuation number parameter 165 for the layer whose TemporalId is TId is set by the following (Equation 5).
[0064] Consecutive number parameter of the layer where TemporalId is TId = number of consecutive B pictures of the layer where TemporalId is TId - sps_max_num_reorder_pics[TId] + 1 (Equation 5)
[0065] The calculated consecutive number parameter 165 is set to sps_max_latency_increase_plus1[TId].
[0066] Next, the limit value setting unit 101 moves to the next tier to be processed by adding 1 to TId (S117). Steps S114 to S117 are repeated until TId reaches the tier number 161, that is, until processing of all tiers is completed (S118).
[0067] Here, the limit value setting unit 101 sets the number of layers 161 and then sets the number of display delay pictures 164 and the number of consecutive B pictures 162 for each layer, but the setting order is not limited to this.
[0068] <Configuration of the encoding unit 102> Fig. 7 is a block diagram showing the internal configuration of the encoding unit 102. As shown in Fig. 7, the encoding unit 102 includes an image rearrangement unit 121, a code block division unit 122, a subtraction unit 123, a transform quantization unit 124, a variable length encoding unit 125, an inverse transform quantization unit 126, an addition unit 127, a frame memory 128, an intra prediction unit 129, an inter prediction unit 130, and a selection unit 131.
[0069] <Operation (encoding)> Next, the encoding process (S102 in FIG. 4) according to this embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart of the encoding process according to this embodiment.
[0070] First, the variable length coding unit 125 performs variable length coding on sps_max_sub_layers_minus1, sps_max_num_reorder_pics[], and sps_max_latency_increase_plus1[] set by the limit value setting unit 101 (S121). sps_max_num_reorder_pics[] and sps_max_latency_increase_plus1[] exist for each layer, but the variable length coding unit 125 codes all of them.
[0071] Next, the image rearrangement unit 121 rearranges the input images 153 in accordance with sps_max_sub_layers_minus1, sps_max_num_reorder_pics[ ], and sps_max_latency_increase_plus1[ ], and determines the picture type of the input images 153 (S122).
[0072] The image rearrangement unit 121 performs this rearrangement using sps_max_num_reorder_pics[sps_max_sub_layers_minus1] and SpsMaxLatencyPictures. SpsMaxLatencyPictures is calculated by the following (Equation 6).
[0073] SpsMaxLatencyPictures=sps_max_num_reorder_pics[sps_max_sub_layers_minus1]+sps_max_latency_increase_plus1[sps_max_sub_layers_minus1]-1 (Formula 6)
[0074] 9A to 9D are diagrams showing this rearrangement. Because the rearrangement shown in FIGS. 9A to 9D is performed, the encoding unit 102 cannot start encoding the input images 153 until a plurality of input images 153 have been input. In other words, a delay occurs between the input of the first input image 153 and the start of output of the code string 155. This delay is the transmission delay time, and the above-mentioned transmission delay time limit value 152 is the limit value of this transmission delay time.
[0075] 9A to 9D show the number of transmission delay pictures corresponding to a coding structure limit value of 154. Fig. 9A shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 1 and SpsMaxLatencyPictures is 2. Fig. 9B shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 2 and SpsMaxLatencyPictures is 3. Fig. 9C shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 3 and SpsMaxLatencyPictures is 7. FIG. 9B shows the number of transmission delay pictures when sps_max_num_reorder_pics[sps_max_sub_layers_minus1] is 4 and SpsMaxLatencyPictures is 7.
[0076] For example, in the case of FIG. 9A , image 0, image 1, image 2, and image 3 are input to the image encoding device 100 in this order, and the image encoding device 100 encodes these images in the order of image 0, image 3, image 1, and image 2. Since the image encoding device 100 needs to transmit the codestream without any gaps, it does not start transmitting the codestream until image 3 is input. Therefore, a transmission delay of three pictures occurs between the input of image 0 and the start of transmission of the codestream. Furthermore, the image rearrangement unit 121 determines the picture type of each image and outputs information indicating which image each image uses as a reference picture to the inter prediction unit 130. Here, the picture types are I-picture, P-picture, and B-picture.
[0077] Next, the code block dividing unit 122 divides the input image 153 into code blocks 171 (S123).
[0078] Next, the intra prediction unit 129 generates a prediction block for intra prediction and calculates the cost of the prediction block (S124). The inter prediction unit 130 generates a prediction block for inter prediction and calculates the cost of the prediction block (S125). The selection unit 131 determines the prediction mode and prediction block 177 to be used using the calculated cost and the like (S126).
[0079] Next, the subtraction unit 123 generates a difference block 172 by calculating the difference between the prediction block 177 and the code block 171 (S127). Next, the transform quantization unit 124 performs frequency transform and quantization on the difference block 172 to generate transform coefficients 173 (S128). Next, the inverse transform quantization unit 126 performs inverse quantization and inverse frequency transform on the transform coefficients 173 to reconstruct a difference block 174 (S129). Next, the addition unit 127 adds the prediction block 177 and the difference block 174 to generate a decoded block 175 (S130). This decoded block 175 is stored in the frame memory 128 and is used in prediction processes by the intra prediction unit 129 and the inter prediction unit 130.
[0080] Next, the variable length coding unit 125 codes prediction information 178 indicating the prediction mode used and the like (S131), and codes the transform coefficients 173 (S132).
[0081] Then, the process moves to the next code block (S133), and the encoding unit 102 repeats steps S124 to S133 until all code blocks in the picture have been processed (S134).
[0082] Then, the encoding unit 102 repeats steps S122 to S134 until processing of all pictures is completed (S135).
[0083] <Effects> As described above, the image coding device 100 according to this embodiment determines the coding structure based on the frame rate 151 and the transmission delay time limit value 152. As a result, when the frame rate 151 is high, the image coding device 100 can deepen the hierarchy without increasing the display delay time of the decoder or the transmission delay time of the encoder, thereby improving temporal scalability. Furthermore, increasing the number of B pictures makes it possible to improve compression performance.
[0084] Furthermore, even in the case of various frame rates, it is possible to prevent the display delay time of the specified decoder and the transmission delay time of the encoder from being exceeded.
[0085] This will be explained in more detail. Fig. 10 is a diagram showing the number of layers 161, the number of display delay pictures 164, and the number of consecutive B pictures 162, which are calculated based on a frame rate 151 and a transmission delay time limit value 152. Fig. 10 also shows an example in which the transmission delay time limit value is set to 4 / 30 seconds.
[0086] 11A to 11D are diagrams showing coding structures based on the conditions of FIG. 10. FIG. 11A shows the structure when the frame rate is 24 fps. FIG. 11B shows the structure when the frame rate is 30 fps. FIG. 11C shows the structure when the frame rate is 60 fps. FIG. 11D shows the structure when the frame rate is 120 fps.
[0087] 9A to 9D show the numbers of pictures with transmission delay in the coded stream when the frame rates are 24 fps, 30 fps, 60 fps, and 120 fps, respectively. Also, FIGS. 12A to 12D show the numbers of pictures with display delay when the frame rates are 24 fps, 30 fps, 60 fps, and 120 fps.
[0088] As shown in Fig. 10, the transmission delay time does not exceed the limit of 4 / 30 seconds at any frame rate. Furthermore, in the current operational regulations (Non-Patent Document 2), the coding structure is limited to three layers as shown in Fig. 11B. That is, the number of display delay pictures is limited to two as shown in Fig. 12B, and the number of transmission delay pictures in a code stream is limited to four as shown in Fig. 9B. Furthermore, in the case of 30 fps, the display delay time is 2 / 30 seconds and the transmission delay time is 4 / 30 seconds. In this embodiment, when the transmission delay time limit value is set to 4 / 30 seconds, even if the number of layers is increased or decreased depending on the frame rate, the transmission delay time does not exceed 4 / 30 seconds and the display delay time does not exceed 2 / 30 seconds.
[0089] Furthermore, in this embodiment, the image coding device 100 determines the coding structure using a limit value for the transmission delay time, rather than the display delay time. By limiting the coding structure using the transmission delay time in this way, it is possible to determine the coding structure so that it does not exceed both the display delay of 2 / 30 seconds and the transmission delay of 4 / 30 seconds specified in the current operational regulations (Non-Patent Document 2). More specifically, if the number of display delay pictures is determined so as not to exceed the display delay of 2 / 30 seconds specified in the current operational regulations (Non-Patent Document 2), the number of display delay pictures at 120 fps is 8 (8 / 120 seconds), and a coding structure of up to 9 layers is permitted. However, if a 9-layer coding structure is used, the number of transmission delay pictures will be 256 (256 / 120 seconds), which far exceeds the transmission delay of 4 / 30 seconds specified in the current operational regulations (Non-Patent Document 2). On the other hand, if we focus on the transmission delay time and determine the number of transmission delay pictures so that the transmission delay time does not exceed 4 / 30 seconds, the number of transmission delay pictures is limited to 16 (16 / 120 seconds) at 120 fps, and the coding structure is limited to 5 layers. In this case, the display delay time does not exceed 2 / 30 seconds. By limiting the transmission delay time in this way, both the transmission delay time and the display delay time can be appropriately limited.
[0090] Furthermore, the image coding device 100 sets a limit value for the coding structure for each layer, which makes it possible to prevent the display delay time from exceeding a specified time even in an image decoding device that decodes only pictures in layers with a small TemporalId.
[0091] In the above description, the image coding device 100 calculates coding structure limit values such as the number of layers using a formula, but the table shown in Fig. 10 may be stored in memory in advance, and the coding structure limit values may be set according to the frame rate 151 and the transmission delay time limit value 152 by referring to the table. The image coding device 100 may also use both a table and a formula. For example, the image coding device 100 may set the coding structure limit value using a table when the frame rate is 24 fps or less, and may set the coding structure limit value using a formula when the frame rate is more than 24 fps.
[0092] In the above description, the image coding device 100 uses the output delay time limit value 152 and the frame rate 151 input from an external source, but this is not limiting. For example, the image coding device 100 may use a predetermined fixed value as at least one of the output delay time limit value 152 and the frame rate 151. Furthermore, the image coding device 100 may determine at least one of the output delay time limit value 152 and the frame rate 151 depending on the internal state of a buffer memory or the like.
[0093] 11A to 11D are merely examples, and the present invention is not limited to these. For example, the arrows indicating reference images are not limited to these, and each picture may use as a reference image any picture with a Temporal Id greater than its own Temporal Id. For example, image B1 shown in FIG. 11B may use image P4 as a reference image.
[0094] Furthermore, the above-mentioned limit values of the coding structure (number of layers, number of consecutive B pictures, and number of display latency pictures) are merely maximum values, and values smaller than the limit values may be used depending on the situation. For example, in FIG. 10, when the frame rate is 30 fps, the number of layers is 3, the number of consecutive B pictures is 3 [2], and the number of display latency pictures is 2 [2], and an coding structure such as that shown in FIG. 11B is shown. However, the number of layers may be 3 or less, and the number of consecutive B pictures and the number of display latency pictures may be values corresponding to the number of layers of 3 or less. For example, the number of layers may be 2, the number of consecutive B pictures may be 2 [1], and the number of display latency pictures may be 2 [2]. In this case, for example, the coding structure shown in FIG. 11A is used. In this case, in the coding structure coding in step S121 shown in FIG. 8, information indicating the coding structure used is coded.
[0095] In the above description, sps_max_num_reorder_pics is set to the number of display delay pictures, but sps_max_num_reorder_pics may be a variable indicating the number of pictures whose order is to be changed. For example, in the example shown in Fig. 9C, input image 8, input image 4, and input image 2 are rearranged and coded so that they are positioned before their positions in the input order (display order). In this case, the number of pictures whose order is to be changed is 3, and this value 3 may be set to sps_max_num_reorder_pics.
[0096] In the above description, the consecutive number parameter is set to sps_max_latency_increase_plus1, and the value of sps_max_num_reorder_pics+sps_max_latency_increase_plus1-1 (SpsMaxLatencyPictures) is treated as the number of consecutive B pictures. However, SpsMaxLatencyPictures may indicate the maximum number of picture decodes, which is the number of pictures decoded after a picture is decoded and stored in the buffer until it becomes displayable. For example, in the case of image P4 in FIG. 12B, after decoding of image P4 is completed, three images, image B2, image B1, and image B3, are decoded, and then image P4 becomes displayable. Furthermore, image B2, image B3, and image P4 are displayed in order. This maximum number of picture decodes, 3, may be set as SpsMaxLatencyPictures.
[0097] In addition, in this embodiment, sps_max_num_reorder_pics and sps_max_latency_increase_plus1 are set for each layer and coded, but this is not limited to this. For example, in a system that does not use temporal scalability, only the values of sps_max_num_reorder_pics and sps_max_latency_increase_plus1 for the deepest layer (the layer with the largest TemporalId) may be set and coded.
[0098] In the above description, the frame rates are 24 fps, 30 fps, 60 fps, and 120 fps, but other frame rates may also be used. The frame rate may also be a numerical value including a decimal point, such as 29.97 fps.
[0099] Furthermore, the processing in this embodiment may be realized by software. This software may be distributed by download or the like. This software may also be recorded on a recording medium such as a CD-ROM and distributed. This also applies to the other embodiments in this specification.
[0100] (Embodiment 2) In this embodiment, an image decoding apparatus corresponding to the image coding apparatus described in the first embodiment will be described.
[0101] <Overall structure> FIG. 13 is a block diagram showing the configuration of an image decoding device 200 according to this embodiment.
[0102] 13 generates an output image 263 by decoding a code string 251. The code string 251 is, for example, the code string 155 generated by the image coding device 100 in Embodiment 1. The image decoding device 200 includes a variable length decoding unit 201, an inverse transform and quantization unit 202, an adder 203, a frame memory 204, an intra-prediction block generation unit 205, an inter-prediction block generation unit 206, a limit value decoding unit 208, an image rearrangement unit 209, and a coding structure confirmation unit 210.
[0103] <Movement (overall)> Next, the image decoding process according to this embodiment will be described with reference to FIG.
[0104] First, the variable length decoding unit 201 decodes the coding structure limiting value 257 from the codestream 251. This coding structure limiting value 257 includes sps_max_sub_layers_minus1, sps_max_num_reorder_pics, and sps_max_latency_increase_plus1. Note that the meaning of these pieces of information is the same as in Embodiment 1. Next, the limiting value decoding unit 208 obtains the number of layers by adding 1 to sps_max_sub_layers_minus1, obtains the number of consecutive B pictures using the formula sps_max_num_reorder_pics+sps_max_latency_increase_plus1-1, and obtains the number of display latency pictures using sps_max_num_reorder_pics (S201). Furthermore, the limit value decoding unit 208 obtains the coding structure 262 (the number of layers, the number of display latency pictures, and the number of consecutive B-pictures) from sps_max_num_reorder_pics and sps_max_latency_increase_plus1 of the layer of TemporalId corresponding to the value of HighestTId 252 input from outside, and outputs the obtained coding structure 262 to the image rearrangement unit 209 and the coding structure confirmation unit 210. Here, HighestTId 252 indicates the TemporalId of the highest layer to be decoded.
[0105] Next, the coding structure checking unit 210 checks whether each value of the coding structure 262 complies with the operating regulations (S202). Specifically, the coding structure checking unit 210 calculates each limit value according to the following (Equation 7) to (Equation 9) using an externally input transmission delay time limit value 253 and a frame rate 256 obtained by variable-length decoding the code string 251, and determines whether the coding structure is equal to or less than the calculated limit value.
[0106] Number of layers = int(log2(transmission delay time limit [s] × frame rate [fps])) + 1 (Formula 7)
[0107] Number of display delay pictures [TId] = int(log2(output delay time limit [s] × frame rate [fps] ÷ 2 (n-TId))) ...(Formula 8)
[0108] Number of consecutive B pictures [TId] = int (transmission delay time limit [s] × frame rate [fps] ÷ 2 (n-TId) -1)...(Formula 9)
[0109] If the coding structure is larger than the limit value (Yes in S203), the coding structure confirmation unit 210 displays an error message to that effect (S204) and ends the decoding process.
[0110] Next, the variable length decoding unit 201 decodes prediction information 255 indicating the prediction mode from the code sequence 251 (S205). If the prediction mode is intra prediction (Yes in S206), the intra prediction block generation unit 205 generates a prediction block 261 by intra prediction (S207). On the other hand, if the prediction mode is inter prediction (No in S206), the inter prediction block generation unit 206 generates a prediction block 261 by inter prediction (S208).
[0111] Next, the variable length decoding unit 201 decodes the transform coefficients 254 from the code string 251 (S209). Next, the inverse transform and quantization unit 202 reconstructs the difference block 258 by performing inverse quantization and inverse frequency transform on the transform coefficients 254 (S210). Next, the adder 203 generates a decoded block 259 by adding the difference block 258 and the prediction block 261 (S211). This decoded block 259 is stored in the frame memory 204 and is used in the prediction block generation process by the intra prediction block generation unit 205 and the inter prediction block generation unit 206.
[0112] Then, the image decoding device 200 moves the process to the next code block (S212), and repeats steps S205 to S212 until all code blocks in the picture have been processed (S213).
[0113] The processes in steps S205 to S212 are performed only on pictures having a TemporalId equal to or less than HighestTId252 that are input from the outside.
[0114] Next, the image rearrangement unit 209 rearranges the decoded pictures in accordance with the coding structure 262 of the layer of HighestTId 252 input from the outside, and outputs the rearranged decoded pictures as output images 263 (S214).
[0115] Then, the image decoding device 200 repeats steps S205 to S214 until processing of all pictures is completed (S215).
[0116] <Effects> As described above, the image decoding device 200 according to this embodiment can decode a code string generated by efficient encoding. Furthermore, the image decoding device 200 can check whether the coding structure complies with the operating rules, and if it does not comply, can stop the decoding process and display an error message.
[0117] In the above description, the image decoding device 200 is configured to decode only pictures in a layer equal to or lower than HighestTId252 in accordance with the externally input HighestTId252, but this is not limited to this. The image decoding device 200 may always decode pictures in all layers. Alternatively, the image decoding device 200 may use a predetermined fixed value as HighestTId252 and always decode only pictures in a layer equal to or lower than the predetermined layer indicated by HighestTId252.
[0118] Furthermore, in the above description, the image decoding device 200 checks whether the coding structure 262 complies with the operating rules, but this function is not essential, and the coding structure 262 does not have to be checked.
[0119] In the above description, the image decoding device 200 uses the output delay time limit value 253 input from outside, but the output delay time limit value 253 may be a predetermined fixed value.
[0120] The rest is the same as in the first embodiment, so a detailed description is omitted.
[0121] The order of each flow is not limited to the above, as with the encoding side.
[0122] As described above in the first and second embodiments, the image coding device 100 according to the first embodiment is an image coding device that generates a code sequence 155 (bit stream) by hierarchically coding an input image 153, and performs the processing shown in FIG. 15.
[0123] First, the image coding device 100 determines the number of layers 161 in hierarchical coding so that the number of layers 161 is equal to or less than a maximum number of layers predetermined according to the frame rate (S301). Here, the maximum number of layers is the number of layers shown in Fig. 10, and is, for example, 2 when the frame rate is 24 fps, 3 when the frame rate is 30 fps, 4 when the frame rate is 60 fps, and 5 when the frame rate is 120 fps. In other words, when the frame rate is 60 fps or more, the maximum number of layers is 4 or more. Also, when the frame rate is 60 fps or less, the maximum number of layers is 4 or less. Also, when the frame rate is higher than 30 fps, the maximum number of layers is greater than 3.
[0124] The image encoding device 100 further determines the picture type of the input image 153 so that the display latency picture count 164 is equal to or less than a maximum picture count predetermined according to the frame rate. Here, the display latency picture count 164 refers to the number of pictures from when the image decoding device starts decoding images until the image is output (displayed) when the image decoding device decodes the code string 155 generated by the image encoding device 100. The picture type refers to an I-picture, a P-picture, or a B-picture. Here, the maximum picture count is the display latency picture count shown in FIG. 10 , and is, for example, 1 when the frame rate is 24 fps, 2 when the frame rate is 30 fps, 3 when the frame rate is 60 fps, and 4 when the frame rate is 120 fps. In other words, when the frame rate is 60 fps or higher, the maximum picture count is 3 or more. Also, when the frame rate is 60 fps or lower, the maximum picture count is 3 or less. Also, if the frame rate is greater than 30 fps, the maximum number of pictures is greater than 2.
[0125] Furthermore, the image coding device 100 determines the picture type of the input image 153 so that the number of consecutive B pictures, that is, the B-picture consecutive number 162, is equal to or less than a maximum consecutive number predetermined depending on the frame rate. Here, the maximum consecutive number is the display delay picture number shown in FIG. 10 , and is, for example, 2 when the frame rate is 24 fps, 3 when the frame rate is 30 fps, 7 when the frame rate is 60 fps, and 15 when the frame rate is 120 fps. In other words, when the frame rate is 60 fps or higher, the maximum consecutive number is 7 or more. When the frame rate is 60 fps or lower, the maximum consecutive number is 7 or less. When the frame rate is higher than 30 fps, the maximum consecutive number is greater than 3.
[0126] Furthermore, the image coding device 100 may determine the maximum number of layers, the maximum number of pictures, and the number of consecutive B pictures depending on the frame rate, as shown in Fig. 10. That is, the image coding device 100 may set the maximum number of layers, the maximum number of pictures, and the number of consecutive B pictures to be larger as the frame rate increases.
[0127] As described above, the number of layers 161, the number of display delay pictures 164, and the number of consecutive B pictures 162 are calculated by the above (Equation 1), (Equation 3), and (Equation 4) using the frame rate 151 and the transmission delay time limit value 152. In other words, the maximum number of pictures, the encoder transmission delay (transmission delay time) which is the time from when an input image 153 is input to the image encoding device 100 until a code string 155 is output, and the frame rate satisfy the following relationship:
[0128] Maximum number of pictures = int(log2(encoder transmission delay [s] × frame rate [fps]))
[0129] Furthermore, the maximum number of consecutive frames, the encoder transmission delay, and the frame rate satisfy the following relationship.
[0130] Maximum number of consecutive frames = int (encoder transmission delay [s] x frame rate [fps] - 1)
[0131] The maximum number of layers, the encoder output delay, and the frame rate satisfy the following relationship.
[0132] Maximum number of layers = int(log2(encoder transmission delay [s] × frame rate [fps])) + 1
[0133] Furthermore, the maximum number of pictures [i] in each layer, the encoder transmission delay, and the frame rate satisfy the following relationship:
[0134] Maximum number of pictures [i] = int(log2(encoder transmission delay [s] × frame rate [fps] / 2 (n-i) ))
[0135] The maximum number of consecutive frames [i] in each layer, the encoder transmission delay, and the frame rate satisfy the following relationship:
[0136] Maximum number of consecutive frames [i] = int (encoder transmission delay [s] × frame rate [fps] / 2 (n-i) -1)
[0137] Here, i is an integer equal to or less than the maximum number of layers and indicates a layer, and n indicates (maximum number of layers - 1).
[0138] Next, the image coding device 100 generates a code string 155 by hierarchically coding the input image 153 using the determined number of layers 161 and picture type (S302). The image coding device 100 also codes first information (sps_max_sub_layers_minus1), second information (sps_max_num_reorder_pics), and third information (sps_max_latency_increase_plus1) indicating the determined number of layers 161, the number of display latency pictures 164, and the number of consecutive B pictures 162.
[0139] Moreover, the image decoding device 200 according to the second embodiment is an image decoding device that generates an output image 263 by decoding a code sequence 251 (bit stream) obtained by hierarchical coding of an image, and performs the processing shown in FIG. 16 .
[0140] First, the image decoding device 200 decodes an image from the codestream 251 (S401).
[0141] Next, the image decoding device 200 decodes, from the codestream 251, first information (sps_max_sub_layers_minus1) indicating the number of layers in hierarchical encoding (S402). For example, this number of layers is equal to or less than the maximum number of layers determined in advance according to the frame rate of the codestream 251.
[0142] Furthermore, the image decoding device 200 decodes second information (sps_max_num_reorder_pics) indicating the number of display latency pictures from the codestream 251. Furthermore, the image decoding device 200 also decodes third information (sps_max_latency_increase_plus1) indicating the number of consecutive B pictures from the codestream 251.
[0143] Next, the image decoding device 200 rearranges and outputs the decoded images using the number of layers indicated by the first information, the number of display delay pictures indicated by the second information, and the number of consecutive B pictures indicated by the third information (S403).
[0144] Note that specific examples and limitations of the maximum number of layers, the maximum number of pictures, and the maximum number of consecutive B-pictures are the same as those in the image coding device 100. The relationships between the maximum number of layers, the maximum number of pictures, and the maximum number of consecutive B-pictures are also the same as those in the image coding device 100.
[0145] Although the image decoding device and image coding device according to the embodiments have been described above, the present invention is not limited to these embodiments.
[0146] Furthermore, each processing unit included in the image decoding device or image coding device according to the above embodiments is typically realized as an LSI, which is an integrated circuit. These units may be individually implemented as single chips, or some or all of them may be integrated into a single chip.
[0147] Furthermore, the integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI fabrication, or reconfigurable processors, which allow the connections and settings of circuit cells within LSIs to be reconfigured, may also be used.
[0148] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may also be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0149] In other words, the image decoding device and the image encoding device include processing circuitry and storage electrically connected to (accessible from) the processing circuitry. The processing circuitry includes at least one of dedicated hardware and a program execution unit. Furthermore, if the processing circuitry includes a program execution unit, the storage unit stores a software program to be executed by the program execution unit. The processing circuitry uses the storage unit to execute the image decoding method or image encoding method according to the above-described embodiment.
[0150] Furthermore, the present invention may be the above-mentioned software program, or a non-transitory computer-readable recording medium on which the above-mentioned program is recorded. Needless to say, the above-mentioned program can be distributed via a transmission medium such as the Internet.
[0151] Furthermore, all the numbers used above are merely examples for the purpose of specifically explaining the present invention, and the present invention is not limited to the numbers used as examples.
[0152] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in time-sharing by a single piece of hardware or software.
[0153] Furthermore, the order in which the steps included in the image decoding method or the image encoding method are executed is merely an example for specifically explaining the present invention, and an order other than the above may be used. Furthermore, some of the steps may be executed simultaneously (in parallel) with other steps.
[0154] Furthermore, the processing described in the above embodiments may be realized by centralized processing using a single device (system), or by distributed processing using multiple devices. Furthermore, the computer that executes the above program may be a single computer or multiple computers. In other words, centralized processing or distributed processing may be performed.
[0155] Furthermore, the present invention is particularly effective in cases where the receiving terminals have diverse capabilities, such as broadcasting to many end users. For example, a signal with the above-described data structure is broadcast. A terminal such as a 4K2K television can decode data at all levels. On the other hand, a smartphone can decode up to two levels. Furthermore, depending on the bandwidth congestion situation, the transmitting device can transmit only the upper levels, rather than the full levels. This enables flexible broadcasting and communication.
[0156] While the image decoding device and the image encoding device according to one or more aspects of the present invention have been described based on the embodiments, the present invention is not limited to these embodiments. As long as they do not deviate from the spirit of the present invention, various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining components of different embodiments may also be included within the scope of one or more aspects of the present invention.
[0157] (Embodiment 3) By recording a program for implementing the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments on a storage medium, it becomes possible to easily perform the processes shown in each of the above embodiments on an independent computer system. The storage medium may be a magnetic disk, optical disk, magneto-optical disk, IC card, semiconductor memory, or any other medium capable of recording a program.
[0158] Furthermore, here, we will explain application examples of the video coding method (image coding method) and video decoding method (image decoding method) shown in each of the above embodiments, and a system using the same. The system is characterized by having an image coding / decoding device consisting of an image coding device using the image coding method and an image decoding device using the image decoding method. Other components of the system can be appropriately changed depending on the situation.
[0159] 17 is a diagram showing the overall configuration of a content supply system ex100 that provides a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.
[0160] This content supply system ex100 is connected to the Internet ex101 via an Internet service provider ex102, a telephone network ex104, and base stations ex106 to ex110, and devices such as a computer ex111, a PDA (Personal Digital Assistant) ex112, a camera ex113, a mobile phone ex114, and a game console ex115.
[0161] However, the content supply system ex100 is not limited to the configuration shown in Fig. 17, and any combination of elements may be connected. Also, each device may be directly connected to the telephone network ex104 without going through base stations ex106 to ex110, which are fixed wireless stations. Also, each device may be directly connected to each other via short-range wireless or the like.
[0162] The camera ex113 is a device capable of shooting moving images, such as a digital video camera, and the camera ex116 is a device capable of shooting still images and moving images, such as a digital camera. The mobile phone ex114 may be any of a GSM (registered trademark) (Global System for Mobile Communications) system, a CDMA (Code Division Multiple Access) system, a W-CDMA (Wideband-Code Division Multiple Access) system, an LTE (Long Term Evolution) system, an HSPA (High Speed Packet Access) mobile phone, or a PHS (Personal Handyphone System) system.
[0163] In the content supply system ex100, a camera ex113 and the like are connected to a streaming server ex103 via a base station ex109 and a telephone network ex104, thereby enabling live streaming and the like. In live streaming, a user shoots content (e.g., video of a live music concert) using the camera ex113, and encodes the content as described in the above embodiments (i.e., functions as an image encoding device according to an aspect of the present invention) and transmits the content to the streaming server ex103. Meanwhile, the streaming server ex103 streams the transmitted content data to a requesting client. Examples of clients include a computer ex111, a PDA ex112, a camera ex113, a mobile phone ex114, a game console ex115, and the like that are capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data (i.e., functions as an image decoding device according to an aspect of the present invention).
[0164] The encoding process of the captured data may be performed by the camera ex113, by the streaming server ex103 that processes the data transmission, or by a mutually shared responsibility. Similarly, the decoding process of the distributed data may be performed by the client, by the streaming server ex103, or by a mutually shared responsibility. Furthermore, still images and / or video data captured by camera ex116, not limited to camera ex113, may be transmitted to the streaming server ex103 via computer ex111. In this case, the encoding process may be performed by the camera ex116, the computer ex111, or the streaming server ex103, or by a mutually shared responsibility.
[0165] Furthermore, these encoding and decoding processes are generally performed by the computer ex111 or an LSIex500 possessed by each device. The LSIex500 may be a single chip or may be configured with multiple chips. It is also possible to embed video encoding and decoding software on some kind of recording medium (CD-ROM, flexible disk, hard disk, etc.) that can be read by the computer ex111, etc., and perform the encoding and decoding processes using that software. Furthermore, if the mobile phone ex114 is equipped with a camera, video data captured by the camera may be transmitted. This video data is data that has been encoded and processed by the LSIex500 possessed by the mobile phone ex114.
[0166] The streaming server ex103 may also be a plurality of servers or computers that process, record, and distribute data in a distributed manner.
[0167] In this way, the content delivery system ex100 allows a client to receive and play back encoded data. In this way, the content delivery system ex100 allows a client to receive, decode, and play back information sent by a user in real time, enabling even users without special rights or equipment to realize personal broadcasting.
[0168] In addition to the example of the content supply system ex100, as shown in FIG. 18, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) according to each of the above embodiments can also be incorporated into a digital broadcasting system ex200. Specifically, a broadcasting station ex201 communicates via radio waves multiplexed data in which music data and the like are multiplexed onto video data, or transmits the multiplexed data to a satellite ex202. This video data is data encoded using the video encoding method described in each of the above embodiments (i.e., data encoded by an image encoding device according to one aspect of the present invention). Receiving this, the broadcasting satellite ex202 transmits broadcasting radio waves, which are received by a home antenna ex204 capable of receiving satellite broadcasts. The received multiplexed data is decoded and played back by a device such as a television (receiver) ex300 or a set-top box (STB) ex217 (i.e., functions as an image decoding device according to one aspect of the present invention).
[0169] The video decoding device or video encoding device described in each of the above embodiments can also be implemented in a reader / recorder ex218 that reads and decodes multiplexed data recorded on a recording medium ex215 such as a DVD or Blu-ray, or encodes a video signal onto the recording medium ex215 and, in some cases, multiplexes it with an audio signal before writing it. In this case, the reproduced video signal is displayed on a monitor ex219, and the video signal can be reproduced in another device or system using the recording medium ex215 on which the multiplexed data is recorded. Alternatively, a video decoding device may be implemented in a set-top box ex217 connected to a cable television cable ex203 or a satellite / terrestrial broadcast antenna ex204, and the video may be displayed on the television monitor ex219. In this case, the video decoding device may be incorporated into the television rather than the set-top box.
[0170] 19 is a diagram showing a television (receiver) ex300 that uses the video decoding method and video encoding method described in each of the above embodiments. The television ex300 includes a tuner ex301 that acquires or outputs multiplexed data in which audio data is multiplexed onto video data via an antenna ex204 that receives the broadcasts or a cable ex203, a modulation / demodulation unit ex302 that demodulates the received multiplexed data or modulates it into multiplexed data to be transmitted externally, and a multiplexing / demultiplexing unit ex303 that separates the demodulated multiplexed data into video data and audio data or multiplexes the video data and audio data encoded by a signal processing unit ex306.
[0171] The television ex300 also has a signal processing unit ex306 having an audio signal processing unit ex304 and a video signal processing unit ex305 (which function as an image encoding device or an image decoding device according to an embodiment of the present invention) that decode the audio data and the video data, respectively, or encode the respective information, and an output unit ex309 having a speaker ex307 that outputs the decoded audio signal and a display unit ex308 such as a display that displays the decoded video signal.The television ex300 also has an interface unit ex317 that has an operation input unit ex312 that accepts user operation input, etc.The television ex300 also has a control unit ex310 that controls each unit overall, and a power supply circuit unit ex311 that supplies power to each unit. In addition to the operation input unit ex312, the interface unit ex317 may have a bridge ex313 connected to an external device such as a reader / recorder ex218, a slot unit ex314 for allowing a recording medium ex216 such as an SD card to be attached, a driver ex315 for connecting to an external recording medium such as a hard disk, a modem ex316 for connecting to a telephone network, etc. The recording medium ex216 is a non-volatile / volatile semiconductor memory element that stores information and allows it to be electrically recorded. The various units of the television ex300 are connected to each other via a synchronous bus.
[0172] First, a configuration in which the television ex300 decodes and plays back multiplexed data acquired from an external source via an antenna ex204 or the like will be described. The television ex300 receives user operation via a remote controller ex220 or the like, and, under the control of a control unit ex310 having a CPU or the like, separates the multiplexed data demodulated by a modulation / demodulation unit ex302 in a multiplexing / separation unit ex303. The television ex300 then decodes the separated audio data in an audio signal processing unit ex304 and decodes the separated video data in a video signal processing unit ex305 using the decoding method described in each of the above embodiments. The decoded audio and video signals are output to the outside from an output unit ex309. When outputting, it is preferable to temporarily store these signals in buffers ex318, ex319, or the like so that the audio and video signals are played back in sync. The television ex300 may also read the multiplexed data from recording media ex215, ex216, such as magnetic / optical discs or SD cards, rather than from broadcasts or the like. Next, a configuration in which the television ex300 encodes audio and video signals and transmits them externally or writes them to a recording medium or the like will be described. The television ex300 receives user operation from a remote controller ex220 or the like, and, under the control of the control unit ex310, encodes the audio signal in the audio signal processing unit ex304 and encodes the video signal in the video signal processing unit ex305 using the encoding method described in each of the above embodiments. The encoded audio and video signals are multiplexed by the multiplexing / demultiplexing unit ex303 and output externally. When multiplexing, these signals may be temporarily stored in buffers ex320, ex321, etc., so that the audio and video signals are synchronized. Note that multiple buffers ex318, ex319, ex320, and ex321 may be provided as shown, or one or more buffers may be shared. Furthermore, data may be stored in buffers other than those shown in the figure, for example, between the modulation / demodulation unit ex302 and the multiplexing / demultiplexing unit ex303, as a buffer to prevent system overflow and underflow.
[0173] Furthermore, in addition to acquiring audio data and video data from broadcasts, recording media, etc., the television ex300 may also be configured to accept AV input from a microphone or camera and perform encoding processing on the data acquired from them. Note that while the television ex300 has been described here as being configured to be capable of the above encoding processing, multiplexing, and external output, it may also be configured not to be able to perform these processes and only be capable of the above reception, decoding processing, and external output.
[0174] Furthermore, when multiplexed data is read from or written to a recording medium using the reader / recorder ex218, the above-mentioned decoding or encoding process may be performed by either the television ex300 or the reader / recorder ex218, or the television ex300 and the reader / recorder ex218 may share the process.
[0175] As an example, Figure 20 shows the configuration of the information reproducing / recording unit ex400 when reading or writing data from an optical disc. The information reproducing / recording unit ex400 includes the following elements ex401, ex402, ex403, ex404, ex405, ex406, and ex407. The optical head ex401 writes information by irradiating a laser spot onto the recording surface of the recording medium ex215, which is an optical disc, and reads the information by detecting the light reflected from the recording surface of the recording medium ex215. The modulation / recording unit ex402 electrically drives the semiconductor laser built into the optical head ex401 and modulates the laser light according to the recorded data. The reproduction / demodulation unit ex403 amplifies the reproduction signal obtained by electrically detecting the light reflected from the recording surface using a photodetector built into the optical head ex401, and separates and demodulates the signal components recorded on the recording medium ex215 to reproduce the required information. The buffer ex404 temporarily stores information to be recorded on the recording medium ex215 and information reproduced from the recording medium ex215. The disk motor ex405 rotates the recording medium ex215. The servo control unit ex406 controls the rotation of the disk motor ex405, moves the optical head ex401 to a specified information track, and performs laser spot tracking. The system control unit ex407 controls the entire information reproduction / recording unit ex400. The system control unit ex407 performs the above read and write processes by using various information stored in the buffer ex404, generating and adding new information as needed, and recording and reproducing information through the optical head ex401 while coordinating the modulation recording unit ex402, reproduction demodulation unit ex403, and servo control unit ex406. The system control unit ex407 is composed of, for example, a microprocessor and performs these processes by executing read and write programs.
[0176] In the above description, the optical head ex401 is described as irradiating a laser spot, but it may be configured to perform higher density recording using near-field light.
[0177] FIG. 21 shows a schematic diagram of recording medium ex215, an optical disc. A spiral guide groove is formed on the recording surface of recording medium ex215, and address information indicating absolute positions on the disc is recorded in advance on information track ex230 by varying the shape of the groove. This address information includes information for identifying the position of recording block ex231, which is the unit of data recording. A recording or playback device can identify a recording block by reproducing information track ex230 and reading the address information. Recording medium ex215 also includes a data recording area ex233, an inner peripheral area ex232, and an outer peripheral area ex234. The data recording area ex233 is the area used for recording user data, while the inner peripheral area ex232 and outer peripheral area ex234, which are located either inner or outer than data recording area ex233, are used for specific purposes other than recording user data. The information reproducing / recording unit ex400 reads and writes encoded audio data, video data, or multiplexed data obtained by multiplexing these data, from the data recording area ex233 of such recording medium ex215.
[0178] The above explanation has been given using examples of optical discs such as single-layer DVDs and BDs, but the present invention is not limited to these and may be an optical disc with a multi-layer structure that allows recording on areas other than the surface. It may also be an optical disc with a structure that allows multidimensional recording / playback, such as recording information using light of various different wavelengths in the same location on the disc or recording different layers of information from various angles.
[0179] In the digital broadcasting system ex200, a car ex210 equipped with an antenna ex205 can receive data from a satellite ex202 or the like, and the video can be played on a display device such as a car navigation system ex211 installed in the car ex210. The car navigation system ex211 can be configured, for example, as shown in Fig. 19, with a GPS receiver added, and similar configurations can be considered for a computer ex111, a mobile phone ex114, and the like.
[0180] 22A is a diagram showing a mobile phone ex114 that uses the video decoding method and video encoding method described in the above embodiment. The mobile phone ex114 includes an antenna ex350 for transmitting and receiving radio waves to and from base station ex110, a camera unit ex365 capable of capturing video and still images, and a display unit ex358 such as an LCD display that displays decoded data of video captured by the camera unit ex365 and video received by the antenna ex350. The mobile phone ex114 further includes a main body unit having an operation key unit ex366, an audio output unit ex357 such as a speaker for outputting audio, an audio input unit ex356 such as a microphone for inputting audio, a memory unit ex367 for storing captured video, still images, recorded audio, or encoded or decoded data of received video, still images, email, etc., or a slot unit ex364 that serves as an interface with a recording medium for similarly storing data.
[0181] Furthermore, a configuration example of mobile phone ex114 will be described with reference to Fig. 22B. Mobile phone ex114 has a main control unit ex360 that comprehensively controls each unit of a main body unit including a display unit ex358 and an operation key unit ex366, and a power supply circuit unit ex361, an operation input control unit ex362, a video signal processing unit ex355, a camera interface unit ex363, an LCD (Liquid Crystal Display) control unit ex359, a modulation / demodulation unit ex352, a multiplexing / demultiplexing unit ex353, an audio signal processing unit ex354, a slot unit ex364, and a memory unit ex367, which are all connected to each other via a bus ex370.
[0182] When the end call and power key is turned on by the user, the power supply circuit unit ex361 starts up the mobile phone ex114 into an operable state by supplying power to each unit from the battery pack.
[0183] Based on the control of a main control unit ex360 having a CPU, ROM, RAM, etc., the mobile phone ex114 converts an audio signal collected by an audio input unit ex356 into a digital audio signal by an audio signal processing unit ex354 in a voice call mode, which undergoes spectrum spread processing by a modulation / demodulation unit ex352, digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit ex351, and then transmits the digital audio signal via an antenna ex350. Furthermore, the mobile phone ex114 amplifies received data received via the antenna ex350 in a voice call mode, performs frequency conversion processing and analog-to-digital conversion processing, performs spectrum despread processing by the modulation / demodulation unit ex352, converts the data into an analog audio signal by the audio signal processing unit ex354, and then outputs the data from an audio output unit ex357.
[0184] Furthermore, when sending an e-mail in data communication mode, the text data of the e-mail entered by operating the operation key unit ex366 or the like of the main unit is sent to the main control unit ex360 via the operation input control unit ex362. The main control unit ex360 performs spectrum spread processing on the text data in the modulation / demodulation unit ex352, performs digital-to-analog conversion processing and frequency conversion processing in the transmission / reception unit ex351, and then transmits the data to the base station ex110 via the antenna ex350. When receiving an e-mail, the received data is subjected to roughly the reverse processing and output to the display unit ex358.
[0185] When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex355 compresses and encodes the video signal supplied from the camera unit ex365 using the video encoding method described in each of the above embodiments (i.e., functions as an image encoding device according to one aspect of the present invention), and sends the encoded video data to the multiplexing / separating unit ex353. In addition, the audio signal processing unit ex354 encodes the audio signal collected by the audio input unit ex356 while the camera unit ex365 is capturing video, still images, etc., and sends the encoded audio data to the multiplexing / separating unit ex353.
[0186] The multiplexing / separation unit ex353 multiplexes the encoded video data supplied from the video signal processing unit ex355 and the encoded audio data supplied from the audio signal processing unit ex354 using a predetermined method, and the resulting multiplexed data is subjected to spectrum spreading processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex352, digital-to-analog conversion processing and frequency conversion processing in the transmission / reception unit ex351, and then transmitted via the antenna ex350.
[0187] When receiving video file data linked to a website or the like in data communication mode, or when receiving an email with video and / or audio attachments, the multiplexer / demultiplexer ex353 decodes the multiplexed data received via the antenna ex350 into a video data bitstream and an audio data bitstream. The multiplexer / demultiplexer ex353 then decodes the multiplexed data into a video data bitstream and an audio data bitstream via a synchronization bus ex370. The video signal processor ex355 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments (i.e., functions as an image decoding device according to one aspect of the present invention). The display unit ex358 displays, via an LCD controller ex359, video and still images included in the video file linked to a website, for example. The audio signal processor ex354 decodes the audio signal, and audio is output from an audio output unit ex357.
[0188] Furthermore, like the television ex300, terminals such as the mobile phone ex114 can be implemented in three ways: a transmitting / receiving terminal with both an encoder and a decoder, a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. Furthermore, in the digital broadcasting system ex200, it has been explained that multiplexed data in which music data and the like are multiplexed onto video data is received and transmitted, but the data may also be multiplexed with text data related to the video in addition to audio data, or it may be video data itself rather than multiplexed data.
[0189] In this way, it is possible to use the video encoding method or video decoding method shown in each of the above embodiments in any of the above-mentioned devices and systems, and by doing so, it is possible to obtain the effects described in each of the above embodiments.
[0190] Furthermore, the present invention is not limited to the above-described embodiment, and various modifications and alterations are possible without departing from the scope of the present invention.
[0191] (Fourth embodiment) It is also possible to generate video data by switching between the video encoding method or device shown in each of the above embodiments and a video encoding method or device conforming to a different standard, such as MPEG-2, MPEG4-AVC, or VC-1, as needed.
[0192] When multiple pieces of video data conforming to different standards are generated, it is necessary to select a decoding method corresponding to each standard when decoding. However, since it is not possible to identify which standard the video data to be decoded conforms to, a problem arises in that it is not possible to select an appropriate decoding method.
[0193] To solve this problem, multiplexed data, which is video data multiplexed with audio data, etc., is configured to include identification information that indicates which standard the video data conforms to. A specific configuration of multiplexed data including video data generated by the video encoding methods or devices described in the above embodiments is described below. The multiplexed data is a digital stream in MPEG-2 transport stream format.
[0194] FIG. 23 shows the structure of multiplexed data. As shown in FIG. 23, the multiplexed data is obtained by multiplexing one or more of a video stream, an audio stream, a presentation graphics stream (PG), and an interactive graphics stream. The video stream represents the main video and secondary video of a movie, the audio stream (IG) represents the main audio portion of the movie and the secondary audio mixed with the main audio, and the presentation graphics stream represents the subtitles of the movie. Here, the main video refers to the normal video displayed on the screen, and the secondary video refers to the video displayed on a small screen within the main video. The interactive graphics stream represents an interactive screen created by arranging GUI components on the screen. The video stream is encoded using the video encoding method or device described in each of the above embodiments or a video encoding method or device conforming to conventional standards such as MPEG-2, MPEG4-AVC, or VC-1. The audio stream is encoded using a format such as Dolby AC-3, Dolby Digital Plus, MLP, DTS, DTS-HD, or Linear PCM.
[0195] Each stream included in the multiplexed data is identified by a PID. For example, 0x1011 is assigned to the video stream used for movie images, 0x1100 to 0x111F to the audio stream, 0x1200 to 0x121F to the presentation graphics, 0x1400 to 0x141F to the interactive graphics stream, 0x1B00 to 0x1B1F to the video stream used for movie secondary video, and 0x1A00 to 0x1A1F to the audio stream used for secondary audio to be mixed with the main audio.
[0196] 24 is a diagram showing how multiplexed data is multiplexed. First, a video stream ex235 consisting of multiple video frames and an audio stream ex238 consisting of multiple audio frames are converted into PES packet sequences ex236 and ex239, respectively, and then converted into TS packets ex237 and ex240. Similarly, presentation graphics stream ex241 and interactive graphics data ex244 are converted into PES packet sequences ex242 and ex245, respectively, and then converted into TS packets ex243 and ex246. Multiplexed data ex247 is constructed by multiplexing these TS packets into a single stream.
[0197] FIG. 25 shows in more detail how a video stream is stored in a PES packet sequence. The first row in FIG. 25 shows a video frame sequence of the video stream. The second row shows a PES packet sequence. As indicated by arrows yy1, yy2, yy3, and yy4 in FIG. 25, I-pictures, B-pictures, and P-pictures, which are multiple Video Presentation Units in the video stream, are divided into individual pictures and stored in the payload of a PES packet. Each PES packet has a PES header, which stores a Presentation Time-Stamp (PTS), which is the display time of the picture, and a Decoding Time-Stamp (DTS), which is the decoding time of the picture.
[0198] Figure 26 shows the format of the TS packet that is ultimately written to the multiplexed data. TS packets are 188-byte fixed-length packets consisting of a 4-byte TS header containing information such as a PID that identifies the stream, and a 184-byte TS payload that stores the data. The PES packets are divided and stored in the TS payload. In the case of BD-ROM, a 4-byte TP_Extra_Header is added to the TS packet, forming a 192-byte source packet that is written to the multiplexed data. The TP_Extra_Header contains information such as an ATS (Arrival Time Stamp). The ATS indicates the start time of the TS packet's transfer to the PID filter of the decoder. As shown in the lower part of Figure 26, source packets are arranged in the multiplexed data, and the number that increments from the beginning of the multiplexed data is called the SPN (Source Packet Number).
[0199] In addition to the individual streams (video, audio, subtitles, etc.), the TS packets contained in the multiplexed data also contain a Program Association Table (PAT), Program Map Table (PMT), and Program Clock Reference (PCR). The PAT indicates the PID of the PMT used in the multiplexed data, and the PAT's own PID is registered as 0. The PMT contains the PIDs of each stream (video, audio, subtitles, etc.) contained in the multiplexed data, as well as attribute information for the streams corresponding to each PID. It also contains various descriptors related to the multiplexed data. The descriptors include copy control information that indicates whether copying of the multiplexed data is permitted or prohibited. The PCR contains information about the Arrival Time Clock (ATC), which is the time axis of the ATS, and the System Time Clock (STC), which is the time axis of the PTS and DTS, and contains information about the STC time corresponding to the ATS at which the PCR packet is transferred to the decoder.
[0200] Figure 27 is a diagram explaining the data structure of a PMT in detail. At the beginning of a PMT is a PMT header that describes the length of the data contained in the PMT, among other things. This is followed by multiple descriptors related to the multiplexed data. The above-mentioned copy control information and other information are written as descriptors. After the descriptors are multiple stream information items related to each stream included in the multiplexed data. The stream information consists of stream descriptors that describe the stream type to identify the stream compression codec, the stream PID, and stream attribute information (frame rate, aspect ratio, etc.). There are as many stream descriptors as there are streams in the multiplexed data.
[0201] When recording on a recording medium, the multiplexed data is recorded together with a multiplexed data information file.
[0202] As shown in FIG. 28, the multiplexed data information file is management information for multiplexed data, has one-to-one correspondence with the multiplexed data, and is composed of multiplexed data information, stream attribute information, and an entry map.
[0203] As shown in Fig. 28, the multiplexed data information consists of a system rate, a playback start time, and a playback end time. The system rate indicates the maximum transfer rate of the multiplexed data to the PID filter of the system target decoder, which will be described later. The interval between ATSs contained in the multiplexed data is set to be equal to or less than the system rate. The playback start time is set to the PTS of the first video frame of the multiplexed data, and the playback end time is set to the PTS of the last video frame of the multiplexed data plus the playback interval of one frame.
[0204] As shown in Figure 29, the stream attribute information for each stream included in the multiplexed data is registered for each PID. The attribute information has different information for each video stream, audio stream, presentation graphics stream, and interactive graphics stream. The video stream attribute information includes information such as the compression codec used to compress the video stream, the resolution of the individual picture data that make up the video stream, the aspect ratio, and the frame rate. The audio stream attribute information includes information such as the compression codec used to compress the audio stream, the number of channels included in the audio stream, the language it supports, and the sampling frequency. This information is used to initialize the decoder before the player starts playback.
[0205] In this embodiment, the stream type included in the PMT of the multiplexed data is used. Furthermore, if multiplexed data is recorded on a recording medium, the video stream attribute information included in the multiplexed data information is used. Specifically, the video coding method or device shown in each of the above embodiments includes a step or means for setting, in the stream type included in the PMT or the video stream attribute information, unique information indicating that the video data is generated by the video coding method or device shown in each of the above embodiments. This configuration makes it possible to distinguish between video data generated by the video coding method or device shown in each of the above embodiments and video data that conforms to other standards.
[0206] FIG. 30 shows the steps of the video decoding method according to this embodiment. In step exS100, the stream type included in the PMT or the video stream attribute information included in the multiplexed data information is obtained from the multiplexed data. Next, in step exS101, it is determined whether the stream type or the video stream attribute information indicates that the multiplexed data was generated by the video coding method or device described in the above embodiments. If it is determined that the stream type or the video stream attribute information was generated by the video coding method or device described in the above embodiments, in step exS102, decoding is performed using the video decoding method described in the above embodiments. If the stream type or the video stream attribute information indicates that the data complies with a conventional standard such as MPEG-2, MPEG4-AVC, or VC-1, decoding is performed using the video decoding method according to the conventional standard in step exS103.
[0207] In this way, by setting a new unique value in the stream type or video stream attribute information, it is possible to determine whether the video decoding method or device shown in each of the above embodiments can decode the data when decoding. Therefore, even when multiplexed data conforming to a different standard is input, an appropriate decoding method or device can be selected, enabling decoding without errors. Furthermore, the video encoding method or device or video decoding method or device shown in this embodiment can be used in any of the above-mentioned devices and systems.
[0208] (Embodiment 5) The video encoding method and device, and video decoding method and device described in each of the above embodiments are typically realized by an LSI, which is an integrated circuit. As an example, FIG. 31 shows the configuration of a single-chip LSI ex500. LSI ex500 includes elements ex501, ex502, ex503, ex504, ex505, ex506, ex507, ex508, and ex509, which are described below, and each element is connected via a bus ex510. When the power supply is on, a power supply circuit unit ex505 supplies power to each unit, thereby activating them into an operable state.
[0209] For example, when performing encoding processing, the LSI ex500 inputs AV signals from the microphone ex117, camera ex113, etc. via the AV I / O ex509 under the control of a control unit ex501 including a CPU ex502, a memory controller ex503, a stream controller ex504, a drive frequency control unit ex512, etc. The input AV signals are temporarily stored in an external memory ex511 such as an SDRAM. Under the control of the control unit ex501, the stored data is divided into multiple batches as appropriate depending on the processing volume and processing speed and sent to the signal processing unit ex507, where the audio signal and / or video signal is encoded. Here, the video signal encoding processing is the encoding processing described in each of the above embodiments. The signal processing unit ex507 may further perform processing such as multiplexing the encoded audio data and the encoded video data, and output the resulting data to the outside from the stream I / O ex506. This output multiplexed data is transmitted to the base station ex107 or written to a recording medium ex215. When multiplexing, it is advisable to temporarily store the data in a buffer ex508 to ensure synchronization.
[0210] Although the memory ex511 has been described above as being external to the LSIex500, it may be included within the LSIex500. The buffer ex508 is not limited to one, and multiple buffers may be provided. Furthermore, the LSIex500 may be formed as a single chip or multiple chips.
[0211] Furthermore, in the above description, the control unit ex501 is described as having a CPU ex502, a memory controller ex503, a stream controller ex504, a drive frequency control unit ex512, etc., but the configuration of the control unit ex501 is not limited to this configuration. For example, the signal processing unit ex507 may further include a CPU. By providing a CPU inside the signal processing unit ex507, it is possible to further improve processing speed. As another example, the CPU ex502 may include the signal processing unit ex507, or a part of the signal processing unit ex507, such as an audio signal processing unit. In such a case, the control unit ex501 is configured to include a CPU ex502 that includes the signal processing unit ex507, or a part of it.
[0212] Although we have referred to it as an LSI here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the level of integration.
[0213] Furthermore, the integrated circuit implementation is not limited to LSIs, but may be realized using dedicated circuits or general-purpose processors. Field programmable gate arrays (FPGAs) that can be programmed after LSI fabrication, or reconfigurable processors that allow the connections and settings of circuit cells within the LSI to be reconfigured, may also be used. Such programmable logic devices can typically execute the video encoding method or video decoding method described in each of the above embodiments by loading or reading from memory a program that constitutes software or firmware.
[0214] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that such technology could be used to integrate functional blocks. The application of biotechnology is also a possibility.
[0215] (Embodiment 6) When decoding video data generated by the video encoding method or device described in each of the above embodiments, the amount of processing is likely to increase compared to when decoding video data conforming to conventional standards such as MPEG-2, MPEG4-AVC, or VC-1. Therefore, it is necessary to set the drive frequency of the LSIex500 to a higher frequency than the drive frequency of the CPUex502 when decoding video data conforming to conventional standards. However, increasing the drive frequency raises the problem of increased power consumption.
[0216] To solve this problem, video decoding devices such as televisions ex300 and LSIs ex500 are configured to identify the standard to which video data conforms and switch the drive frequency according to the standard. FIG. 32 shows a configuration ex800 in this embodiment. If the video data was generated using the video encoding method or device described in each of the above embodiments, a drive frequency switching unit ex803 sets a high drive frequency. The unit then instructs a decoding processing unit ex801, which executes the video decoding method described in each of the above embodiments, to decode the video data. On the other hand, if the video data conforms to a conventional standard, the unit sets a low drive frequency compared to when the video data was generated using the video encoding method or device described in each of the above embodiments. The unit then instructs a decoding processing unit ex802, which conforms to the conventional standard, to decode the video data.
[0217] More specifically, the drive frequency switching unit ex803 is composed of the CPU ex502 and drive frequency control unit ex512 in FIG. 31. The decoding processing unit ex801 that executes the video decoding method described in each of the above embodiments and the decoding processing unit ex802 that complies with the conventional standard correspond to the signal processing unit ex507 in FIG. 31. The CPU ex502 identifies the standard to which the video data conforms. The driving frequency control unit ex512 sets the drive frequency based on the signal from the CPU ex502. The signal processing unit ex507 decodes the video data based on the signal from the CPU ex502. Here, the video data can be identified using, for example, the identification information described in the fourth embodiment. The identification information is not limited to that described in the fourth embodiment, and may be any information that can identify the standard to which the video data conforms. For example, if it is possible to identify the standard to which the video data conforms based on an external signal that identifies whether the video data is for use on a television or a disc, then the identification may be based on such an external signal. Furthermore, the selection of the drive frequency in the CPUex502 can be performed based on a lookup table that associates the video data standard with the drive frequency, as shown in Fig. 34. The lookup table is stored in the buffer ex508 or the internal memory of the LSI, and the CPUex502 can select the drive frequency by referring to this lookup table.
[0218] FIG. 33 shows steps for implementing the method of this embodiment. First, in step exS200, the signal processing unit ex507 acquires identification information from the multiplexed data. Next, in step exS201, the CPU ex502 identifies, based on the identification information, whether the video data was generated by the encoding method or device described in any of the above embodiments. If the video data was generated by the encoding method or device described in any of the above embodiments, in step exS202, the CPU ex502 sends a signal to the driving frequency control unit ex512 to set the driving frequency to a high level. The driving frequency control unit ex512 then sets the driving frequency to a high level. On the other hand, if the video data indicates that the video data complies with a conventional standard such as MPEG-2, MPEG4-AVC, or VC-1, in step exS203, the CPU ex502 sends a signal to the driving frequency control unit ex512 to set the driving frequency to a low level. The driving frequency control unit ex512 then sets the driving frequency to a lower level than when the video data was generated by the encoding method or device described in any of the above embodiments.
[0219] Furthermore, by changing the voltage applied to the LSIex500 or a device including the LSIex500 in conjunction with switching the drive frequency, it is possible to further enhance the power saving effect. For example, when the drive frequency is set low, it is conceivable to set the voltage applied to the LSIex500 or a device including the LSIex500 lower in response to this change than when the drive frequency is set high.
[0220] Furthermore, the method of setting the drive frequency is not limited to the above-described setting method, and may be such that a high drive frequency is set when the decoding processing volume is large, and a low drive frequency is set when the decoding processing volume is small. For example, if the processing volume required to decode video data conforming to the MPEG4-AVC standard is larger than the processing volume required to decode video data generated by the video encoding method or device described in each of the above-described embodiments, the drive frequency may be set in the opposite way to the above-described setting method.
[0221] Furthermore, the method of setting the drive frequency is not limited to a configuration that lowers the drive frequency. For example, if the identification information indicates that the video data is generated by the video encoding method or device described in each of the above embodiments, the voltage applied to the LSIex500 or a device including the LSIex500 can be set high. If the identification information indicates that the video data complies with conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, the voltage applied to the LSIex500 or a device including the LSIex500 can be set low. As another example, if the identification information indicates that the video data is generated by the video encoding method or device described in each of the above embodiments, the drive of the CPUex502 can be suspended without stopping. If the identification information indicates that the video data complies with conventional standards such as MPEG-2, MPEG4-AVC, or VC-1, the drive of the CPUex502 can be suspended because there is sufficient processing capacity. Even if the identification information indicates that the video data is generated by the video encoding method or device described in each of the above embodiments, the drive of the CPUex502 can be suspended if there is sufficient processing capacity. In this case, it is conceivable to set the stop time shorter than when the video data indicates that it is video data that complies with conventional standards such as MPEG-2, MPEG4-AVC, and VC-1.
[0222] In this way, by switching the drive frequency depending on the standard to which the video data conforms, it is possible to achieve power savings. Furthermore, if the LSIex500 or a device including the LSIex500 is driven by a battery, the power savings can also extend the battery life.
[0223] (Embodiment 7) The above-mentioned devices and systems, such as televisions and mobile phones, may receive multiple inputs of video data conforming to different standards. To ensure that the signal processing unit ex507 of the LSIex500 can decode such inputs, the signal processing unit ex507 must support multiple standards. However, using separate signal processing units ex507 for each standard increases the circuit size of the LSIex500 and increases costs.
[0224] To solve this problem, a configuration is provided in which a decoding processing unit for executing the video decoding method described in each of the above embodiments is partially shared with a decoding processing unit conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. An example of this configuration is shown in ex900 of FIG. 35A. For example, the video decoding method described in each of the above embodiments and a video decoding method conforming to the MPEG4-AVC standard share some of the processing content, such as entropy coding, inverse quantization, deblocking filtering, and motion compensation. A possible configuration is to share a decoding processing unit ex902 conforming to the MPEG4-AVC standard for the common processing content, and use a dedicated decoding processing unit ex901 for other processing content unique to one aspect of the present invention that does not conform to the MPEG4-AVC standard. In particular, since one aspect of the present invention is characterized by hierarchical coding, it is possible to use a dedicated decoding processing unit ex901 for inverse quantization, and share a decoding processing unit for any or all of the other processing, such as entropy decoding, inverse quantization, deblocking filtering, and motion compensation. Regarding the sharing of the decoding processing unit, for common processing content, the decoding processing unit for executing the video decoding method shown in each of the above embodiments may be shared, and for processing content specific to the MPEG4-AVC standard, a dedicated decoding processing unit may be used.
[0225] Another example of partially sharing processing is shown in ex1000 in FIG. 35B. In this example, a dedicated decoding processing unit ex1001 corresponding to processing content specific to one aspect of the present invention, a dedicated decoding processing unit ex1002 corresponding to processing content specific to another conventional standard, and a shared decoding processing unit ex1003 corresponding to processing content common to the video decoding method according to one aspect of the present invention and the video decoding method of another conventional standard are used. Here, the dedicated decoding processing units ex1001 and ex1002 are not necessarily specialized for processing content specific to one aspect of the present invention or another conventional standard, and may be capable of performing other general-purpose processing. The configuration of this embodiment can also be implemented using an LSI ex500.
[0226] In this way, by sharing a decoding processing unit for processing content that is common between a video decoding method according to one embodiment of the present invention and a video decoding method of a conventional standard, it is possible to reduce the circuit size of the LSI and reduce costs. [Industrial Applicability]
[0227] The present invention can be applied to an image decoding method and apparatus, or an image encoding method and apparatus, and can also be used in high-resolution information display devices or imaging devices that include an image decoding device, such as televisions, digital video recorders, car navigation systems, mobile phones, digital cameras, and digital video cameras. [Explanation of symbols]
[0228] 100 Image encoding device 101 Limit value setting section 102 Encoding section 111 Hierarchy number setting section 112 Layer number parameter setting section 113 Display delay picture number setting unit 114 B-picture consecutive number setting section 115 Continuous number parameter setting section 121, 209 Image sorting section 122 Code block division part 123 Subtraction section 124 Transformation and Quantization Unit 125 Variable-length coding section 126, 202 Inverse transformation and quantization unit 127, 203 Addition section 128, 204 frame memory 129 Intra Prediction Unit 130 Inter Prediction Unit 131 Selection Section 151, 256 frame rates 152, 253 Transmission delay time limit 153 input images 154, 257 coding structure limit value 155, 251 code string 161 hierarchical levels 162 consecutive B pictures 163 Layer Count Parameter 164 display delay pictures 165 consecutive number parameters 171 Code Blocks 172, 174, 258 differential blocks 173, 254 conversion factor 175, 259 decoding blocks 177, 261 predicted blocks 178, 255 Forecast information 200 Image decoding device 201 Variable length decoding unit 205 Intra prediction block generation unit 206 Inter-prediction block generation unit 208 Limit Value Decoding Unit 210 Encoding structure confirmation section 252 Highest TId 262 encoding structure 263 output image
Claims
1. a generating step of generating a bitstream including the video encoded by hierarchically encoding a plurality of images included in the video into a hierarchical structure having one or more layers, and control information; a transmitting step of transmitting the bitstream and the control information; the lowest layer of the hierarchical structure includes I pictures and P pictures; a layer other than the lowest layer of the hierarchical structure includes a B-picture; the control information includes information about a frame rate of the moving image, the number of layers is predetermined based on a frame rate of the moving image; The number of consecutive B pictures in display order is equal to or less than a predetermined maximum number of consecutive B pictures. Sending method.
2. a processing circuit; a storage device accessible by the processing circuitry; The processing circuitry uses the storage device to: Execute the transmission method according to claim 1 Transmitting device.
Citation Information
Patent Citations
Device for encoding animation
JP2000324486A
Video transmission system, transmitting device, and repeating apparatus
JP2011216986A
Systems and methods for error resilient scheme for low latency h.264 video coding
US20120082226A1
Method and apparatus for decoding / encoding of a video signal
WO2008030068A1
Method and apparatus for video coding and decoding
WO2013004911A1