Receiving apparatus and receiving method

By generating and encoding both basic and extended video streams with identification information, the technology ensures efficient transmission and decoding of high-quality format image data, accommodating various display capabilities.

JP7711823B2Active Publication Date: 2025-07-23SONY GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024151962
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-07-31
Filing Date
2024-09-04
Publication Date
2025-07-23
Estimated Expiration
2035-07-09

AI Technical Summary

Technical Problem

Existing technologies face challenges in satisfactorily transmitting a predetermined number of high-quality format image data together with basic format image data, particularly in scenarios where selective use of high-quality format image data is required on the receiving side.

Method used

An image encoding unit generates a basic video stream from basic format image data and a predetermined number of extended video streams from high-quality format image data, with identification information inserted into the container to facilitate selective decoding based on display capability.

Benefits of technology

Enables efficient transmission and decoding of high-quality format image data alongside basic format data, allowing for seamless adaptation to the receiving device's display capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711823000001
    Figure 0007711823000001
  • Figure 0007711823000002
    Figure 0007711823000002
  • Figure 0007711823000003
    Figure 0007711823000003
Patent Text Reader

Abstract

To successfully transmit basic format image data and a predetermined number of pieces of high-quality format image data.SOLUTION: A basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding respectively the predetermined number of pieces of high-quality format image data are generated. A container in a predetermined format including each of the video streams is transmitted. Identification information in a high-quality format corresponding to each of the predetermined number of extended video streams is inserted into a layer of the container and / or the video stream.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to , subject to a transmission device and a receiving party the law .

Background Art

[0002] Conventionally, it has been known to transmit high-quality format image data together with basic format image data and selectively use the basic format image data or the high-quality format image data on the receiving side. For example, Patent Document 1 describes performing scalable media encoding to generate a base layer stream for a low-resolution video service and an enhancement layer stream for a high-resolution video service, and transmitting a broadcast signal including these. Note that in addition to high resolution, high-quality formats include high frame frequency, high dynamic range, wide color gamut, high bit length, and the like.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of this technology is to satisfactorily transmit a predetermined number of high-quality format image data together with basic format image data.

Means for Solving the Problems

[0005] The concept of this technology is an image encoding unit that generates a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, A transmitting unit that transmits a container in a predetermined format including the basic video stream and the predetermined number of extended video streams generated by the image encoding unit; An identification information insertion unit that inserts identification information of a high-quality format corresponding to each of the predetermined number of extended video streams into the container and / or the layer of the video stream. It is in the transmission device.

[0006] In this technology, the image encoding unit generates a basic video stream and a predetermined number of extended video streams. Here, the basic video stream is obtained by encoding basic format image data. Also, each of the predetermined number of extended video streams is obtained by encoding a predetermined number of high-quality format image data.

[0007] For example, the image encoding unit may be configured to perform predictive encoding processing within the basic format image data to generate a basic video stream for the basic format image data, and for the high-quality format image data, selectively perform predictive encoding processing within the high-quality format image data or predictive encoding processing between the basic format image data or other high-quality format image data to generate an extended video stream.

[0008] The transmitting unit transmits a container in a predetermined format including the basic video stream and the predetermined number of extended video streams generated by the image encoding unit. For example, the container may be a transport stream (MPEG-2 TS) adopted in the digital broadcast standard. Also, for example, the container may be an MP4 used for Internet distribution or a container in another format.

[0009] The identification information insertion unit inserts identification information in a high-quality format corresponding to a predetermined number of extended video streams into the layer of the container or the video stream. For example, the container is MPEG2-TS, and when the identification information insertion unit inserts the identification information into the layer of the container, the identification information may be inserted into each video elementary stream loop corresponding to a predetermined number of extended video streams existing under the program map table (PMT). Also, for example, the video stream has a NAL (Network Abstraction Layer) unit structure, and the identification information insertion unit may insert the identification information into the header of the NAL unit.

[0010] Thus, in this technology, identification information in a high-quality format corresponding to a predetermined number of extended video streams is inserted into and transmitted from the layer of the container or the video stream. Therefore, on the receiving side, it becomes easy to selectively perform decoding processing on a predetermined video stream based on the identification information to obtain image data according to the display capability.

[0011] In this technology, for example, information indicating whether identification information inserted into the layer of the container is generated by performing predictive coding processing between a predetermined number of extended video streams and basic format image data or between high-quality format image data may be added. In this case, on the receiving side, it becomes possible to easily recognize whether basic format image data or other high-quality format image data is referenced in the predictive coding processing when generating each of the predetermined number of extended video streams.

[0012] In the present technology, for example, the identification information inserted into the layer of the container may be such that information indicating a video stream corresponding to the image data referred to in the prediction encoding process between the basic format image data or other high-quality format image data performed when generating a predetermined number of extended video streams is added. In this case, on the receiving side, it becomes possible to easily recognize which video stream corresponds to the image data referred to in the prediction encoding process when generating each of the predetermined number of extended video streams.

[0013] Another concept of the present technology is a receiving unit that receives a container in a predetermined format including a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, in the layer of the container and / or the layer of the video stream, identification information in a high-quality format corresponding to each of the predetermined number of extended video streams is inserted, and further includes a processing unit that processes each of the video streams included in the received container based on the identification information in a receiving device.

[0014] In the present technology, a container including a basic video stream and a predetermined number of extended video streams is received by the receiving unit. Here, the basic video stream is obtained by encoding basic format image data. Also, each of the predetermined number of extended video streams is obtained by encoding a predetermined number of high-quality format image data. In the layer of the container or the video stream, identification information in a high-quality format corresponding to each of the predetermined number of extended video streams is inserted.

[0015] For example, the basic video stream is generated by performing predictive encoding processing on the basic format image data within the basic format image data, and the extended video stream is generated by selectively performing predictive encoding processing on the high-quality format image data within the high-quality format image data or on the predictive encoding processing between the basic format image data and other high-quality format image data. This may be done.

[0016] By the processing unit, each video stream included in the received container is processed based on the identification information. For example, the processing unit may perform decoding processing on the basic video stream and a predetermined extended video stream based on the identification information and the display capability information to obtain image data corresponding to the display capability.

[0017] Thus, in this technology, a predetermined number of extended video streams inserted into the container or layer of the video stream and sent are each used to process each video stream based on the identification information of the corresponding high-quality format. Therefore, it becomes easy to selectively perform decoding processing on a predetermined video stream to obtain image data corresponding to the reception capability.

Advantages of the Invention

[0018] According to this technology, a predetermined number of high-quality format image data can be transmitted well together with the basic format image data. Note that the effects described here are not necessarily limited, and any effect described in the present disclosure may be applicable.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0020] Hereinafter, embodiments for carrying out the invention (hereinafter referred to as "embodiments") will be described. The description will be made in the following order. 1. Embodiments 2. Modification Examples

[0021] <1. Embodiments> [Transmission / Reception System] FIG. 1 shows a configuration example of a transmission / reception system 10 as an embodiment. This transmission / reception system 10 is configured to include a transmission device 100 and a reception device 200.

[0022] The transmission device 100 transmits a transport stream TS as a container by mounting it on a broadcast wave or a network packet. This transport stream TS includes a basic video stream and a predetermined number of extended video streams.

[0023] The basic video stream is generated by encoding basic format image data with, for example, H.264 / AVC, H.265 / HEVC, etc. Here, regarding the basic format image data, predictive encoding processing within this basic format image data is performed to generate the basic video stream.

[0024] The predetermined number of extended video streams are each generated by encoding a predetermined number of high-quality image data with, for example, H.264 / AVC, H.265 / HEVC, etc. Here, regarding the high-quality format image data, predictive encoding processing within this high-quality format image data, or predictive encoding processing between the basic format image data or other such high-quality format image data is selectively performed to generate the extended video stream.

[0025] Identification information of a corresponding high-quality format is inserted into the layer of the container for each of the predetermined number of extended video streams. On the receiving side, based on this identification information, it becomes possible to easily grasp the corresponding high-quality format for each of the predetermined number of extended video streams at the layer of the container. In this embodiment, the identification information is inserted into each video elementary stream loop corresponding to the predetermined number of extended video streams existing under the program map table.

[0026] This identification information is appended with information indicating whether a predetermined number of extended video streams are each generated by performing predictive encoding processing with respect to basic format image data or by performing predictive encoding processing with respect to high-quality format image data. On the receiving side, based on this information, it becomes possible to easily recognize in the container layer whether basic format image data or other high-quality format image data is referenced in the predictive encoding processing when generating each of the predetermined number of extended video streams.

[0027] In addition, this identification information is appended with information indicating the video stream corresponding to the image data referenced in the predictive encoding processing between the basic format image data or other high-quality format image data performed when generating each of the predetermined number of extended video streams. On the receiving side, based on this information, it becomes possible to easily recognize in the container layer which video stream corresponds to the image data referenced in the predictive encoding processing when generating each of the predetermined number of extended video streams.

[0028] Identification information of a high-quality format corresponding to each of the predetermined number of extended video streams is inserted into the layer of the video stream. On the receiving side, based on this identification information, it becomes possible to easily grasp the high-quality format corresponding to each of the predetermined number of extended video streams. In this embodiment, the identification information is inserted into the header of the NAL unit.

[0029] The receiving device 200 receives the above-described transport stream TS transmitted on a broadcast wave or a packet on the network from the transmitting device 100. As described above, identification information of a high-quality format corresponding to each of the predetermined number of extended video streams included in the transport stream TS is inserted into the layers of the container and the video stream. The receiving device 200 processes each video stream included in the transport stream TS based on this identification information and acquires image data according to the display capability.

[0030] "Configuration of the Transmitting Device" Figure 2 shows a configuration example of the transmission device 100. This transmission device 100 handles, as transmission image data, basic format image data Vb and three high-quality format image data Vh1, Vh2, and Vh3. Here, the basic format image data Vb is LDR (Low Dynamic Range) image data with a frame frequency of 50 Hz. The high-quality format image data Vh1 is LDR image data with a frame frequency of 100 Hz. The LDR image data has a luminance range of 0% to 100% with respect to the brightness of the white peak of a conventional LDR image.

[0031] The high-quality format image data Vh2 is HDR (High Dynamic Range) image data with a frame frequency of 50 Hz. The high-quality format image data Vh3 is HDR image data with a frame frequency of 100 Hz. This HDR image data has a luminance range of 0 to 100% * N, for example, 0 to 400% or 0 to 800%, etc., when the brightness of the white peak of a conventional LDR image is set to 100%.

[0032] Figure 3 shows a configuration example of the image data generation unit 150 that generates the basic format image data Vb and the three high-quality format image data Vh1, Vh2, and Vh3. This image data generation unit 150 includes an HDR camera 151, a frame rate conversion unit 152, a dynamic range conversion unit 153, and a frame rate conversion unit 154.

[0033] The HDR camera 151 captures an object and outputs HDR image data with a frame frequency of 100 Hz, that is, high-quality format image data Vh3. The frame rate conversion unit 152 performs a process of converting the frame frequency of the high-quality format image data Vh3 output from the HDR camera 151 from 100 Hz to 50 Hz, and outputs HDR image data with a frame frequency of 50 Hz, that is, high-quality format image data Vh2.

[0034] The dynamic range conversion unit 153 performs a process of converting from HDR to LDR on the high-quality format image data Vh3 output from the HDR camera 151, and outputs LDR image data with a frame frequency of 100 Hz, that is, high-quality format image data Vh1. The frame rate conversion unit 154 performs a process of converting the frame frequency from 100 Hz to 50 Hz on the high-quality format image data Vh1 output from the dynamic range conversion unit 153, and outputs LDR image data with a frame frequency of 50 Hz, that is, basic format image data Vb.

[0035] Returning to FIG. 2, the transmission device 100 includes a control unit 101, LDR photoelectric conversion units 102 and 103, HDR photoelectric conversion units 104 and 105, a video encoder 106, a system encoder 107, and a transmission unit 108. The control unit 101 is configured with a CPU (Central Processing Unit) and controls the operations of each part of the transmission device 100 based on a control program.

[0036] The LDR photoelectric conversion unit 102 applies photoelectric conversion characteristics for LDR images (LDR OETF curve) to the basic format image data Vb to obtain basic format image data Vb' for transmission. The LDR photoelectric conversion unit 103 applies photoelectric conversion characteristics for LDR images to the high-quality format image data Vh1 to obtain high-quality format image data Vh1' for transmission.

[0037] The HDR photoelectric conversion unit 104 applies photoelectric conversion characteristics for HDR images (HDR OETF curve) to the high-quality format image data Vh2 to obtain high-quality format image data Vh2' for transmission. The HDR photoelectric conversion unit 105 applies photoelectric conversion characteristics for HDR images to the high-quality format image data Vh3 to obtain high-quality format image data Vh3' for transmission.

[0038] The video encoder 106 has four encoding units 106-0, 106-1, 106-2, and 106-3. The encoding unit 106-0 performs predictive encoding processing such as H.264 / AVC and H.265 / HEVC on the basic format image data Vb' for transmission to generate a basic video stream STb. In this case, the encoding unit 106-0 performs prediction within the image data Vb'.

[0039] The encoding unit 106-1 performs predictive encoding processing such as H.264 / AVC and H.265 / HEVC on the high-quality format image data Vh1' for transmission to generate an extended video stream STe1. In this case, the encoding unit 106-1 selectively performs prediction within the image data Vh1' or prediction between the image data Vh1' and the image data Vb' for each encoding block in order to reduce the prediction residual.

[0040] The encoding unit 106-2 performs predictive encoding processing such as H.264 / AVC and H.265 / HEVC on the high-quality format image data Vh2' for transmission to generate an extended video stream STe2. In this case, the encoding unit 106-2 selectively performs prediction within the image data Vh2' or prediction between the image data Vh2' and the image data Vb' for each encoding block in order to reduce the prediction residual.

[0041] The encoding unit 106-3 performs predictive encoding processing such as H.264 / AVC and H.265 / HEVC on the high-quality format image data Vh3' for transmission to generate an extended video stream STe3. In this case, the encoding unit 106-3 selectively performs prediction within the image data Vh3' or prediction between the image data Vh3' and the image data Vh2' for each encoding block in order to reduce the prediction residual.

[0042] Figure 4 shows a configuration example of the main part of the encoding unit 160. This encoding unit 160 can be applied to the encoding units 106-1, 106-2, and 106-3. This encoding unit 160 includes an intra-layer prediction unit 161, an inter-layer prediction unit 162, a prediction adjustment unit 163, a selection unit 164, and an encoding function unit 165.

[0043] The intra-layer prediction unit 161 performs prediction (intra-layer prediction) within the image data V1 to be encoded on the image data V1 to obtain prediction residual data. The inter-layer prediction unit 162 performs prediction (inter-layer prediction) between the image data V1 to be encoded and the reference image data V2 to obtain prediction residual data.

[0044] The prediction adjustment unit 163 performs the following processing according to the type of scalable extension of the image data V1 with respect to the image data V2 in order to efficiently perform the inter-layer prediction in the inter-layer prediction unit 162. In the case of dynamic range extension, level adjustment for conversion from LDR to HDR is performed. In the case of spatial scalable extension, the block is enlarged to a predetermined size. In the case of frame rate extension, it is bypassed. In the case of color gamut extension, mapping is performed for each of the luminance and chrominance differences. In the case of bit length extension, conversion for aligning the MSBs of the pixels is performed.

[0045] For example, in the case of the encoding unit 106-1, the image data V1 is high-quality format image data Vh1´ (100 Hz, LDR), the image data V2 is basic format image data Vb´ (50 Hz, LDR), and the type of scalable extension corresponds to frame rate extension. Therefore, in the prediction adjustment unit 163, the image data Vb´ is directly bypassed.

[0046] Also, for example, in the case of the encoding unit 106-2, the image data V1 is high-quality format image data Vh2´ (50 Hz, HDR), the image data V2 is basic format image data Vb´ (50 Hz, LDR), and the type of scalable extension corresponds to dynamic range extension. Therefore, in the prediction adjustment unit 163, level adjustment for converting the image data Vb´ from LDR to HDR is performed.

[0047] Also, for example, in the case of the encoding unit 106-3, the image data V1 is high-quality format image data Vh3´ (100 Hz, HDR), the image data V2 is high-quality format image data Vh2´ (50 Hz, HDR), and the type of scalable extension corresponds to frame rate extension. Therefore, in the prediction adjustment unit 163, the image data Vb´ is bypassed as it is.

[0048] The selection unit 164 selectively extracts the prediction residual data obtained by the intra-layer prediction unit 161 or the prediction residual data obtained by the inter-layer prediction unit 162 for each coding block and sends it to the encoding function unit 165. In this case, in the selection unit 164, for example, the one with the smaller prediction residual is extracted. The encoding function unit 165 performs encoding processes such as transform coding, quantization, and entropy coding on the prediction residual data extracted from the selection unit 164 to obtain the video stream ST.

[0049] Returning to FIG. 2, the video encoder 106 inserts identification information in the high-quality format corresponding to each of the extended video streams STe1, STe2, STe3 into their layers. The video encoder 106 inserts this identification information, for example, into the header of the NAL unit.

[0050] Figure 5(a) shows an example of the structure (Syntax) of the NAL unit header, and Figure 5(b) shows the content (Semantics) of the main parameters in this example of the structure. The 1-bit field of "Forbidden_zero_bit" must be 0. The 6-bit field of "nal_unit_type" indicates the NAL unit type. The 6-bit field of "Nuh_layer_id" is an ID indicating the layer extension type of the stream. The 3-bit field of "nuh_temporal_id_plus1" indicates the temporal_id (0 to 6) and takes a value (1 to 7) obtained by adding 1.

[0051] In this embodiment, the 6-bit field of "nuh_layer_id" indicates the identification information (stream extension category information) of the high-quality format corresponding to each extended video stream. For example, "0" indicates the base stream. "1 to 4" indicate the spatial extension stream. "5 to 8" indicate the frame rate extension stream. "9 to 12" indicate the dynamic range extension stream. "13 to 16" indicate the color gamut extension stream. "17 to 20" indicate the bit length extension stream. "21 to 24" indicate the spatial extension and frame rate extension. "25 to 28" indicate the frame rate extension and dynamic range extension.

[0052] For example, the basic video stream STb corresponds to the base stream. Therefore, "nuh_layer_id" in the header of the NAL unit constituting this basic video stream STb is set to "0". Also, for example, the extended video stream STe1 corresponds to the frame rate extension stream. Therefore, "nuh_layer_id" in the header of the NAL unit constituting this extended video stream STe1 is set to any value in the range of "5 to 8".

[0053] Also, for example, the extended video stream STe2 corresponds to a dynamic range extended stream. Therefore, the "nuh_layer_id" in the header of the NAL unit constituting this extended video stream STe2 is set to be any value in the range of "9 to 12". Also, for example, the extended video stream STe3 corresponds to a stream with frame rate extension and dynamic range extension. Therefore, the "nuh_layer_id" in the header of the NAL unit constituting this extended video stream STe3 is set to be any value in the range of "25 to 28".

[0054] Figure 6 shows a configuration example of the basic video stream STb and the extended video streams STe1, STe2, and STe3. The horizontal axis represents the display order (POC: picture order of composition), with the left side having an earlier display time and the right side having a later display time. Each of the rectangular frames represents a picture, and the solid arrows indicate the reference relationship of the pictures in predictive coding.

[0055] The basic video stream STb is composed of encoded image data of pictures such as "00", "01", ···. The extended video stream STe1 is composed of encoded image data of pictures such as "10", "11", ··· located between each picture of the basic video stream STb. The extended video stream STe2 is composed of encoded image data of pictures such as "20", "21", ··· at the same position as each picture of the basic video stream STb. And the extended video stream STe3 is composed of encoded image data of pictures such as "30", "31", ··· located between each picture of the extended video stream STe2.

[0056] Returning to Figure 2, the system encoder 107 generates a transport stream TS including the basic video stream STb and the extended video streams STe1, STe2, and STe3 generated by the video encoder 106. Then, the transmission unit 108 places this transport stream TS on broadcast waves or network packets and transmits it to the receiving device 200.

[0057] At this time, the system encoder 107 inserts identification information in a high-quality format corresponding to each of the extended video streams STe1, STe2, and STe3 into the layer of the container (transport stream). In this embodiment, for example, a scalable extension descriptor including the identification information is inserted into the video elementary stream loop corresponding to each extended video stream existing under the PMT (Program Map Table).

[0058] FIG. 7 shows an example of the structure (Syntax) of this scalable extension descriptor. FIG. 8 shows the content (Semantics) of the main information in the structure example shown in FIG. 7. The 8-bit field of "descriptor_tag" indicates the descriptor type, and here it indicates that it is a scalable extension descriptor. The 8-bit field of "descriptor_length" indicates the length (size) of the descriptor and indicates the number of bytes hereafter as the length of the descriptor.

[0059] The 4-bit field of "type of enhancement" indicates the identification information (stream extension category information) of the high-quality format corresponding to each extended video stream. For example, "1" indicates spatial scalable extension. "2" indicates frame rate scalable extension. "3" indicates dynamic range scalable extension. "4" indicates color gamut scalable extension. "5" indicates bit length scalable extension. "6" indicates spatial / frame rate scalable extension. "7" indicates frame rate / dynamic range scalable extension.

[0060] For example, the extended video stream STe1 corresponds to a frame rate scalable extension. Therefore, the "type of enhancement" of the scalable extension descriptor corresponding to this extended video stream STe1 is set to "2".

[0061] Also, for example, the extended video stream STe2 corresponds to a dynamic range scalable extension. Therefore, the "type of enhancement" of the scalable extension descriptor corresponding to this extended video stream STe2 is set to "3".

[0062] Also, for example, the extended video stream STe3 corresponds to a frame rate / dynamic range scalable extension. Therefore, the "type of enhancement" of the scalable extension descriptor corresponding to this extended video stream STe3 is set to "7".

[0063] FIG. 9 shows the correspondence between the value of the "type of enhancement" field and the value of the "nuh_layer_id" field of the NAL unit header described above. Thus, it can be seen that the identification information (stream extension category information) of the high-quality format corresponding to each extended video stream can be grasped in the same way from any field.

[0064] Returning to FIG. 7, the 4-bit field of "scalable_priority" indicates the priority within the same extension category of each extended video stream. That is, this field indicates whether each extended video stream is generated by performing predictive coding processing with the basic format image data or with the high-quality format image data.

[0065] For example, "0" indicates that it is the first priority stream referring to the basic stream, that is, it is generated by performing predictive coding processing with the basic format image data. Also, for example, "1" indicates that it is the second priority stream referring to the first priority stream, that is, it is generated by performing predictive coding processing with the high-quality format image data.

[0066] For example, the extended video stream STe1 is related to the encoding of high-quality format image data Vh1´ and is generated by performing predictive encoding processing with the basic format image data Vb´. Therefore, the "scalable_priority" of the scalable extension descriptor corresponding to this extended video stream STe1 is set to "0".

[0067] Also, for example, the extended video stream STe2 is related to the encoding of high-quality format image data Vh2´ and is generated by performing predictive encoding processing with the basic format image data Vb´. Therefore, the "scalable_priority" of the scalable extension descriptor corresponding to this extended video stream STe2 is set to "0".

[0068] Also, for example, the extended video stream STe3 is related to the encoding of high-quality format image data Vh3´ and is generated by performing predictive encoding processing with the high-quality format image data Vh2´. Therefore, the "scalable_priority" of the scalable extension descriptor corresponding to this extended video stream STe3 is set to "1".

[0069] The 32-bit field of "enhancement reference PID" indicates the PID value of the reference stream. That is, this field indicates the PID value of the video stream corresponding to the image data referenced in the predictive encoding process between the basic format image data or other high-quality format image data performed when generating each extended video stream.

[0070] For example, the extended video stream STe1 is related to the encoding of high-quality format image data Vh1´, and is generated by performing predictive encoding processing with the basic format image data Vb´. Therefore, the "enhancement reference PID" of the scalable extension descriptor corresponding to this extended video stream STe1 indicates the PID value of the basic video stream STb.

[0071] Also, for example, the extended video stream STe2 is related to the encoding of high-quality format image data Vh2´, and is generated by performing predictive encoding processing with the basic format image data Vb´. Therefore, the "enhancement reference PID" of the scalable extension descriptor corresponding to this extended video stream STe2 indicates the PID value of the basic video stream STb.

[0072] Also, for example, the extended video stream STe3 is related to the encoding of high-quality format image data Vh3´, and is generated by performing predictive encoding processing with the high-quality format image data Vh2´. Therefore, the "enhancement reference PID" of the scalable extension descriptor corresponding to this extended video stream STe3 indicates the PID value of the extended video stream STe2.

[0073] [Configuration of Transport Stream TS] Figure 10 shows a configuration example of the transport stream TS. This transport stream TS includes four video streams: the basic video stream STb and the extended video streams STe1, STe2, and STe3. In this configuration example, there are PES packets "video PES" for each video stream.

[0074] The packet identifier (PID) of the basic video stream STb is, for example, PID1. In the encoded image data of each picture of this video stream, there are NAL units such as AUD, VPS, SPS, PPS, PSEI, SLICE, SSEI, and EOS. The "nuh_layer_id" in the headers of these NAL units is set to "0", indicating that it is the basic video stream (see Figure 9).

[0075] Also, the packet identifier (PID) of the extended video stream STe1 is, for example, PID2. In the encoded image data of each picture of this video stream, there are NAL units such as AUD, SPS, PPS, PSEI, SLICE, SSEI, and EOS. The "nuh_layer_id" in the headers of these NAL units is, for example, set to "5", indicating that it is a frame rate extended stream (see Figure 9).

[0076] Also, the packet identifier (PID) of the extended video stream STe2 is, for example, PID3. In the encoded image data of each picture of this video stream, there are NAL units such as AUD, SPS, PPS, PSEI, SLICE, SSEI, and EOS. The "nuh_layer_id" in the headers of these NAL units is, for example, set to "9", indicating that it is a dynamic range extended stream (see Figure 9).

[0077] Furthermore, the packet identifier (PID) of the extended video stream STe3 is, for example, PID4. In the encoded image data of each picture of this video stream, there are NAL units such as AUD, SPS, PPS, PSEI, SLICE, SSEI, and EOS. The "nuh_layer_id" in the headers of these NAL units is, for example, set to "25", indicating that it is a stream with both frame rate extension and dynamic range extension (see Figure 9).

[0078] In addition, a Program Map Table (PMT) is included in a transport stream TS as Program Specific Information (PSI). This PSI is information that describes to which program each elementary stream included in the transport stream belongs.

[0079] In the PMT, there is a program loop that describes information related to the entire program. Also, in the PMT, there is an elementary stream loop that has information related to each elementary stream. In this configuration example, there are four video elementary stream loops corresponding to the four video streams of a basic video stream STb and extended video streams STe1, STe2, and STe3. In the video elementary stream loop corresponding to the basic video stream STb, information such as a stream type (ST0) and a packet identifier (PID1) is arranged.

[0080] Also, in the video elementary stream loop corresponding to the extended video stream STe1, information such as a stream type (ST1) and a packet identifier (PID2) is arranged, and a descriptor that describes information related to this extended video stream STe1 is also arranged. As one of these descriptors, the above-described Scalable extension descriptor is inserted.

[0081] In this descriptor, "type of enhancement" is set to "2", indicating that it is a frame rate extension stream (frame rate scalable extension) (see Figure 9). Also, "scalable_priority" in this descriptor is set to "0", indicating that it is a first-priority stream that references the basic stream. Further, "enhancement reference PID" in this descriptor is set to "PID1", indicating that it references the basic video stream STb.

[0082] Also, in the video elementary stream loop corresponding to the extended video stream STe2, information such as stream type (ST2) and packet identifier (PID3) is arranged, and a descriptor for describing information related to this extended video stream STe2 is also arranged. As one of these descriptors, the scalable extension descriptor described above is inserted.

[0083] In this descriptor, "type of enhancement" is set to "3", indicating that it is a dynamic range extension stream (dynamic range scalable extension) (see Figure 9). Also, "scalable_priority" in this descriptor is set to "0", indicating that it is a first-priority stream that references the basic stream. Further, "enhancement reference PID" in this descriptor is set to "PID1", indicating that it references the basic video stream STb.

[0084] Also, in the video elementary stream loop corresponding to the extended video stream STe3, information such as stream type (ST3) and packet identifier (PID4) is arranged, and a descriptor for describing information related to this extended video stream STe3 is also arranged. As one of these descriptors, the scalable extension descriptor described above is inserted.

[0085] The "type of enhancement" in this descriptor is set to "7", indicating that it is a stream of frame rate enhancement and dynamic range enhancement (frame rate / dynamic range scalable enhancement) (see Figure 9). Also, the "scalable_priority" in this descriptor is set to "1", indicating that it is a second priority stream referring to the first priority stream. Further, the "enhancement reference PID" in this descriptor is set to "PID3", indicating that it refers to the enhanced video stream STe2.

[0086] The operation of the transmission device 100 shown in Figure 2 will be briefly described. The basic format image data Vb, which is LDR image data with a frame frequency of 50 Hz, is supplied to the LDR photoelectric conversion unit 102. In this LDR photoelectric conversion unit 102, the photoelectric conversion characteristics for LDR images (LDR OETF curve) are applied to the basic format image data Vb to obtain the basic format image data Vb' for transmission. This basic format image data Vb' is supplied to the encoding units 106-0, 106-1, 106-2 of the video encoder 106.

[0087] Also, the high-quality format image data Vh1, which is LDR image data with a frame frequency of 100 Hz, is supplied to the LDR photoelectric conversion unit 103. In this LDR photoelectric conversion unit 103, the photoelectric conversion characteristics for LDR images (LDR OETF curve) are applied to the high-quality format image data Vh1 to obtain the high-quality format image data Vh1' for transmission. This high-quality format image data Vh1' is supplied to the encoding unit 106-1 of the video encoder 106.

[0088] In addition, high-quality format image data Vh2, which is HDR image data with a frame frequency of 50 Hz, is supplied to the HDR photoelectric conversion unit 104. In this HDR photoelectric conversion unit 104, for the high-quality format image data Vh2, the photoelectric conversion characteristics for HDR images (HDR OETF curve) are applied to obtain high-quality format image data Vh2' for transmission. This high-quality format image data Vh2' is supplied to the encoding units 106-2 and 106-3 of the video encoder 106.

[0089] In addition, high-quality format image data Vh3, which is HDR image data with a frame frequency of 100 Hz, is supplied to the HDR photoelectric conversion unit 105. In this HDR photoelectric conversion unit 105, for the high-quality format image data Vh3, the photoelectric conversion characteristics for HDR images (HDR OETF curve) are applied to obtain high-quality format image data Vh3' for transmission. This high-quality format image data Vh3' is supplied to the encoding unit 106-3 of the video encoder 106.

[0090] In the video encoder 106, encoding processing is performed on each of the basic format image data Vb', high-quality format image data Vh1', Vh2', and Vh3' to generate a video stream. That is, in the encoding unit 106-0, predictive encoding processing such as H.264 / AVC and H.265 / HEVC is performed on the basic format image data Vb' for transmission, and a basic video stream STb including the encoded image data of each picture is generated. In this case, in the encoding unit 106-0, prediction is performed within the image data Vb'.

[0091] In addition, in the encoding unit 106-1, predictive encoding processing such as H.264 / AVC and H.265 / HEVC is performed on the high-quality format image data Vh1' for transmission, and an extended video stream STe1 including the encoded image data of each picture is generated. In this case, in the encoding unit 106-1, in order to reduce the prediction residual, prediction within the image data Vh1' or prediction between the image data Vh1' and the image data Vb' is selectively performed for each encoding block.

[0092] Also, in the encoding unit 106-2, predictive encoding processing such as H.264 / AVC and H.265 / HEVC is performed on the high-quality format image data Vh2' for transmission, and an extended video stream STe2 including the encoded image data of each picture is generated. In this case, in the encoding unit 106-2, in order to reduce the prediction residual, prediction within the image data Vh2' or prediction between the image data Vb' is selectively performed for each encoding block.

[0093] Also, in the encoding unit 106-3, predictive encoding processing such as H.264 / AVC and H.265 / HEVC is performed on the high-quality format image data Vh3' for transmission, and an extended video stream STe3 including the encoded image data of each picture is generated. In this case, in the encoding unit 106-3, in order to reduce the prediction residual, prediction within the image data Vh3' or prediction between the image data Vh2' is selectively performed for each encoding block.

[0094] Also, in the video encoder 106, identification information of the corresponding high-quality format is inserted into the layers of the extended video streams STe1, STe2, and STe3. That is, in the video encoder 106, the identification information (stream extension category information) of the high-quality format corresponding to each extended video stream is set in the "nuh_layer_id" field of the NAL unit header (see FIGS. 5 and 9).

[0095] The basic video stream STb, the extended video streams STe1, STe2, and STe generated by the video encoder 106 are supplied to the system encoder 107. In this system encoder 107, a transport stream TS including each video stream is generated.

[0096] In this system encoder 107, identification information in a high-quality format corresponding to each of the extended video streams STe1, STe2, and STe3 is inserted into the layer of the container (transport stream). That is, in the system encoder 107, a scalable extension descriptor including identification information (extended category information of the stream) is inserted into the video elementary stream loop corresponding to each extended video stream existing under the PMT (see FIGS. 7 and 9).

[0097] The transport stream TS generated by the system encoder 107 is sent to the transmission unit 108. In the transmission unit 108, this transport stream TS is placed on a broadcast wave or a packet on the network and transmitted to the receiving device 200.

[0098] "Configuration of the receiving device" FIG. 11 shows a configuration example of the receiving device 200. This receiving device 200 corresponds to the configuration example of the transmitting device 100 in FIG. 2. This receiving device 200 includes a control unit 201, a receiving unit 202, a system decoder 203, a video decoder 204, LDR electro-optical conversion units 205 and 206L, HDR electro-optical conversion units 207 and 208, and a display unit (display device) 209. The control unit 201 is configured to include a CPU (Central Processing Unit) and controls the operations of each part of the receiving device 200 based on a control program stored in a storage (not shown).

[0099] The receiving unit 202 receives the transport stream TS sent on a broadcast wave or a packet on the network from the transmitting device 100. The system decoder 203 extracts the basic video stream STb and the extended video streams STe1, STe2, and STe3 from this transport stream TS.

[0100] In addition, the system decoder 203 extracts various information inserted into the layer of the container (transport stream) and sends it to the control unit 201. This information also includes the scalable extension descriptor described above. The control unit 201 can grasp the identification information (stream extension category information) of the high-quality format corresponding to each of the extended video streams STe1, STe2, and STe3 from the "type of enhancement" field of this descriptor.

[0101] In addition, the control unit 201 can grasp, from the "scalable_priority" field of this descriptor, the priority within the same extension category for each of the extended video streams STe1, STe2, and STe3, that is, whether it is the first-priority stream that references the basic stream or the second-priority stream that references the first-priority stream. Furthermore, the control unit 201 can grasp the PID value of the video stream referenced by each of the extended video streams STe1, STe2, and STe3 from the "enhancement reference PID" field of this descriptor.

[0102] The video decoder 204 has four decoding units 204-0, 204-1, 204-2, and 204-3. The decoding unit 204-0 performs decoding processing on the basic video stream STb to generate basic format image data Vb'. In this case, the decoding unit 204-0 performs prediction compensation within the image data Vb'.

[0103] The decoding unit 204-1 performs decoding processing on the extended video stream STe1 to generate high-quality format image data Vh1'. In this case, the decoding unit 204-1 performs prediction compensation within the image data Vh1' or prediction compensation with the image data Vb' for each encoding block in correspondence with the prediction at the time of encoding.

[0104] The decoding unit 204-2 performs a decoding process on the extended video stream STe2 to generate high-quality format image data Vh2´. In this case, the decoding unit 204-2 performs prediction compensation within the image data Vh2´ or prediction compensation between the image data Vh2´ and the image data Vb´ for each encoded block in correspondence with the prediction during encoding.

[0105] The decoding unit 204-3 performs a decoding process on the extended video stream STe3 to generate high-quality format image data Vh3´. In this case, the decoding unit 204-3 performs prediction compensation within the image data Vh3´ or prediction compensation between the image data Vh3´ and the image data Vh2´ for each encoded block in correspondence with the prediction during encoding.

[0106] FIG. 12 shows a configuration example of the main part of the decoding unit 240. This decoding unit 240 can be applied to the decoding units 204-1, 204-2, and 204-3. This decoding unit 240 performs a process reverse to that of the encoding unit 165 in FIG. 4. This decoding unit 240 includes a decoding function unit 241, an intra-layer prediction compensation unit 242, an inter-layer prediction compensation unit 243, a prediction adjustment unit 244, and a selection unit 245.

[0107] The decoding function unit 241 performs a decoding process other than prediction compensation on the video stream ST to obtain prediction residual data. The intra-layer prediction compensation unit 242 performs prediction compensation (intra-layer prediction compensation) within the image data V1 on the prediction residual data to obtain the image data V1. The inter-layer prediction compensation unit 243 performs prediction compensation (inter-layer prediction compensation) between the prediction residual data and the reference target image data V2 to obtain the image data V1.

[0108] Although detailed description is omitted, the prediction adjustment unit 244 performs processing on the image data V1 according to the type of scalable extension of the image data V1 with respect to the image data V2, similarly to the prediction adjustment unit 163 of the encoding unit 160 in FIG. 4. The selection unit 245 selectively extracts the image data V1 obtained by the intra-layer prediction compensation unit 242 or the image data V1 obtained by the inter-layer prediction compensation unit 243 for each encoding block in correspondence with the prediction at the time of encoding, and outputs the same.

[0109] Returning to FIG. 11, the video decoder 204 sends the header information of the NAL unit of each video stream to the control unit 201. The control unit 201 can grasp the identification information (stream extension category information) of the high-quality format corresponding to each of the extended video streams STe1, STe2, and STe3 from the "nuh_layer_id" field of this header information.

[0110] The LDR electro-optical conversion unit 205 performs electro-optical conversion with characteristics opposite to those of the LDR photoelectric conversion unit 102 in the transmission device 100 described above on the basic format image data Vb´ obtained by the decoding unit 204-0 to obtain the basic format image data Vb. This basic format image data is LDR image data with a frame frequency of 50 Hz.

[0111] Further, the LDR electro-optical conversion unit 206 performs electro-optical conversion with characteristics opposite to those of the LDR photoelectric conversion unit 103 in the transmission device 100 described above on the high-quality format image data Vh1´ obtained by the decoding unit 204-1 to obtain the high-quality format image data Vh1. This high-quality format image data Vh1 is LDR image data with a frame frequency of 100 Hz.

[0112] Further, the HDR electro-optical conversion unit 207 performs electro-optical conversion with characteristics opposite to those of the HDR photoelectric conversion unit 104 in the transmission device 100 described above on the high-quality format image data Vh2´ obtained by the decoding unit 204-2 to obtain the high-quality format image data Vh2. This high-quality format image data Vh2 is HDR image data with a frame frequency of 50 Hz.

[0113] Further, the HDR electro-optical conversion unit 208 performs electro-optical conversion with characteristics opposite to those of the HDR photoelectric conversion unit 105 in the transmission device 100 described above on the high-quality format image data Vh3´ obtained by the decoding unit 204-3 to obtain high-quality format image data Vh3. This high-quality format image data Vh3 is HDR image data with a frame frequency of 100 Hz.

[0114] The display unit 209 is composed of, for example, an LCD (Liquid Crystal Display), an organic EL (Organic Electro-Luminescence) panel, or the like. The display unit 209 displays an image based on any one of the basic format image data Vb, high-quality format image data Vh1, Vh2, and Vh3 according to its display capability.

[0115] In this case, the control unit 201 controls the image data to be supplied to the display unit 209. This control is performed based on the identification information (stream extension category information) in the high-quality format corresponding to each of the extended video streams STe1, STe2, and STe3 grasped by the control unit 201 as described above, and the display capability information of the display unit 209.

[0116] That is, when the display unit 209 is unable to perform high frame frequency display or high dynamic range display, it is controlled so that the basic format image data Vb related to the decoding of the basic video stream STb is supplied to the display unit 209. In this case, the control unit 201 controls the decoding unit 204-0 to decode the basic video stream STb and the LDR electro-optical conversion unit 205 to output the basic format image data Vb.

[0117] Also, when the display unit 209 can display at a high frame frequency but cannot display with a high dynamic range, control is performed so that high-quality format image data Vh1 related to the decoding of the extended video stream STe1 is supplied to the display unit 209. In this case, the control unit 201 controls the decoding unit 204-0 to decode the basic video stream STb, the decoding unit 204-1 to decode the extended video stream STe1, and the LDR electro-optical conversion unit 206 to output the high-quality format image data Vh1.

[0118] Also, when the display unit 209 cannot display at a high frame frequency but can display with a high dynamic range, control is performed so that high-quality format image data Vh2 related to the decoding of the extended video stream STe2 is supplied to the display unit 209. In this case, the control unit 201 controls the decoding unit 204-0 to decode the basic video stream STb, the decoding unit 204-2 to decode the extended video stream STe2, and the HDR electro-optical conversion unit 207 to output the high-quality format image data Vh2.

[0119] Also, when the display unit 209 can display both at a high frame frequency and with a high dynamic range, control is performed so that high-quality format image data Vh3 related to the decoding of the extended video stream STe3 is supplied to the display unit 209. In this case, the control unit 201 controls the decoding unit 204-0 to decode the basic video stream STb, the decoding unit 204-2 to decode the extended video stream STe2, the decoding unit 204-3 to decode the extended video stream STe3, and the HDR electro-optical conversion unit 208 to output the high-quality format image data Vh3.

[0120] The operation of the receiving apparatus 200 shown in FIG. 11 will be briefly described. In the receiving unit 202, a transport stream TS sent on a broadcast wave or a packet on the network from the transmitting apparatus 100 is received. This transport stream TS is supplied to the system decoder 203. In the system decoder 203, a basic video stream STb and extended video streams STe1, STe2, STe3 are extracted from this transport stream TS.

[0121] Also, in the system decoder 203, various information inserted in the layer of the container (transport stream) is extracted and sent to the control unit 201. This information includes a scalable extension descriptor. In the control unit 201, identification information (stream extension category information) of a high-quality format corresponding to each of the extended video streams STe1, STe2, STe3 is grasped from the "type of enhancement" field of this descriptor.

[0122] When the display unit 209 cannot perform high frame frequency display or high dynamic range display, basic format image data Vb is supplied from the LDR electro-optical conversion unit 205 to the display unit 209. In the display unit 209, this basic format image data Vb, that is, an image with a frame frequency of 50 Hz and LDR image data is displayed.

[0123] In this case, the basic video stream STb extracted by the system decoder 203 is supplied to the decoding unit 204-0. In the decoding unit 204-0, a decoding process is performed on the basic video stream STb, and basic format image data Vb' is generated. Here, in the decoding unit 204-0, it is possible to confirm that the supplied video stream is the basic video stream STb from the "nuh_layer_id" field of the NAL unit header.

[0124] The basic format image data Vb' generated by the decoding unit 204-0 is supplied to the LDR electro-optical conversion unit 205. In the LDR electro-optical conversion unit 205, electro-optical conversion is performed on this basic format image data Vb', and the basic format image data Vb is obtained and supplied to the display unit 209.

[0125] Also, when the display unit 209 can display at a high frame frequency but cannot display with a high dynamic range, high-quality format image data Vh1 is supplied from the LDR electro-optical conversion unit 206 to the display unit 209. On the display unit 209, this high-quality format image data Vh1, that is, an image based on LDR image data with a frame frequency of 100 Hz is displayed.

[0126] In this case, the basic video stream STb extracted by the system decoder 203 is supplied to the decoding unit 204-0. In the decoding unit 204-0, decoding processing is performed on the basic video stream STb, and basic format image data Vb' is generated. Also, the extended video stream STe1 extracted by the system decoder 203 is supplied to the decoding unit 204-1. In the decoding unit 204-1, with reference to the basic format image data Vb', decoding processing is performed on the extended video stream STe1, and high-quality format image data Vh1' is generated.

[0127] Here, in the decoding unit 204-0, it is possible to confirm from the "nuh_layer_id" field of the NAL unit header that the supplied video stream is the basic video stream STb. Also, in the decoding unit 204-1, it is possible to confirm from the "nuh_layer_id" field of the NAL unit header that the supplied video stream is the extended video stream STe1.

[0128] The high-quality format image data Vh1' generated by the decoding unit 204-1 is supplied to the LDR electro-optical conversion unit 206. In the LDR electro-optical conversion unit 206, electro-optical conversion is performed on this high-quality format image data Vh1', and high-quality format image data Vh1 is obtained and supplied to the display unit 209.

[0129] Also, when the display unit 209 cannot perform high frame rate display but can perform high dynamic range display, high-quality format image data Vh2 is supplied from the HDR electro-optical conversion unit 207 to the display unit 209. In the display unit 209, this high-quality format image data Vh2, that is, an image based on HDR image data with a frame frequency of 50 Hz is displayed.

[0130] In this case, the basic video stream STb extracted by the system decoder 203 is supplied to the decoding unit 204-0. In the decoding unit 204-0, decoding processing is performed on the basic video stream STb, and basic format image data Vb' is generated. Also, the extended video stream STe2 extracted by the system decoder 203 is supplied to the decoding unit 204-2. In the decoding unit 204-2, decoding processing is performed with reference to the basic format image data Vb' on the extended video stream STe2, and high-quality format image data Vh2' is generated.

[0131] Here, in the decoding unit 204-0, it is possible to confirm from the "nuh_layer_id" field of the NAL unit header that the supplied video stream is the basic video stream STb. Also, in the decoding unit 204-2, it is possible to confirm from the "nuh_layer_id" field of the NAL unit header that the supplied video stream is the extended video stream STe2.

[0132] The high-quality format image data Vh2´ generated by the decoding unit 204-2 is supplied to the HDR electro-optical conversion unit 207. In the HDR electro-optical conversion unit 207, the high-quality format image data Vh2´ is subjected to electro-optical conversion to obtain high-quality format image data Vh2, which is then supplied to the display unit 209.

[0133] Also, when the display unit 209 is capable of both high frame rate display and high dynamic range display, high-quality format image data Vh3 is supplied from the HDR electro-optical conversion unit 208 to the display unit 209. On the display unit 209, this high-quality format image data Vh3, that is, an image based on HDR image data with a frame frequency of 100 Hz is displayed.

[0134] In this case, the basic video stream STb extracted by the system decoder 203 is supplied to the decoding unit 204-0. In the decoding unit 204-0, a decoding process is performed on the basic video stream STb to generate basic format image data Vb´. Also, the extended video stream STe2 extracted by the system decoder 203 is supplied to the decoding unit 204-2. In the decoding unit 204-2, with reference to the basic format image data Vb´, a decoding process is performed on the extended video stream STe2 to generate high-quality format image data Vh2´.

[0135] Furthermore, the extended video stream STe3 extracted by the system decoder 203 is supplied to the decoding unit 204-3. In the decoding unit 204-3, with reference to the high-quality format image data Vh2´, a decoding process is performed on the extended video stream STe3 to generate high-quality format image data Vh3´.

[0136] Here, in the decoding unit 204-0, it is possible to confirm that the supplied video stream is the basic video stream STb from the "nuh_layer_id" field in the header of the NAL unit. Also, in the decoding unit 204-2, it is possible to confirm that the supplied video stream is the extended video stream STe2 from the "nuh_layer_id" field in the header of the NAL unit. Further, in the decoding unit 204-3, it is possible to confirm that the supplied video stream is the extended video stream STe3 from the "nuh_layer_id" field in the header of the NAL unit.

[0137] The high-quality format image data Vh3' generated in the decoding unit 204-3 is supplied to the HDR electro-optical conversion unit 208. In the HDR electro-optical conversion unit 208, electro-optical conversion is performed on this high-quality format image data Vh3', and high-quality format image data Vh3 is obtained and supplied to the display unit 209.

[0138] As described above, in the transmission / reception system 10 shown in FIG. 1, in the transmission device 100, a predetermined number of extended video streams included in the transport stream TS are each inserted with corresponding high-quality format identification information (stream extension category information) into the container and the layer of the video stream and then transmitted. Therefore, on the receiving side, based on this identification information, it becomes easy to selectively perform decoding processing on a predetermined video stream to obtain image data corresponding to the display ability.

[0139] <2. Modified Example> Note that in the above-described embodiment, an example is shown in which the corresponding high-quality format identification information (stream extension category information) of a predetermined number of extended video streams included in the transport stream TS is inserted into both the container and the layer of the video stream and then transmitted. However, it is also conceivable to insert this identification information only into the layer of the container or only into the layer of the video stream.

[0140] Also, instead of transmitting an ID indicating the layer extension type of the stream, the extension category of the stream, and information indicating the priority within the extension category, it is also possible to indicate their combined state with the value of "stream_type". For example, as shown in FIG. 10, the basic stream is "Stream_type = ST0", the first stream of the frame rate scalable extension is "Stream_type = ST1", the first stream of the dynamic range scalable extension is "Stream_type = ST2", and the stream of the frame rate / dynamic range scalable extension (the second extension stream) is "Stream_type = ST3", and so on.

[0141] Also, in the above-described embodiment, the transmission / reception system 10 including the transmission device 100 and the reception device 200 is shown, but the configuration of the transmission / reception system to which this technology can be applied is not limited to this. For example, the part of the reception device 200 may be configured as a set-top box and a monitor connected by a digital interface such as HDMI (High-Definition Multimedia Interface). In this case, the set-top box can obtain display capability information by acquiring EDID (Extended display identification data) from the monitor. Note that "HDMI" is a registered trademark.

[0142] Also, in the above-described embodiment, an example where the container is a transport stream (MPEG-2 TS) is shown. However, this technology can be similarly applied to a system configured to be distributed to a receiving terminal using a network such as the Internet. In Internet distribution, it is often distributed in a container of MP4 or other formats. That is, as the container, various formats of containers such as the transport stream (MPEG-2 TS) adopted in the digital broadcast standard and MP4 used in Internet distribution are applicable.

[0143] In addition, the present technology can also be configured as follows. (1) An image encoding unit that generates a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively; A transmission unit that transmits a container in a predetermined format including the basic video stream and the predetermined number of extended video streams generated by the image encoding unit; An identification information insertion unit that inserts identification information in a high-quality format corresponding to each of the predetermined number of extended video streams into a layer of the container. Transmission device. (2) The image encoding unit: Regarding the basic format image data, perform predictive encoding processing within the basic format image data to generate the basic video stream; Regarding the high-quality format image data, selectively perform predictive encoding processing within the high-quality format image data or predictive encoding processing between the basic format image data or other high-quality format image data to generate the extended video stream. The transmission device according to (1) above. (3) The identification information inserted into the layer of the container includes: Information indicating whether each of the predetermined number of extended video streams is generated by performing predictive encoding processing with the basic format image data or by performing predictive encoding processing with the high-quality format image data is added. The transmission device according to (2) above. (4) The identification information inserted into the layer of the container includes: Information indicating a video stream corresponding to the image data referred to in the predictive encoding processing between the basic format image data or other high-quality format image data performed when generating each of the predetermined number of extended video streams is added. The transmission device according to (2) or (3) above. (5) The container is MPEG2-TS, The identification information insertion unit, inserts the identification information into each video elementary stream loop corresponding to the predetermined number of extended video streams existing under the program map table. The transmission device according to any one of (1) to (4) above. (6) The identification information insertion unit, further inserts the identification information of the high-quality format corresponding to each of the predetermined number of extended video streams into the layer of the video stream. The transmission device according to any one of (1) to (5) above. (7) The video stream has an NAL unit structure, The identification information insertion unit, inserts the identification information into the header of the NAL unit. The transmission device according to (6) above. (8) An image encoding step of generating a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, A transmission step of transmitting, by a transmission unit, a container in a predetermined format including the basic video stream and the predetermined number of extended video streams generated in the image encoding step, and an identification information insertion step of inserting the identification information of the high-quality format corresponding to each of the predetermined number of extended video streams into the layer of the container. Transmission method. (9) An image encoding unit that generates a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, A transmission unit that transmits a container in a predetermined format including the basic video stream and the predetermined number of extended video streams generated by the image encoding unit, An identification information insertion unit that inserts the identification information of the corresponding high-quality format into each of the extended video streams of the predetermined number into the layer of the video stream. Transmission device. (10) The image encoding unit Regarding the basic format image data, performs predictive encoding processing within the basic format image data to generate the basic video stream, Regarding the high-quality format image data, selectively performs predictive encoding processing within the high-quality format image data or predictive encoding processing between the basic format image data or other high-quality format image data to generate the extended video stream. The transmission device according to (9) above. (11) The video stream has an NAL unit structure, The identification information insertion unit Inserts the identification information into the header of the NAL unit. The transmission device according to (9) or (10) above. (12) An image encoding step of generating a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, A transmission step of transmitting, by a transmission unit, a container of a predetermined format including the basic video stream and the predetermined number of extended video streams generated in the image encoding step, An identification information insertion step of inserting the identification information of the corresponding high-quality format into each of the predetermined number of extended video streams into the layer of the video stream. Transmission method. (13) A receiving unit that receives a container of a predetermined format including a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively, In the layer of the container, identification information of a high-quality format corresponding to each of the predetermined number of extended video streams is inserted. The receiving apparatus further includes a processing unit that processes each of the video streams included in the received container based on the identification information. Receiving apparatus. (14) The processing unit performs decoding processing on the basic video stream and the predetermined extended video streams based on the identification information and the display capability information to obtain image data corresponding to the display capability. The receiving apparatus according to (13) above. (15) The basic video stream is generated by performing predictive encoding processing on the basic format image data within the basic format image data. The extended video stream is generated by selectively performing predictive encoding processing on the high-quality format image data within the high-quality format image data or predictive encoding processing between the basic format image data or other high-quality format image data. The receiving apparatus according to (13) or (14) above. (16) A receiving step of receiving a container in a predetermined format including a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively by a receiving unit, In the layer of the container, identification information of a high-quality format corresponding to each of the predetermined number of extended video streams is inserted. The receiving method further includes a processing step of processing each of the video streams included in the received container based on the identification information. Receiving method. (17) A receiving unit that receives a container in a predetermined format including a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively. In the layer of the video stream, identification information of a high-quality format corresponding to each of the predetermined number of extended video streams is inserted, The receiving apparatus further includes a processing unit that processes each of the video streams included in the received container based on the identification information. Receiving apparatus. (18) The processing unit Performs decoding processing on the basic video stream and the predetermined extended video streams based on the identification information and display capability information to obtain image data corresponding to the display capability. The receiving apparatus according to (17) above. (19) The basic video stream is generated by performing predictive coding processing on the basic format image data within the basic format image data, The extended video stream is generated by selectively performing predictive coding processing on the high-quality format image data within the high-quality format image data or predictive coding processing between the basic format image data or other high-quality format image data. The receiving apparatus according to (17) or (18) above. (20) A receiving step of receiving a container of a predetermined format including a basic video stream obtained by encoding basic format image data and a predetermined number of extended video streams obtained by encoding a predetermined number of high-quality format image data respectively by a receiving unit, In the layer of the video stream, identification information of a high-quality format corresponding to each of the predetermined number of extended video streams is inserted, The receiving method further includes a processing step of processing each of the video streams included in the received container based on the identification information. Receiving method.

[0144] The main feature of this technology is that a predetermined number of extended video streams included in the transport stream TS each insert identification information (stream extended category information) in a corresponding high-quality format into a container or a layer of the video stream and transmit it, making it easier for the receiving side to obtain image data according to the display capabilities (see Figure 10).

[0145] 10 ··· Transmission and reception system 100 ··· Transmitter 101 ··· Control unit 102, 103 ··· LDR photoelectric conversion unit 104, 105 ··· HDR photoelectric conversion unit 106 ··· Video encoder 106-0, 106-1, 106-1, 106-1 ··· Encoding unit 107 ··· System encoder 108 ··· Transmitting unit 150 ··· Image data generation unit 151 ··· HDR camera 152, 154 ··· Frame rate conversion unit 153 ··· Dynamic range conversion unit 160 ··· Encoding unit 161 ··· Intra-layer prediction unit 162 ··· Inter-layer prediction unit 163 ··· Prediction adjustment unit 164 ··· Selection unit 165 ··· Encoding function unit 200 ··· Receiver 201 ··· Control unit 202 ··· Receiving unit 203 ··· System decoder 204 ··· Video decoder 204-0, 204-1, 204-1, 204-1 ··· Decoding unit 205, 206 ··· LDR electro-optical conversion unit 207, 208 ··· HDR electro-optical conversion unit 209 ··· Display unit 240 ··· Decoding unit 241 ··· Decoding function unit 242 ··· Intra-layer prediction compensation unit 243 ··· Inter-layer prediction compensation unit 244 ··· Prediction adjustment unit 245 ··· Selection unit

Claims

1. A receiving unit that receives a container including a basic video stream obtained by encoding image data in a basic format, a predetermined number of extended video streams obtained by encoding image data in a predetermined number of high-quality formats corresponding to scalable extensions of a predetermined number of types, and type information indicating the types of scalable extensions corresponding to the predetermined number of extended video streams; A processing unit that processes the video streams included in the received container based on the type information, wherein the basic video stream and the predetermined number of extended video streams include NAL units, wherein the header of the NAL unit includes identification information corresponding to whether the NAL unit is included in the basic video stream or the extended video stream, and the processing unit controls the image data supplied to the display unit based on the identification information and the display capability information of the display unit A receiving device.

2. A receiving step of receiving a container including a basic video stream obtained by encoding image data in a basic format, a predetermined number of extended video streams obtained by encoding image data in a predetermined number of high-quality formats corresponding to scalable extensions of a predetermined number of types, and type information indicating the types of scalable extensions corresponding to the predetermined number of extended video streams; A processing step of processing the video streams included in the received container based on the type information, wherein the basic video stream and the predetermined number of extended video streams include NAL units, wherein the header of the NAL unit includes identification information corresponding to whether the NAL unit is included in the basic video stream or the extended video stream, and in the processing step, the image data supplied to the display unit is controlled based on the identification information and the display capability information of the display unit A receiving method.

Citation Information

Patent Citations

  • Method and apparatus for hierarchical transmission and reception in digital broadcasting

    JP2008543142A

  • Multiplexing device for hierarchized elementary stream, demultiplexing device, multiplexing method, and program

    JP2009267537A

  • Level signaling for layered video coding

    WO2013151814A1

  • Image data transmission device, image data transmission method, image data reception device, and image data reception method

    WO2013161442A1

  • Transmission device, transmission method, reception device, and reception method

    WO2014034463A1