Transmitting method

The transmission method addresses the challenge of selecting hierarchically encoded data by generating and transmitting packets with distinct asset IDs and decoding information, facilitating easy selection and decoding on the receiving side.

JP2025085801AActive Publication Date: 2025-06-05PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025047238
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-06-18
Filing Date
2025-03-21
Publication Date
2025-06-05
Estimated Expiration
2034-06-04

AI Technical Summary

Technical Problem

Existing methods for transmitting hierarchically encoded data, such as those using HEVC, do not facilitate easy selection of hierarchically coded data on the receiving side, particularly in terms of frame rate selection.

Method used

A transmission method that generates and transmits packets for base and enhancement layers with distinct asset IDs, including information for decoding and identifying the number of layers, allowing for easy selection and filtering of hierarchically encoded data.

Benefits of technology

Enables easy selection and decoding of hierarchically encoded data on the receiving side, allowing for flexible frame rate adjustments and efficient data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085801000001_ABST
    Figure 2025085801000001_ABST
Patent Text Reader

Abstract

To provide a transmitting method for transmitting encoded data that can easily select hierarchically encoded data.SOLUTION: A transmission method includes a transmitting step of transmitting a packet containing a base layer level and a packet containing an enhanced layer level in an encoded stream, and transmitting a packet containing information. Different packet IDs are assigned to the packet containing the base layer level of encoded data, the packet containing the enhanced layer level of the encoded data, and the packet containing the information. The information includes information indicating an asset ID of the base layer level used for decoding the enhanced layer level, information indicating the correspondence between the asset ID and the packet ID assigned to the packet, and information for specifying the number of layers.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a transmission method for transmitting hierarchically encoded data. [Background technology]

[0002] Conventionally, a technique for transmitting encoded data by a predetermined multiplexing method is known. The encoded data is generated by encoding content including video data and audio data based on a video encoding standard such as High Efficiency Video Coding (HEVC).

[0003] The predetermined multiplexing method is, for example, MPEG-2 TS (Moving Picture Examples include MPEG (MPEG Media Transport) and MMT (Experts Group-2 Transport Stream) (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Information technology - High efficiency coding and media delivery in heterogeneous environment - Part1:MPEG media transport(MMT), ISO / IEC DIS 23008-1 Summary of the Invention [Problem to be solved by the invention]

[0005] By the way, HEVC allows hierarchical encoding. On the receiving side, it is possible to select the frame rate of the video by selecting the hierarchically encoded data according to the layer level.

[0006] The present invention provides a method for transmitting coded data, which allows easy selection of hierarchically coded data. [Means for solving the problem]

[0007] A transmission method according to one embodiment of the present invention is a transmission method for encoded data in which video is encoded in a time hierarchical manner and the encoding order and display order are different, the method including: a generation step of generating packets in which the encoded data is packetized into different packets for a base layer level and an enhancement layer level, which are layer levels of the encoded data, the packets being associated with different asset IDs depending on at least the base layer level and the enhancement layer level, and the packets including information indicating the asset ID of the base layer level used for decoding the enhancement layer level, information indicating the correspondence between the asset ID and a packet ID assigned to the packet, and information for identifying the number of layers; and a transmission step of transmitting the packets including the base layer level and the packets including the enhancement layer level in an encoded stream, and transmitting the packets including the generated information, wherein different packet IDs are assigned to the packets including the base layer level of the encoded data, the packets including the enhancement layer level of the encoded data, and the packets including the information.

[0008] These comprehensive or specific aspects may be realized by a system, an apparatus, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. These comprehensive or specific aspects may be realized by any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium. Effect of the Invention

[0009] According to the present invention, hierarchically encoded data can be easily selected. [Brief description of the drawings]

[0010] [Figure 1]FIG. 1 is a diagram for explaining coded data that has been coded in a time scalable manner. [Diagram 2] FIG. 2 is a first diagram for explaining the data structure of a coded stream in MMT. [Diagram 3] FIG. 3 is a second diagram for explaining the data structure of a coded stream in MMT. [Figure 4] FIG. 4 is a diagram showing a correspondence relationship between packet IDs and data (assets) in a coded stream according to the first embodiment. [Diagram 5] FIG. 5 is a block diagram showing a configuration of a transmission device according to the first embodiment. [Figure 6] FIG. 6 is a flowchart of the transmission method according to the first embodiment. [Figure 7] FIG. 7 is a block diagram showing a configuration of a receiving device according to the first embodiment. [Figure 8] FIG. 8 is a flowchart of the receiving method according to the first embodiment. [Figure 9] FIG. 9 is a diagram conceptually showing a receiving method according to the first embodiment. [Figure 10] FIG. 10 is a block diagram showing a configuration of a receiving device according to the second embodiment. [Figure 11] FIG. 11 is a first diagram for explaining an overview of a transmission / reception method according to the second embodiment. [Figure 12] FIG. 12 is a second diagram for explaining the outline of the transmission and reception method according to the second embodiment. [Figure 13] FIG. 13 is a first diagram for explaining an example of packetizing encoded data in units of fragmented MFUs. [Figure 14] FIG. 14 is a second diagram for explaining an example of packetizing encoded data in units of fragmented MFUs. [Figure 15] FIG. 15 is a third diagram for explaining an example of packetizing encoded data in units of fragmented MFUs. [Figure 16]FIG. 16 is a diagram showing an example in which encoded data is arranged in the MP4 data in the original order. [Figure 17] FIG. 17 is a diagram showing a first example of arranging encoded data in MP4 data for each hierarchical level. [Figure 18] FIG. 18 is a diagram illustrating a second example of arranging encoded data in MP4 data for each hierarchical level. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] (Findings on which the present invention is based) The video coding method HEVC (High Efficiency Video Coding) supports time scalability, and can play back, for example, a 120 fps video as a 60 fps video. Fig. 1 is a diagram for explaining coded data coded in a time scalable manner.

[0012] A temporal ID is assigned to each layer of the encoded data that is encoded in a time scalable manner. In Fig. 1, for example, if a picture (I0, P4) with a temporal ID of 0 and a picture (B2) with a temporal ID of 1 are displayed, the video is displayed at 60 fps, and if a picture (B1, B3) with a temporal ID of 2 is additionally displayed, the video is displayed at 120 fps.

[0013] In the example of FIG. 1, coded data with a Temporal ID of 0 or 1 is the base layer (base layer level), and data with a Temporal ID of 2 is the enhancement layer (enhancement layer level).

[0014] A picture in the base layer can be coded independently or can be decoded using other pictures in the base layer. In contrast, a picture in the enhancement layer cannot be decoded independently, and can only be decoded after the reference picture at the start point of the arrow in Figure 1 is decoded. Therefore, a picture in the base layer that is a reference picture for a picture in the enhancement layer must be decoded before the picture in the enhancement layer.

[0015] Note that the decoding order is different from the image display order. In the example of Fig. 1, the image display order is (I0, B1, B2, B3, P4), whereas the decoding order is (I0, P4, B2, B1, B3). The image display order is determined based on a PTS (Presentation Time Stamp) assigned to each picture, and the decoding order is determined based on a DTS (Decode Time Stamp) assigned to each picture.

[0016] In the case of not only temporal scalable coding but also spatial scalable coding and SNR scalable coding, when pictures are divided into a base layer and an enhancement layer, pictures belonging to the enhancement layer cannot be decoded independently, and must be decoded together with pictures belonging to the base layer.

[0017] It is desirable that scalably coded (hierarchically coded) coded data can be easily selected on the receiving side (decoding side).

[0018] Therefore, a transmission method according to one embodiment of the present invention is a method for transmitting encoded data in which video is hierarchically encoded, and includes a generation step of generating an encoded stream including packets in which the encoded data has been packetized, the packets being assigned packet IDs that differ depending on at least the hierarchical level of the encoded data, and information indicating a correspondence between the packet IDs and the hierarchical level, and a transmission step of transmitting the generated encoded stream and the generated information indicating the correspondence.

[0019] This allows coded data to be sorted by hierarchical level by filtering the packet ID, which means that coded data can be easily sorted on the receiving side.

[0020] In addition, the hierarchical levels may include a base hierarchical level and an enhanced hierarchical level, and the encoded data of the base hierarchical level may be decodable independently or by referring to data obtained after decoding of other encoded data of the base hierarchical level, and the encoded data of the enhanced hierarchical level may be decodable by referring to data obtained after decoding of the encoded data of the base hierarchical level.

[0021] In addition, the generating step may generate a first encoded stream, which is an encoded stream including the packetized packets of the encoded data at the base layer level and not including the packetized packets of the encoded data at the enhanced layer level, and a second encoded stream, which is an encoded stream including the packetized packets of the encoded data at the enhanced layer level and not including the packetized packets of the encoded data at the base layer level, and the transmitting step may transmit the first encoded stream using a first transmission path and transmit the second encoded stream using a second transmission path different from the first transmission path.

[0022] In addition, in the generating step, the first encoded stream and the second encoded stream may be generated in accordance with different multiplexing methods.

[0023] In addition, in the generating step, one of the first encoded stream and the second encoded stream may be generated in accordance with MPEG2-TS (Moving Picture Experts Group-2 Transport Stream), and the other of the first encoded stream and the second encoded stream may be generated in accordance with MMT (MPEG Media Transport).

[0024] Furthermore, one of the first transmission path and the second transmission path may be a transmission path for broadcasting, and the other of the first transmission path and the second transmission path may be a transmission path for communication.

[0025] In addition, the generating step may generate the encoded stream including information indicating the correspondence relationship, and the transmitting step may transmit the encoded stream including the information indicating the correspondence relationship.

[0026] In addition, the information indicating the correspondence may include either information indicating that the encoded stream can be decoded independently, or information indicating other encoded streams necessary for decoding the encoded stream.

[0027] These comprehensive or specific aspects may be realized by a system, an apparatus, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. These comprehensive or specific aspects may be realized by any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium.

[0028] Hereinafter, the embodiment will be specifically described with reference to the drawings.

[0029] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.

[0030] (Embodiment 1) [How to send] Hereinafter, a transmission method (transmission device) according to embodiment 1 will be described with reference to the drawings. In embodiment 1, a transmission method for transmitting encoded data in accordance with MMT will be described as an example.

[0031] First, the data structure of an encoded stream in MMT will be described. Figures 2 and 3 are diagrams for explaining the data structure of an encoded stream in MMT.

[0032] As shown in Fig. 2, the coded data is composed of a plurality of AUs (Access Units). The coded data is, for example, AV data coded based on a video coding standard such as HEVC. Specifically, the coded data includes video data, audio data, and metadata, still images, files, etc. associated therewith. When the coded data is video data, one AU is a unit equivalent to one picture (one frame).

[0033] In MMT, encoded data is converted into MP4 data (with an MP4 header) in units of GOPs (Group Of Pictures) according to the MP4 file format. The MP4 header included in the MP4 data describes the relative values ​​of the presentation time of the AU (the PTS mentioned above) and the decoding time (the DTS mentioned above). The MP4 header also describes the sequence number of the MP4 data. Note that MP4 data (MP4 file) is an example of an MPU (Media Processing Unit), which is a data unit defined in the MMT standard.

[0034] In the following, an example will be described in which MP4 data (file) is transmitted, but the transmitted data does not have to be MP4 data. For example, data in a file format different from an MP4 file may be transmitted, and as long as the encoded data and information required for decoding the encoded data (for example, information included in the MP4 header) are transmitted, the receiving side can decode the encoded data.

[0035] 3, an encoded stream 10 in MMT includes program information 11, time offset information 12, and a plurality of MMT packets 13. In other words, the encoded stream 10 is a packet sequence of the MMT packets 13.

[0036] The coded stream 10 (MMT stream) is one of one or more streams that make up one MMT package. The MMT package corresponds to, for example, one broadcast program content.

[0037] The program information 11 includes information indicating that the coded stream 10 is a scalable coded stream (a stream including both a base layer and an enhancement layer), and information on the type of scalable coding and the number of hierarchical levels (number of layers). Here, the type of scalable coding is time scalability, spatial scalability, SNR scalability, etc., and the number of hierarchical levels is the number of layers such as the base layer and the enhancement layer. Note that the program information 11 does not need to include all of the above information, and may include at least one of the information.

[0038] The program information 11 also includes, for example, information indicating the correspondence between a plurality of assets and packet IDs. An asset is a data entity including data with the same transport characteristics, and is, for example, either video data or audio data. The program information 11 may also include a descriptor indicating the hierarchical relationship of each packet ID (or asset).

[0039] Specifically, the program information 11 is CI (Composition Information) and MPT (MMT Package Table) in MMT. Note that the program information 11 is PMT (Program Map Table) in MPEG2-TS and MPD (Media Presentation Description) in MPEG-DASH.

[0040] The time offset information 12 is time information for determining the PTS or DTS of an AU. Specifically, the time offset information 12 is, for example, the absolute PTS or DTS of the first AU belonging to the base layer.

[0041] The MMT packet 13 is data in which MP4 data has been packetized. In the first embodiment, one MMT packet 13 includes one MP4 data (MPU). As shown in Fig. 3, the MMT packet 13 includes a header 13a (MTT packet header, or TS packet header in the case of MPEG2-TS) and a payload 13b.

[0042] MP4 data is stored in the payload 13b. Note that the payload 13b may store divided MP4 data.

[0043] The header 13a is additional information related to the payload 13b, for example, the header 13a includes a packet ID.

[0044] The packet ID is an identification number that indicates an asset of data included in the MMT packet 13 (payload 13b). The packet ID is an identification number that is unique to each asset that constitutes an MMT package.

[0045] The encoded stream 10 is characterized in that the video data of the base layer and the video data of the enhancement layer are treated as different assets. That is, different packet IDs are assigned to the MMT packets 13 of the encoded stream 10 according to the layer level of the encoded data stored therein. Fig. 4 is a diagram showing the correspondence between the packet IDs and data (assets) in the encoded stream 10. Note that Fig. 4 is an example of the correspondence.

[0046] As shown in Fig. 4, in the first embodiment, a packet ID "1" is assigned to MMT packet 13 obtained by packetizing video data of the base layer (encoded data at the base layer level). That is, the packet ID "1" is described in header 13a. And a packet ID "2" is assigned to MMT packet 13 obtained by packetizing video data of the enhancement layer (encoded data at the enhancement layer level). That is, the packet ID "2" is described in header 13a.

[0047] Similarly, the MMT packet 13 obtained by packetizing the audio data is assigned a packet ID of "3," and the MMT packet 13 obtained by packetizing the time offset information 12 is assigned a packet ID of "4." The MMT packet 13 obtained by packetizing the program information 11 is assigned a packet ID of "5."

[0048] 4 is described in the program information 11 of the coded stream 10. The above correspondence includes information indicating that the MMT packet 13 to which the packet ID "1" is assigned and the MMT packet 13 to which the packet ID "2" is assigned are a pair and are used for scalability.

[0049] The above-described method (transmission device) for transmitting the coded stream 10 according to the first embodiment will be described. Fig. 5 is a block diagram showing the configuration of the transmission device according to the first embodiment. Fig. 6 is a flowchart of the transmission method according to the first embodiment.

[0050] 5, the transmitting device 15 includes an encoding unit 16, a multiplexing unit 17, and a transmitting unit 18. Specifically, the components of the transmitting device 15 are realized by a microcomputer, a processor, or a dedicated circuit.

[0051] In the method of transmitting an encoded stream 10 according to the first embodiment, first, an encoded stream 10 is generated, which includes an MMT packet 13 to which a packet ID is assigned and information indicating the correspondence between the packet ID and the layer level (S11).

[0052] Specifically, when packetizing the encoded data output from the encoding unit 16, the multiplexing unit 17 determines (selects) a packet ID according to the layer level of the encoded data. Then, the multiplexing unit 17 generates an MMT packet 13 including the determined packet ID. Meanwhile, the multiplexing unit 17 generates information indicating the above-mentioned correspondence relationship. Then, the multiplexing unit 17 generates an encoded stream 10 including the generated MMT packet 13 and the generated information indicating the correspondence relationship.

[0053] The generated coded stream 10 is transmitted by the transmitter 18 using a transmission path (S12).

[0054] In this way, if an encoded stream 10 including MMT packets 13 to which different packet IDs are assigned according to the hierarchical level of the encoded data is transmitted, the encoded data can be easily selected on the receiving side by utilizing a conventional packet filter mechanism.

[0055] The information indicating the correspondence between the packet ID and the layer level may not be included in the encoded stream 10 and may be transmitted separately from the encoded stream 10. Also, if the correspondence between the packet ID and the layer level is already known on the receiving side, the information indicating the correspondence between the packet ID and the layer level does not need to be transmitted.

[0056] For example, the information indicating the above-mentioned correspondence may be included in program information that is repeatedly inserted into a continuous signal such as a broadcast signal, or may be obtained from a communication server before the start of decoding.

[0057] [How to receive] Hereinafter, a receiving method (receiving device) according to the first embodiment will be described with reference to the drawings. Fig. 7 is a block diagram showing the configuration of the receiving device according to the first embodiment. Fig. 8 is a flowchart of the receiving method according to the first embodiment.

[0058] In the following description, the base layer may be referred to as layer level A, and the enhancement layer may be referred to as layer level B.

[0059] 7, the receiving device 20 includes a packet filter 21, a program information analysis unit 22, a control unit 23, a packet buffer 24, a decoding unit 25, and a presentation unit 26. Among the components of the receiving device 20, the components other than the packet buffer 24 and the presentation unit 26 are specifically realized by a microcomputer, a processor, or a dedicated circuit. The packet buffer 24 is, for example, a storage device such as a semiconductor memory. The presentation unit 26 is, for example, a display device such as a liquid crystal panel.

[0060] 8, first, packet filter 21 separates MMT packets 13 included in encoded stream 10 (S21), and outputs program information 11 to program information analysis unit 22. Here, packet filter 21 recognizes the packet ID of MMT packet 13 including program information 11 in advance (or can acquire the packet ID of MMT packet 13 including program information 11 from other control information), and therefore can separate MMT packet 13 including program information 11 from encoded stream 10.

[0061] Next, the program information analysis unit 22 analyzes (S22) the program information 11. The program information 11 includes the correspondence between the packet IDs and the assets as described above.

[0062] Meanwhile, the control unit 23 determines which layer level of encoded data (MMT packets 13) to extract (S23). This determination may be made based on a user input received by an input receiving unit (not shown in FIG. 7), or may be made according to the specifications of the presentation unit 26 (for example, a frame rate supported by the presentation unit 26, etc.).

[0063] Then, the packet filter 21 extracts (filters) the encoded data (MMT packets 13) of the determined layer level under the control of the control unit 23 (S24). The control unit 23 recognizes the packet ID for each layer level through the analysis by the program information analysis unit 22, and therefore can cause the packet filter 21 to extract the encoded data of the determined layer level.

[0064] Next, the packet buffer 24 buffers the coded data extracted by the packet filter 21, and outputs it to the decoding unit 25 at the timing of the DTS (S25). The timing of the DTS is calculated based on the program information 11, the time offset information 12, and, for example, the time information transmitted in the MP4 header. Note that, in cases where the same DTS is assigned to the coded data of the base layer and the coded data of the enhancement layer due to spatial scalability or the like, the decoding order may be rearranged so that the coded data of the base layer is decoded prior to the coded data of the enhancement layer.

[0065] The encoded data buffered by the packet buffer 24 is decoded by the decoding unit 25 and presented (displayed) by the presenting unit 26 at the timing of the PTS (S26). The timing of the PTS is calculated based on the program information 11, the time offset information 12, and the time information in the MP4 header.

[0066] Such a receiving method will be further explained with reference to Fig. 9. Fig. 9 is a diagram conceptually showing the receiving method according to the first embodiment.

[0067] 9, for example, when it is determined that the layer level is layer level A (the extraction target is only the encoded data of the base layer), the packet filter 21 extracts all MMT packets 13 with the packet ID "1" assigned, and does not extract MMT packets 13 with the packet ID "2." As a result, a video with a low frame rate (for example, 60 fps) is displayed by the presentation unit 26.

[0068] Also, for example, when it is determined that the layer level is layer level A+B (the extraction target is both the coded data of the base layer and the coded data of the enhancement layer), the packet filter 21 extracts all MMT packets 13 to which the packet ID "1" or "2" is assigned. As a result, a video with a high frame rate (for example, 120 fps) is displayed by the presentation unit 26.

[0069] In this way, in the receiving device 20, the packet filter 21 can be used to easily distinguish between the coded data at the base layer level and the coded data at the enhancement layer level.

[0070] (Embodiment 2) [Send / receive method] A transmission method and a reception method (reception device) according to the second embodiment will be described below with reference to the drawings. Fig. 10 is a block diagram showing the configuration of a reception device according to the second embodiment. Figs. 11 and 12 are diagrams for explaining an outline of the transmission method according to the second embodiment. Note that the block diagram of the transmission device and the flowcharts of the reception method and the transmission method are almost the same as those explained in the first embodiment except for the use of a layer level ID, and therefore explanations thereof will be omitted.

[0071] As shown in FIG. 10, a receiving device 20a according to the second embodiment differs from the receiving device 20 in that a hierarchical filter 27 is provided.

[0072] As shown in (1) of Figure 11, in the encoded stream transmitted in the transmission method of embodiment 2, the same packet ID (packet ID: 1 in Figure 10) is assigned to each of MMT packets 13 of the base layer and MMT packets 13 of the enhancement layer.

[0073] Then, a layer level ID, which is an identifier related to a layer level, is assigned separately from the packet ID to the MMT packets 13 to which the same packet ID is assigned. In the example of Fig. 10, for example, the layer level ID is assigned A to the base layer and B to the enhancement layer. The packet ID and layer level ID are described, for example, in the header 13a (MTT packet header) corresponding to the MTT packet.

[0074] The hierarchical level ID may be defined as a new identifier or may be realized using private user data or other identifiers.

[0075] When the TS packet header is used, the layer level ID may be defined as a new identifier or may be realized using an existing identifier. For example, a function equivalent to the layer level ID can be realized by using either or both of the transport priority identifier and the elementary stream priority identifier.

[0076] As shown in (2) of Fig. 11 and (2) of Fig. 12, the transmitted coded stream is packet-filtered by the packet filter 21 of the receiving device 20a. That is, the transmitted coded stream is filtered based on the packet ID added to the packet header.

[0077] As shown in (3) of Figure 11 and (3) of Figure 12, the packet-filtered MMT packet 13 is further layer-level filtered by the layer filter 27 based on the layer level ID. The filtered encoded data is temporarily buffered in the packet buffer 24, and then decoded by the decoder 25 at the timing of the DTS. Then, as shown in (4) of Figure 11 and (4) of Figure 12, the decoded data is presented by the presenter 26 at the timing of the PTS.

[0078] Here, to obtain a video in which only the base layer is decoded (for example, a 60 fps video), it is sufficient to decode only the MMT packets 13 (encoded data) of the layer level ID "A" of the low layer. Therefore, in the layer level filtering, only the MMT packets 13 of the layer level ID "A" are extracted.

[0079] On the other hand, to obtain a video in which the base layer and the enhancement layer have been decoded (e.g., a 120 fps video), it is necessary to decode both the MMT packets 13 with layer level ID "A" in the low layer and the MMT packets 13 with layer level ID "B" in the high layer. Therefore, in the layer level filtering, both the MMT packets 13 with layer level ID "A" and the MMT packets 13 with layer level ID "B" are extracted.

[0080] In this way, the receiving method (receiving device 20a) of embodiment 2 has two series: a series that filters only MMT packets 13 with a layer level ID of "A" and decodes and presents video of only the base layer, and a series that filters MMT packets 13 with a layer level ID of "A" or "B" and decodes and presents video of the base layer + enhancement layer.

[0081] In addition, which packet ID or layer level ID to filter in packet filtering and layer level filtering is determined taking into consideration the type of scalable encoding described in program information 11, information on the number of layers, and which layer of encoded data the receiving device 20a will decode and display.

[0082] Such a determination is made by the receiving device 20a depending on, for example, the processing capability of the receiving device 20a. The transmitting device may further transmit, as signaling information, information related to the capability of the receiving device 20a required for decoding and displaying the content. In such a case, the receiving device 20a makes the above determination by comparing the signaling information with the capability of the receiving device 20a.

[0083] It is also possible to provide a filter section that combines the packet filter 21 and the layer level filter 27 into one, and the filter section may perform filtering collectively based on the packet ID and the layer level ID.

[0084] As described above, according to the transmission / reception method of the second embodiment, coded data can be selected for each layer level by filtering the layer level ID. That is, coded data can be easily selected on the receiving side. In addition, since the packet ID and the layer level ID are assigned separately, the coded data of the base layer and the coded data of the enhancement layer can be treated as the same stream in the packet filtering.

[0085] In addition, by assigning a layer level ID to the packet, it becomes possible to extract the desired hierarchically encoded data by a filtering operation alone, making reassembly unnecessary.

[0086] Furthermore, by being able to extract desired hierarchically encoded data using a hierarchical filter, a receiving device capable of decoding only the base layer can reduce the memory required to buffer data packets of the enhancement layer.

[0087] [Example 1] In MMT, an MPU made up of MP4 data can be fragmented into MFUs (Media Fragment Units), and a header 13a can be added in units of MFUs to generate an MMT packet 13. Here, an MFU can be fragmented down to the minimum NAL unit unit.

[0088] An example of packetizing encoded data in fragmented MFU units will be described below as a specific example 1 of the second embodiment. Figures 13, 14, and 15 are diagrams for explaining an example of packetizing encoded data in fragmented MFU units. Note that in Figures 13, 14, and 15, white AUs indicate AUs in the base layer, and hatched AUs indicate AUs in the enhancement layer (the same applies to the following Figures 16 to 18).

[0089] When fragmented MFUs are packetized, the same packet ID is assigned to the packet ID of the MMT packet header, and further, a layer level ID is assigned to the MMT packet header. Furthermore, an ID indicating that it is common is assigned to the MMT packet header of common data (common information) unrelated to the layer level among 'ftyp', 'moov', 'moof', etc. In Fig. 13, layer level A is the base layer, layer level B is the enhancement layer, and layer level Z is the common information. However, the layer levels of the base layer and the common information may be the same.

[0090] In such a configuration, the coded data of layer level B is treated as one asset. In the receiving device 20a, after filtering based on the packet ID is performed, layer level filtering becomes possible.

[0091] When it is desired to decode both the coded data of the base layer and the coded data of the enhancement layer, the receiving device 20a performs filtering based on the packet ID, and then extracts all layer level IDs by filtering based on the layer level ID. That is, in the layer level filtering, all of the layer level A: base layer, layer level B: enhancement layer, and layer level Z: common information are extracted. The extracted data is as shown in FIG. 14.

[0092] When it is desired to decode only the coded data of the base layer, the receiving device 20a performs filtering based on the packet ID, and then filters and extracts the layer level A: base layer and the layer level Z: common information. The extracted data is as shown in (a) of FIG.

[0093] At this time, since the AUs of the enhancement layer are removed, the AUs of the base layer are acquired in a packed state in the decoding unit 25 as shown in (b) of Fig. 15. However, the time offset information and data size of the sample (AU) described in 'moof' are information generated with the enhancement layer included. For this reason, a mismatch occurs between the information described in the header and the actual data.

[0094] Therefore, it is necessary to store information required for reconstructing the MP4 data, such as separately storing the size and offset information of the removed AUs.

[0095] Therefore, when acquiring the AU of the base layer, or the DTS and PTS, the decoding unit 25 may perform the decoding process taking into account that the AU of the enhancement layer (removed by filtering) does not exist in 'mdat' in the header information such as 'moof'.

[0096] For example, the offset information of the access unit (sample) in 'moof' is set as if the AU of the enhancement layer exists. Therefore, when the decoding unit 25 acquires only the base layer, the size of the removed AU is subtracted from the offset information. The data obtained by subtraction is typically as shown in (c) of FIG. 15.

[0097] Similarly, the DTS and PTS are calculated based on the sample_duration (the difference in DTS between consecutive access units) and sample_composition_time_offset (the difference between the DTS and PTS in an access unit) corresponding to the AU of the removed enhancement layer.

[0098] Instead of the subtraction as described above, header information for decoding data extracted from only the base layer (header information for AUs with only the base layer) may be described in advance in the MP4 header. Furthermore, the MP4 header may be described with information for distinguishing between header information for decoding only the base layer and header information for decoding both the base layer and the enhancement layer.

[0099] [Example 2] As a second concrete example of the second embodiment, an example in which the MPU is not fragmented but packetized on an MPU-by-MPU basis will be described below.

[0100] First, an example of arranging coded data in MP4 data in its original order will be described. Fig. 16 is a diagram showing an example of arranging coded data in MP4 data in its original order (when AUs of different layer levels are multiplexed simultaneously).

[0101] When the encoded data is placed directly in the MP4 data, base layer AUs and enhancement layer AUs are mixed in one track in the 'mdat' box. In this case, a layer level ID is assigned to each AU. The layer level ID for each AU (sample) is described in 'moov' or 'moof'. Note that it is preferable to place corresponding base layer AUs and enhancement layer AUs in the same 'mdat' box. Note that when a transport header is added to the MP4 data and packetized, the same packet ID is assigned.

[0102] In the above configuration, filtering on a packet basis is not possible because a layer level ID cannot be added to the packet header. Filtering is possible by analyzing the MP4 data.

[0103] As another method, the base layer and the enhancement layer of the encoded data are divided into tracks, and the correspondence is described in the header.

[0104] The data packetized in this manner is packet filtered in the receiving device 20a, after which the layer level of the AU is determined when the MP4 data is analyzed, and the AU of the desired layer is extracted and decoded.

[0105] Next, a first example of arranging encoded data in MP4 data for each hierarchical level will be described. Fig. 17 is a diagram showing a first example of arranging encoded data in MP4 data for each hierarchical level.

[0106] When coding data is placed in MP4 data for each hierarchical level, the coding data is separated for each hierarchical level and placed in the 'mdat' box of the fragment for each hierarchical level. In this case, the hierarchical level ID is described in 'moof'. The common header is assigned a hierarchical level ID to indicate that the information is common regardless of the layer.

[0107] In addition, the same packet ID is assigned to the packet header. In this case, filtering on a packet basis is also not possible.

[0108] The data packetized in this manner is packet filtered in the receiving device 20a, after which the layer level of the fragment is determined when the MP4 data is analyzed, and the fragment of the desired layer is extracted and decoded.

[0109] Finally, a second example of arranging encoded data in MP4 data for each hierarchical level will be described. Fig. 18 is a diagram for explaining a second example of arranging encoded data in MP4 data for each hierarchical level.

[0110] In this example, the coded data is separated for each hierarchical level and placed in an 'mdat' box for each hierarchical level.

[0111] MP4 data in which the AUs of the base layer are stored and MP4 data in which the AUs of the enhancement layer are stored are generated.

[0112] The layer level ID is written in either or both of the MP4 data header and the transport packet header. Here, the layer level ID indicates the hierarchical relationship between MP4 data or transport packets. The same packet ID is assigned to the packet header.

[0113] The data packetized in this manner is packet filtered in the receiving device 20a, and then packets of a desired layer are extracted and decoded based on the layer level ID in the packet header.

[0114] [Variations] In the above embodiment 2, it has been described that the packet ID and the layer level ID are assigned separately, but the layer level ID may be assigned using some bits of the packet ID, or new bits may be assigned as an extended packet ID. Note that assigning the layer level ID using some bits of the packet ID is equivalent to assigning a different packet ID for each layer level based on the rule that the same ID is assigned except for the bit indicating the layer level ID.

[0115] In addition, in the above-described second embodiment, when both the coded data of the base layer and the coded data of the enhancement layer are decoded, the data of the base layer and the enhancement layer are extracted by filtering using packet filtering or layer level filtering. However, the coded data may be reconstructed after being once separated into the base layer and the enhancement layer by layer level filtering.

[0116] (Other embodiments) The present invention is not limited to the above-described embodiment.

[0117] In the above first and second embodiments, the coded stream multiplexed according to MMT has been described. However, the coded stream may be multiplexed according to other multiplexing methods such as MPEG2-TS or RTP (Real Transport Protocol). Also, the MMT packets may be transmitted according to MPEG-TS2. In either case, the receiving side can easily select the coded data.

[0118] In the above-mentioned first and second embodiments, the coded data of the base layer and the coded data of the enhancement layer are mixed in one coded stream. However, a first coded stream consisting of the coded data of the base layer and a second coded stream consisting of the coded data of the enhancement layer may be generated separately. Here, the first coded stream is, more specifically, a coded stream including packets in which the coded data of the base layer is packetized, and including no packets in which the coded data of the enhancement layer is packetized. The second coded stream is a coded stream including packets in which the coded data of the enhancement layer is packetized, and including no packets in which the coded data of the base layer is packetized.

[0119] In such a case, the first encoded stream and the second encoded stream may be generated according to different multiplexing methods. For example, one of the first encoded stream and the second encoded stream may be generated according to MPEG2-TS, and the other of the first encoded stream and the second encoded stream may be generated according to MMT.

[0120] When the two encoded streams are generated according to different multiplexing methods, the packet ID or layer level ID is assigned according to each multiplexing method.In this case, the correspondence between the packet ID and the layer level in each encoded stream or the correspondence between the layer level ID and the layer level is described in the common program information.Specifically, only one of the two encoded streams includes the common program information, or both of the two encoded streams include the common program information.

[0121] In the receiving device, packet filtering and layer level filtering are performed based on the correspondence described in the program information, and the coded data of the desired layer level is extracted and decoded. That is, one image is displayed from two coded streams.

[0122] Also, the first encoded stream and the second encoded stream may be transmitted using (physically) different transmission paths. Specifically, for example, one of the first encoded stream and the second encoded stream may be transmitted using a transmission path for broadcasting, and the other of the first encoded stream and the second encoded stream may be transmitted using a transmission path for communication. Such transmission is assumed, for example, in the case of hierarchical transmission or bulk transmission across channels. In this case, the correspondence between each packet ID and layer level ID is described in common program information. Note that the program information does not have to be common. It is sufficient that the receiving device recognizes the correspondence between the packet ID and the layer level, or the correspondence between the layer level ID and the layer level.

[0123] In the above first and second embodiments, the coded stream is described as being composed of two layers, that is, one base layer and one enhancement layer, but the coded stream may be composed of three or more layer levels by configuring the enhancement layers in multiple stages. In this case, a different packet ID (or a different layer level ID) is assigned to each of the three layer levels.

[0124] In the above-mentioned first and second embodiments, the transmitting device 15 includes the encoding unit 16, but the transmitting device does not have to have an encoding function. In this case, a coding device having an encoding function is provided separately from the transmitting device 15.

[0125] Similarly, in the above-mentioned first and second embodiments, the receiving devices 20 and 20a include the decoding unit 25, but the receiving devices 20 and 20a may not have the decoding function. In this case, a decoding device having the decoding function is provided separately from the receiving devices 20 and 20a.

[0126] In the above first and second embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0127] In addition, in the above-mentioned embodiment 1, the processes executed by a specific processing unit may be executed by another processing unit. Also, the order of multiple processes may be changed, and multiple processes may be executed in parallel.

[0128] In addition, a comprehensive or specific aspect of the present invention may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. In addition, a comprehensive or specific aspect of the present invention may be realized as any combination of a system, a method, an integrated circuit, a computer program, or a recording medium.

[0129] The present invention is not limited to these embodiments or their modifications. As long as the modifications do not deviate from the spirit of the present invention, the scope of the present invention also includes various modifications conceivable by those skilled in the art to the present embodiments or their modifications, or configurations constructed by combining components of different embodiments or their modifications.

[0130] A first transmitting device according to one aspect of the present invention transmits time-scalable encoded data. Here, the time-scalable encoded data includes base layer data that can be decoded using data included in the layer, and enhancement layer data that cannot be decoded alone and must be decoded together with the base layer data. Here, the base layer data is data for decoding, for example, 60p video, and the enhancement layer data is data for decoding, for example, 120p video by using together with the base layer data.

[0131] The base layer data is transmitted as a first asset to which a first packet ID is assigned, and the enhancement layer data is transmitted as a second asset to which a second packet ID is assigned. The packet ID is described in the header of a packet in which the data is stored. The transmitting device multiplexes and transmits the base layer data, the enhancement layer data, and the program information. Here, the program information may include, for example, an identifier indicating a hierarchical relationship of each packet ID (or an asset corresponding to each packet ID). Here, the information indicating the hierarchical relationship includes, for example, information indicating that the data of the first packet ID (first asset) can be decoded independently, and information indicating that the data of the second packet ID (second asset) cannot be decoded independently and needs to be decoded using the data of the first packet ID (first asset).

[0132] In addition, the program information may include at least one of information indicating that the stream constituting the program is a scalable coded stream (a stream including both a base layer and an enhancement layer), information indicating the type of scalable coding, information indicating the number of layers, and information related to the layer level.

[0133] Also, a first receiving device according to an aspect of the present invention receives time-scalable encoded data. Here, the time-scalable encoded data includes base layer data that can be decoded using data included in the layer, and enhancement layer data that cannot be decoded alone and must be decoded together with the base layer data. Here, the base layer data is data for decoding, for example, 60p video, and the enhancement layer data is data for decoding, for example, 120p video by using together with the base layer data.

[0134] The base layer data is transmitted as a first asset to which a first packet ID is assigned, and the enhancement layer data is transmitted as a second asset to which a second packet ID is assigned. The packet ID is described in the header of a packet in which the data is stored. The transmitting device multiplexes and transmits the base layer data, the enhancement layer data, and the program information. Here, the program information may include, for example, an identifier indicating a hierarchical relationship of each packet ID (or an asset corresponding to each packet ID). Here, the information indicating the hierarchical relationship includes, for example, information indicating that the data of the first packet ID (first asset) can be decoded independently, and information indicating that the data of the second packet ID (second asset) cannot be decoded independently and needs to be decoded using the data of the first packet ID (first asset).

[0135] In addition, the program information may include at least one of information indicating that the stream constituting the program is a scalable coded stream (a stream including both a base layer and an enhancement layer), information indicating the type of scalable coding, information indicating the number of layers, and information related to the layer level.

[0136] According to the above-mentioned first transmitting device and second receiving device, the receiving side can obtain a packet ID (asset) required for decoding data of a selected packet ID (asset) from the program information and perform filtering based on the packet ID (asset). For example, when playing a first asset to which a first packet ID is assigned, the data of the first packet ID (first asset) can be decoded independently, so the data of the first packet ID (first asset) is obtained by filtering. On the other hand, when playing a second asset to which a second packet ID is assigned, the data of the second packet ID (second asset) cannot be decoded independently and needs to be decoded using the data of the first packet ID (first asset), so the data of the first packet ID (first asset) and the data of the second packet ID (second asset) are obtained by filtering.

[0137] In the above configuration, the program information describes information indicating data of packet IDs (asset) other than the packet ID (asset) that is necessary for decoding data of each packet ID (asset). According to this configuration, even if the scalable encoded data includes three layers, namely, layer A, layer B that is decoded together with layer A, and layer C that is decoded together with layer A, when the layer to be played is selected, the receiving device can identify the packet ID (asset) that transmits the data required for decoding without performing complex judgment.

[0138] In particular, considering that in the future there may be cases where the hierarchical depth is three or more, or where data encoded using multiple types of scalable encoding is multiplexed and transmitted, the above configuration is useful as it can identify the packet ID (asset) that transmits the data required for decoding without making complex judgments. [Industrial Applicability]

[0139] The present invention can be applied to television broadcasting, video distribution, and the like, as a method for transmitting coded data that allows the receiving side to easily select hierarchically coded data. [Explanation of symbols]

[0140] 10 Encoded Stream 11 Program Information 12 Time offset information 13 MMT Packets 13a Header 13b Payload 15 Transmitting device 16 Encoding section 17 Multiplexer 18 Transmitter 20, 20a Receiving device 21 Packet Filters 22 Program Information Analysis Unit 23 Control Unit 24 Packet Buffers 25 Decoding section 26 Presentation section

Claims

[Claim 1] A method for transmitting coded data in which video is coded in a time hierarchical manner and the coding order and the display order are different, comprising the steps of: a generating step of generating packets in which the encoded data is packetized into packets that differ depending on a base layer level and an enhancement layer level, which are layer levels of the encoded data, the packets being associated with different asset IDs depending on at least the base layer level and the enhancement layer level, and the packets including information indicating the asset ID of the base layer level used for decoding the enhancement layer level and information for identifying the number of layers of the base layer level; a transmitting step of transmitting the packet including the base layer level and the packet including the enhancement layer level in an encoded stream, and transmitting the packet including the generated information; Different packet IDs are assigned to the packet containing the base layer level of the encoded data, the packet containing the enhancement layer level of the encoded data, and the packet containing the information. Transmission method.

Citation Information

Patent Citations

  • Stream transmission system, transmitter, receiver, stream transmission method, and program

    JP2012095053A

  • Image encoding method, image decoding method, memory management method, image encoding device, image decoding device, memory management device, and image encoding / decoding device

    WO2012096186A1