Transmission method, reception method, transmission device, and reception device

By converting variable-length packet headers and payloads into fixed-length packets and managing their transmission to avoid consecutive minimum sizes, the method ensures stable buffer operation for ultra-high definition video content, addressing decoding challenges in existing transmission methods.

JP2026015490AActive Publication Date: 2026-01-29PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025194193
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2015-08-03
Filing Date
2025-11-13
Publication Date
2026-01-29
Estimated Expiration
2036-07-12

AI Technical Summary

Technical Problem

Existing methods for transmitting ultra-high definition video content, such as 8K and 4K, face challenges in ensuring buffer operation at the receiving device due to the large processing load during decoding, particularly when using multiplexing methods like MMT, which may not guarantee buffer operation.

Method used

A transmission method that converts variable-length packet headers and payloads into fixed-length packets, ensuring they are stored and transmitted in a manner that avoids consecutive minimum packet sizes, using a reception buffer model to manage the packet flow and maintain buffer stability.

Benefits of technology

This approach guarantees buffer operation at the receiving device, enabling efficient and stable decoding of ultra-high definition video content by managing packet storage and transmission to prevent buffer underflow or overflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015490000001_ABST
    Figure 2026015490000001_ABST
Patent Text Reader

Abstract

To provide a transmission method and the like capable of guaranteeing a buffer operation of a reception device when transmitting data using an MMT system.SOLUTION: A transmission method for transmitting a plurality of transmission packets in a state where a specification by a reception buffer model is satisfied, the transmission packet including a variable-length packet header and a variable-length payload, the reception buffer model including a buffer configured to receive the transmission packet, convert a first packet configured by the variable-length packet header and the variable-length payload stored in the received transmission packet into a second packet having a fixed-length packet header subjected to header decompression, and output the second packet obtained by the conversion at a predetermined extraction rate, the plurality of transmission packets equal to or less than a predetermined number of packets are transmitted per unit time (S2301), and the predetermined number of packets is set so that the transmission packets of the minimum packet size do not continue.SELECTED DRAWING: Figure 86
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a transmission method, a reception method, a transmission device, and a reception device. [Background technology]

[0002] As broadcasting and communication services become more sophisticated, the introduction of ultra-high definition video content such as 8K (7680 x 4320 pixels, hereinafter also referred to as 8K4K) and 4K (3840 x 2160 pixels, hereinafter also referred to as 4K2K) is being considered. A receiving device needs to decode and display the received encoded data of ultra-high definition video in real time. However, video with a resolution such as 8K imposes a large processing load during decoding, making it difficult to decode such video in real time using a single decoder. Therefore, methods are being considered for achieving real-time processing by using multiple decoders to parallelize the decoding process, thereby reducing the processing load per decoder.

[0003] Furthermore, the encoded data is multiplexed based on a multiplexing method such as MPEG-2 TS (Transport Stream) or MMT (MPEG Media Transport) and then transmitted. For example, Non-Patent Document 1 discloses a technique for transmitting encoded media data packet by packet in accordance with MMT. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Information technology - High efficiency coding and media delivery in heterogeneous environment - Part1:MPEG media transport(MMT), ISO / IEC DIS 23008-1 Summary of the Invention [Problem to be solved by the invention]

[0005] Meanwhile, with the advancement of broadcasting and communication services, the introduction of ultra-high-definition video content such as 8K and 4K (3840x2160 pixels) is being considered. In multiplexing methods such as MMT (MPEG Media Transport) and MPEG-DASH, media data such as video, audio, and files are multiplexed into ISOBMFF (ISO based media file format) (MP4 format) files, and after being packetized into MMT packets, they are transmitted over broadcast or communication transmission paths.

[0006] However, in MMT transmission, it may not be possible to guarantee the buffer operation of the receiving device.

[0007] The present invention provides a transmission method that can guarantee the buffer operation of a receiving device when transmitting data using a method such as MMT. [Means for solving the problem]

[0008] In order to achieve the above object, a transmission method according to one aspect of the present invention is a transmission method for transmitting a plurality of transmission packets, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and a reception buffer model for receiving the plurality of transmission packets to be transmitted includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, stores the plurality of transmission packets up to a predetermined number of packets in a predetermined data unit, and transmits the plurality of transmission packets, and the plurality of transmission packets are actual packets different from invalid packets. and a second transmission packet storing the invalid packet, wherein, when storing the first transmission packets in the predetermined data unit, the number of the first transmission packets is counted as a first packet number, and when storing the second transmission packets in the predetermined data unit, the packet size of the second transmission packet is divided by the maximum packet size of the first transmission packet, and the quotient is obtained by adding 1 to the quotient, and the integer value is counted as a second packet number, and the plurality of transmission packets are stored in the predetermined data unit so that the sum of the first packet number and the second packet number obtained by counting is equal to or less than the predetermined packet number, and the predetermined packet number is a value set so that the transmission packets of the minimum packet size do not occur consecutively.

[0009] In order to achieve the above object, one aspect of the present invention provides a receiving method in a receiving device having a buffer, the receiving method including: receiving a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload; converting, using the buffer, first packets composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets into second packets having headers of a fixed length packet header that have been expanded; outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer being larger than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; and storing actual packets different from invalid packets in the plurality of transmission packets. The plurality of transmission packets stored in a predetermined data unit include a first transmission packet storing an invalid packet and a second transmission packet storing the invalid packet, and the plurality of transmission packets stored in a predetermined data unit are counted as a first packet count when the first transmission packet is stored in the predetermined data unit, and are counted as a second packet count when the second transmission packet is stored in the predetermined data unit, such that the sum of the first packet count and the second packet count obtained by counting is equal to or less than the predetermined packet count, and the predetermined packet count is a value set so that the transmission packets of the minimum packet size do not occur consecutively.

[0010] In order to achieve the above object, a transmission method according to one aspect of the present invention is a transmission method for transmitting a plurality of transmission packets, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and a reception buffer model for receiving the plurality of transmission packets includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, stores the plurality of transmission packets up to a predetermined number of packets in a predetermined data unit, and transmits the predetermined data unit in each of a plurality of consecutive unit transmission periods that do not overlap each other, thereby obtaining a predetermined number of packets per unit time. a first transmission packet storing unit configured to store the first transmission packets in the predetermined data unit, the first transmission packet storing an actual packet different from an invalid packet, and a second transmission packet storing the invalid packet; when storing the first transmission packets in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; when storing the second transmission packets in the predetermined data unit, the packet size of the second transmission packet is divided by the maximum packet size of the first transmission packet, and the quotient is obtained by adding 1 to the quotient; the second packet number is counted as an integer value; and the plurality of transmission packets are stored in the predetermined data unit such that the sum of the first packet number and the second packet number obtained by counting is not more than the predetermined packet number; the predetermined packet number is a value set so that the transmission packets of the minimum packet size do not occur consecutively.

[0011] Furthermore, a receiving method according to one aspect of the present invention is a receiving method in a receiving device having a buffer, the receiving method including the steps of: receiving a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload; converting, using the buffer, first packets each composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets into second packets each having a header-expanded fixed-length packet header; outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer being greater than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; and storing and transmitting the plurality of transmission packets in a predetermined data unit in each of a plurality of consecutive unit transmission periods over a period of unit time that does not overlap with one another, thereby the plurality of transmission packets stored in the predetermined data unit are packets that are not more than a predetermined number of packets transmitted per unit time, and include first transmission packets that store actual packets different from invalid packets and second transmission packets that store the invalid packets; when the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as the first packet number; when the second transmission packets are stored in the predetermined data unit, the number of the second transmission packets is counted as the second packet number, where the quotient is an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet; and the sum of the first packet number and the second packet number obtained by counting is not more than the predetermined packet number; the predetermined number of packets is a value set so that the transmission packets of the minimum packet size do not occur consecutively.

[0012] In order to achieve the above object, a transmission method according to one aspect of the present invention is a transmission method for transmitting a plurality of transmission packets, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and a reception buffer model for receiving the plurality of transmission packets includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, stores the plurality of transmission packets of a predetermined number of packets or less in a predetermined data unit, and transmits the predetermined data unit in each of a plurality of consecutive unit transmission periods that do not overlap each other, thereby transmitting the predetermined number of packets or less per unit time. the plurality of transmission packets include first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets, and in the storing, when storing the first transmission packets in the predetermined data unit, count the number of the first transmission packets as a first packet number, and when storing the second transmission packets in the predetermined data unit, count an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packets as a second packet number, and store the plurality of transmission packets in the predetermined data unit so that the sum of the first packet number and the second packet number obtained by counting is equal to or less than the predetermined packet number, the predetermined data unit is a transmission frame, and the unit transmission period is a transmission frame length corresponding to the transmission frame.

[0013] Furthermore, a receiving method according to one aspect of the present invention is a receiving method in a receiving device having a buffer, the receiving method including: receiving a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload; converting, using the buffer, first packets composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets into second packets having headers expanded to fixed-length packet headers; and outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer is greater than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; and storing and transmitting the plurality of transmission packets in a predetermined data unit in each of a plurality of consecutive unit transmission periods over a period of unit time that does not overlap with one another, thereby The transmission packets are not more than a predetermined number of transmitted packets, and include first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets. The plurality of transmission packets stored in the predetermined data unit are counted as the first packet number when the first transmission packets are stored in the predetermined data unit, and are counted as the second packet number when the second transmission packets are stored in the predetermined data unit by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packets, and are stored in the predetermined data unit so that the sum of the first packet number and the second packet number obtained by counting is not more than the predetermined packet number. The predetermined data unit is a transmission frame, and the unit transmission period is a transmission frame length corresponding to the transmission frame.

[0014] In order to achieve the above object, a transmission method according to one aspect of the present invention is a transmission method for transmitting a plurality of transmission packets, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and a reception buffer model for receiving the plurality of transmission packets includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, stores the plurality of transmission packets up to a predetermined number of packets in a predetermined data unit, and transmits the predetermined data unit in each of a plurality of consecutive unit transmission periods that do not overlap each other, thereby obtaining the plurality of transmission packets per unit time. a plurality of transmission packets equal to or less than a predetermined number of packets are transmitted, the plurality of transmission packets including first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets; when storing the first transmission packets in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; when storing the second transmission packets in the predetermined data unit, the packet size of the second transmission packet is divided by the maximum packet size of the first transmission packet, and the quotient obtained by adding 1 to the quotient is counted as a second packet number; and the plurality of transmission packets are stored in the predetermined data unit such that the sum of the first packet number and the second packet number obtained by counting is equal to or less than the predetermined packet number; and the maximum packet size of the first transmission packet is 1500 B.

[0015] Furthermore, a receiving method according to one aspect of the present invention is a receiving method in a receiving device having a buffer, the receiving method including: receiving a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload; converting, using the buffer, first packets each composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets into second packets each having a header-expanded fixed-length packet header; outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer being greater than a maximum buffer occupancy calculated using the transmission rate of the plurality of transmission packets, the average packet length of the transmission packets, and the extraction rate; and storing and transmitting the plurality of transmission packets in a predetermined data unit in each of a plurality of consecutive unit transmission periods over a unit time period that do not overlap with one another, The plurality of transmission packets stored in the predetermined data unit are transmission packets transmitted per unit time equal to or less than the predetermined number of packets, and include first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets, and the plurality of transmission packets are packets stored in the predetermined data unit such that, when the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as the first number of packets, and when the second transmission packets are stored in the predetermined data unit, the number of the second transmission packets is counted as the second number of packets, and an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet is counted as the second number of packets, and the sum of the first number of packets and the second number of packets obtained by counting is equal to or less than the predetermined number of packets, and the maximum packet size of the first transmission packet is 1500 B.

[0016] In order to achieve the above object, a transmission method according to one aspect of the present invention is a transmission method for transmitting a plurality of transmission packets, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and a reception buffer model for receiving the plurality of transmission packets includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, stores the plurality of transmission packets up to a predetermined number of packets in a predetermined data unit, and transmits the predetermined data unit in each of a plurality of consecutive unit transmission periods that do not overlap each other, thereby the plurality of transmission packets including first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets; when storing the first transmission packets in the predetermined data unit, the number of the first transmission packets is counted as the first packet number; when storing the second transmission packets in the predetermined data unit, the packet size of the second transmission packet is divided by the maximum packet size of the first transmission packets, and the quotient is obtained by adding 1 to the quotient; the plurality of transmission packets are stored in the predetermined data unit such that the sum of the first packet number and the second packet number obtained by counting is equal to or less than the predetermined packet number; the invalid packets are NULL packets.

[0017] Furthermore, a receiving method according to one aspect of the present invention is a receiving method in a receiving device having a buffer, the receiving method including: receiving a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload; converting, using the buffer, first packets composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets into second packets having headers expanded to fixed-length packet headers; outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer being greater than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; and transmitting the plurality of transmission packets in a unit time period that does not overlap with one another and for each of a plurality of consecutive unit transmission periods, the plurality of transmission packets being stored in a predetermined data unit in the unit transmission period and transmitted. When the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as the first number of packets, and when the second transmission packets are stored in the predetermined data unit, the number of the second transmission packets is counted as the second number of packets, where the number of the first transmission packets is the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet and adding 1 to the quotient. The invalid packets are NULL packets.

[0018] In order to achieve the above object, a transmission method according to one embodiment of the present invention is a transmission method for transmitting a plurality of transmission packets while satisfying the requirements of a predetermined reception buffer model to ensure buffer operation of a receiving device, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and the reception buffer model includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, and transmits the plurality of transmission packets in a predetermined number of packets or less per unit time.

[0019] Furthermore, a receiving method according to one embodiment of the present invention is a receiving method in a receiving device having a buffer, which receives a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload, converts first packets, each of which is composed of a variable-length packet header and a variable-length payload and is stored in the received plurality of transmission packets, into second packets having headers of fixed lengths that have been expanded using the buffer, and outputs the second packets obtained by the conversion from the buffer at a predetermined extraction rate, wherein the buffer size of the buffer is greater than a maximum buffer occupancy calculated using the transmission rate of the plurality of transmission packets, the average packet length of the transmission packets, and the extraction rate.

[0020] These general or specific aspects may be realized as a system, device, integrated circuit, computer program, or computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, device, integrated circuit, computer program, and recording medium. [Effects of the Invention]

[0021] The present invention can guarantee the buffer operation of the receiving device when transmitting data using a method such as MMT. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 is a diagram showing an example of dividing a picture into slice segments. [Figure 2] FIG. 2 is a diagram showing an example of a PES packet sequence in which picture data is stored. [Figure 3] FIG. 3 is a diagram showing an example of division of a picture according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of dividing a picture according to a comparative example of the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of data of an access unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a transmission device according to the first embodiment. [Figure 7] FIG. 7 is a block diagram of a receiving device according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of an MMT packet according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating another example of an MMT packet according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of data input to each decoding unit according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 12] FIG. 12 is a diagram showing another example of data input to each decoding unit according to the first embodiment. [Figure 13] FIG. 13 is a diagram showing an example of dividing a picture according to the first embodiment. [Figure 14] FIG. 14 is a flowchart of a transmission method according to the first embodiment. [Figure 15] FIG. 15 is a block diagram of a receiving device according to the first embodiment. [Figure 16] FIG. 16 is a flowchart of a receiving method according to the first embodiment. [Figure 17] FIG. 17 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 19] FIG. 19 is a diagram illustrating the configuration of the MPU. [Figure 20] FIG. 20 is a diagram showing the structure of MF metadata. [Figure 21] FIG. 21 is a diagram for explaining the data transmission order. [Figure 22] FIG. 22 is a diagram showing an example of a method for performing decoding without using header information. [Figure 23] FIG. 23 is a block diagram of a transmission device according to the second embodiment. [Figure 24] FIG. 24 is a flowchart of a transmission method according to the second embodiment. [Figure 25] FIG. 25 is a block diagram of a receiving device according to the second embodiment. [Figure 26] FIG. 26 is a flowchart of the operation for identifying the MPU start position and the NAL unit position. [Figure 27] FIG. 27 is a flowchart of an operation of obtaining initialization information based on a transmission order type and decoding media data based on the initialization information. [Figure 28] FIG. 28 is a flowchart of the operation of a receiving device when a low-delay presentation mode is provided. [Figure 29] FIG. 29 is a diagram illustrating an example of the transmission order of MMT packets when auxiliary data is transmitted. [Figure 30] FIG. 30 is a diagram illustrating an example in which a transmitting device generates auxiliary data based on the configuration of moof. [Figure 31] FIG. 31 is a diagram for explaining reception of auxiliary data. [Figure 32] FIG. 32 is a flowchart of a receiving operation using auxiliary data. [Figure 33]FIG. 33 shows the structure of an MPU that is made up of multiple movie fragments. [Figure 34] FIG. 34 is a diagram for explaining the transmission order of MMT packets when the MPU having the configuration of FIG. 33 is transmitted. [Figure 35] FIG. 35 is a first diagram illustrating an example of the operation of a receiving device when one MPU is made up of multiple movie fragments. [Figure 36] FIG. 36 is a second diagram illustrating an example of the operation of a receiving device when one MPU is made up of multiple movie fragments. [Figure 37] FIG. 37 is a flowchart of the operation of the receiving method described with reference to FIGS. [Figure 38] FIG. 38 is a diagram showing a case where non-VCL NAL units are aggregated as individual data units. [Figure 39] FIG. 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. [Figure 40] FIG. 40 is a flowchart showing the operation of the receiving device when a packet loss occurs. [Figure 41] FIG. 41 is a flowchart of the receiving operation when the MPU is divided into multiple movie fragments. [Figure 42] FIG. 42 is a diagram showing an example of a prediction structure of a picture for each TemporalId when temporal scalability is realized. [Figure 43] FIG. 43 shows the relationship between the decoding time (DTS) and the display time (PTS) for each picture in FIG. [Figure 44] FIG. 44 is a diagram showing an example of a prediction structure of a picture that requires picture delay processing and reorder processing. [Figure 45] Figure 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and an MMTP packet. [Figure 46]FIG. 46 is a diagram for explaining the calculation method and problems of the PTS and DTS. [Figure 47] FIG. 47 is a flowchart of the receiving operation when the DTS is calculated using information for DTS calculation. [Figure 48] FIG. 48 is a diagram for explaining a method of storing a data unit in a payload in MMT. [Figure 49] FIG. 49 is a flow chart showing the operation of the transmitting device according to the third embodiment. [Figure 50] FIG. 50 shows an operation flow of the receiving device according to the third embodiment. [Figure 51] FIG. 51 is a diagram illustrating an example of a specific configuration of a transmission device according to the third embodiment. [Figure 52] FIG. 52 is a diagram illustrating an example of a specific configuration of a receiving device according to the third embodiment. [Figure 53] Figure 53 shows how non-timed media is stored in the MPU and how it is transmitted in MMTP packets. [Figure 54] FIG. 54 shows an example in which a file is divided into a plurality of divided data pieces, each of which is packetized and transmitted. [Figure 55] FIG. 55 shows another example in which a plurality of divided data pieces obtained by dividing a file are packetized and transmitted. [Figure 56] FIG. 56 is a diagram showing the syntax of a loop for each file in the asset management table. [Figure 57] FIG. 57 shows the operational flow for identifying the divided data number in the receiving device. [Figure 58] FIG. 58 shows the operational flow for identifying the number of divided data pieces in the receiving device. [Figure 59] FIG. 59 shows an operational flow for determining whether to operate a fragment counter in a transmitting device. [Figure 60] FIG. 60 is a diagram for explaining a method for identifying the number of divided data pieces and divided data numbers (when a fragment counter is used). [Figure 61] Figure 61 shows the operational flow of a transmitting device when utilizing a fragment counter. [Figure 62] FIG. 62 shows the operational flow of a receiving device when utilizing a fragment counter. [Figure 63] FIG. 63 shows a service configuration in which the same program is transmitted using multiple IP data flows. [Figure 64] FIG. 64 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 65] FIG. 65 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 66] FIG. 66 shows the operational flow of the transmitting device. [Figure 67] FIG. 67 shows the operational flow of the receiving device. [Figure 68] FIG. 68 shows a reception buffer model based on the reception buffer model defined in ARIB STD B-60, particularly when only a broadcast transmission channel is used. [Figure 69] FIG. 69 is a diagram showing an example in which multiple data units are aggregated and stored in one payload. [Figure 70] FIG. 70 shows an example in which a plurality of data units are aggregated and stored in one payload, where a video signal in NAL size format is used as one data unit. [Figure 71] Figure 71 shows the structure of the payload of an MMTP packet in which the data unit length is not indicated. [Figure 72] FIG. 72 shows an example of the extend area assigned to each packet. [Figure 73] FIG. 73 shows the operational flow of the receiving device. [Figure 74] FIG. 74 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 75] FIG. 75 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 76] FIG. 76 shows the operational flow of the transmitting device. [Figure 77] FIG. 77 shows the operational flow of the receiving device. [Figure 78] FIG. 78 is a diagram showing a protocol stack of the MMT / TLV method defined in ARIB STD-B60. [Figure 79] FIG. 79 is a diagram showing the structure of a TLV packet. [Figure 80] FIG. 80 is a diagram illustrating an example of a block diagram of a receiving device. [Figure 81] FIG. 81 is a diagram showing the configuration of a transmission slot. [Figure 82] FIG. 82 is a diagram showing a transmission frame and one TLV stream (TLV packet sequence) having a specific tlv_stream_id stored in the transmission frame. [Figure 83] FIG. 83 is a flowchart of a transmission method when the buffer size of the TLV packet buffer and the maximum number of packets per transmission frame are specified. [Figure 84] FIG. 84 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 85] FIG. 85 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 86] FIG. 86 shows the operational flow of the transmitting device. [Figure 87] FIG. 87 shows the operational flow of the receiving device. DETAILED DESCRIPTION OF THE INVENTION

[0023] A transmission method according to one embodiment of the present invention is a transmission method for transmitting a plurality of transmission packets while satisfying the requirements of a predetermined reception buffer model to ensure buffer operation of a receiving device, wherein the transmission packets are composed of a variable-length packet header and a variable-length payload, and the reception buffer model includes a buffer that receives the transmission packets, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packets into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate, and transmits the plurality of transmission packets at a predetermined number of packets or less per unit time.

[0024] Such a transmitting device can guarantee the buffer operation of the receiving device when transmitting data using a method such as MMT.

[0025] For example, the plurality of transmission packets, each of which has a predetermined number of packets or less, may be stored in a predetermined data unit, and in the transmission, the predetermined data unit may be transmitted in each of a plurality of consecutive unit transmission periods during the unit time that do not overlap each other, thereby transmitting the plurality of transmission packets, each of which has a number of packets or less per unit time.

[0026] For example, the specified number of packets may include multiple packet numbers that are different from each other and are set according to the transmission rate of the specified data unit, and in the storage, the multiple transmission packets may be stored in the specified data unit in a number equal to or less than the number of packets that corresponds to the transmission rate of the specified data unit.

[0027] For example, the plurality of transmission packets may include first transmission packets storing actual packets different from invalid packets and second transmission packets storing the invalid packets, and in the storing, when the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets may be counted as a first packet number, and when the second transmission packets are stored in the predetermined data unit, an integer value obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packets may be counted as a second packet number, and the plurality of transmission packets may be stored in the predetermined data unit so that the sum of the first packet number and the second packet number obtained by counting is equal to or less than the predetermined packet number.

[0028] For example, the invalid packet may be a NULL packet.

[0029] For example, the maximum value of the packet size of the first transmission packet may be 1500B.

[0030] For example, the predetermined data unit may be a transmission frame, and the unit transmission period may be a transmission frame length corresponding to the transmission frame.

[0031] For example, the predetermined number of packets may be a value set so that the transmission packets of the minimum packet size are not consecutive.

[0032] A receiving method according to one embodiment of the present invention is a receiving method in a receiving device having a buffer, which receives a plurality of transmission packets, each of which is composed of a variable-length packet header and a variable-length payload, converts first packets composed of a variable-length packet header and a variable-length payload stored in the received plurality of transmission packets using the buffer into second packets having header-expanded fixed-length packet headers, and outputs the second packets obtained by the conversion from the buffer at a predetermined extraction rate, wherein the buffer size of the buffer is greater than a maximum buffer occupancy calculated using the transmission rate of the plurality of transmission packets, the average packet length of the transmission packets, and the extraction rate.

[0033] These comprehensive or specific aspects may be realized as a system, device, integrated circuit, computer program, or computer-readable recording medium such as a CD-ROM, or as any combination of a system, device, integrated circuit, computer program, or recording medium.

[0034] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0035] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.

[0036] (Findings that form the basis of the present invention) In recent years, the resolution of displays such as TVs, smartphones, and tablet devices has been increasing. In particular, 8K4K (8K x 4K resolution) services are scheduled for broadcasting in Japan in 2020. Since it is difficult to decode ultra-high resolution moving images such as 8K4K in real time using a single decoder, methods for decoding images in parallel using multiple decoders are being investigated.

[0037] Since the coded data is multiplexed and transmitted based on a multiplexing method such as MPEG-2 TS or MMT, the receiving device must separate the coded data of the video from the multiplexed data before decoding. Hereinafter, the process of separating the coded data from the multiplexed data will be referred to as demultiplexing.

[0038] When parallelizing the decoding process, it is necessary to allocate the coded data to be decoded to each decoder. When allocating the coded data, the coded data itself needs to be analyzed, and the processing load for analysis is large, especially for content such as 8K, where the bit rate is very high. Therefore, the demultiplexing part becomes a bottleneck, making real-time playback impossible.

[0039] In video coding standards such as H.264 and H.265 standardized by MPEG and ITU, a transmitting device divides a picture into multiple regions called slices or slice segments and encodes the picture so that each region can be decoded independently. Therefore, in the case of H.265, for example, a receiving device that receives a broadcast can separate data for each slice segment from received data and output the data for each slice segment to a separate decoder, thereby achieving parallel decoding.

[0040] 1 is a diagram showing an example in which one picture is divided into four slice segments in HEVC. For example, a receiving device includes four decoders, and each decoder decodes one of the four slice segments.

[0041] In conventional broadcasting, a transmitting device stores one picture (access unit in the MPEG system standard) in one PES packet and multiplexes the PES packet into a sequence of TS packets. Therefore, a receiving device must separate the payload of the PES packet, analyze the data of the access unit stored in the payload, separate each slice segment, and output the data of each separated slice segment to a decoder.

[0042] However, the present inventors have found that there is a problem in that it is difficult to perform this processing in real time because the amount of processing required to analyze the data of an access unit and separate the slice segments is large.

[0043] FIG. 2 is a diagram showing an example in which data of a picture divided into slice segments is stored in the payload of a PES packet.

[0044] 2, for example, data of a plurality of slice segments (slice segments 1 to 4) is stored in the payload of one PES packet, and the PES packets are multiplexed into a sequence of TS packets.

[0045] (Embodiment 1) In the following, an example will be described in which H.265 is used as the video encoding method, but this embodiment can also be applied to cases in which other encoding methods such as H.264 are used.

[0046] 3 is a diagram showing an example of dividing an access unit (picture) into division units in this embodiment. The access unit is divided into two equal parts horizontally and vertically, into a total of four tiles, by a function called tile introduced by H.265. Furthermore, there is a one-to-one correspondence between slice segments and tiles.

[0047] The reason for dividing the data into two equal parts horizontally and vertically will be explained below. First, when decoding, a line memory is generally required to store one horizontal line of data. However, when it comes to ultra-high resolutions such as 8K4K, the horizontal size increases, so the size of the line memory also increases. When implementing a receiving device, it is desirable to be able to reduce the size of the line memory. In order to reduce the size of the line memory, vertical division is necessary. Vertical division requires a data structure called a tile. For these reasons, tiles are used.

[0048] On the other hand, since images generally have high correlation in the horizontal direction, being able to refer to a wider range in the horizontal direction improves coding efficiency. Therefore, from the viewpoint of coding efficiency, it is desirable to divide the access unit horizontally.

[0049] By dividing the access unit into two equal parts horizontally and vertically, these two characteristics can be achieved simultaneously, and both implementation and coding efficiency can be taken into consideration.If a single decoder can decode 4K2K video in real time, by dividing an 8K4K image into four equal parts and dividing each slice segment into 4K2K, the receiving device can decode the 8K4K image in real time.

[0050] Next, we will explain why tiles obtained by dividing an access unit in the horizontal and vertical directions correspond one-to-one to slice segments. In H.265, an access unit is composed of multiple units called NAL (Network Adaptation Layer) units.

[0051] The payload of an NAL unit stores any of the following: an access unit delimiter indicating the start position of an access unit; an SPS (Sequence Parameter Set) which is initialization information used in common for each sequence at the time of decoding; a PPS (Picture Parameter Set) which is initialization information used in common for each picture at the time of decoding; SEI (Supplemental Enhancement Information) which is not required for the decoding process itself but is required for processing and displaying the decoding results; and coded data of a slice segment. The header of an NAL unit contains type information for identifying the data stored in the payload.

[0052] Here, the transmitting device transmits the coded data in MPEG-2 TS, MMT (MPEG Media Transport), MPEG DASH (Dynamic Adaptive Streaming) format. When multiplexing using a multiplexing format such as Streaming over HTTP (Streaming over HTTP) or RTP (Real-time Transport Protocol), the basic unit can be set to an NAL unit. In order to store one slice segment in one NAL unit, it is desirable to divide the access unit into regions by slice segment units. For this reason, the transmitting device associates tiles with slice segments in a one-to-one correspondence.

[0053] As shown in Fig. 4, the transmitting device can also set tiles 1 to 4 together in one slice segment. In this case, however, all tiles are stored in one NAL unit, making it difficult for the receiving device to separate the tiles in the multiplexing layer.

[0054] Note that there are two types of slice segments: independent slice segments that can be decoded independently, and reference slice segments that refer to independent slice segments. Here, a case where an independent slice segment is used will be described.

[0055] Fig. 5 is a diagram showing an example of data of an access unit divided so that the boundaries between tiles and slice segments coincide as shown in Fig. 3. The data of the access unit includes a NAL unit in which an access unit delimiter placed at the beginning is stored, followed by NAL units of SPS, PPS, and SEI, and data of slice segments in which data from tiles 1 to 4 is stored. Note that the data of the access unit does not have to include some or all of the NAL units of SPS, PPS, and SEI.

[0056] Next, a configuration of transmitting apparatus 100 according to this embodiment will be described. Fig. 6 is a block diagram showing an example configuration of transmitting apparatus 100 according to this embodiment. This transmitting apparatus 100 includes an encoding unit 101, a multiplexing unit 102, a modulating unit 103, and a transmitting unit 104.

[0057] The encoding unit 101 generates encoded data by encoding an input image according to, for example, H.265. Furthermore, the encoding unit 101 divides an access unit into four slice segments (tiles) and encodes each slice segment, for example, as shown in FIG. 3 .

[0058] The multiplexing unit 102 multiplexes the coded data generated by the coding unit 101. The modulation unit 103 modulates the data obtained by the multiplexing. The transmission unit 104 transmits the modulated data as a broadcast signal.

[0059] Next, a configuration of receiving apparatus 200 according to this embodiment will be described. Fig. 7 is a block diagram showing an example configuration of receiving apparatus 200 according to this embodiment. Receiving apparatus 200 includes tuner 201, demodulation unit 202, demultiplexing unit 203, a plurality of decoding units 204A to 204D, and display unit 205.

[0060] The tuner 201 receives a broadcast signal. The demodulation unit 202 demodulates the received broadcast signal. The demodulated data is input to the demultiplexing unit 203.

[0061] The demultiplexing unit 203 separates the demodulated data into division units and outputs the data for each division unit to the decoding units 204A to 204D. Here, a division unit is a divided area obtained by dividing an access unit, such as a slice segment in H.265. Also, here, an 8K4K image is divided into four 4K2K images. Therefore, there are four decoding units 204A to 204D.

[0062] The decoding units 204A to 204D operate in synchronization with one another based on a predetermined reference clock. Each decoding unit decodes the coded data in division units in accordance with the DTS (Decoding Time Stamp) of the access unit, and outputs the decoding result to the display unit 205.

[0063] The display unit 205 generates an 8K4K output image by integrating the multiple decoding results output from the multiple decoding units 204A to 204D. The display unit 205 displays the generated output image in accordance with the PTS (Presentation Time Stamp) of the access unit acquired separately. When integrating the decoding results, the display unit 205 may perform filtering such as a deblocking filter on boundary regions between adjacent division units, such as tile boundaries, so that the boundaries become less noticeable visually.

[0064] In the above description, the transmitting device 100 and the receiving device 200 that transmit or receive broadcasts are used as examples, but content may be transmitted and received via a communication network. When the receiving device 200 receives content via a communication network, the receiving device 200 separates the multiplexed data from IP packets received via a network such as Ethernet.

[0065] In broadcasting, the transmission path delay from when a broadcast signal is transmitted until it reaches the receiving device 200 is constant. On the other hand, in communication networks such as the Internet, due to the effects of congestion, the transmission path delay from when data transmitted from a server reaches the receiving device 200 is not constant. Therefore, the receiving device 200 often does not perform strictly synchronized playback based on a reference clock such as PCR in the broadcast MPEG-2 TS. Therefore, the receiving device 200 may display the 8K4K output image on the display device according to the PTS without strictly synchronizing each decoding unit.

[0066] Furthermore, due to congestion in the communication network, etc., the decoding process for all division units may not be completed by the time indicated by the PTS of the access unit. In this case, the receiving device 200 skips displaying the access unit or delays displaying it until decoding of at least four division units is completed and generation of the 8K4K image is completed.

[0067] The content may be transmitted and received by combining broadcasting and communication. The present method can also be applied to playing back multiplexed data stored on a recording medium such as a hard disk or memory.

[0068] Next, a method for multiplexing access units divided into slice segments when MMT is used as the multiplexing method will be described.

[0069] 8 is a diagram showing an example of packetizing data of an access unit of HEVC into MMT. Although the SPS, PPS, SEI, and the like do not necessarily need to be included in the access unit, the case where they are present is illustrated here.

[0070] NAL units that are located before the first slice segment in an access unit, such as the access unit delimiter, SPS, PPS, and SEI, are stored together in MMT packet #1. Subsequent slice segments are stored in separate MMT packets for each slice segment.

[0071] As shown in FIG. 9, an NAL unit that is arranged before the first slice segment in an access unit may be stored in the same MMT packet as the first slice segment.

[0072] Furthermore, when NAL units such as End-of-Sequence or End-of-Bitstream, which indicate the end of a sequence or stream, are added after the last slice segment, they are stored in the same MMT packet as the last slice segment. However, since NAL units such as End-of-Sequence or End-of-Bitstream are inserted at the end point of the decoding process or the connection point of two streams, it may be desirable for receiving device 200 to easily obtain these NAL units in the multiplexing layer. In this case, these NAL units may be stored in an MMT packet separate from the slice segment. This allows receiving device 200 to easily separate these NAL units in the multiplexing layer.

[0073] Note that TS, DASH, RTP, or the like may be used as the multiplexing method. In these methods, the transmitting device 100 stores different slice segments in different packets. This ensures that the receiving device 200 can separate the slice segments in the multiplexing layer.

[0074] For example, when TS is used, coded data is packetized in slice segment units as PES packets. When RTP is used, coded data is packetized in slice segment units as RTP packets. Even in these cases, NAL units placed before slice segments and slice segments may be packetized separately, as in MMT packet #1 shown in FIG. 8.

[0075] When TS is used, the transmitting device 100 indicates the unit of data stored in a PES packet by using a data alignment descriptor, etc. Furthermore, since DASH is a method of downloading MP4 format data units called segments via HTTP, the transmitting device 100 does not packetize the encoded data for transmission. For this reason, the transmitting device 100 may create subsamples in slice segment units and store information indicating the storage locations of the subsamples in the MP4 header so that the receiving device 200 can detect slice segments in the multiplexing layer in MP4.

[0076] The MMT packetization of slice segments will be described in detail below.

[0077] As shown in Fig. 8, by packetizing the encoded data, data commonly referenced when decoding all slice segments in an access unit, such as SPS and PPS, is stored in MMT packet #1. In this case, receiving device 200 concatenates the payload data of MMT packet #1 with the data of each slice segment and outputs the obtained data to the decoding unit. In this way, receiving device 200 can easily generate input data for the decoding unit by concatenating the payloads of multiple MMT packets.

[0078] FIG. 10 is a diagram showing an example in which input data to decoding units 204A to 204D is generated from the MMT packets shown in FIG. 8. Demultiplexing unit 203 generates data necessary for decoding unit 204A to decode slice segment 1 by concatenating payload data of MMT packet #1 and MMT packet #2. Demultiplexing unit 203 similarly generates input data for decoding units 204B to 204D. That is, demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #3 to generate input data for decoding unit 204B. Demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #4 to generate input data for decoding unit 204C. Demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #5 to generate input data for decoding unit 204D.

[0079] In addition, the demultiplexing unit 203 may remove NAL units that are not necessary for the decoding process, such as the access unit delimiter and SEI, from the payload data of MMT packet #1, and separate only the NAL units of SPS and PPS that are necessary for the decoding process and add them to the data of the slice segment.

[0080] 9, when encoded data is packetized, demultiplexing unit 203 outputs MMT packet #1 including the start data of an access unit in the multiplexing layer to the first decoding unit 204A. In addition, demultiplexing unit 203 analyzes the MMT packet including the start data of an access unit in the multiplexing layer, separates the SPS and PPS NAL units, and adds the separated SPS and PPS NAL units to each of the data of the second and subsequent slice segments to generate input data for each of the second and subsequent decoding units.

[0081] Furthermore, it is desirable that receiving device 200 can identify the type of data stored in the MMT payload and the index number of the slice segment within an access unit when a slice segment is stored in the payload, using information included in the header of the MMT packet. Here, the data type refers to either pre-slice segment data (a collective term for NAL units arranged before the first slice segment in an access unit) or slice segment data. When storing a unit obtained by fragmenting an MPU, such as a slice segment, in an MMT packet, a mode for storing MFUs (Media Fragment Units) is used. When using this mode, transmitting device 100 can, for example, set the Data Unit, which is the basic unit of data in an MFU, to a sample (a data unit in MMT, equivalent to an access unit) or a subsample (a unit obtained by dividing a sample).

[0082] At this time, the header of the MMT packet includes a field called a fragmentation indicator and a field called a fragment counter.

[0083] The Fragmentation indicator indicates whether the data stored in the payload of an MMT packet is a fragment of a Data unit, and if so, whether the fragment is the first or last fragment of the Data unit, or a fragment that is neither the first nor the last. In other words, the Fragmentation indicator included in the header of a certain packet is identification information that indicates whether (1) only the packet in question is included in the Data unit, which is the basic data unit, (2) the Data unit is divided into multiple packets and stored, and the packet in question is the first packet of the Data unit, (3) the Data unit is divided into multiple packets and stored, and the packet in question is a packet other than the first or last packet of the Data unit, or (4) the Data unit is divided into multiple packets and stored, and the packet in question is the last packet of the Data unit.

[0084] The fragment counter is an index number that indicates which fragment in the data unit the data stored in the MMT packet corresponds to.

[0085] Therefore, by transmitting device 100 setting samples in MMT to Data units and setting data before a slice segment and each slice segment to fragment units of Data units, receiving device 200 can identify the type of data stored in the payload using information included in the header of the MMT packet. That is, demultiplexing unit 203 can generate input data for each of decoding units 204A to 204D by referring to the header of the MMT packet.

[0086] FIG. 11 is a diagram showing an example in which a sample is set to a data unit, and data before a slice segment and a slice segment are packetized as fragments of the data unit.

[0087] The data before the slice segment and the slice segment are divided into five fragments, fragment #1 to fragment #5. Each fragment is stored in a separate MMT packet. At this time, the values ​​of the fragmentation indicator and fragment counter included in the header of the MMT packet are as shown in the figure.

[0088] For example, the Fragment indicator is a 2-bit binary value. The Fragment indicator of MMT packet #1, which is the head of the Data unit, the Fragment indicator of MMT packet #5, which is the last packet, and the Fragment indicators of MMT packet #2 to MMT packet #4, which are the packets in between, are all set to different values. Specifically, the Fragment indicator of MMT packet #1, which is the head of the Data unit, is set to 01, the Fragment indicator of MMT packet #5, which is the last packet, is set to 11, and the Fragment indicators of MMT packet #2 to MMT packet #4, which are the packets in between, are set to 10. Note that if a Data unit contains only one MMT packet, the Fragment indicator is set to 00.

[0089] In addition, the fragment counter is 4 in MMT packet #1, which is the total number of fragments, 5, minus 1, and is decremented by 1 in subsequent packets, reaching 0 in the final MMT packet #5.

[0090] Therefore, receiving apparatus 200 can identify an MMT packet that stores pre-slice segment data by using either the fragment indicator or the fragment counter. Also, receiving apparatus 200 can identify an MMT packet that stores the N-th slice segment by referring to the fragment counter.

[0091] The header of the MMT packet also includes the sequence number within the MPU of the Movie Fragment to which the Data Unit belongs, the sequence number of the MPU itself, and the sequence number within the Movie Fragment of the sample to which the Data Unit belongs. By referencing these, the demultiplexer 203 can uniquely determine the sample to which the Data Unit belongs.

[0092] Furthermore, since demultiplexing unit 203 can determine the index number of a fragment within a data unit from a fragment counter or the like, it can uniquely identify the slice segment stored in the fragment even if packet loss occurs. For example, even if fragment #4 shown in FIG. 11 cannot be acquired due to packet loss, demultiplexing unit 203 can determine that the fragment received next after fragment #3 is fragment #5, and can therefore correctly output slice segment 4 stored in fragment #5 to decoding unit 204D rather than decoding unit 204C.

[0093] Note that, when a transmission path that guarantees no packet loss is used, demultiplexing unit 203 can periodically process the arrived packets without referring to the header of the MMT packet to determine the type of data stored in the MMT packet or the index number of the slice segment. For example, when an access unit is transmitted using five MMT packets, including pre-slice data and four slice segments, receiving device 200 can sequentially acquire the pre-slice data and the data of the four slice segments by sequentially processing the received MMT packets after determining the pre-slice data of the access unit from which decoding is to be started.

[0094] A variation of packetization will now be described.

[0095] Slice segments do not necessarily have to be divided both horizontally and vertically within the plane of an access unit; as shown in Figure 1, they may be divided only horizontally or only vertically within the plane of an access unit.

[0096] Also, if the access unit is divided only horizontally, tiles do not need to be used.

[0097] Furthermore, the number of divisions within an access unit is arbitrary and is not limited to 4. However, the area sizes of slice segments and tiles must be equal to or greater than the lower limit of encoding standards such as H.265.

[0098] Transmitting device 100 may store identification information indicating the division method for an access unit in an MMT message, a TS descriptor, or the like. For example, information indicating the number of divisions in the horizontal and vertical directions within the plane may be stored. Alternatively, unique identification information may be assigned to the division method, such as dividing the access unit into two equal parts in each of the horizontal and vertical directions as shown in FIG. 3, or dividing the access unit into four equal parts in the horizontal direction as shown in FIG. 1. For example, if the access unit is divided as shown in FIG. 3, the identification information indicates mode 1, and if the access unit is divided as shown in FIG. 1, the identification information indicates mode 1.

[0099] Furthermore, information indicating constraints on coding conditions related to the intra-plane division method may be included in the multiplexing layer. For example, information indicating that one slice segment is composed of one tile may be used. Alternatively, information indicating that reference blocks when performing motion compensation during decoding of a slice segment or tile are limited to slice segments or tiles at the same position within the screen, or limited to blocks within a predetermined range in adjacent slice segments may be used.

[0100] Furthermore, the transmitting device 100 may switch whether to divide an access unit into a plurality of slice segments depending on the resolution of the video. For example, the transmitting device 100 may not perform intra-frame division when the video to be processed has a resolution of 4K2K, but may divide the access unit into four when the video to be processed has a resolution of 8K4K. By predefining the division method for 8K4K video, the receiving device 200 can determine whether to perform intra-frame division and the division method by acquiring the resolution of the video to be received, and switch the decoding operation.

[0101] Furthermore, receiving device 200 can detect whether or not a frame is fragmented by referring to the header of the MMT packet. For example, if an access unit is not fragmented, if the MMT Data unit is set to Sample, the Data unit is not fragmented. Therefore, receiving device 200 can determine that an access unit is not fragmented if the value of the Fragment counter included in the header of the MMT packet is always zero. Alternatively, receiving device 200 may detect whether the value of the Fragmentation indicator is always 01. Receiving device 200 can also determine that an access unit is not fragmented if the value of the Fragmentation indicator is always 01.

[0102] Furthermore, the receiving device 200 can also handle a case where the number of divisions in a plane of an access unit does not match the number of decoding units. For example, if the receiving device 200 includes two decoding units 204A and 204B that can decode 8K2K encoded data in real time, the demultiplexing unit 203 outputs two of the four slice segments that make up the 8K4K encoded data to the decoding unit 204A.

[0103] Fig. 12 is a diagram showing an example of operation when MMT packetized data as shown in Fig. 8 is input to two decoding units 204A and 204B. Here, it is desirable that receiving device 200 be able to integrate and output the decoding results of decoding units 204A and 204B as they are. Therefore, demultiplexing unit 203 selects slice segments to output to each of decoding units 204A and 204B so that the decoding results of each of decoding units 204A and 204B are spatially continuous.

[0104] Furthermore, the demultiplexing unit 203 may select a decoding unit to use depending on the resolution or frame rate of the encoded data of the video. For example, if the receiving device 200 is equipped with four 4K2K decoding units, and the resolution of the input image is 8K4K, the receiving device 200 performs the decoding process using all four decoding units. Furthermore, if the resolution of the input image is 4K2K, the receiving device 200 performs the decoding process using only one decoding unit. Alternatively, even if the plane is divided into four, if 8K4K can be decoded in real time by a single decoding unit, the demultiplexing unit 203 integrates all the division units and outputs them to a single decoding unit.

[0105] Furthermore, the receiving device 200 may determine the decoding unit to be used taking the frame rate into consideration. For example, if the receiving device 200 is equipped with two decoding units, each of which has an upper limit of 60 fps for the frame rate that can be decoded in real time when the resolution is 8K4K, there may be a case where 120 fps coded data in 8K4K is input. In this case, if the plane is configured with four division units, slice segment 1 and slice segment 2 are input to the decoding unit 204A, and slice segment 3 and slice segment 4 are input to the decoding unit 204B, as in the example of FIG. 12. Each of the decoding units 204A and 204B can decode up to 120 fps in real time at 8K2K (resolution half that of 8K4K), and therefore the decoding process is performed by these two decoding units 204A and 204B.

[0106] Furthermore, even if the resolution and frame rate are the same, the processing amount differs depending on the profile or level of the encoding method, or the encoding method itself, such as H.264 or H.265. Therefore, the receiving device 200 may select a decoder to use based on this information. Note that if the receiving device 200 is unable to decode all of the encoded data received via broadcasting or communication, or if it is unable to decode all of the slice segments or tiles constituting a region selected by the user, the receiving device 200 may automatically determine slice segments or tiles that can be decoded within the processing range of the decoding device. Alternatively, the receiving device 200 may provide a user interface that allows the user to select a region to decode. In this case, the receiving device 200 may display a warning message indicating that all regions cannot be decoded, or may display information indicating the number of regions, slice segments, or tiles that can be decoded.

[0107] The above method can also be applied to cases where MMT packets storing slice segments of the same coded data are transmitted and received using multiple transmission paths such as broadcasting and communication.

[0108] Furthermore, the transmitting device 100 may perform encoding so that the slice segments overlap each other to make the boundaries between division units less noticeable. In the example shown in FIG. 13, an 8K4K picture is divided into four slice segments 1 to 4. Each of slice segments 1 to 3 is, for example, 8K x 1.1K, and slice segment 4 is 8K x 1K. Adjacent slice segments overlap each other. This allows efficient motion compensation during encoding at the boundaries of the four divisions indicated by the dotted lines, improving the image quality of the boundary areas. In this way, image quality degradation at the boundary areas is reduced.

[0109] In this case, the display unit 205 extracts an 8K×1K region from the 8K×1.1K region and integrates the obtained regions. Note that the transmitting device 100 may include information indicating whether the slice segments are coded with overlapping and the extent of the overlap in the multiplexing layer or coded data and transmit the information separately.

[0110] Note that a similar approach can be applied when tiles are used.

[0111] The following describes the flow of operations of the transmission device 100. FIG.

[0112] First, the encoding unit 101 divides a picture (access unit) into a plurality of slice segments (tiles), which are a plurality of regions (S101). Next, the encoding unit 101 generates encoded data corresponding to each of the plurality of slice segments by encoding each of the plurality of slice segments so that each of the plurality of slice segments can be decoded independently (S102). Note that the encoding unit 101 may encode the plurality of slice segments using a single encoding unit, or may encode the plurality of slice segments in parallel using a plurality of encoding units.

[0113] Next, multiplexing unit 102 multiplexes the plurality of coded data generated by coding unit 101 by storing the plurality of coded data in a plurality of MMT packets (S103). Specifically, as shown in FIGS. 8 and 9, multiplexing unit 102 stores the plurality of coded data in a plurality of MMT packets so that coded data corresponding to different slice segments are not stored in one MMT packet. Furthermore, as shown in FIG. 8, multiplexing unit 102 stores control information commonly used for all decoding units in a picture in MMT packet #1, which is different from the plurality of MMT packets #2 to #5 in which the plurality of coded data are stored. Here, the control information includes at least one of an access unit delimiter, an SPS, a PPS, and an SEI.

[0114] Note that multiplexing unit 102 may store the control information in the same MMT packet as one of the plurality of MMT packets in which the plurality of encoded data are stored. For example, as shown in Fig. 9, multiplexing unit 102 may store the control information in the first MMT packet (MMT packet #1 in Fig. 9) of the plurality of MMT packets in which the plurality of encoded data are stored.

[0115] Finally, the transmitting device 100 transmits the multiple MMT packets. Specifically, the modulating unit 103 modulates the data obtained by multiplexing, and the transmitting unit 104 transmits the modulated data (S104).

[0116] Fig. 15 is a block diagram showing an example of the configuration of receiving device 200, and is a diagram showing in detail the configuration of demultiplexing unit 203 and the subsequent stages shown in Fig. 7. As shown in Fig. 15, receiving device 200 further includes a decoding command unit 206. In addition, demultiplexing unit 203 includes a type discrimination unit 211, a control information acquisition unit 212, a slice information acquisition unit 213, and a decoded data generation unit 214.

[0117] The following describes the flow of operation of receiving apparatus 200. Fig. 16 is a flowchart showing an example of operation of receiving apparatus 200. Here, the operation for one access unit is shown. When decoding processing for multiple access units is executed, the processing of this flowchart is repeated.

[0118] First, the receiving device 200 receives, for example, a plurality of packets (MMT packets) generated by the transmitting device 100 (S201).

[0119] Next, the type determination unit 211 analyzes the header of the received packet to obtain the type of encoded data stored in the received packet (S202).

[0120] Next, the type determination unit 211 determines whether the data stored in the received packet is pre-slice segment data or slice segment data, based on the type of the acquired coded data (S203).

[0121] If the data stored in the received packet is pre-slice segment data (Yes in S203), the control information acquisition unit 212 acquires the pre-slice segment data of the access unit to be processed from the payload of the received packet and stores the pre-slice segment data in memory (S204).

[0122] On the other hand, if the data stored in the received packet is data of a slice segment (No in S203), the receiving device 200 uses the header information of the received packet to determine which of the multiple areas the data stored in the received packet is encoded data for. Specifically, the slice information acquisition unit 213 analyzes the header of the received packet to acquire the index number Idx of the slice segment stored in the received packet (S205). Specifically, the index number Idx is the index number within the Movie Fragment of the access unit (sample in MMT).

[0123] The process of step S205 may be performed together with step S202.

[0124] Next, the decoding data generation unit 214 determines a decoding unit that will decode the slice segment (S206). Specifically, the index number Idx and a plurality of decoding units are associated in advance, and the decoding data generation unit 214 determines the decoding unit that will decode the slice segment as the decoding unit that corresponds to the index number Idx acquired in step S205.

[0125] 12, the decoding data generation unit 214 may determine the decoding unit that decodes the slice segment based on at least one of the resolution of the access unit (picture), the division method of the access unit into multiple slice segments (tiles), and the processing capabilities of the multiple decoding units included in the receiving device 200. For example, the decoding data generation unit 214 determines the division method of the access unit based on identification information in a descriptor such as an MMT message or a TS section.

[0126] Next, the decoded data generation unit 214 generates multiple pieces of input data (combined data) to be input to multiple decoding units by combining control information, which is included in one of the multiple packets and is used commonly for all decoding units in the picture, with each of the multiple pieces of encoded data for the multiple slice segments. Specifically, the decoded data generation unit 214 acquires slice segment data from the payload of the received packet. The decoded data generation unit 214 generates input data to the decoding unit determined in step S206 by combining the pre-slice segment data stored in memory in step S204 with the acquired slice segment data (S207).

[0127] After step S204 or S207, if the data of the received packet is not the final data of the access unit (No in S208), the processes from step S201 onward are performed again. That is, the above processes are repeated until input data for the multiple decoding units 204A to 204D corresponding to all slice segments included in the access unit are generated.

[0128] The timing at which the packets are received is not limited to the timing shown in FIG. 16, and a plurality of packets may be received in advance or sequentially and stored in a memory or the like.

[0129] On the other hand, if the data of the received packet is the final data of the access unit (Yes in S208), the decoding command unit 206 outputs the plurality of input data generated in step S207 to the corresponding decoding units 204A to 204D (S209).

[0130] Next, the multiple decoding units 204A to 204D decode the multiple pieces of input data in parallel in accordance with the DTS of the access unit, thereby generating multiple decoded images (S210).

[0131] Finally, the display unit 205 generates a display image by combining the decoded images generated by the decoding units 204A to 204D, and displays the display image in accordance with the PTS of the access unit (S211).

[0132] The receiving device 200 acquires the DTS and PTS of an access unit by analyzing the header information of an MPU or the payload data of an MMT packet that stores the header information of a Movie Fragment. Furthermore, when TS is used as the multiplexing method, the receiving device 200 acquires the DTS and PTS of an access unit from the header of a PES packet. When RTP is used as the multiplexing method, the receiving device 200 acquires the DTS and PTS of an access unit from the header of an RTP packet.

[0133] Furthermore, when integrating the decoding results of the multiple decoders, the display unit 205 may perform filtering, such as deblocking filtering, at the boundaries between adjacent division units. Note that, since filtering is not necessary when displaying the decoding results of a single decoder, the display unit 205 may switch the filtering process depending on whether or not to perform filtering at the boundaries between the decoding results of the multiple decoders. Whether filtering is necessary may be determined in advance depending on whether or not division is performed. Alternatively, information indicating whether filtering is necessary may be separately stored in the multiplexing layer. Information necessary for filtering, such as filter coefficients, may be stored in the SPS, PPS, SEI, or slice segment. The decoding units 204A to 204D or the demultiplexing unit 203 acquire this information by analyzing the SEI and output the acquired information to the display unit 205. The display unit 205 performs filtering using this information. Note that, if this information is stored in the slice segment, it is preferable that the decoding units 204A to 204D acquire this information.

[0134] In the above description, an example was shown in which the types of data stored in a fragment are two types: pre-slice segment data and slice segments. However, the types of data may be three or more. In this case, in step S203, the cases are classified according to the types.

[0135] Furthermore, when the data size of a slice segment is large, transmitting device 100 may fragment the slice segment and store the fragmented slice segment in an MMT packet. That is, transmitting device 100 may fragment the data before the slice segment and the slice segment. In this case, if the access unit and the data unit are set equal as in the packetization example shown in FIG. 11, the following problem occurs.

[0136] For example, if slice segment 1 is divided into three fragments, slice segment 1 is divided and transmitted as three packets with fragment counter values ​​of 1 to 3. Furthermore, from slice segment 2 onwards, the fragment counter value becomes 4 or greater, and it becomes impossible to associate the fragment counter value with the data stored in the payload. Therefore, receiving device 200 cannot identify the packet that stores the first data of the slice segment from the information in the header of the MMT packet.

[0137] In such a case, receiving apparatus 200 may analyze data in the payload of the MMT packet to identify the start position of the slice segment. Here, there are two types of formats for storing NAL units in the multiplexing layer in H.264 or H.265: a format called a byte stream format in which a start code consisting of a specific bit string is added immediately before the NAL unit header, and a format called an NAL size format in which a field indicating the size of the NAL unit is added.

[0138] The byte stream format is used in MPEG-2 systems and RTP, etc. The NAL size format is used in MP4, and DASH and MMT, which use MP4, etc.

[0139] When the byte stream format is used, the receiving device 200 analyzes whether the leading data of the packet matches the start code. If the leading data of the packet matches the start code, the receiving device 200 can detect whether the data included in the packet is data of a slice segment by obtaining the type of NAL unit from the NAL unit header that follows.

[0140] On the other hand, in the case of the NAL size format, the receiving device 200 cannot detect the start position of the NAL unit based on the bit string. Therefore, in order to obtain the start position of the NAL unit, the receiving device 200 needs to shift the pointer by reading data by the size of the NAL unit, starting from the first NAL unit of the access unit.

[0141] However, if the size of the subsample unit is indicated in the header of an MPU or Movie Fragment in MMT, and the subsample corresponds to pre-slice data or a slice segment, receiving device 200 can identify the start position of each NAL unit based on the size information of the subsample. Therefore, transmitting device 100 may include information indicating whether subsample unit information exists in an MPU or Movie Fragment in information that receiving device 200 acquires when starting to receive data, such as an MPT in MMT.

[0142] Note that MPU data is an extension of the MP4 format. MP4 has a mode in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and a mode in which they cannot be stored. Information for identifying this mode is indicated as the entry name of SampleEntry. When a mode in which parameter sets can be stored is used and a parameter set is included in a sample, receiving device 200 acquires the parameter set by the method described above.

[0143] On the other hand, when a mode that cannot store parameter sets is used, the parameter sets are stored as Decoder Specific Information in SampleEntry, or are stored using a stream for the parameter sets. Here, since streams for parameter sets are not generally used, it is desirable that the transmitting device 100 stores the parameter sets in Decoder Specific Information. In this case, the receiving device 200 analyzes the SampleEntry transmitted as metadata of the MPU or metadata of the Movie Fragment in the MMT packet, and acquires the parameter sets referenced by the access unit.

[0144] When a parameter set is stored as sample data, the receiving device 200 can acquire the parameter set required for decoding by referring only to the sample data without referring to the SampleEntry. In this case, the transmitting device 100 does not need to store the parameter set in the SampleEntry. This allows the transmitting device 100 to use the same SampleEntry in different MPUs, thereby reducing the processing load on the transmitting device 100 when generating an MPU. Another advantage is that the receiving device 200 does not need to refer to the parameter set in the SampleEntry.

[0145] Alternatively, transmitting device 100 may store one default parameter set in SampleEntry, and store the parameter set referenced by the access unit in the sample data. In conventional MP4, parameter sets were typically stored in SampleEntry, so there was a possibility that some receiving devices would stop playback if no parameter set existed in SampleEntry. This problem can be solved by using the above method.

[0146] Alternatively, transmitting device 100 may store a parameter set in sample data only when a parameter set different from the default parameter set is used.

[0147] In both modes, it is possible to store parameter sets in SampleEntry, so transmitting device 100 may always store parameter sets in VisualSampleEntry, and receiving device 200 may always acquire parameter sets from VisualSampleEntry.

[0148] In the MMT standard, MP4 header information such as Moov and Moof is transmitted as MPU metadata or movie fragment metadata, but the transmitting device 100 does not necessarily have to transmit MPU metadata and movie fragment metadata. Furthermore, the receiving device 200 can also determine whether an SPS and a PPS are stored in the sample data based on the service, asset type, or whether MPU meta is transmitted in the ARIB (Association of Radio Industries and Businesses) standard.

[0149] FIG. 17 is a diagram showing an example in which data before a slice segment and each slice segment are set to different data units.

[0150] 17, the data sizes of the data before the slice segment and slice segments 1 to 4 are Length #1 to Length #5, respectively. The field values ​​of the Fragmentation indicator, Fragment counter, and Offset included in the header of the MMT packet are as shown in the figure.

[0151] Here, Offset is offset information indicating the bit length (offset) from the beginning of the coded data of the sample (access unit or picture) to which the payload data belongs to to the first byte of the payload data (coded data) included in the MMT packet. Note that although the value of the Fragment counter is explained as starting from a value obtained by subtracting 1 from the total number of fragments, it may start from another value.

[0152] Fig. 18 is a diagram showing an example of when a Data unit is fragmented. In the example shown in Fig. 18, slice segment 1 is divided into three fragments, which are stored in MMT packets #2 to #4, respectively. In this case, if the data size of each fragment is Length #2_1 to Length #2_3, respectively, the values ​​of each field are as shown in the diagram.

[0153] In this way, when a data unit such as a slice segment is set to Data unit, the start of the access unit and the start of the slice segment can be determined as follows based on the field values ​​of the MMT packet header.

[0154] The start of the payload in a packet with an Offset value of 0 is the start of the access unit.

[0155] The start of the payload of a packet in which the Offset value is a value other than 0 and the Fragmentation indicator value is 00 or 01 is the start of the slice segment.

[0156] In addition, if no fragmentation of the data unit occurs and no packet loss occurs, the receiving device 200 can identify the index number of the slice segment to be stored in the MMT packet based on the number of slice segments obtained after detecting the beginning of the access unit.

[0157] Similarly, even when the data unit of the data before the slice segment is fragmented, the receiving device 200 can detect the beginning of the access unit and the slice segment.

[0158] Furthermore, even when packet loss occurs or when the SPS, PPS, and SEI included in the data before the slice segment are set in different Data units, receiving device 200 can identify the MMT packet that stores the start data of the slice segment based on the analysis result of the MMT header, and then analyze the header of the slice segment to identify the start position of the slice segment or tile within the picture (access unit). The amount of processing involved in analyzing the slice header is small, so the processing load is not a problem.

[0159] In this way, each of the plurality of coded data of the plurality of slice segments is in one-to-one correspondence with a basic data unit, which is a unit of data stored in one or more packets. Also, each of the plurality of coded data is stored in one or more MMT packets.

[0160] The header information of each MMT packet includes a fragmentation indicator (identification information) and an offset (offset information).

[0161] Receiving device 200 determines that the start of payload data included in a packet having header information including a Fragmentation indicator whose value is 00 or 01 is the start of the encoded data of each slice segment. Specifically, receiving device 200 determines that the start of payload data included in a packet having header information including an Offset whose value is not 0 and a Fragmentation indicator whose value is 00 or 01 is the start of the encoded data of each slice segment.

[0162] 17, the start of a Data Unit is either the start of an access unit or the start of a slice segment, and the value of the Fragmentation indicator is 00 or 01. Furthermore, receiving device 200 can also detect the start of an access unit or the start of a slice segment without referring to the Offset by referring to the type of NAL unit and determining whether the start of a Data Unit is an access unit delimiter or a slice segment.

[0163] In this way, by transmitting apparatus 100 packetizing data so that the beginning of an NAL unit always starts at the beginning of the payload of an MMT packet, receiving apparatus 200 can detect the beginning of an access unit or a slice segment by analyzing the fragmentation indicator and the NAL unit header, even when the data before the slice segment is divided into multiple data units. The type of the NAL unit is present in the first byte of the NAL unit header. Therefore, when analyzing the header portion of the MMT packet, receiving apparatus 200 can obtain the type of the NAL unit by analyzing one additional byte of data. In the case of audio, receiving apparatus 200 only needs to detect the beginning of the access unit, and can make a determination based on whether the value of the fragmentation indicator is 00 or 01.

[0164] Furthermore, as described above, when storing coded data that has been coded so that it can be divided and decoded in PES packets of MPEG-2 TS, the transmitting device 100 can use a data alignment descriptor. An example of a method for storing coded data in PES packets will be described in detail below.

[0165] For example, in HEVC, the transmission device 100 can indicate whether the data stored in the PES packet is an access unit, a slice segment, or a tile by using a data alignment descriptor. The alignment types in HEVC are specified as follows:

[0166] Alignment type=8 indicates an HEVC slice segment. Alignment type=9 indicates an HEVC slice segment or access unit. Alignment type=12 indicates an HEVC slice segment or tile.

[0167] Therefore, the transmitting device 100 can indicate that the data of the PES packet is either a slice segment or pre-slice segment data by using, for example, type 9. Since a type indicating a slice rather than a slice segment is also separately defined, the transmitting device 100 may use a type indicating a slice rather than a slice segment.

[0168] Furthermore, the DTS and PTS included in the header of a PES packet are set only in the PES packet that contains the first data of an access unit. Therefore, if the type is 9 and the PES packet contains a DTS or PTS field, receiving device 200 can determine that the PES packet stores the entire access unit or the first division unit of the access unit.

[0169] Furthermore, the transmitting device 100 may enable the receiving device 200 to distinguish the data contained in the packet using a field such as transport_priority, which indicates the priority of a TS packet that stores a PES packet containing the first data of an access unit. The receiving device 200 may also determine the data contained in the packet by analyzing whether the payload of the PES packet is an access unit delimiter. Furthermore, the data_alignment_indicator in the PES packet header indicates whether data is stored in the PES packet according to these types. If this flag (data_alignment_indicator) is set to 1, it is guaranteed that the data stored in the PES packet complies with the type indicated in the data alignment descriptor.

[0170] Furthermore, the transmitting device 100 may use the data alignment descriptor only when PES packetizing is performed in units that can be divided and decoded, such as slice segments. As a result, if the data alignment descriptor is present, the receiving device 200 can determine that the coded data has been PES packetized in units that can be divided and decoded, and if the data alignment descriptor is not present, the receiving device 200 can determine that the coded data has been PES packetized in units of access units. Note that if the data_alignment_indicator is set to 1 and the data alignment descriptor is not present, the MPEG-2 TS standard specifies that the unit of PES packetization is the access unit.

[0171] If the PMT includes a data alignment descriptor, the receiving device 200 determines that the PES packetization is performed in units that can be divided and decoded, and can generate input data for each decoder based on the packetized units. If the PMT does not include a data alignment descriptor and the receiving device 200 determines that parallel decoding of the encoded data is necessary based on program information or other descriptor information, the receiving device 200 generates input data for each decoder by analyzing the slice header of the slice segment, etc. If the encoded data can be decoded by a single decoder, the receiving device 200 decodes the data of the entire access unit by that decoder. If information indicating whether the encoded data is composed of units that can be divided and decoded, such as slice segments or tiles, is separately indicated by a descriptor in the PMT, the receiving device 200 may determine whether the encoded data can be parallel decoded based on the analysis result of the descriptor.

[0172] Furthermore, since the DTS and PTS included in the header of a PES packet are set only in the PES packet containing the first data of an access unit, when an access unit is divided and packetized as PES packets, the second and subsequent PES packets do not contain information indicating the DTS and PTS of the access unit. Therefore, when decoding processes are performed in parallel, each of the decoding units 204A to 204D and the display unit 205 uses the DTS and PTS stored in the header of the PES packet containing the first data of an access unit.

[0173] (Embodiment 2) In the second embodiment, a method for storing data in the NAL size format in an MPU based on the MP4 format in MMT will be described. Note that, although a storage method in an MPU used in MMT will be described below as an example, such a storage method can also be applied to DASH, which is also based on the MP4 format.

[0174] [Storage method in MPU] In the MP4 format, multiple access units are stored together in a single MP4 file. The MPU used in MMT stores data for each media in a single MP4 file, and the data can contain any number of access units. Since an MPU is a unit that can be decoded independently, for example, an MPU stores access units in units of GOPs.

[0175] 19 is a diagram showing the structure of an MPU. At the beginning of an MPU are ftyp, mmpu, and moov, which are collectively defined as MPU metadata. moov stores initialization information common to the file and an MMT hint track.

[0176] Additionally, moof stores initialization information and size for each sample and subsample, information (sample_duration, sample_size, sample_composition_time_offset) that can identify the presentation time (PTS) and decoding time (DTS), and data_offset that indicates the position of the data.

[0177] Furthermore, each of the multiple access units is stored as a sample in mdat (mdat box). Data in moof and mdat excluding samples is defined as movie fragment metadata (hereinafter referred to as MF metadata), and sample data in mdat is defined as media data.

[0178] Fig. 20 is a diagram showing the structure of MF metadata. As shown in Fig. 20, the MF metadata is more specifically made up of the type, length, and data of a moof box (moof), and the type and length of an mdat box (mdat).

[0179] When storing access units in MP4 data, there are two modes: one in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and one in which they cannot be stored.

[0180] In the non-storable mode, the parameter set is stored in the Decoder Specific Information of the SampleEntry in the moov, and in the storable mode, the parameter set is included in the sample.

[0181] The MPU metadata, MF metadata, and media data are each stored in an MMT payload, and a fragment type (FT) is stored in the header of the MMT payload as an identifier for identifying these data. FT=0 indicates MPU metadata, FT=1 indicates MF metadata, and FT=2 indicates media data.

[0182] Note that, although Fig. 19 illustrates an example in which MPU metadata units and MF metadata units are stored as data units in the MMT payload, units such as ftyp, mmpu, moov, and moof may also be stored as data units in the MMT payload in data unit units. Similarly, Fig. 19 illustrates an example in which sample units are stored as data units in the MMT payload. However, data units may be configured in sample units or NAL unit units, and such data units may be stored in the MMT payload in data unit units. Such data units may also be further fragmented and stored in the MMT payload.

[0183] [Conventional transmission methods and issues] Conventionally, when multiple access units are encapsulated in the MP4 format, moov and moof are created when all samples to be stored in the MP4 are available.

[0184] When transmitting MP4 format in real time via broadcasting, for example, if the samples stored in one MP4 file are in GOP units, delays occur due to encapsulation because the moov and moof are created after the GOP-unit time samples are accumulated. This encapsulation on the transmitting side always increases the end-to-end delay by the GOP unit time. This makes it difficult to provide services in real time, and leads to degradation of the service for viewers, especially when transmitting live content.

[0185] Fig. 21 is a diagram for explaining the data transmission order. When MMT is applied to broadcasting, as shown in Fig. 21(a), if data is loaded onto MMT packets and transmitted in the order of the MPU configuration (transmitting MMT packets #1, #2, #3, #4, #5, and #6 in that order), a delay occurs in the transmission of the MMT packets due to encapsulation.

[0186] To prevent this delay due to encapsulation, a method has been proposed in which MPU header information such as MPU metadata and MF metadata is not sent (packets #1 and #2 are not sent, and packets #3 to #6 are sent in this order), as shown in (b) of Figure 21. Another possible method is to send media data first without waiting for the creation of MPU header information, and then send the MPU header information after the media data has been sent (sending packets #3 to #6, #1, and #2 in that order), as shown in (c) of Figure 20.

[0187] If the MPU header information is not transmitted, the receiving device decodes without using the MPU header information. If the MPU header information is delayed relative to the media data, the receiving device waits until it obtains the MPU header information before decoding.

[0188] However, conventional MP4-compliant receiving devices are not guaranteed to be able to decode without using MPU header information. Furthermore, if a receiving device performs special processing to decode without using the MPU header, using conventional transmission methods can complicate the decoding process, potentially making real-time decoding difficult. Furthermore, if a receiving device waits for MPU header information before decoding, media data must be buffered until the receiving device acquires the header information. However, because no buffer model is specified, decoding is not guaranteed.

[0189] Therefore, the transmitting device according to the second embodiment stores only common information in the MPU metadata, as shown in (d) of Fig. 20, so that the MPU metadata is transmitted before the media data.The transmitting device according to the second embodiment then transmits the MF metadata, which is generated with a delay, after the media data.This provides a transmitting method or receiving method that can guarantee decoding of the media data.

[0190] The following describes the reception method when using each of the transmission methods (a) to (d) in FIG.

[0191] In each transmission method shown in FIG. 21, first, MPU data is configured in the order of MPU metadata, MFU metadata, and media data.

[0192] After constructing the MPU data, if the transmitting device transmits data in the order of MPU metadata, MF metadata, and media data, as shown in (a) of Figure 21, the receiving device can perform decoding using either of the following methods (A-1) and (A-2).

[0193] (A-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data using the MPU header information.

[0194] (A-2) The receiving device decodes the media data without using the MPU header information.

[0195] Although these methods all incur delays due to encapsulation on the transmitting side, they have the advantage that the receiving device does not need to buffer the media data to obtain the MPU header. If buffering is not performed, there is no need to install memory for buffering, and buffering delays do not occur. Furthermore, method (A-1) is applicable to conventional receiving devices because it uses MPU header information for decoding.

[0196] When the transmitting device transmits only media data as shown in (b) of FIG. 21, the receiving device can perform decoding using the following method (B-1).

[0197] (B-1) The receiving device decodes the media data without using the MPU header information.

[0198] Although not shown, if MPU metadata is transmitted before the transmission of the media data in (b) of FIG. 21, decoding can be performed using the following method (B-2).

[0199] (B-2) The receiving device decodes the media data using the MPU metadata.

[0200] The advantages of both methods (B-1) and (B-2) above are that there is no delay due to encapsulation on the sending side, and there is no need to buffer media data to obtain the MPU header. However, since neither method (B-1) nor (B-2) performs decoding using MPU header information, special processing may be required for decoding.

[0201] When the transmitting device transmits data in the order of media data, MPU metadata, and MF metadata, as shown in (c) of Figure 21, the receiving device can perform decoding using either of the following methods (C-1) and (C-2).

[0202] (C-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data.

[0203] (C-2) The receiving device decodes the media data without using the MPU header information.

[0204] When the above method (C-1) is used, it is necessary to buffer the media data in order to obtain the MPU header information. On the other hand, when the above method (C-2) is used, it is not necessary to buffer the media data in order to obtain the MPU header information.

[0205] In addition, neither method (C-1) nor (C-2) causes delays due to encapsulation on the sending side. Furthermore, method (C-2) may require special processing because it does not use MPU header information.

[0206] When the transmitting device transmits data in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Figure 21, the receiving device can perform decoding using either of the following methods (D-1) and (D-2).

[0207] (D-1) After acquiring the MPU metadata, the receiving device further acquires the MF metadata, and then decodes the media data.

[0208] (D-2) After acquiring the MPU metadata, the receiving device decodes the media data without using the MF metadata.

[0209] When the above method (D-1) is used, it is necessary to buffer the media data in order to obtain the MF metadata, but when the above method (D-2) is used, it is not necessary to buffer the media data in order to obtain the MF metadata.

[0210] The above method (D-2) does not perform decoding using MF metadata, and therefore may require special processing.

[0211] As described above, when decoding is possible using MPU metadata and MF metadata, there is an advantage that decoding can also be performed by a conventional MP4 receiving device.

[0212] 21, the MPU data is structured in the order of MPU metadata, MFU metadata, and media data, and in moof, position information (offset) for each sample and subsample is determined based on this structure. Also, the MF metadata includes data other than the media data in the mdat box (box size and type).

[0213] Therefore, when a receiving device identifies media data based on MF metadata, the receiving device reconstructs the data in the order in which the MPU data was constructed, regardless of the order in which the data was transmitted, and then decodes it using the moov in the MPU metadata or the moof in the MF metadata.

[0214] In FIG. 21, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data, but the MPU data may be configured in an order different from that shown in FIG. 21, and the position information (offset) may be determined.

[0215] For example, MPU data may be configured in the order of MPU metadata, media data, and MF metadata, and negative position information (offset) may be indicated in the MF metadata. In this case, regardless of the order in which the data is transmitted, the receiving device reconstructs the data in the order in which the MPU data was configured on the transmitting side, and then performs decoding using moov or moof.

[0216] The transmitting device may signal information indicating the order in which MPU data is constructed, and the receiving device may reconstruct the data based on the signaled information.

[0217] As described above, the receiving device receives packetized MPU metadata, packetized media data (sample data), and packetized MF metadata in this order, as shown in (d) of Fig. 21. Here, the MPU metadata is an example of first metadata, and the MF metadata is an example of second metadata.

[0218] Next, the receiving device reconstructs MPU data (an MP4 format file) including the received MPU metadata, the received MF metadata, and the received sample data. Then, the receiving device decodes the sample data included in the reconstructed MPU data using the MPU metadata and MF metadata. The MF metadata is metadata including data that can be generated only after the sample data is generated on the transmitting side (for example, the length stored in the mbox).

[0219] The operation of the receiving device is more specifically performed by each component of the receiving device. For example, the receiving device includes a receiving unit that receives the data, a reconstructing unit that reconstructs the MPU data, and a decoding unit that decodes the MPU data. The receiving unit, generating unit, and decoding unit are each realized by a microcomputer, a processor, a dedicated circuit, etc.

[0220] [Method of decrypting without using header information] Next, a method for decoding without using header information will be described. Here, a method for decoding without using header information in a receiving device will be described, regardless of whether header information is sent on the transmitting side or not. That is, this method is applicable to any of the transmission methods described with reference to FIG. 21. However, some decoding methods are applicable only to specific transmission methods.

[0221] Figure 22 is a diagram showing an example of a method for decoding without using header information. Figure 22 shows only MMT payloads and MMT packets containing only media data, and does not show MMT payloads and MMT packets containing MPU metadata or MF metadata. In the following description of Figure 22, it is assumed that media data belonging to the same MPU are transmitted continuously. In addition, although an example will be described in which samples are stored in the payload as media data, in the following description of Figure 22, it goes without saying that NAL units or fragmented NAL units may be stored.

[0222] To decode media data, a receiving device must first obtain initialization information necessary for decoding. If the media is video, the receiving device must obtain initialization information for each sample, identify the start position of the MPU (a random access unit), and obtain the start positions of the samples and NAL units. The receiving device must also identify the decoding time (DTS) and presentation time (PTS) of each sample.

[0223] Therefore, the receiving device can perform decoding without using header information, for example, by using the following method: Note that when NAL unit units or units obtained by fragmenting NAL units are stored in the payload, "sample" in the following description can be read as "NAL unit in a sample."

[0224] <Random access (=identify the first sample of the MPU)> When header information is not transmitted, the receiving device can identify the first sample of the MPU using the following methods 1 and 2. Note that when header information is transmitted, method 3 can be used.

[0225] [Method 1] The receiving device acquires samples contained in MMT packets with 'RAP_flag=1' in the MMT packet header.

[0226] [Method 2] The receiving device acquires the sample with 'sample number=0' in the MMT payload header.

[0227] [Method 3] When at least one of MPU metadata and MF metadata is transmitted before or after the media data, the receiving device acquires samples contained in the MMT payload whose fragment type (FT) in the MMT payload header has been switched to media data.

[0228] In Methods 1 and 2, if a single payload contains a mixture of multiple samples belonging to different MPUs, it is impossible to determine which NAL unit is a random access point (RAP_flag = 1 or sample number = 0). For this reason, it is necessary to impose a constraint such as not mixing samples from different MPUs in a single payload, or, if a single payload contains a mixture of samples from different MPUs, to impose a constraint such as setting RAP_flag to 1 if the last (or first) sample is a random access point.

[0229] Furthermore, in order for the receiving device to obtain the start position of the NAL unit, it is necessary to shift the data read pointer by the size of the NAL unit, starting from the first NAL unit of the sample.

[0230] If the data is fragmented, the receiving device can identify the data unit by referring to the fragment_indicator and fragment_number.

[0231] <Determining the DTS of a sample> There are two methods for determining the DTS of a sample: Method 1 and Method 2 below.

[0232] [Method 1] The receiver determines the DTS of the first sample based on the prediction structure. However, this method requires analysis of the coded data, which may make real-time decoding difficult. Therefore, the following method 2 is preferable.

[0233] [Method 2] The receiving device separately transmits the DTS for the first sample and acquires the transmitted DTS for the first sample. Examples of methods for transmitting the DTS for the first sample include transmitting the DTS for the MPU's first sample using MMT-SI, or transmitting a DTS for each sample using the MMT packet header extension field. The DTS may be an absolute value or a relative value to the PTS. The transmitting side may also signal whether the DTS for the first sample is included.

[0234] In both Method 1 and Method 2, the DTS of subsequent samples is calculated assuming a fixed frame rate.

[0235] In addition to using an extension field, a method for storing the DTS for each sample in a packet header also includes storing the DTS of the sample contained in the MMT packet in the 32-bit NTP timestamp field in the MMT packet header. If the DTS cannot be expressed using the number of bits in one packet header (32 bits), the DTS may be expressed using multiple packet headers. Alternatively, the DTS may be expressed by combining the NTP timestamp field and extension field in the packet header. If DTS information is not included, a known value (e.g., ALL 0) is used.

[0236] <Determining the PTS of a sample> The receiving device obtains the PTS of the first sample from the MPU timestamp descriptor for each asset included in the MPU. The receiving device calculates the PTS of subsequent samples based on a fixed frame rate, using parameters such as POC that indicate the display order of the samples. In this way, transmission at a fixed frame rate is essential to calculate the DTS and PTS without using header information.

[0237] In addition, when MF metadata is transmitted, the receiving device can calculate the absolute values ​​of the DTS and PTS from the relative time information of the DTS and PTS from the first sample indicated in the MF metadata and the absolute value of the timestamp of the MPU first sample indicated in the MPU timestamp descriptor.

[0238] When analyzing the coded data and calculating the DTS and PTS, the receiving device may use SEI information included in the access unit.

[0239] <Initialization information (parameter set)> [For video] In the case of video, the parameter set is stored in the sample data. Furthermore, if the MPU metadata and MF metadata are not transmitted, it is guaranteed that the parameter set required for decoding can be obtained by referring only to the sample data.

[0240] Also, when MPU metadata is transmitted before media data, as in (a) and (d) of Figure 21, it may be specified that parameter sets are not stored in SampleEntry. In this case, the receiving device does not refer to the parameter sets in SampleEntry, but only refers to the parameter sets in the sample.

[0241] Furthermore, when MPU metadata is transmitted before media data, SampleEntry stores a parameter set common to the MPU or a default parameter set, and a receiving device may refer to the parameter set in SampleEntry and the parameter set in the sample. Storing a parameter set in SampleEntry enables decoding even on conventional receiving devices that cannot play back data unless a parameter set is present in SampleEntry.

[0242] [For audio] For audio, an LATM header is required for decoding, and in MP4, the LATM header must be included in the sample entry. However, if the header information is not transmitted, it is difficult for the receiving device to obtain the LATM header, so the LATM header is separately included in control information such as SI. The LATM header may also be included in a message, table, or descriptor. The LATM header may also be included in the sample.

[0243] The receiving device acquires the LATM header from the SI or the like before starting decoding, and starts decoding the audio. Alternatively, as shown in (a) and (d) of Figure 21, if the MPU metadata is transmitted before the media data, the receiving device can receive the LATM header before the media data. Therefore, if the MPU metadata is transmitted before the media data, decoding can be performed even using a conventional receiving device.

[0244] <Other> The transmission order and the type of transmission order may be notified as control information such as an MMT packet header, a payload header, or an MPT or other table, message, descriptor, etc. Note that the type of transmission order here refers to, for example, the four types of transmission order shown in (a) to (d) of Fig. 21, and an identifier for identifying each type may be stored in a location that can be obtained before decoding begins.

[0245] Furthermore, different transmission order types may be used for audio and video, or a common type may be used for audio and video. Specifically, for example, audio may be transmitted in the order of MPU metadata, MF metadata, and media data, as shown in (a) of Fig. 21, and video may be transmitted in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Fig. 21.

[0246] The above-described method allows a receiving device to decode without using header information. Also, if the MPU metadata is transmitted before the media data (see (a) and (d) in FIG. 21), decoding becomes possible even with a conventional receiving device.

[0247] In particular, by transmitting the MF metadata after the media data ((d) in FIG. 21), delay due to encapsulation does not occur, and decoding can be performed even by a conventional receiving device.

[0248] [Configuration and operation of transmitting device] Next, the configuration and operation of the transmission device will be described. Fig. 23 is a block diagram of the transmission device according to the second embodiment, and Fig. 24 is a flowchart of the transmission method according to the second embodiment.

[0249] As shown in FIG. 23, the transmitting device 15 includes an encoding unit 16, a multiplexing unit 17, and a transmitting unit .

[0250] The encoding unit 16 generates encoded data by encoding the video or audio to be encoded according to, for example, H.265 (S10).

[0251] The multiplexing unit 17 multiplexes (packetizes) the encoded data generated by the encoding unit 16 (S11). Specifically, the multiplexing unit 17 packetizes each of the sample data, MPU metadata, and MF metadata that make up an MP4 format file. The sample data is data obtained by encoding a video signal or an audio signal, the MPU metadata is an example of first metadata, and the MF metadata is an example of second metadata. Both the first metadata and the second metadata are metadata used to decode the sample data, but the difference between them is that the second metadata includes data that can be generated only after the sample data is generated.

[0252] Here, the data that can be generated only after the sample data is generated is, for example, data other than the sample data stored in mdat in MP4 format (data in the header of mdat, i.e., type and length shown in Fig. 20). Here, the second metadata only needs to include the length, which is at least a part of this data.

[0253] The transmitting unit 18 transmits the packetized MP4 format file (S12). The transmitting unit 18 transmits the MP4 format file, for example, by the method shown in (d) of Fig. 21. That is, the transmitting unit 18 transmits the packetized MPU metadata, packetized sample data, and packetized MF metadata in this order.

[0254] Each of the encoding unit 16, multiplexing unit 17, and transmitting unit 18 is realized by a microcomputer, a processor, a dedicated circuit, or the like.

[0255] [Configuration of receiving device] Next, the configuration and operation of the receiving device will be described. Fig. 25 is a block diagram of the receiving device according to the second embodiment.

[0256] As shown in FIG. 25, the receiving device 20 includes a packet filtering unit 21, a transmission order type discrimination unit 22, a random access unit 23, a control information acquisition unit 24, a data acquisition unit 25, a PTS, DTS calculation unit 26, an initialization information acquisition unit 27, a decoding command unit 28, a decoding unit 29, and a presentation unit 30.

[0257] [Receiver operation 1] First, an operation of the receiving device 20 for identifying the MPU start position and the NAL unit position when the media is video will be described. Fig. 26 is a flowchart of such an operation of the receiving device 20. Note that it is assumed here that the transmission order type of the MPU data is stored in the SI information by the transmitting device 15 (multiplexing unit 17).

[0258] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering and acquires the transmission order type of the MPU data (S21).

[0259] Next, the transmission order type determination unit 22 determines (discriminates) whether or not MPU header information (at least one of MPU metadata and MF metadata) is included in the data after packet filtering (S22). If the MPU header information is included (Yes in S22), the random access unit 23 detects that the fragment type of the MMT payload header has switched to media data, thereby identifying the MPU first sample (S23).

[0260] On the other hand, if the MPU header information is not included (No in S22), the random access unit 23 identifies the MPU first sample based on the RAP_flag in the MMT packet header or the sample number in the MMT payload header (S24).

[0261] Furthermore, the transmission order type determination unit 22 determines whether or not MF metadata is included in the packet-filtered data (S25). If it is determined that MF metadata is included (Yes in S25), the data acquisition unit 25 acquires NAL units by reading the NAL units based on the sample, subsample offset, and size information included in the MF metadata (S26). On the other hand, if it is determined that MF metadata is not included (No in S25), the data acquisition unit 25 acquires NAL units by reading data of the size of the NAL units in order from the first NAL unit of the sample (S27).

[0262] Note that even if it is determined in step S22 that MPU header information is included, receiving device 20 may identify the MPU first sample using the process of step S24 instead of step S23. Furthermore, if it is determined that MPU header information is included, the process of step S23 and the process of step S24 may be used together.

[0263] Furthermore, even if it is determined in step S25 that MF metadata is included, the receiving device 20 may acquire the NAL unit using the process of step S27 without using the process of step S26. Furthermore, if it is determined that MF metadata is included, the process of step S23 and the process of step S24 may be used in combination.

[0264] Furthermore, when it is determined in step S25 that MF metadata is included, it is assumed that the MF data is transmitted after the media data. In this case, the receiving device 20 may buffer the media data and wait until the MF metadata is acquired before performing the process of step S26, or the receiving device 20 may determine whether to perform the process of step S27 without waiting for the MF metadata to be acquired.

[0265] For example, receiving device 20 may determine whether to wait for acquisition of MF metadata based on whether it has a buffer with a buffer size capable of buffering media data. Also, receiving device 20 may determine whether to wait for acquisition of MF metadata based on whether the end-to-end delay will be small. Also, receiving device 20 may perform the decoding process mainly using the process of step S26, and use the process of step S27 in the case of a processing mode when packet loss or the like occurs.

[0266] In addition, if the transmission order type is predetermined, steps S22 and S26 may be omitted, and in this case, the receiving device 20 may determine the method for identifying the MPU first sample and the method for identifying the NAL unit, taking into account the buffer size and end-to-end delay.

[0267] If the transmission order type is known in advance, the transmission order type discriminator 22 in the receiving device 20 is not necessary.

[0268] 26, a decoding instruction unit 28 outputs the data acquired by the data acquisition unit to a decoding unit 29 based on the PTS and DTS calculated by the PTS / DTS calculation unit 26 and the initialization information acquired by the initialization information acquisition unit 27. The decoding unit 29 decodes the data, and a presentation unit 30 presents the decoded data.

[0269] [Receiver operation 2] Next, a description will be given of an operation of the receiving device 20 to obtain initialization information based on the transmission order type and decode media data based on the initialization information. Fig. 27 is a flowchart of such an operation.

[0270] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering and acquires the transmission order type (S301).

[0271] Next, the transmission order type determination unit 22 determines whether or not MPU metadata has been transmitted (S302). If it is determined that MPU metadata has been transmitted (Yes in S302), the transmission order type determination unit 22 determines, based on the analysis result of step S301, whether or not the MPU metadata has been transmitted before the media data (S303). If the MPU metadata has been transmitted before the media data (Yes in S303), the initialization information acquisition unit 27 decodes the media data based on the common initialization information included in the MPU metadata and the initialization information of the sample data (S304).

[0272] On the other hand, if it is determined that the MPU metadata was transmitted after the media data (No in S303), the data acquisition unit 25 buffers the media data until the MPU metadata is acquired (S305), and performs the processing of step S304 after the MPU metadata is acquired.

[0273] Furthermore, if it is determined in step S302 that the MPU metadata has not been transmitted (No in S302), the initialization information acquisition unit 27 decodes the media data based only on the initialization information of the sample data (S306).

[0274] If decoding of the media data is guaranteed only based on the initialization information of the sample data on the transmitting side, the processes based on the determinations of steps S302 and S303 are not performed, and the process of step S306 is used.

[0275] Furthermore, receiving device 20 may determine whether or not to buffer the media data before step S305. In this case, if receiving device 20 determines to buffer the media data, it proceeds to the process of step S305, and if receiving device 20 determines not to buffer the media data, it proceeds to the process of step S306. The determination of whether or not to buffer the media data may be made based on the buffer size and occupancy of receiving device 20, or may be made taking into account end-to-end delay, for example, by selecting the buffer with the smaller end-to-end delay.

[0276] [Receiver operation 3] Here, we will explain in detail the transmission method and reception method when MF metadata is transmitted after media data ((c) of FIG. 21 and (d) of FIG. 21). The following explains the case of (d) of FIG. 21 as an example. Note that in transmission, only the method of (d) of FIG. 21 is used, and signaling of the transmission order type is not performed.

[0277] As mentioned above, when data is transmitted in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Figure 21, (D-1) The receiving device 20 acquires the MPU metadata, and then acquires the MF metadata, and then decodes the media data. (D-2) After acquiring the MPU metadata, the receiving device 20 decodes the media data without using the MF metadata. There are two possible decoding methods:

[0278] Here, D-1 requires buffering of media data to acquire MF metadata, but since decoding can be performed using MPU header information, it can be decoded by a conventional MP4-compliant receiving device.D-2 does not require buffering of media data to acquire MF metadata, but since decoding cannot be performed using MF metadata, special processing is required for decoding.

[0279] Furthermore, the method of FIG. 21(d) has the advantage that the MF metadata is transmitted after the media data, so no delay occurs due to encapsulation, and the end-to-end delay can be reduced.

[0280] The receiving device 20 can select one of the two decoding methods described above depending on the capabilities of the receiving device 20 and the quality of service that the receiving device 20 provides.

[0281] The transmitting device 15 must ensure that decoding can be performed with reduced occurrence of buffer overflow and underflow during the decoding operation in the receiving device 20. For example, the following parameters can be used as elements for defining the decoder model when decoding using the D-1 method.

[0282] Buffer size for reconfiguring the MPU (MPU buffer) For example, buffer size = maximum rate × maximum MPU time × α, where the maximum rate is the upper limit rate of the profile and level of the encoded data + MPU header overhead. The maximum MPU time is the maximum time length of a GOP when 1 MPU = 1 GOP (video).

[0283] Here, audio may be in the GOP unit common to video, or in a different unit. α is a margin to prevent overflow, and may be multiplied or added to the maximum rate x maximum MPU time. When multiplied, α≧1, and when added, α≧0.

[0284] Upper limit of the decoding delay time from when data is input to the MPU buffer until it is decoded. (TSTD_delay in MPEG-TS STD) For example, at the time of transmission, the DTS is set so that the time when acquisition of the MPU data at the receiver is completed<=DTS, taking into consideration the maximum MPU time and the upper limit of the decoding delay time.

[0285] Furthermore, the transmitting device 15 may assign a DTS and a PTS according to a decoder model for decoding using the D-1 method, thereby ensuring that decoding is possible for a receiving device that performs decoding using the D-1 method, and may also transmit auxiliary information required for decoding using the D-2 method.

[0286] For example, the transmitting device 15 can guarantee the operation of a receiving device that decodes using the D-2 method by signaling the pre-buffering time in the decoder buffer when decoding using the D-2 method.

[0287] The pre-buffering time may be included in SI control information such as a message, table, or descriptor, or may be included in the header of an MMT packet or MMT payload. Alternatively, the SEI in the encoded data may be overwritten. The DTS and PTS for decoding using the D-1 method may be stored in an MPU timestamp descriptor or SampleEntry, and the DTS and PTS for decoding using the D-2 method or the pre-buffering time may be described in the SEI.

[0288] If the receiving device 20 only supports MP4-compliant decoding operations using an MPU header, it may select decoding method D-1, and if it supports both D-1 and D-2, it may select either one.

[0289] The transmitting device 15 may assign a DTS and a PTS to one of the streams (D-1 in this example) so as to ensure the decoding operation, and may also transmit auxiliary information to assist the decoding operation of the other stream.

[0290] Furthermore, when the D-2 method is used, compared to when the D-1 method is used, there is a high possibility that the end-to-end delay will be larger due to the delay caused by pre-buffering of the MF metadata. Therefore, the receiving device 20 may select the D-2 method for decoding when it is desired to reduce the end-to-end delay. For example, the receiving device 20 may always use the D-2 method when it is desired to always reduce the end-to-end delay. Furthermore, the receiving device 20 may use the D-2 method only when operating in a low-delay presentation mode in which it is desired to present live content, channel selection, zapping, etc. with low delay.

[0291] FIG. 28 is a flowchart of such a receiving method.

[0292] First, the receiving device 20 receives an MMT packet and acquires MPU data (S401). Then, the receiving device 20 (transmission order type determination unit 22) determines whether to present the program in low-delay presentation mode (S402).

[0293] If the program is not presented in low-latency presentation mode (No in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using the header information (S405). Also, the receiving device 20 (PTS, DTS calculation unit 26, decoding instruction unit 28, decoding unit 29, presentation unit 30) performs decoding and presentation processing based on the PTS and DTS assigned by the transmitting side (S406).

[0294] On the other hand, when the program is presented in low-latency presentation mode (Yes in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using a decoding method that does not use header information (S403). The receiving device 20 also performs decoding and presentation processing based on auxiliary information for decoding without using PTS, DTS, and header information that is assigned by the transmitting side (S404). Note that in steps S403 and S404, processing may be performed using MPU metadata.

[0295] [Transmission and reception method using auxiliary data] The above has described the transmission and reception operations in the case where MF metadata is transmitted after media data (cases (c) and (d) in Figure 21). Next, a method will be described in which the transmitting device 15 transmits auxiliary data having some of the functions of the MF metadata, thereby enabling decoding to begin earlier and reducing end-to-end delay. Here, an example will be described in which auxiliary data is further transmitted based on the transmission method shown in (d) in Figure 21, but the method using auxiliary data is also applicable to the transmission methods shown in (a) to (c) in Figure 21.

[0296] Figure 29(a) is a diagram showing an MMT packet transmitted using the method shown in Figure 21(d). That is, data is transmitted in the order of MPU metadata, media data, and MF metadata.

[0297] Here, sample #1, sample #2, sample #3, and sample #4 are samples included in the media data. Note that, although an example in which media data is stored in MMT packets in sample units is described here, the media data may be stored in MMT packets in NAL unit units, or in units obtained by dividing an NAL unit. Note that there are also cases in which multiple NAL units are aggregated and stored in an MMT packet.

[0298] As explained in D-1 above, in the case of the method shown in (d) of Fig. 21, that is, when data is transmitted in the order of MPU metadata, media data, and MF metadata, there is a method in which the MPU metadata is acquired, then the MF metadata is acquired, and then the media data is decoded. This method of D-1 requires buffering of the media data to acquire the MF metadata, but has the advantage that the method of D-1 can be applied to conventional MP4-compliant receiving devices because decoding is performed using MPU header information. On the other hand, it has the disadvantage that the receiving device 20 must wait to start decoding until the MF metadata is acquired.

[0299] In contrast, as shown in (b) of FIG. 29, in the method using auxiliary data, the auxiliary data is transmitted before the MF metadata.

[0300] MF metadata contains information indicating the DTS, PTS, offsets, and sizes of all samples contained in a movie fragment, while ancillary data contains information indicating the DTS, PTS, offsets, and sizes of some of the samples contained in a movie fragment.

[0301] For example, the MF metadata includes information on all samples (samples #1-#4), whereas the auxiliary data includes information on some samples (samples #1-#2).

[0302] In the case shown in (b) of Figure 29, the auxiliary data enables decoding of sample #1 and sample #2, so the end-to-end delay is smaller than in the transmission method of D-1. Note that the auxiliary data may include any combination of sample information, and the auxiliary data may be transmitted repeatedly.

[0303] 29(c), when transmitting auxiliary information at timing A, transmitting device 15 includes information on sample #1 in the auxiliary information, and when transmitting auxiliary information at timing B, transmitting device 15 includes information on sample #1 and sample #2 in the auxiliary information. When transmitting auxiliary information at timing C, transmitting device 15 includes information on sample #1, sample #2, and sample #3 in the auxiliary information.

[0304] The MF metadata includes information on sample #1, sample #2, sample #3, and sample #4 (information on all samples in the movie fragment).

[0305] The auxiliary data does not necessarily have to be transmitted immediately after it is generated.

[0306] In addition, in the header of an MMT packet or an MMT payload, a type indicating that auxiliary data is stored is specified.

[0307] For example, when auxiliary data is stored in the MMT payload using the MPU mode, a data type indicating that it is auxiliary data is specified as the fragment_type field value (e.g., FT=3). The auxiliary data may be data based on the moof structure or may have other structures.

[0308] When auxiliary data is stored as a control signal (descriptor, table, message) in the MMT payload, a descriptor tag, table ID, message ID, etc. that indicate that it is auxiliary data are specified.

[0309] In addition, the PTS or DTS may be stored in the header of the MMT packet or MMT payload.

[0310] [Example of generating auxiliary data] An example in which a transmission device generates auxiliary data based on the configuration of moof will be described below. Fig. 30 is a diagram for explaining an example in which a transmission device generates auxiliary data based on the configuration of moof.

[0311] In a normal MP4, a moof is created for a movie fragment, as shown in Fig. 20. The moof contains information indicating the DTS, PTS, offset, and size of the samples included in the movie fragment.

[0312] Here, the transmitting device 15 composes an MP4 file (MP4 file) using only a portion of the sample data that composes the MPU, and generates auxiliary data.

[0313] For example, as shown in (a) of Figure 30, the transmitting device 15 generates an MP4 using only sample #1 of samples #1-#4 that make up the MPU, and the header of moof+mdat is used as auxiliary data.

[0314] Next, as shown in (b) of Figure 30, the transmitting device 15 generates an MP4 using samples #1 and #2 of samples #1-#4 that make up the MPU, and the header of moof+mdat is used as the next auxiliary data.

[0315] Next, as shown in (c) of Figure 30, the transmitting device 15 generates an MP4 using samples #1, #2, and #3 of samples #1-#4 that make up the MPU, and the header of moof+mdat is set as the next auxiliary data.

[0316] Next, as shown in (d) of Figure 30, the transmission device 15 generates all MP4s from samples #1-#4 that make up the MPU, and the header of moof+mdat among them becomes movie fragment metadata.

[0317] Although transmitting device 15 generates auxiliary data for each sample here, it may generate auxiliary data for every N samples. The value of N is an arbitrary number, and for example, if auxiliary data is transmitted M times when transmitting one MPU, N may be set to total samples / M.

[0318] The information indicating the offset of a sample in moof may be an offset value after the sample entry area for the subsequent number of samples is secured as a NULL area.

[0319] The auxiliary data may be generated so as to have a configuration in which the MF metadata is fragmented.

[0320] [Example of reception using auxiliary data] The following describes reception of auxiliary data generated as described in Fig. 30. Fig. 31 is a diagram for explaining reception of auxiliary data. In Fig. 31(a), the number of samples constituting the MPU is 30, and auxiliary data is generated and transmitted every 10 samples.

[0321] In FIG. 30(a), auxiliary data #1 includes sample information for samples #1-#10, auxiliary data #2 includes sample information for samples #1-#20, and MF metadata includes sample information for samples #1-#30.

[0322] Although samples #1-#10, samples #11-#20, and samples #21-#30 are stored in one MMT payload, they may also be stored in sample units or NAL units, or in fragment or aggregate units.

[0323] The receiving device 20 receives the packets of MPU meta, samples, MF meta, and auxiliary data, respectively.

[0324] The receiving device 20 concatenates the sample data in the order in which they are received (to the end), and after receiving the latest auxiliary data, updates the previous auxiliary data.Furthermore, the receiving device 20 can configure a complete MPU by finally replacing the auxiliary data with MF metadata.

[0325] Upon receiving auxiliary data #1, receiving device 20 concatenates the data to form an MP4, as shown in the upper part of (b) of Fig. 31. This allows receiving device 20 to parse samples #1-#10 using the MPU metadata and information in auxiliary data #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0326] Furthermore, upon receiving auxiliary data #2, receiving device 20 concatenates the data to form an MP4, as shown in the middle part of (b) of Fig. 31. This allows receiving device 20 to parse samples #1-#20 using the MPU metadata and information in auxiliary data #2, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0327] Furthermore, upon receiving the MF metadata, the receiving device 20 concatenates the data as shown in the lower part of (b) of Fig. 31 to construct an MP4 file. This enables the receiving device 20 to parse samples #1-#30 using the MPU metadata and MF metadata, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF metadata.

[0328] In the absence of auxiliary data, receiving device 20 could only acquire sample information after receiving MF metadata, and therefore had to start decoding after receiving MF metadata. However, by transmitting device 15 generating and transmitting auxiliary data, receiving device 20 can acquire sample information using the auxiliary data without waiting for reception of MF metadata, thereby speeding up the decoding start time. Furthermore, by transmitting device 15 generating auxiliary data based on the moof described with reference to Figure 30, receiving device 20 can parse using a conventional MP4 parser as is.

[0329] Furthermore, newly generated auxiliary data and MF metadata contain sample information that overlaps with previously transmitted auxiliary data. Therefore, even if previous auxiliary data cannot be obtained due to packet loss, etc., it is possible to reconstruct the MP4 and obtain sample information (PTS, DTS, size, and offset) by using the newly obtained auxiliary data and MF metadata.

[0330] It should be noted that the auxiliary data does not necessarily have to include information on past sample data. For example, auxiliary data #1 may correspond to sample data #1-#10, and auxiliary data #2 may correspond to sample data #11-#20. For example, as shown in (c) of Figure 31, transmitting device 15 may use complete MF metadata as a data unit and sequentially transmit fragments of the data unit as auxiliary data.

[0331] Furthermore, the transmitting device 15 may repeatedly transmit the auxiliary data or the MF metadata in order to deal with packet loss.

[0332] The MMT packet and MMT payload in which auxiliary data is stored contain an MPU sequence number and an asset ID, as well as MPU metadata, MF metadata, and sample data.

[0333] The above-described receiving operation using auxiliary data will be described with reference to the flowchart in Fig. 32. Fig. 32 is a flowchart of the receiving operation using auxiliary data.

[0334] First, receiving device 20 receives an MMT packet and analyzes the packet header and payload header (S501). Next, receiving device 20 analyzes whether the fragment type is auxiliary data or MF metadata (S502). If the fragment type is auxiliary data, receiving device 20 overwrites and updates the previous auxiliary data (S503). At this time, if there is no previous auxiliary data for the same MPU, receiving device 20 uses the received auxiliary data as new auxiliary data. Then, receiving device 20 acquires samples based on the MPU metadata, auxiliary data, and sample data, and performs decoding (S507).

[0335] On the other hand, if the fragment type is MF metadata, the receiving device 20 overwrites the previous auxiliary data with the MF metadata in step S505 (S505).Then, the receiving device 20 obtains the sample in the form of a complete MPU based on the MPU metadata, MF metadata, and sample data, and performs decoding (S506).

[0336] Although not shown in Figure 32, in step S502, if the fragment type is MPU metadata, the receiving device 20 stores the data in a buffer, and if the fragment type is sample data, it stores the data concatenated at the end for each sample in a buffer.

[0337] If the auxiliary data cannot be obtained due to packet loss, the receiving device 20 can either overwrite the sample with the latest auxiliary data or decode the sample using the previous auxiliary data.

[0338] The transmission cycle and the number of transmissions of the auxiliary data may be predetermined values. Information on the transmission cycle and the number of times (count, countdown) may be transmitted together with the data. For example, the transmission cycle, the number of times of transmission, and a timestamp such as initial_cpb_removal_delay may be stored in the data unit header.

[0339] By transmitting ancillary data including information on the first sample of the MPU at least once before the initial_cpb_removal_delay, it is possible to comply with the CPB buffer model. In this case, the MPU timestamp descriptor is set to a value based on the picture timing SEI.

[0340] The transmission method for receiving operations using such auxiliary data is not limited to the MMT method, but can also be applied to streaming transmission of packets configured in ISOBMFF file format, such as MPEG-DASH.

[0341] [Transmission method when one MPU consists of multiple movie fragments] In the explanation from Fig. 19 onwards, one MPU is composed of one movie fragment, but here we will explain the case where one MPU is composed of multiple movie fragments. Fig. 33 shows the configuration of an MPU composed of multiple movie fragments.

[0342] In Figure 33, samples (#1-#6) stored in one MPU are divided into two movie fragments. The first movie fragment is generated based on samples #1-#3, and a corresponding moof box is generated. The second movie fragment is generated based on samples #4-#6, and a corresponding moof box is generated.

[0343] The headers of the moof box and mdat box in the first movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #1. Meanwhile, the headers of the moof box and mdat box in the second movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #2. Note that in Figure 33, the MMT payload in which movie fragment metadata is stored is hatched.

[0344] The number of samples constituting an MPU and the number of samples constituting a movie fragment are arbitrary. For example, the number of samples constituting an MPU may be the number of samples in a GOP unit, and two movie fragments may be composed by using half the number of samples in a GOP unit as movie fragments.

[0345] Note that, although an example is shown here in which one MPU contains two movie fragments (a moof box and an mdat box), one MPU may contain three or more movie fragments instead of two. Also, the samples stored in a movie fragment do not have to be divided equally, but may be divided into any number of samples.

[0346] 33, the MPU metadata unit and the MF metadata unit are each stored as a data unit in the MMT payload. However, transmitting device 15 may store units such as ftyp, mmpu, moov, and moof as data units in the MMT payload in data unit units, or may store data units in the MMT payload in fragmented units. Furthermore, transmitting device 15 may store data units in the MMT payload in aggregated units.

[0347] 33, samples are stored in the MMT payload in sample units. However, transmitting device 15 may configure data units in NAL unit units or units aggregating multiple NAL units instead of sample units, and store the data units in the MMT payload. Also, transmitting device 15 may store data units in fragmented units or aggregated units in the MMT payload.

[0348] In Fig. 33, the MPU is configured in the order of moof#1, mdat#1, moof#2, mdat#2, and an offset is assigned to moof#1, assuming that the corresponding mdat#1 is attached after it. However, an offset may also be assigned to mdat#1, assuming that it is attached before moof#1. In this case, however, movie fragment metadata cannot be generated in the form of moof+mdat, and the headers of moof and mdat are transmitted separately.

[0349] Next, a description will be given of the transmission order of MMT packets when transmitting an MPU having the configuration described in Fig. 33. Fig. 34 is a diagram for explaining the transmission order of MMT packets.

[0350] Figure 34(a) shows the transmission order when MMT packets are transmitted in the configuration order of the MPUs shown in Figure 33. Figure 34(a) specifically shows an example in which MPU meta, MF meta #1, media data #1 (samples #1-#3), MF meta #2, and media data #2 (samples #4-#6) are transmitted in this order.

[0351] FIG. 34(b) shows an example in which MPU meta, media data #1 (samples #1-#3), MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in this order.

[0352] FIG. 34(c) shows an example in which media data #1 (samples #1-#3), MPU meta, MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in this order.

[0353] MF meta #1 is generated using samples #1-#3, and MF meta #2 is generated using samples #4-#6. Therefore, when the transmission method of Figure 34(a) is used, a delay occurs in the transmission of sample data due to encapsulation.

[0354] In contrast, when the transmission methods of Figures 34(b) and 34(c) are used, samples can be transmitted without waiting for the MF meta to be generated, so no delay due to encapsulation occurs and end-to-end delay can be reduced.

[0355] Also, in the transmission order (a) of Figure 34, one MPU is divided into multiple movie fragments, and the number of samples stored in the MF meta is smaller than in the case of Figure 19, so the amount of delay due to encapsulation can be reduced compared to the case of Figure 19.

[0356] In addition to the method shown here, for example, transmitting device 15 may concatenate MF meta #1 and MF meta #2 and transmit them together at the end of the MPU. In this case, MF meta of different movie fragments may be aggregated and stored in one MMT payload. Also, MF meta of different MPUs may be aggregated and stored in an MMT payload.

[0357] [How to receive when one MPU consists of multiple movie fragments] Here, a description will be given of an example of operation of receiving device 20 that receives and decodes MMT packets transmitted in the transmission order described in (b) of Fig. 34. Figs. 35 and 36 are diagrams for explaining such an example of operation.

[0358] The receiving device 20 receives each of the MMT packets including the MPU meta, samples, and MF meta transmitted in the transmission order shown in Fig. 35. The sample data is concatenated in the order in which it is received.

[0359] At T1, which is the time when MF meta #1 is received, receiving device 20 concatenates the data as shown in (1) of Fig. 36 to construct an MP4. This allows receiving device 20 to obtain samples #1-#3 based on the MPU metadata and information on MF meta #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF meta.

[0360] Furthermore, receiving device 20 concatenates the data as shown in (2) of Fig. 36 at T2, which is the time when MF meta #2 is received, to construct an MP4. This allows receiving device 20 to acquire samples #4-#6 based on the MPU metadata and information in MF meta #2, and to perform decoding based on the PTS, DTS, offset, and size information in the MF meta. Receiving device 20 may also acquire samples #1-#6 based on the information in MF meta #1 and MF meta #2 by concatenating the data as shown in (3) of Fig. 36 to construct an MP4.

[0361] By dividing a single MPU into multiple movie fragments, the time it takes for the MPU to acquire the initial MF meta is shortened, which allows for earlier decoding start time and reduces the buffer size for storing samples before decoding.

[0362] The transmitting device 15 may set the division unit of the movie fragment so that the time from transmitting (or receiving) the first sample in the movie fragment to transmitting (or receiving) the MF meta corresponding to the movie fragment is shorter than the initial_cpb_removal_delay specified by the encoder. By setting it in this way, the receiving buffer can follow the cpb buffer, realizing low-delay decoding. In this case, absolute times based on the initial_cpb_removal_delay can be used for the PTS and DTS.

[0363] Alternatively, the transmitting device 15 may divide the movie fragments at equal intervals, or divide subsequent movie fragments at shorter intervals than the previous movie fragments, which allows the receiving device 20 to always receive MF meta containing information about a sample before decoding that sample, enabling continuous decoding.

[0364] The absolute time of the PTS and DTS can be calculated using the following two methods.

[0365] (1) The absolute times of the PTS and DTS are determined based on the reception time (T1 or T2) of the MF meta #1 or MF meta #2 and the relative times of the PTS and DTS included in the MF meta.

[0366] (2) The absolute time of the PTS and DTS is determined based on the absolute time signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative time of the PTS and DTS included in the MF meta.

[0367] Also, (2-A) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on the initial_cpb_removal_delay specified by the encoder.

[0368] Also, (2-B) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on a predicted value of the reception time of the MF meta.

[0369] Note that MF meta #1 and MF meta #2 may be transmitted repeatedly. By repeatedly transmitting MF meta #1 and MF meta #2, the receiving device 20 can acquire the MF meta again even if it was unable to acquire it due to packet loss or the like.

[0370] The payload header of an MFU containing samples that constitute a movie fragment can store an identifier indicating the order of the movie fragment. On the other hand, an identifier indicating the order of the MF meta that constitutes a movie fragment is not included in the MMT payload. Therefore, the receiving device 20 identifies the order of the MF meta by the packet_sequence_number. Alternatively, the transmitting device 15 may store and signal an identifier indicating the ordinal number of the movie fragment to which the MF meta belongs in control information (message, table, descriptor), the MMT header, the MMT payload header, or the data unit header.

[0371] The transmitting device 15 may transmit the MPU meta, MF meta, and samples in a predetermined transmission order, and the receiving device 20 may perform the receiving process based on the predetermined transmission order. Alternatively, the transmitting device 15 may signal the transmission order, and the receiving device 20 may select (determine) the receiving process based on the signaling information.

[0372] The above-described receiving method will be explained using Fig. 37. Fig. 37 is a flowchart of the operation of the receiving method explained in Figs.

[0373] First, the receiving device 20 determines (identifies) whether the data included in the payload is MPU metadata, MF metadata, or sample data (MFU) based on the fragment type indicated in the MMT payload (S601, S602). If the data is sample data, the receiving device 20 buffers the sample and waits for reception of MF metadata corresponding to the sample and for the start of decoding (S603).

[0374] On the other hand, in step S602, if the data is MF metadata, the receiving device 20 obtains sample information (PTS, DTS, position information, and size) from the MF metadata, obtains a sample based on the obtained sample information, and decodes and presents the sample based on the PTS and DTS (S604).

[0375] Although not shown, if the data is MPU metadata, the MPU metadata contains initialization information necessary for decoding, which the receiving device 20 stores and uses to decode the sample data in step S604.

[0376] When the receiving device 20 stores the received MPU data (MPU metadata, MF metadata, and sample data) in a storage device, it stores the data after rearranging it into the MPU configuration described in Figure 19 or Figure 33.

[0377] On the transmitting side, packet sequence numbers are assigned to MMT packets with the same packet ID. At this time, packet sequence numbers may be assigned after MMT packets containing MPU metadata, MF metadata, and sample data are rearranged in transmission order, or packet sequence numbers may be assigned in the order before rearrangement.

[0378] If packet sequence numbers are assigned in the order before rearrangement, reception device 20 can rearrange the data in the order configured in the MPU based on the packet sequence numbers, facilitating storage.

[0379] [Method for detecting the beginning of an access unit and the beginning of a slice segment] A method for detecting the beginning of an access unit or a slice segment based on information in the MMT packet header and the MMT payload header will be described.

[0380] Here, two examples are shown: one where non-VCL NAL units (such as access unit delimiters, VPS, SPS, PPS, and SEI) are collectively stored as data units in an MMT payload, and one where each non-VCL NAL unit is treated as a data unit, and the data units are aggregated and stored in a single MMT payload.

[0381] FIG. 38 is a diagram showing a case where non-VCL NAL units are aggregated as individual data units.

[0382] 38, the start of the access unit is an MMT packet whose fragment_type value is MFU, and is the start data of an MMT payload that includes a data unit whose aggregation_flag value is 1 and whose offset value is 0. In this case, the Fragmentation_indicator value is 0.

[0383] Also, in the case of Figure 38, the beginning of the slice segment is an MMT packet whose fragment_type value is MFU, and is the beginning data of an MMT payload whose aggregation_flag value is 0 and whose fragmentation_indicator value is 00 or 01.

[0384] 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. Note that the field values ​​of the packet header are as shown in FIG. 17 (or FIG. 18).

[0385] In the case of FIG. 39, the head of the access unit is the head data of the payload in the packet with an Offset value of 0.

[0386] In addition, in the case of FIG. 39, the start of a slice segment is the first data of the payload of a packet whose Offset value is a value other than 0 and whose fragmentation indicator value is 00 or 01.

[0387] [Reception process when packet loss occurs] Generally, when transmitting MP4 format data in an environment where packet loss occurs, the receiving device 20 restores packets using ALFEC (Application Layer FEC), packet retransmission control, or the like.

[0388] However, if packet loss occurs in streaming such as broadcasting when AL-FEC cannot be used, the packets cannot be restored.

[0389] After data is lost due to packet loss, the receiving device 20 needs to resume decoding of video and audio. To do this, the receiving device 20 needs to detect the beginning of an access unit or NAL unit and start decoding from the beginning of the access unit or NAL unit.

[0390] However, since there is no start code at the beginning of an NAL unit in MP4 format, the receiving device 20 cannot detect the beginning of an access unit or an NAL unit even if it analyzes the stream.

[0391] FIG. 40 is a flowchart of the operation of the receiving device 20 when a packet loss occurs.

[0392] The receiving device 20 detects packet loss using the packet sequence number, packet counter, fragment counter, etc. in the header of the MMT packet or MMT payload (S701), and determines which packet has been lost based on the context (S702).

[0393] If it is determined that no packet loss has occurred (No in S702), the receiving device 20 constructs an MP4 file and decodes the access units or NAL units (S703).

[0394] If it is determined that a packet loss has occurred (Yes in S702), the receiving device 20 generates a NAL unit corresponding to the NAL unit that has experienced the packet loss using dummy data, and constructs an MP4 file (S704). When inserting dummy data into a NAL unit, the receiving device 20 indicates that the NAL unit type is dummy data.

[0395] In addition, the receiving device 20 can resume decoding by detecting the beginning of the next access unit or NAL unit and inputting the beginning data into the decoder based on the methods described in Figures 17, 18, 38, and 39 (S705).

[0396] In addition, if packet loss occurs, the receiving device 20 may resume decoding from the beginning of the access unit and NAL unit based on information detected based on the packet header, or may resume decoding from the beginning of the access unit and NAL unit based on header information of the reconstructed MP4 file, which includes a dummy data NAL unit.

[0397] When storing an MP4 file (MPU), the receiving device 20 may separately acquire packet data (NAL units, etc.) lost due to packet loss and store (replace) it from broadcasting or communication.

[0398] At this time, when receiving device 20 acquires the lost packet from the communication, it notifies the server of information about the lost packet (packet ID, MPU sequence number, packet sequence number, IP data flow number, IP address, etc.) and acquires the packet. The receiving device 20 is not limited to acquiring only the lost packet, but may also acquire a group of packets before and after the lost packet at the same time.

[0399] [How to compose a movie fragment] Here we will explain in detail how to configure movie fragments.

[0400] As described in Fig. 33, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are arbitrary. For example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU may be a fixed predetermined number, or may be dynamically determined.

[0401] Here, by configuring the movie fragments on the transmitting side (transmitting device 15) so as to satisfy the following conditions, low-delay decoding in the receiving device 20 can be guaranteed.

[0402] The conditions are as follows:

[0403] The transmitting device 15 generates and transmits MF meta as movie fragments, which are units obtained by dividing sample data, so that the receiving device 20 can always receive MF meta containing information about any sample (Sample(i)) before the decoding time (DTS(i)) of that sample.

[0404] Specifically, the transmitting device 15 constructs a movie fragment using samples (including the i-th sample) that have been coded before DTS(i).

[0405] To ensure low-latency decoding, the following method, for example, is used to dynamically determine the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU.

[0406] (1) At the start of decoding, the decoding time DTS(0) of the first sample Sample(0) of the GOP is based on initial_cpb_removal_delay. The transmitting device constructs a first movie fragment using samples that have already been coded at a time before DTS(0). The transmitting device 15 also generates MF metadata corresponding to the first movie fragment and transmits it at a time before DTS(0).

[0407] (2) The transmitting device 15 constructs movie fragments so that the above conditions are met for subsequent samples as well.

[0408] For example, if the first sample of a movie fragment is the kth sample, the MF meta of the movie fragment including the kth sample is transmitted by the decoding time DTS(k) of the kth sample. If the encoding completion time of the lth sample is before DTS(k) and the encoding completion time of the (l+1)th sample is after DTS(k), the transmitting device 15 constructs a movie fragment using the kth sample to the lth sample.

[0409] In addition, the transmitting device 15 may construct a movie fragment using the kth sample to a sample less than the lth sample.

[0410] (3) After completing the encoding of the last sample in the MPU, the transmitting device 15 constructs a movie fragment using the remaining samples, generates MF metadata corresponding to the movie fragment, and transmits it.

[0411] It should be noted that the transmitting device 15 may construct a movie fragment using only a portion of the samples that have been coded, rather than using all of the samples that have been coded.

[0412] In the above example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are dynamically determined based on the above conditions to ensure low-latency decoding. However, the method for determining the number of samples and the number of movie fragments is not limited to this method. For example, the number of movie fragments constituting one MPU may be fixed to a predetermined value, and the number of samples may be determined to satisfy the above conditions. Furthermore, the number of movie fragments constituting one MPU and the time at which the movie fragments are divided (or the code amount of the movie fragments) may be fixed to predetermined values, and the number of samples may be determined to satisfy the above conditions.

[0413] In addition, if the MPU is divided into multiple movie fragments, information indicating whether the MPU is divided into multiple movie fragments, attributes of the divided movie fragments, or MF meta attributes for the divided movie fragments may be transmitted.

[0414] Here, the attribute of a movie fragment is information indicating whether the movie fragment is the first movie fragment of an MPU, the last movie fragment of an MPU, or some other movie fragment.

[0415] In addition, the attributes of the MF meta are information that indicates whether the MF meta corresponds to the first movie fragment of the MPU, the last movie fragment of the MPU, or the other movie fragment.

[0416] The transmitting device 15 may store and transmit the number of samples that make up a movie fragment and the number of movie fragments that make up one MPU as control information.

[0417] [Operation of receiving device] The operation of the receiving device 20 based on the movie fragment configured as above will now be described.

[0418] The receiving device 20 determines the absolute times of the PTS and DTS based on the absolute times signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative times of the PTS and DTS included in the MF meta.

[0419] Based on information on whether the MPU is divided into multiple movie fragments, the receiving device 20 performs the following processing based on the attributes of the divided movie fragments if the MPU is divided.

[0420] (1) When the movie fragment is the first movie fragment of the MPU, the receiving device 20 generates the absolute time of the PTS and DTS using the absolute time of the PTS of the first sample included in the MPU timestamp descriptor and the relative times of the PTS and DTS included in the MF meta.

[0421] (2) If the movie fragment is not the first movie fragment of the MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the relative times of the PTS and DTS included in the MF meta, without using the information in the MPU timestamp descriptor.

[0422] (3) If the movie fragment is the last movie fragment in the MPU, the receiving device 20 calculates the absolute times of the PTS and DTS of all samples and then resets the PTS and DTS calculation process (addition process of relative times). Note that the reset process may also be performed on the first movie fragment in the MPU.

[0423] The receiving device 20 may determine whether a movie fragment is divided as follows: The receiving device 20 may also acquire attribute information of the movie fragment as follows.

[0424] For example, the receiving device 20 may determine whether the movie fragment has been divided based on the identifier movie_fragment_sequence_number field value that indicates the order of the movie fragments indicated in the MMTP (MMT Protocol) payload header.

[0425] Specifically, the receiving device 20 may determine that an MPU is divided into multiple movie fragments if the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 1, and there is a value of 2 or greater in the field value.

[0426] In addition, the receiving device 20 may determine that an MPU is divided into multiple movie fragments if the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 0, and there is a value other than 0 in the field value.

[0427] The attribute information of the movie fragment may also be determined based on the movie_fragment_sequence_number.

[0428] In addition, without using the movie_fragment_sequence_number, it is also possible to determine whether a movie fragment is divided and the attribute information of the movie fragment by counting the transmission of movie fragments and MF meta contained in one MPU.

[0429] With the above-described configurations of the transmitting device 15 and the receiving device 20, the receiving device 20 can receive movie fragment metadata at intervals shorter than that of an MPU, enabling decoding to start with low latency. Also, decoding with low latency can be performed using a decoding process based on the MP4 parsing method.

[0430] The reception operation when the MPU is divided into multiple movie fragments as described above will be explained using a flowchart. Figure 41 is a flowchart of the reception operation when the MPU is divided into multiple movie fragments. Note that this flowchart illustrates the operation of step S604 in Figure 37 in more detail.

[0431] First, based on the data type indicated in the MMTP payload header, if the data type is MF meta, the receiving device 20 acquires the MF meta data (S801).

[0432] Next, the receiving device 20 determines whether the MPU is divided into multiple movie fragments (S802), and if the MPU is divided into multiple movie fragments (Yes in S802), determines whether the received MF metadata is the first metadata in the MPU (S803). If the received MF metadata is the first MF metadata in the MPU (Yes in S803), the receiving device 20 calculates the absolute times of the PTS and DTS from the absolute time of the PTS indicated in the MPU timestamp descriptor and the relative times of the PTS and DTS indicated in the MF metadata (S804), and determines whether the metadata is the last metadata in the MPU (S805).

[0433] On the other hand, if the received MF metadata is not the MF metadata at the beginning of the MPU (No in S803), the receiving device 20 calculates the absolute times of the PTS and DTS using the relative times of the PTS and DTS indicated in the MF metadata without using the information in the MPU timestamp descriptor (S808), and proceeds to processing in step S805.

[0434] If it is determined in step S805 that this is the last MF metadata of the MPU (Yes in S805), the receiving device 20 calculates the absolute times of the PTS and DTS of all samples and then resets the PTS and DTS calculation process.If it is determined in step S805 that this is not the last MF metadata of the MPU (No in S805), the receiving device 20 ends the process.

[0435] Also, if it is determined in step S802 that the MPU is not divided into multiple movie fragments (No in S802), the receiving device 20 acquires sample data based on the MF metadata transmitted after the MPU and determines the PTS and DTS (S807).

[0436] Finally, although not shown, the receiving device 20 performs decoding and presentation processes based on the determined PTS and DTS.

[0437] [Issues that arise when splitting movie fragments and their solutions] So far, we have explained how to reduce end-to-end delay by dividing movie fragments. From here, we will explain the new issues that arise when dividing movie fragments and how to solve them.

[0438] First, as background, the picture structure in coded data will be described. Figure 42 is a diagram showing an example of a prediction structure of a picture for each TemporalId when implementing temporal scalability.

[0439] In coding formats such as MPEG-4 AVC and HEVC (High Efficiency Video Coding), temporal scalability can be achieved by using B-pictures (bidirectional reference predictive pictures) that can be referenced from other pictures.

[0440] TemporalId shown in (a) of Figure 42 is an identifier for a layer in the coding structure, with larger TemporalId values ​​indicating deeper layers. Square blocks indicate pictures, with Ix in each block indicating an I-picture (intra-picture predicted picture), Px indicating a P-picture (forward reference predicted picture), and Bx and bx indicating B-pictures (bidirectional reference predicted pictures). The x in Ix / Px / Bx indicates the display order, indicating the order in which the pictures are displayed. Arrows between pictures indicate reference relationships; for example, picture B4 indicates that a predicted image is generated using I0 and B8 as reference images. It is prohibited for a picture to use as a reference image another picture with a TemporalId higher than its own. The layers are specified to allow for temporal scalability. For example, in Figure 42, decoding all pictures results in 120 fps (frames per second) video, but decoding only layers with TemporalId from 0 to 3 results in 60 fps video.

[0441] Figure 43 shows the relationship between the decoding time (DTS) and display time (PTS) for each picture in Figure 42. For example, picture I0 shown in Figure 43 is displayed after the decoding of B4 is completed so that no gap occurs in the decoding and display.

[0442] As shown in Figure 43, when the prediction structure includes a B picture, the decoding order and the display order are different, so after decoding the picture in the receiving device 20, picture delay processing and picture reordering processing are required.

[0443] Although examples of picture prediction structures in temporal scalability have been described above, even when temporal scalability is not used, picture delay processing and reordering processing may be required depending on the prediction structure. Figure 44 is a diagram showing an example of a prediction structure of a picture that requires picture delay processing and reordering processing. Note that the numbers in Figure 44 indicate the decoding order.

[0444] As shown in Figure 44, depending on the prediction structure, the first sample in decoding order may differ from the first sample in presentation order, and in Figure 44, the first sample in presentation order is the fourth sample in decoding order. Note that Figure 44 shows an example of a prediction structure, and the prediction structure is not limited to this structure. In other prediction structures, the first sample in decoding order may differ from the first sample in presentation order.

[0445] Like FIG. 33, FIG. 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and MMTP packets. Note that the number of samples constituting an MPU and the number of samples constituting a movie fragment are arbitrary. For example, the number of samples constituting an MPU may be the number of samples per GOP, and two movie fragments may be constituted by using half the number of samples per GOP. One sample may be treated as one movie fragment, or the samples constituting an MPU may not be divided.

[0446] While FIG. 45 shows an example in which one MPU contains two movie fragments (a moof box and an mdat box), the number of movie fragments contained in one MPU does not have to be two. The number of movie fragments contained in one MPU may be three or more, or may be the number of samples contained in the MPU. Furthermore, the samples stored in a movie fragment do not have to be divided equally, but may be divided into any number of samples.

[0447] Movie fragment metadata (MF metadata) contains information on the PTS, DTS, offset, and size of the samples contained in the movie fragment, and when the receiving device 20 decodes a sample, it extracts the PTS and DTS from the MF meta containing information about the sample and determines the decoding timing and presentation timing.

[0448] Hereinafter, for the sake of detailed explanation, the absolute value of the decoding time of the i sample will be referred to as DTS(i), and the absolute value of the presentation time will be referred to as PTS(i).

[0449] The information of the i-th sample among the timestamp information stored in the moof in the MF meta is specifically the relative value of the decoding time of the i-th sample and the (i+1)-th sample, and the relative value of the decoding time of the i-th sample and the presentation time, which will be referred to as DT(i) and CT(i) hereafter.

[0450] Movie fragment metadata #1 contains DT(i) and CT(i) for samples #1-#3, and movie fragment metadata #2 contains DT(i) and CT(i) for samples #4-#6.

[0451] The absolute PTS value of the access unit at the beginning of the MPU is stored in the MPU timestamp descriptor or the like, and the receiving device 20 calculates the PTS and DTS based on the PTS_MPU of the access unit at the beginning of the MPU, the CT, and the DT.

[0452] FIG. 46 is a diagram for explaining a method of calculating PTS and DTS and problems involved when an MPU is configured using samples #1 to #10.

[0453] (a) of Figure 46 shows an example where the MPU is not divided into movie fragments, (b) of Figure 46 shows an example where the MPU is divided into two movie fragments of 5 sample units, and (c) of Figure 46 shows an example where the MPU is divided into 10 movie fragments of sample units.

[0454] As explained in Figure 45, when the PTS and DTS are calculated using the MPU timestamp descriptor and the timestamp information (CT and DT) in the MP4, the first sample in the presentation order in Figure 44 is the fourth sample in decoding order. Therefore, the PTS stored in the MPU timestamp descriptor is the PTS (absolute value) of the fourth sample in decoding order. Note that, hereinafter, this sample will be referred to as sample A. Furthermore, the first sample in decoding order will be referred to as sample B.

[0455] Because the only absolute time information related to the timestamp is the information in the MPU timestamp descriptor, the receiving device 20 cannot calculate the PTS (absolute time) and DTS (absolute time) of other samples until the arrival of sample A. The receiving device 20 also cannot calculate the PTS and DTS of sample B.

[0456] 46(a), sample A is included in the same movie fragment as sample B and is stored in one MF meta, so receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.

[0457] 46(b), sample A is included in the same movie fragment as sample B and is stored in one MF meta. Therefore, receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.

[0458] In the example of (c) in Figure 46, sample A is included in a different movie fragment from sample B. Therefore, receiving device 20 cannot determine the DTS of sample B until it receives MF meta including the CT and DT of the movie fragment that includes sample A.

[0459] Therefore, in the example of FIG. 46(c), the receiving device 20 cannot start decoding immediately after the arrival of the B sample.

[0460] In this way, if a movie fragment containing a B sample does not contain an A sample, the receiving device 20 cannot start decoding the B sample until it has received the MF meta for the movie fragment containing the A sample.

[0461] This issue occurs when the first sample in presentation order does not match the first sample in decoding order, and the movie fragment is split to the point where sample A and sample B are no longer stored in the same movie fragment. This issue also occurs regardless of whether the MF meta is forward or backward.

[0462] In this way, if the first sample in presentation order does not match the first sample in decoding order, and if sample A and sample B are not stored in the same movie fragment, the DTS cannot be determined immediately after receiving sample B. Therefore, transmitting device 15 separately transmits the DTS (absolute value) of sample B or information that allows the receiving side to calculate the DTS (absolute value) of sample B. Such information may be transmitted using control information, a packet header, etc.

[0463] Using this information, the receiving device 20 calculates the DTS (absolute value) of sample B. Fig. 47 is a flowchart of the receiving operation when the DTS is calculated using this information.

[0464] The receiving device 20 receives the movie fragment at the beginning of the MPU (S901), and determines whether the A sample and the B sample are stored in the same movie fragment (S902). If they are stored in the same movie fragment (Yes in S902), the receiving device 20 calculates the DTS using only the MF meta information, without using the DTS (absolute time) of the B sample, and starts decoding (S904). Note that in step S904, the receiving device 20 may determine the DTS using the DTS of the B sample.

[0465] On the other hand, if sample A and sample B are not stored in the same movie fragment in step S902 (No in S902), receiving device 20 obtains the DTS (absolute time stamp) of sample B, determines the DTS, and starts decoding (S903).

[0466] In the above description, an example has been described in which the absolute value of the decoding time and the absolute value of the presentation time of each sample are calculated using MF meta (timestamp information stored in the moof in MP4 format) in the MMT standard, but it goes without saying that the MF meta may be replaced with any control information that can be used to calculate the absolute value of the decoding time and the absolute value of the presentation time of each sample. Examples of such control information include control information in which the above-mentioned relative value CT(i) of the decoding time between the i-th sample and the (i+1)-th sample is replaced with the relative value of the presentation time between the i-th sample and the (i+1)-th sample, and control information that includes both the relative value CT(i) of the decoding time between the i-th sample and the (i+1)-th sample and the relative value of the presentation time between the i-th sample and the (i+1)-th sample.

[0467] (Embodiment 3) [overview] In the third embodiment, a content transmission method and data structure when transmitting content such as video, audio, subtitles, and data broadcasting via broadcasting will be described. That is, a content transmission method and data structure specialized for playing back broadcast streams will be described.

[0468] In the third embodiment, an example will be described in which the MMT method (hereinafter also simply referred to as MMT) is used as the multiplexing method, but other multiplexing methods such as MPEG-DASH or RTP may also be used.

[0469] First, a method for storing a data unit (DU) in a payload in MMT will be described in detail. Fig. 48 is a diagram for explaining a method for storing a data unit in a payload in MMT.

[0470] In MMT, a transmitting device stores part of the data that constitutes an MPU in an MMTP payload as a data unit, and transmits it with a header. The header includes an MMTP payload header and an MMTP packet header. The data unit may be in units of NAL units or samples.

[0471] (a) of Fig. 48 shows an example in which a transmitting device aggregates multiple data units and stores them in one payload. In the example of (a) of Fig. 48, a data unit header (DUH) and a data unit length (DUL) are added to the beginning of each of the multiple data units, and multiple data units with the data unit header and data unit length added are stored together in the payload.

[0472] Figure 48(b) shows an example in which one data unit is stored in one payload. In the example of Figure 48(b), a data unit header is added to the beginning of the data unit and stored in the payload. Figure 48(c) shows an example in which one data unit is divided, and the divided data units are added with data unit headers and stored in the payload.

[0473] There are various types of data units, such as timed-MFU, which is media including information related to synchronization of video, audio, or subtitles, non-timed-MFU, which is media including no information related to synchronization such as files, MPU metadata, and MF metadata, and a data unit header is defined depending on the type of data unit. Note that MPU metadata and MF metadata do not have a data unit header.

[0474] Furthermore, although a transmitting device cannot in principle aggregate different types of data units, it may be specified to be able to aggregate different types of data units. For example, when the size of MF metadata is small, such as when it is divided into movie fragments for each sample, aggregating the MF metadata and media data can reduce the number of packets and also the transmission capacity.

[0475] If the data unit is an MFU, some information about the MPU, such as information for configuring the MPU (MP4), is stored as a header.

[0476] For example, the header of a timed-MFU includes movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter, while the header of a non-timed-MFU includes item_iD. The meaning of each field is specified in standards such as ISO / IEC23008-1 or ARIB STD-B60. The meaning of each field specified in such standards will be explained below.

[0477] The movie_fragment_sequence_number indicates the sequence number of the movie fragment to which the MFU belongs, and is also specified in ISO / IEC14496-12.

[0478] The sample_number indicates the sample number to which the MFU belongs, and is also specified in ISO / IEC14496-12.

[0479] The offset indicates the offset amount of the MFU in the sample to which the MFU belongs, in bytes.

[0480] The priority indicates the relative importance of the MFU in the MPU to which the MFU belongs, and an MFU with a larger priority number is more important than an MFU with a smaller priority number.

[0481] The dependency_counter indicates the number of MFUs whose decoding process depends on the MFU (i.e., the number of MFUs whose decoding process cannot be performed unless the MFU is decoded). For example, when the MFU is HEVC and a B picture or a P picture refers to an I picture, the B picture or the P picture cannot be decoded unless the I picture is decoded.

[0482] Therefore, when the MFU is in sample units, the dependency_counter in the MFU of an I-picture indicates the number of pictures that reference the I-picture. When the MFU is in NAL unit units, the dependency_counter in the MFU belonging to the I-picture indicates the number of NAL units that belong to the picture that references the I-picture. Furthermore, in the case of a video signal that has been temporally hierarchically coded, the MFU of the enhancement layer depends on the MFU of the base layer, so the dependency_counter in the MFU of the base layer indicates the number of MFUs of the enhancement layer. This field can only be generated after the number of dependent MFUs has been determined.

[0483] The item_iD indicates an identifier that uniquely identifies the item.

[0484] [MP4 non-support mode] As explained in Figures 19 and 21, the transmitting device can transmit the MPU in MMT by transmitting MPU metadata or MF metadata before or after the media data, or by transmitting only the media data. In addition, the receiving device can decode using a receiving device or method that complies with MP4, or by decoding without using a header.

[0485] As a method of transmitting data specialized for broadcast stream playback, there is, for example, a transmission method that does not support MP4 reconstruction in a receiving device.

[0486] An example of a transmission method that does not support MP4 reconstruction in a receiving device is a method that does not transmit metadata (MPU metadata and MF metadata), as shown in (b) of Figure 21. In this case, the field value of the fragment type (information indicating the type of data unit) included in the MMTP packet is fixed to 2 (=MFU).

[0487] If metadata is not transmitted, as explained above, an MP4-compliant receiving device cannot decode the received data as MP4, but it can decode it without using the metadata (header).

[0488] Therefore, metadata is not necessarily essential information for decoding and playing back a broadcast stream. Similarly, the information in the data unit header in timed-MFU, as explained in Fig. 48, is information for reconstructing MP4 in a receiving device. Since there is no need to reconstruct MP4 for broadcast stream playback, the information in the data unit header in timed-MFU (hereinafter also referred to as timed-MFU header) is not necessarily information required for broadcast stream playback.

[0489] A receiving device can easily reconstruct an MP4 file by using the metadata and the information for reconstructing an MP4 file in the data unit header (hereinafter also referred to as MP4 configuration information). However, a receiving device cannot reconstruct an MP4 file even if only one of the metadata and the MP4 configuration information in the data unit header is transmitted. There is little benefit to transmitting only one of the metadata and the information for reconstructing an MP4 file, and generating and transmitting unnecessary information increases processing and reduces transmission efficiency.

[0490] Therefore, the transmitting device controls the data structure and transmission of the MP4 configuration information using the following method. The transmitting device determines whether to indicate the MP4 configuration information in the data unit header based on whether metadata is transmitted. Specifically, if metadata is transmitted, the transmitting device indicates the MP4 configuration information in the data unit header, and if metadata is not transmitted, the transmitting device does not indicate the MP4 configuration information in the data unit header.

[0491] As a method for not indicating MP4 configuration information in the data unit header, for example, the following method can be used.

[0492] 1. The transmitting device sets the MP4 configuration information as reserved and does not use it. This reduces the amount of processing on the sending side (the amount of processing on the transmitting device) that generates the MP4 configuration information.

[0493] 2. The transmitting device deletes the MP4 configuration information and compresses the header, which reduces the amount of processing on the sending side that generates the MP4 configuration information and also reduces transmission capacity.

[0494] When the transmitting device deletes the MP4 configuration information and compresses the header, the transmitting device may indicate a flag indicating that the MP4 configuration information has been deleted (compressed). The flag is indicated in the header (MMTP packet header, MMTP payload header, data unit header) or control information.

[0495] Furthermore, information as to whether metadata is transmitted may be determined in advance, or may be separately signaled in the header or control information and transmitted to the receiving device.

[0496] For example, the MFU header may store information indicating whether metadata corresponding to the MFU has been transmitted.

[0497] On the other hand, the receiving device can determine whether MP4 configuration information is indicated based on whether metadata is transmitted.

[0498] Here, if the data transmission order (for example, an order such as MPU metadata, MF metadata, and media data) is fixed, the receiving device may make a determination based on whether the metadata is received before the media data.

[0499] If MP4 configuration information is indicated, the receiving device can use the MP4 configuration information to reconstruct the MP4, or the receiving device can use the MP4 configuration information to detect the beginning of other access units or NAL units.

[0500] The MP4 configuration information may be the entire timed-MFU header or a part of it.

[0501] Similarly, the transmitting device determines whether metadata is transmitted in the non-timed-MFU header. You may decide whether to show the id.

[0502] The transmitting device may indicate MP4 configuration information in only one of timed-MFU and non-timed-MFU. If the transmitting device indicates MP4 configuration information in only one of timed-MFU and non-timed-MFU, the transmitting device determines whether to indicate MP4 configuration information based on whether metadata is transmitted and whether the MFU is timed or non-timed. The receiving device can determine whether to indicate MP4 configuration information based on whether metadata is transmitted and the timed / non-timed flag.

[0503] In the above description, the transmitting device determines whether to indicate the MP4 configuration information based on whether the metadata (both the MPU metadata and the MF metadata) is transmitted. However, the transmitting device may not indicate the MP4 configuration information if some of the metadata (either the MPU metadata or the MF metadata) is not transmitted.

[0504] The transmitting device may also determine whether to indicate MP4 configuration information based on information other than metadata.

[0505] For example, modes such as an MP4 support mode / an MP4 non-support mode may be defined, and the transmitting device may indicate MP4 configuration information in a data unit header in the MP4 support mode, and not indicate MP4 configuration information in the data unit header in the MP4 non-support mode. Also, the transmitting device may transmit metadata and indicate MP4 configuration information in a data unit header in the MP4 support mode, and not transmit metadata and not indicate MP4 configuration information in the data unit header in the MP4 non-support mode.

[0506] [Transmitter operation flow] Next, the operation flow of the transmitting device will be described with reference to Figure 49.

[0507] The transmitting device first determines whether to transmit metadata (S1001). If the transmitting device determines to transmit metadata (Yes in S1002), the transmitting device proceeds to step S1003, generates MP4 configuration information, stores it in a header, and transmits it (S1003). In this case, the transmitting device also generates and transmits metadata.

[0508] On the other hand, if the transmitting device determines not to transmit metadata (No in S1002), it transmits the MP4 configuration information without generating it and storing it in the header (S1004). In this case, the transmitting device does not generate or transmit metadata.

[0509] Note that whether or not to transmit metadata in step S1001 may be determined in advance, or may be determined based on whether metadata has been generated within the transmitting device or whether metadata is being transmitted within the transmitting device.

[0510] [Operation flow of receiving device] Next, the operation flow of the receiving device will be explained, as shown in Figure 50.

[0511] The receiving device first determines whether metadata is being transmitted (S1101). Whether metadata is being transmitted can be determined by monitoring the fragment type in the MMTP packet payload. Alternatively, whether metadata is being transmitted may be determined in advance.

[0512] If the receiving device determines that metadata has been transmitted (Yes in S1102), it reconstructs the MP4 and executes a decoding process using the MP4 configuration information (S1103). On the other hand, if the receiving device determines that metadata has not been transmitted (No in S1102), it does not reconstruct the MP4 and executes a decoding process without using the MP4 configuration information (S1104).

[0513] In addition, using the methods described above, the receiving device can detect random access points, the beginning of access units, the beginning of NAL units, etc. without using MP4 configuration information, and can perform decoding processes, packet loss detection, and recovery from packet loss.

[0514] For example, the beginning of an access unit is the beginning data of an MMT payload in which the aggregation_flag value is 1. In this case, the fragmentation_indicator value is 0.

[0515] In addition, the beginning of a slice segment is the beginning data of an MMT payload in which the aggregation_flag value is 0 and the fragmentation_indicator value is 00 or 01.

[0516] Based on the above information, the receiving device can detect the beginning of an access unit and a slice segment.

[0517] In addition, the receiving device may analyze the NAL unit header in a packet containing the beginning of a data unit whose fragmentation_indicator value is 00 or 01, and detect that the type of NAL unit is an AU delimiter and that the type of NAL unit is a slice segment.

[0518] [Broadcast Simple Mode] So far, we have described a method for transmitting data specialized for broadcast stream playback that does not support MP4 configuration information in the receiving device, but the method for transmitting data specialized for broadcast stream playback is not limited to this.

[0519] As a method for transmitting data specialized for broadcast stream playback, for example, the following method may be used.

[0520] A transmitting device does not need to use AL-FEC in a fixed broadcast reception environment. If AL-FEC is not used, the FEC_type in the MMTP packet header is always fixed to 0.

[0521] The transmitting device may always use AL-FEC in a mobile broadcast reception environment and in the communication UDP transmission mode. When AL-FEC is used, the FEC_type in the MMTP packet header is always 0 or 1.

[0522] The transmitting device may not transmit assets in bulk. If assets are not transmitted in bulk, location_infolocation, which indicates the number of asset transmission locations within the MPT, may be fixed to 1.

[0523] · The transmitting device does not have to perform hybrid transmission of assets, programs, and messages.

[0524] Also, for example, if a broadcast simple mode is defined, the transmitting device may set the MP4 non-support mode when in the broadcast simple mode, or may use the data transmission method specialized for broadcast stream playback described above. Whether the mode is the broadcast simple mode may be determined in advance, or the transmitting device may store a flag indicating the broadcast simple mode as control information and transmit it to the receiving device.

[0525] In addition, the transmitting device may, based on whether the MP4 non-support mode is in effect (whether metadata is transmitted) as described in Figure 49, use the data transmission method specialized for broadcast stream playback shown above as broadcast simple mode if the MP4 non-support mode is in effect.

[0526] When the receiving device is in broadcast simple mode, it is considered to be in MP4 non-support mode and can perform decoding processing without reconstructing the MP4.

[0527] Furthermore, when the receiving device is in the broadcast simple mode, it determines that the function is specialized for broadcasting, and can perform reception processing specialized for broadcasting.

[0528] As a result, when in broadcast simple mode, by using only functions specialized for broadcasting, not only can unnecessary processing be reduced for the transmitting device and receiving device, but transmission overhead can also be reduced by not compressing and transmitting unnecessary information.

[0529] When the MP4 non-support mode is used, hint information supporting a storage method other than the MP4 format may be indicated.

[0530] Storage methods other than MP4 configuration include, for example, directly storing MMT packets or IP packets, or converting MMT packets into MPEG-2 TS packets.

[0531] In the case of an MP4 non-support mode, a format that does not conform to the MP4 structure may be used.

[0532] For example, in the case of an MP4 non-support mode, the data stored in the MFU may be in a format with a byte start code at the beginning of the NAL unit, rather than in the MP4 format where the size of the NAL unit is added to the beginning of the NAL unit.

[0533] In MMT, the asset type indicating the type of asset is described in 4CC registered in MP4REG (http: / / www.mp4ra.org), and when HEVC is used as the video signal, 'HEV1' or 'HVC1' is used. 'HVC1' is a format that may include parameter sets in samples, while 'HEV1' is a format that does not include parameter sets in samples but includes parameter sets in the sample entries in the MPU metadata.

[0534] In the case of broadcast simple mode or MP4 non-support mode, if MPU metadata and MF metadata are not transmitted, it may be specified that a parameter set must be included in the sample. Also, whether 'HEV1' or 'HVC1' is indicated in the asset type, it may be specified that the format must be 'HVC1'.

[0535] [Supplement 1: Transmitter] As described above, when metadata is not transmitted, the MP4 configuration information is set to reserved, and a transmitting device that is not in operation can also be configured as shown in Fig. 51. Fig. 51 is a diagram showing an example of a specific configuration of a transmitting device.

[0536] The transmitting device 300 includes an encoding unit 301, an assigning unit 302, and a transmitting unit 303. Each of the encoding unit 301, the assigning unit 302, and the transmitting unit 303 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0537] The encoding unit 301 encodes a video signal or an audio signal to generate sample data, which is specifically a data unit.

[0538] The adding unit 302 adds header information including MP4 configuration information to sample data, which is data obtained by encoding a video signal or an audio signal. The MP4 configuration information is information for reconstructing the sample data as an MP4 format file on the receiving side, and the content of the information varies depending on whether the presentation time of the sample data is specified.

[0539] As described above, the attachment unit 302 includes MP4 configuration information such as movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter in the header (header information) of timed-MFU, which is an example of sample data with a defined presentation time (sample data including information about synchronization).

[0540] On the other hand, the attachment unit 302 includes MP4 configuration information such as item_id in the header (header information) of timed-MFU, which is an example of sample data for which the presentation time is not specified (sample data that does not include information about synchronization).

[0541] Then, when the transmitting unit 303 does not transmit metadata corresponding to the sample data (for example, in the case of (b) in Figure 21), the attaching unit 302 attaches header information that does not include MP4 configuration information to the sample data, depending on whether the presentation time of the sample data is specified or not.

[0542] Specifically, when the presentation time of the sample data is determined, the attachment unit 302 attaches header information that does not include the first MP4 configuration information to the sample data, and when the presentation time of the sample data is not determined, the attachment unit 302 attaches header information that includes the second MP4 configuration information to the sample data.

[0543] 49, when the transmitting unit 303 does not transmit metadata corresponding to the sample data, the adding unit 302 sets the MP4 configuration information to reserved (a fixed value), thereby not substantially generating the MP4 configuration information and not substantially storing the MP4 configuration information in the header (header information). Note that the metadata includes MPU metadata and movie fragment metadata.

[0544] The transmitting unit 303 transmits the sample data to which the header information has been added. More specifically, the transmitting unit 303 packetizes the sample data to which the header information has been added in accordance with the MMT method and transmits the packetized data.

[0545] As described above, in the transmission method and reception method specialized for playing back a broadcast stream, the receiving device does not need to reconstruct data units into MP4. If the receiving device does not need to reconstruct data into MP4, the processing load of the transmitting device is reduced by not generating unnecessary information such as MP4 configuration information.

[0546] On the other hand, the transmitting device must transmit necessary information, but must maintain compatibility with the standard so that it does not have to transmit unnecessary additional information separately.

[0547] With a configuration such as that of the transmitting device 300, by setting the area in which the MP4 configuration information is stored to a fixed value, the MP4 configuration information is not transmitted, and only necessary information is transmitted based on the standard, thereby eliminating the need to transmit unnecessary additional information. In other words, the configuration of the transmitting device and the amount of processing performed by the transmitting device can be reduced. Furthermore, since unnecessary data is not transmitted, transmission efficiency can be improved.

[0548] [Supplement 2: Receiving device] Furthermore, a receiving device corresponding to transmitting device 300 may be configured, for example, as shown in Fig. 52. Fig. 52 is a diagram showing another example of the configuration of a receiving device.

[0549] The receiving device 400 includes a receiving unit 401 and a decoding unit 402. The receiving unit 401 and the decoding unit 402 are realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0550] The receiving unit 401 receives sample data, which is data in which a video signal or an audio signal has been encoded, and which is provided with header information including MP4 configuration information for reconstructing the sample data as a file in MP4 format.

[0551] If the receiving unit does not receive metadata corresponding to the sample data and the presentation time of the sample data is determined, the decoding unit 402 decodes the sample data without using the MP4 configuration information.

[0552] For example, as shown in step S1104 in FIG. 50, if the receiving unit 401 does not receive metadata corresponding to the sample data, the decoding unit 402 executes the decoding process without using the MP4 configuration information.

[0553] This allows the configuration of the receiving device 400 and the amount of processing in the receiving device 400 to be reduced.

[0554] (Fourth embodiment) [overview] In the fourth embodiment, a method for storing asynchronous (non-timed) media that does not contain information related to synchronization, such as files, in an MPU and a method for transmitting it in an MMTP packet will be described. Note that in the fourth embodiment, an MPU in MMT will be used as an example, but the method can also be applied to DASH, which is also MP4-based.

[0555] First, the details of how non-timed media (hereinafter also referred to as "asynchronous media data") is stored in the MPU will be explained using Figure 53. Figure 53 is a diagram showing how non-timed media is stored in the MPU and how it is transmitted in MMTP packets.

[0556] An MPU that stores non-timed media consists of boxes such as ftyp, mmpu, moov, and meta, and stores information about the files stored in the MPU. Multiple idat boxes can be stored in a meta box, and an idat box stores one file as an item.

[0557] A part of the ftyp, mmpu, moov, and meta boxes constitute one data unit as MPU metadata, and the item or idat box constitutes a data unit as MFU.

[0558] After the data units are aggregated or fragmented, they are given a data unit header, an MMTP payload header, and an MMTP packet header, and then transmitted as an MMTP packet.

[0559] Note that Figure 53 shows an example in which File #1 and File #2 are stored in one MPU. The MPU metadata is not divided, and the MFU is divided and stored in an MMTP packet, but this is not limited to this, and data units may be aggregated or fragmented depending on the size of the data unit. Also, the MPU metadata does not have to be transmitted, in which case only the MFU is transmitted.

[0560] Header information such as the data unit header indicates an itemID (an identifier that uniquely identifies an item), and the MMTP payload header and MMTP packet header include a packet sequence number (a sequence number for each packet) and an MPU sequence number (a sequence number for the MPU, a number unique within an asset).

[0561] Note that the data structures of the MMTP payload header and the MMTP packet header other than the data unit header are the same as those of the timed media (hereinafter also referred to as "synchronized media data") described so far, and include an aggregation_flag, a fragmentation_indicator, a fragment_counter, etc.

[0562] Next, a specific example of the header information when a file (= Item = MFU) is split and packetized will be described using FIGS. 54 and 55.

[0563] FIGS. 54 and 55 are diagrams showing examples of packetizing and transmitting for each of a plurality of split data obtained by splitting a file. FIGS. 54 and 55 specifically show information (packet sequence number, fragment counter, fragmentation indicator, MPU sequence number, item ID) included in any of the data unit header, the MMTP payload header, and the MMTP packet header, which is the header information for each split MMTP packet. Note that FIG. 54 shows an example in which File #1 is split into M (M <= 256) parts, and FIG. 55 shows an example in which File #2 is split into N (256 <N) parts.

[0564] The split data number indicates the index of the split data from the beginning of the file, and this information is not transmitted. That is, the split data number is not included in the header information. Also, the split data number is a number assigned to each packet corresponding to each of the plurality of split data obtained by splitting the file, and is a number incremented by 1 in ascending order from the first packet.

[0565] The packet sequence number is the sequence number of packets having the same packet ID. In FIGS. 54 and 55, assuming that the split data at the beginning of the file is A, consecutive numbers are assigned up to the split data at the end of the file. The packet sequence number is a number incremented by 1 in ascending order from the split data at the beginning of the file, and is a number corresponding to the split data number.

[0566] The fragment counter indicates the number of split data pieces that come after the split data piece in question among the multiple split data pieces obtained by splitting a single file. Furthermore, if the number of split data pieces, which is the number of split data pieces obtained by splitting a single file, exceeds 256, the fragment counter indicates the remainder when the number of split data pieces is divided by 256. In the example of Figure 54, the number of split data pieces is 256 or less, so the field value of the fragment counter is (M-split data number). On the other hand, in the example of Figure 55, the number of split data pieces exceeds 256, so the value obtained by dividing (N-split data number) by 256 is ((N-split data number)%256).

[0567] The fragmentation indicator indicates the fragmentation state of the data stored in the MMTP packet, and is a value indicating whether the fragment is the first fragment of a fragmented data unit, the last fragment, any other fragment, or one or more unfragmented data units. Specifically, the fragmentation indicator is "01" for the first fragment, "11" for the last fragment, "10" for the remaining fragments, and "00" for unfragmented data units.

[0568] In this embodiment, when the number of divided data pieces exceeds 256, it is explained as indicating the remainder when the number of divided data pieces is divided by 256, but the number of divided data pieces is not limited to 256 and may be another number (predetermined number).

[0569] 54 and 55, when a file is divided and conventional header information is added to each of the multiple data segments obtained by dividing the file and transmitted, the receiving device does not have information that can be used to determine the position of the data segment in the original file (data segment number) that the data stored in the received MMTP packet is, the number of data segments in the file, or the data segment numbers and number of data segments. For this reason, with conventional transmission methods, even if an MMTP packet is received, it is not possible to uniquely detect the data segment number or number of data segments for the data stored in the received MMTP packet.

[0570] For example, if the number of divided data pieces is 256 or less, as shown in Figure 54, and it is known in advance that the number of divided data pieces is 256 or less, it is possible to identify the divided data piece numbers and the number of divided data pieces by referencing the fragment counter. However, if the number of divided data pieces is 256 or more, it is not possible to identify the divided data piece numbers and the number of divided data pieces.

[0571] Furthermore, if the number of divided data segments of a file is limited to 256 or less, and the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed, and in this case the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes]. Therefore, if you want to transmit a file that is larger than 1 [Mbytes], you cannot limit the number of divided data segments of the file to 256 or less.

[0572] Furthermore, for example, since the first and last fragments of a file can be detected by referencing the fragmentation indicator, it is possible to count the number of MMTP packets until the MMTP packet containing the last fragment of the file is received, or to calculate the fragment number and number of fragments by combining it with the packet sequence number after receiving the MMTP packet containing the last fragment of the file. Therefore, the fragment number and number of fragments may be signaled by combining the fragmentation indicator and the packet sequence number. However, if reception starts from an MMTP packet containing fragment data in the middle of a file (i.e., fragment data that is neither the first nor the last fragment of the file), the fragment number and number of fragments of that fragment cannot be identified. The fragment number and number of fragments of that fragment can only be identified after receiving the MMTP packet containing the last fragment of the file.

[0573] To address the issue described in Figures 54 and 55, that is, to uniquely determine the data segment number and number of data segments of a file when a packet containing the file's data segment is received midway, the following method is used.

[0574] First, the divided data number will be explained.

[0575] For the divided data number, the packet sequence number in the divided data at the beginning of the file (item) is signaled.

[0576] As a signaling method, it is stored in the control information that manages the file. Specifically, in Figures 54 and 55, the packet sequence number A of the divided data at the beginning of the file is stored in the control information. The receiving device obtains the value of A from the control information and calculates the divided data number from the packet sequence number indicated in the packet header.

[0577] The divided data number of the divided data is obtained by subtracting the packet sequence number A of the first divided data from the packet sequence number of the divided data.

[0578] An example of control information for managing files is the asset management table specified in ARIB STD-B60. The asset management table indicates the file size, version information, etc. for each file, and is stored in a data transmission message for transmission. Figure 56 shows the syntax of a loop for each file in the asset management table.

[0579] If the area of ​​the existing asset management table cannot be expanded, signaling may be performed using a 32-bit area in part of the item_info_byte field that indicates item information. A flag indicating whether the packet sequence number in the first divided data of the file (item) is indicated in part of the item_info_byte area may be indicated in, for example, a reserved_future_use field of the control information.

[0580] When a file is repeatedly transmitted, such as in a data carousel, multiple packet sequence numbers may be indicated, or the packet sequence number of the beginning of the file to be transmitted immediately after may be indicated.

[0581] The packet sequence number is not limited to the packet sequence number of the divided data at the beginning of the file, and may be any information that links the divided data number of the file with the packet sequence number.

[0582] Next, the number of divided data items will be described.

[0583] The order of loops for each file included in the asset management table may be defined as the transmission order of the files. This allows the first packet sequence numbers of two consecutive files in transmission order to be known, and the number of divided data pieces of the previously transmitted file can be determined by subtracting the first packet sequence number of the previously transmitted file from the first packet sequence number of the later transmitted file. That is, for example, if File #1 shown in Figure 54 and File #2 shown in Figure 55 are consecutive files in this order, the last packet sequence number of File #1 and the first packet sequence number of File #2 are assigned consecutive numbers.

[0584] The number of divided data pieces of a file may also be specified by specifying a file division method. For example, if the number of divided data pieces is N, the size of each of the 1st to (N-1)th divided data pieces is set to L, and the size of the Nth divided data piece is specified as a fraction (item_size-L*(N-1)), so that the number of divided data pieces can be calculated backward from the item_size shown in the asset management table. In this case, the integer value obtained by rounding up (item_size / L) becomes the number of divided data pieces. However, the file division method is not limited to this.

[0585] The number of divided data items may also be stored directly in the asset management table.

[0586] By using the above method, the receiving device receives the control information and calculates the number of divided data pieces based on the control information. It can also calculate a packet sequence number corresponding to the divided data piece number of the file based on the control information. If the timing of receiving the divided data packets is earlier than the timing of receiving the control information, the divided data piece number and the number of divided data pieces may be calculated at the timing of receiving the control information.

[0587] When the fragment data number or the number of fragment data is signaled using the above method, the fragment data number or the number of fragment data is not identified based on the fragment counter, and the fragment counter becomes unnecessary data. Therefore, in the transmission of asynchronous media, when information that can identify the fragment data number and the number of fragment data is signaled using the above method, the fragment counter may not be used, or header compression may be performed. This reduces the processing load of the transmitting device and receiving device, and also improves transmission efficiency. In other words, when transmitting asynchronous media, the fragment counter may be reserved (disabled). Specifically, the value of the fragment counter may be set to a fixed value, for example, "0." Furthermore, when receiving asynchronous media, the fragment counter may be ignored.

[0588] When storing synchronous media such as video and audio, the order in which MMTP packets are sent at the sending device matches the order in which they arrive at the receiving device, and packets are not retransmitted. In such a case, if there is no need to detect packet loss and reconstruct packets, the fragment counter may not be used. In other words, in this case, the fragment counter may be reserved (disabled).

[0589] In addition, it is possible to detect random access points, the beginning of access units, the beginning of NAL units, etc. without using a fragment counter, and it is possible to perform decoding processes, detect packet loss, and recover from packet loss.

[0590] Furthermore, transmission of real-time content such as live broadcasts requires even lower latency, and requires that data be packetized and transmitted sequentially starting from the data that has been completely encoded. However, in the transmission of real-time content, conventional fragment counters cannot determine the number of divided data pieces when transmitting the first divided data piece, so the first divided data piece is transmitted after all the encoding of the data unit has been completed and the number of divided data pieces has been determined, resulting in a delay. Even in such cases, this delay can be reduced by using the above method and not operating a fragment counter.

[0591] FIG. 57 shows the operational flow for identifying the divided data number in the receiving device.

[0592] The receiving device acquires control information that describes file information (S1201). The receiving device determines whether the control information indicates the packet sequence number of the beginning of the file (S1202). If the control information indicates the packet sequence number of the beginning of the file (Yes in S1202), the receiving device calculates a packet sequence number that corresponds to the divided data number of the divided data of the file (S1203). Then, after acquiring an MMTP packet that stores the divided data, the receiving device identifies the divided data number of the file from the packet sequence number stored in the packet header of the acquired MMTP packet (S1204). On the other hand, if the control information does not indicate the packet sequence number of the beginning of the file (No in S1202), the receiving device acquires an MMTP packet that includes the last divided data of the file, and then identifies the divided data number using the fragment indicator and packet sequence number stored in the packet header of the acquired MMTP packet (S1205).

[0593] FIG. 58 shows the operational flow for identifying the number of divided data pieces in the receiving device.

[0594] The receiving device acquires control information that describes file information (S1301). The receiving device determines whether the control information includes information that allows the number of divided data pieces of the file to be calculated (S1302). If it determines that the control information includes information that allows the number of divided data pieces to be calculated (Yes in S1302), it calculates the number of divided data pieces based on the information included in the control information (S1303). On the other hand, if it determines that the number of divided data pieces cannot be calculated (No in S1302), the receiving device acquires an MMTP packet that includes the last divided data piece of the file, and then identifies the number of divided data pieces using the fragment indicator and packet sequence number stored in the packet header of the acquired MMTP packet (S1304).

[0595] FIG. 59 shows an operational flow for determining whether to operate a fragment counter in a transmitting device.

[0596] First, the transmitting device determines whether the media to be transmitted (hereinafter also referred to as "media data") is synchronous media or asynchronous media (S1401).

[0597] If the result of the determination in step S1401 is synchronous media (Synchronous Media in S1402), the transmitting device determines whether the order of MMTP packets sent and received matches in the environment in which the synchronous media is transmitted, and whether packet reassembly is unnecessary in the event of packet loss (S1403). If the transmitting device determines that it is unnecessary (Yes in S1403), it does not operate a fragment counter (S1404). On the other hand, if the transmitting device determines that it is not unnecessary (No in S1403), it operates a fragment counter (S1405).

[0598] If the result of the determination in step S1401 is asynchronous media (Asynchronous Media in S1402), the transmitting device determines whether or not to operate a fragment counter based on whether the fragment data number and the number of fragment data are signaled using the method described above. Specifically, if the fragment data number and the number of fragment data are signaled (Yes in S1406), the transmitting device does not operate a fragment counter (S1404). On the other hand, if the fragment data number and the number of fragment data are not signaled (No in S1406), the transmitting device operates a fragment counter (S1405).

[0599] In addition, if the transmitting device does not operate a fragment counter, it may set the value of the fragment counter to reserved or may perform header compression.

[0600] In addition, the transmitting device may determine whether to signal the above-mentioned fragment data number and number of fragment data based on whether or not a fragment counter is operated.

[0601] If the synchronous media does not operate a fragment counter, the transmitting device may signal the divided data number and the number of divided data using the method described above for the asynchronous media. Conversely, the operation of the synchronous media may be determined based on whether the asynchronous media operates a fragment counter. In this case, the synchronous media and the asynchronous media can be operated in the same manner regarding whether or not to operate fragments.

[0602] Next, a method for identifying the number of divided data pieces and the divided data numbers (when a fragment counter is used) will be described. Figure 60 is a diagram for explaining a method for identifying the number of divided data pieces and the divided data numbers (when a fragment counter is used).

[0603] As explained using Figure 54, if the number of divided data pieces is 256 or less and it is known in advance that the number of divided data pieces is 256 or less, it is possible to identify the divided data piece number and the number of divided data pieces by referring to the fragment counter.

[0604] If the number of divided data segments in a file is limited to 256 or less, and the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed, and in this case the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes].

[0605] If the file size exceeds the maximum size of a transmittable file, the file is split in advance so that each split file is no larger than x*256 bytes. Each of the multiple split files obtained by splitting the file is treated as a single file (item) and is further split into no more than 256 files. Each of the split data obtained by further splitting is stored in an MMTP packet and transmitted.

[0606] Note that information indicating that the item is a split file, the number of split files, and the sequence numbers of the split files may be stored in the control information and transmitted to the receiving device. This information may also be stored in the asset management table, or may be indicated using part of the existing field item_info_byte.

[0607] When an item is one of multiple split files obtained by splitting a single file, the receiving device can identify the other split files and reconstruct the original file. Furthermore, the receiving device can uniquely identify the number of split data pieces and the split data numbers by using the number of split files, the split file index, and the fragment counter in the control information. Furthermore, the number of split data pieces and the split data numbers can be uniquely identified without using packet sequence numbers, etc.

[0608] Here, it is desirable that the item_id of each of the split files obtained by splitting one file is the same. If a different item_id is assigned, the item_id of the first split file may be indicated in order to uniquely refer to the file from other control information, etc.

[0609] Alternatively, multiple split files may always belong to the same MPU. When multiple files are stored in an MPU, files of different types may not be stored, and files that are split from a single file may always be stored. A receiving device can detect file updates by checking the version information for each MPU, without checking the version information for each item.

[0610] Figure 61 shows the operational flow of a transmitting device when utilizing a fragment counter.

[0611] First, the transmitting device checks the size of the file to be transmitted (S1501). Next, the transmitting device determines whether the file size exceeds x*256 [bytes] (x is the data size that can be transmitted in one packet, for example, the MTU size) (S1502). If the file size exceeds x*256 [bytes] (Yes in S1502), the transmitting device divides the file so that the size of each divided file is less than x*256 [bytes] (S1503). Then, the transmitting device transmits the divided files as items, and transmits information about the divided files (for example, the fact that they are divided files, the sequence numbers in the divided files, etc.) in control information (S1504). On the other hand, if the file size is less than x*256 [bytes] (No in S1502), the transmitting device transmits the file as an item as usual (S1505).

[0612] FIG. 62 shows the operational flow of a receiving device when utilizing a fragment counter.

[0613] First, the receiving device acquires and analyzes control information related to file transmission, such as an asset management table (S1601). Next, the receiving device determines whether the desired item is a split file (S1602). If the receiving device determines that the desired file is a split file (Yes in S1602), it acquires information for reconstructing the file, such as the split file and its index, from the control information (S1603). Then, the receiving device acquires the items that make up the split file and reconstructs the original file (S1604). On the other hand, if the receiving device determines that the desired file is not a split file (No in S1602), it acquires the file as usual (S1605).

[0614] In short, the transmitting device signals the packet sequence number of the divided data at the beginning of the file. The transmitting device also signals information that can identify the number of divided data. Alternatively, the transmitting device defines a fragmentation rule that can identify the number of divided data. The transmitting device also performs reserved or header compression without using a fragment counter.

[0615] When the packet sequence number of the data at the beginning of the file is signaled, the receiving device determines the divided data number and the number of divided data from the packet sequence number of the divided data at the beginning of the file and the packet sequence number of the MMTP packet.

[0616] From another perspective, the transmitting device divides a file, divides data into individual divided files, and transmits the divided files by signaling information linking the divided files (such as a sequence number and the number of divisions).

[0617] The receiving device identifies the divided data number and the number of divided data pieces based on the fragment counter and the sequence number of the divided file.

[0618] This allows the divided data number and divided data to be uniquely identified. In addition, since the divided data number of a divided data item can be identified when the divided data item is received, waiting time and memory usage can be reduced.

[0619] Furthermore, by not using a fragment counter, the configuration of the transmitting / receiving device can reduce the amount of processing and improve transmission efficiency.

[0620] Figure 63 shows a service configuration in which the same program is transmitted over multiple IP data flows. This example shows a case in which part of the data (video and audio) of a program with service ID = 2 is transmitted over an IP data flow using the MMT method, and data with the same service ID but different from the part of the data is transmitted over an IP data flow using the advanced BS data transmission method (in this example, the file transmission protocols are different, but may be the same protocol).

[0621] The transmitting device multiplexes the IP data so as to ensure that data consisting of multiple IP data flows is ready by the time of decoding at the receiving device.

[0622] The receiving device can realize guaranteed receiver operation by processing data consisting of multiple IP data flows based on the decoding time.

[0623] [Supplementary information: Transmitting and receiving devices] As described above, a transmitting device that transmits data without operating a fragment counter can also be configured as shown in Figure 64. Also, a receiving device that receives data without operating a fragment counter can also be configured as shown in Figure 65. Figure 64 is a diagram showing an example of a specific configuration of a transmitting device. Figure 65 is a diagram showing an example of a specific configuration of a receiving device.

[0624] The transmitting device 500 includes a dividing unit 501, a composing unit 502, and a transmitting unit 503. Each of the dividing unit 501, the composing unit 502, and the transmitting unit 503 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0625] The receiving device 600 includes a receiving unit 601, a determining unit 602, and a configuration unit 603. The receiving unit 601, the determining unit 602, and the configuration unit 603 are each realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0626] Detailed explanations of each component of the transmitting device 500 and the receiving device 600 will be given in the explanations of the transmitting method and the receiving method, respectively.

[0627] First, the transmission method will be described with reference to Fig. 66. Fig. 66 shows the operation flow (transmission method) performed by the transmission device.

[0628] First, the division unit 501 of the transmission device 500 divides data into a plurality of divided data (S1701).

[0629] Next, the configuration unit 502 of the transmitting device 500 configures a plurality of packets by adding header information to each of the plurality of divided data and packetizing the data (S1702).

[0630] Then, the transmitting unit 503 of the transmitting device 500 transmits the configured plurality of packets (S1703). The transmitting unit 503 transmits the divided data information and the value of the invalidated fragment counter. The divided data information is information for identifying the divided data number and the number of divided data. The divided data number is a number indicating the ordinal number of the divided data among the plurality of divided data. The number of divided data is the number of the plurality of divided data.

[0631] This allows the amount of processing by the transmission device 500 to be reduced.

[0632] Next, the receiving method will be described with reference to Fig. 67. Fig. 67 shows the operation flow (receiving method) of the receiving device.

[0633] First, the receiving unit 601 of the receiving device 600 receives a plurality of packets (S1801).

[0634] Next, the determination unit 602 of the receiving device 600 determines whether or not divided data information has been acquired from the received plurality of packets (S1802).

[0635] Then, if the determination unit 602 determines that the divided data information has been acquired (Yes in S1802), the construction unit 603 of the receiving device 600 constructs data from the multiple received packets without using the value of the fragment counter included in the header information (S1803).

[0636] On the other hand, if the judgment unit 602 determines that the divided data information has not been acquired (No in S1802), the construction unit 603 may construct data from the multiple received packets using the value of the fragment counter included in the header information (S1804).

[0637] This allows the amount of processing by the receiving device 600 to be reduced.

[0638] (Embodiment 5) [overview] In the fifth embodiment, a method for transmitting transport packets (TLV packets) when NAL units are stored in the multiplexing layer in the NAL size format will be described.

[0639] As described in the first embodiment, there are two types of formats for storing H.264 or H.265 NAL units in the multiplexing layer. One is a format called byte stream format, in which a start code consisting of a specific bit string is added immediately before the NAL unit header. The other is a format called NAL size format, in which a field indicating the size of the NAL unit is added. The byte stream format is used in MPEG-2 systems and RTP, and the NAL size format is used in MP4, or DASH and MMT that use MP4.

[0640] In the byte stream format, the start code consists of three bytes, and an optional byte (a byte whose value is 0) can be added.

[0641] On the other hand, in the general NAL size format in MP4, size information is indicated as one byte, two bytes, or four bytes. This size information is indicated by the lengthSizeMinusOne field in the HEVC sample entry. A value of "0" in this field indicates one byte, "1" indicates two bytes, and "3" indicates four bytes.

[0642] In ARIB STD-B60 "Media Transport Method Using MMT in Digital Broadcasting," standardized in July 2014, when storing NAL units in the multiplexing layer, if the output of the HEVC encoder is a byte stream, the byte start code is removed and the size of the NAL unit in bytes, expressed as a 32-bit (unsigned integer), is added immediately before the NAL unit as length information. Note that MPU metadata including HEVC sample entries is not transmitted, and the size information is fixed at 32 bits (4 bytes).

[0643] In addition, ARIB STD-B60 "Media Transport Method Using MMT in Digital Broadcasting" specifies that in the reception buffer model that a transmitting device takes into account during transmission to ensure buffer operation in a receiving device, the pre-decoding buffer for video signals is CPB.

[0644] However, there are the following issues: CPB in the MPEG-2 system and HRD in HEVC are specified on the assumption that the video signal is in byte stream format. Therefore, for example, if transmission packet rate control is performed on the assumption that the video signal is in byte stream format with a 3-byte start code, a receiving device that receives transmission packets in the NAL size format with a 4-byte size field may not be able to satisfy the receiving buffer model in ARIB STD-B60. Furthermore, the receiving buffer model in ARIB STD-B60 does not specify a specific buffer size or extraction rate, making it difficult to guarantee buffer operation in the receiving device.

[0645] Therefore, in order to solve the above problem, a receiving buffer model for guaranteeing buffer operation in the receiver is defined as follows.

[0646] FIG. 68 shows a reception buffer model based on the reception buffer model defined in ARIB STD B-60, particularly when only a broadcast transmission channel is used.

[0647] The receiving buffer model includes a TLV packet buffer (first buffer), an IP packet buffer (second buffer), an MMTP buffer (third buffer), and a pre-decoding buffer (fourth buffer). Note that a de-jitter buffer and a buffer for FEC are not required for broadcast transmission paths, so they are omitted.

[0648] The TLV packet buffer receives TLV packets (transmission packets) from the broadcast transmission path, converts the IP packets, which consist of the variable-length packet headers (IP packet headers, full headers when the IP packets are compressed, and compressed headers when the IP packets are compressed) stored in the received TLV packets and the variable-length payloads, into IP packets (first packets) with header-expanded fixed-length IP packet headers, and outputs the IP packets obtained by the conversion at a constant bit rate.

[0649] The IP packet buffer converts IP packets into MMTP packets (second packets) with a packet header and a variable-length payload, and outputs the MMTP packets obtained by the conversion at a constant bit rate. Note that the IP packet buffer may be merged with the MMTP buffer.

[0650] The MMTP buffer converts the output MMTP packets into NAL units, and outputs the NAL units obtained by the conversion at a constant bit rate.

[0651] The pre-decoding buffer sequentially stores the output NAL units, generates access units from the stored NAL units, and outputs the generated access units to the decoder at the decoding time corresponding to the access units.

[0652] The receive buffer model shown in Figure 68 is characterized in that the MMTP buffer and pre-decoding buffer, which are buffers other than the upstream TLV packet buffer and IP packet buffer, follow the receive buffer model in MPEG-2 TS.

[0653] For example, the MMTP buffer for video (MMTP B1) is composed of buffers equivalent to the transport buffer (TB) and multiplexing buffer (MB) in MPEG-2 TS, and the MMTP buffer for audio (MMTP Bn) is composed of buffers equivalent to the transport buffer (TB) in MPEG-2 TS.

[0654] The buffer size of the transport buffer is the same as that of MPEG-2 TS and is a fixed value, for example, n times the MTU size (n can be a decimal or an integer, and is 1 or greater).

[0655] The MMTP packet size is also specified so that the overhead rate of the MMTP packet header is smaller than the overhead rate of the PES packet header, which allows the transport buffer extraction rates RX1, RXn, and RXs in MPEG-2 TS to be applied as is to the transport buffer extraction rates.

[0656] The size of the multiplexing buffer and the extraction rate are the same as those of MPEG-2 The MB size in TS and RBX1.

[0657] In addition to the above receive buffer model, the following constraints are set to solve the problem.

[0658] The HEVC HRD specification assumes a byte stream format, and MMT uses a NAL size format that adds a 4-byte size field to the beginning of the NAL unit. Therefore, during encoding, rate control is performed in the NAL size format to satisfy the HRD.

[0659] That is, the transmitting device controls the rate of transmission packets based on the above-mentioned receiving buffer model and constraints.

[0660] In the receiving device, by performing receiving processing using the above signals, decoding operations can be performed without underflow or overflow.

[0661] The extraction rate of the TLV packet buffer (the bit rate at which the TLV packet buffer outputs IP packets) is set taking into consideration the transmission rate after the IP header is expanded.

[0662] That is, the transmission rate of the output IP packet is taken into account after inputting a TLV packet with a variable data size, removing the TLV header, and expanding (restoring) the IP header. In other words, the amount of header increase or decrease is taken into account relative to the input transmission rate.

[0663] Specifically, the transmission rate of output IP packets is not unique because the data size is variable, packets with compressed IP headers are mixed with packets without compressed IP headers, and the IP header size differs depending on the packet type, such as IPv4 or IPv6. For this reason, the average packet length of variable-length data sizes is determined, and the transmission rate of IP packets output from TLV packets is determined.

[0664] Here, in order to define the maximum transmission speed after the IP header is expanded, the transmission rate is determined assuming that the IP header is always compressed.

[0665] In addition, when IPv4 and IPv6 packet types are mixed, or when specifying without distinguishing between packet types, the transmission rate is determined assuming IPv6 packets, which have a large header size and a large growth rate after header expansion.

[0666] For example, if the average packet length of TLV packets input to the TLV packet buffer is S, and all IP packets stored in the TLV packets are IPv6 packets and are header-compressed, the maximum output transmission rate after removing the TLV header and expanding the IP header is: Input rate × {S / (S + IPv6 header compression amount)} This becomes:

[0667] More specifically, the average packet length S of a TLV packet is defined as S = 0.75 x 1500 (1500 is the maximum MTU size assumed) is set as the standard, Amount of IPv6 header compression = TLV header length - IPv6 header length - UDP header length =3-40-8 In this case, the maximum output transmission rate after removing the TLV header and expanding the IP header is Input rate × 1.0417 ≒ Input rate × 1.05 This becomes:

[0668] FIG. 69 is a diagram showing an example in which multiple data units are aggregated and stored in one payload.

[0669] In the MMT method, when data units are aggregated, a data unit length and a data unit header are added before the data units, as shown in FIG.

[0670] However, for example, when a video signal in the NAL size format is stored as one data unit, as shown in Figure 70, there are two fields indicating the size for one data unit, and the information is redundant. Figure 70 shows an example of aggregating and storing multiple data units into one payload, where a video signal in the NAL size format is treated as one data unit. Specifically, the first size field in the NAL size format (hereinafter referred to as the "size field") and the data unit length field located before the data unit header in the MMTP payload header are both fields indicating the size, and the information is redundant. For example, if the length of the NAL unit is L bytes, the size field indicates L bytes, and the data unit length field indicates L bytes + "length of the size field" (bytes). Although the values ​​indicated in the size field and the data unit length field do not exactly match, they can be said to be redundant because one value can be easily calculated from the other.

[0671] In this way, when data containing data size information is stored as a data unit and multiple such data units are aggregated and stored in a single payload, there is a problem that the size information is duplicated, resulting in large overhead and poor transmission efficiency.

[0672] Therefore, in a transmitting device, when data containing data size information is stored as a data unit and multiple such data units are aggregated and stored in a single payload, it is possible to store them as shown in Figures 71 and 72.

[0673] As shown in Figure 71, it is conceivable to store a NAL unit including a size field as a data unit, and not indicate the data unit length that is conventionally included in the MMTP payload header. Figure 71 shows the structure of the payload of an MMTP packet in which the data unit length is not indicated.

[0674] Also, as shown in Figure 72, a flag indicating whether the data unit length is indicated and information indicating the length of the size field may be newly stored in the header. The location where the flag and information indicating the length of the size field are stored may be indicated on a data unit basis, such as in a data unit header, or may be indicated on a unit where multiple data units are aggregated (packet basis). Figure 72 shows an example of the extend field assigned on a packet basis. Note that the storage location of the above newly indicated information is not limited to this, and may also be the MMTP payload header, MMTP packet header, or control information.

[0675] On the receiving side, if the flag indicating whether the data unit length is compressed indicates that the data unit length is compressed, the length information of the size area inside the data unit is obtained, and the size area is obtained based on the length information of the size area, and the data unit length can be calculated using the length information of the obtained size area and the size area.

[0676] By using the above method, the amount of data can be reduced on the sending side, and transmission efficiency can be improved.

[0677] Note that overhead may be reduced by reducing the size field instead of reducing the data unit length. When reducing the size field, information indicating whether the size field has been reduced or the length of the data unit length field may be stored.

[0678] The MMTP payload header also contains length information.

[0679] When a NAL unit containing a size field is stored as a data unit, the payload size field in the MMTP payload header may be reduced regardless of whether aggregation is performed or not.

[0680] Also, even when data that does not include a size field is stored as a data unit, if it is aggregated and the data unit length is indicated, the payload size field in the MMTP payload header may be reduced.

[0681] When the payload size area is reduced, a flag indicating whether or not the reduction has been made and length information of the reduced size field may be indicated, as described above.

[0682] FIG. 73 shows the operation flow of the receiving device.

[0683] As described above, the transmitting device stores NAL units containing size fields as data units, and the data unit length contained in the MMTP payload header is not indicated in the MMTP packet.

[0684] In the following, we will explain an example in which whether the data unit length is indicated is indicated by a flag or the length information in the size field in the MMTP packet.

[0685] The receiving device determines whether the data unit includes a size field and whether the data unit length has been reduced based on information transmitted from the transmitting side (S1901).

[0686] If it is determined that the data unit length has been reduced (Yes in S1902), the length information of the size field inside the data unit is obtained, and then the size field inside the data unit is analyzed and the data unit length is calculated (S1903).

[0687] On the other hand, if it is determined that the data unit length has not been reduced (No in S1902), the data unit length is calculated as usual from either the data unit length or the size field inside the data unit (S1904).

[0688] [Supplementary information: Transmitting and receiving devices] As described above, a transmitting device that performs rate control so as to satisfy the requirements of the receiving buffer model during encoding can also be configured as shown in Fig. 74. Also, a receiving device that receives and decodes transmission packets transmitted from the transmitting device can also be configured as shown in Fig. 75. Fig. 74 is a diagram showing an example of a specific configuration of a transmitting device. Fig. 75 is a diagram showing an example of a specific configuration of a receiving device.

[0689] The transmission device 700 includes a generation unit 701 and a transmission unit 702. Each of the generation unit 701 and the transmission unit 702 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0690] The receiving device 800 includes a receiving unit 801, a first buffer 802, a second buffer 803, a third buffer 804, a fourth buffer 805, and a decoding unit 806. Each of the receiving unit 801, the first buffer 802, the second buffer 803, the third buffer 804, the fourth buffer 805, and the decoding unit (decoder) 806 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0691] Detailed explanations of each component of the transmitting device 700 and the receiving device 800 will be given in the explanations of the transmitting method and the receiving method, respectively.

[0692] First, the transmission method will be described with reference to Fig. 76. Fig. 76 shows the operation flow (transmission method) performed by the transmission device.

[0693] First, the generating unit 701 of the transmitting device 700 generates a coded stream by performing rate control so as to satisfy the regulations of a predetermined receiving buffer model in order to guarantee the buffer operation of the receiving device (S2001).

[0694] Next, the transmitting unit 702 of the transmitting device 700 packetizes the generated coded stream and transmits the transmission packets obtained by the packetization (S2002).

[0695] The receiving buffer model used in the transmitting device 700 has the same configuration as the receiving device 800, including the first to fourth buffers 802 to 805, and therefore a description thereof will be omitted.

[0696] This allows the transmitting device 700 to guarantee the buffer operation of the receiving device 800 when transmitting data using a method such as MMT.

[0697] Next, the receiving method will be described with reference to Fig. 77. Fig. 77 shows the operation flow (receiving method) of the receiving device.

[0698] First, the receiving unit 801 of the receiving device 800 receives a transmission packet made up of a variable-length packet header and a variable-length payload (S2101).

[0699] Next, the first buffer 802 of the receiving device 800 converts the packet consisting of a variable-length packet header and a variable-length payload stored in the received transmission packet into a first packet having a header-expanded fixed-length packet header, and outputs the first packet obtained by the conversion at a constant bit rate (S2102).

[0700] Next, the second buffer 803 of the receiving device 800 converts the first packet obtained by the conversion into a second packet consisting of a packet header and a variable-length payload, and outputs the second packet obtained by the conversion at a constant bit rate (S2103).

[0701] Next, the third buffer 804 of the receiving device 800 converts the output second packets into NAL units, and outputs the NAL units obtained by the conversion at a constant bit rate (S2104).

[0702] Next, the fourth buffer 805 of the receiving device 800 sequentially accumulates the output NAL units, generates an access unit from the accumulated multiple NAL units, and outputs the generated access unit to the decoder at the decoding time corresponding to the access unit (S2105).

[0703] Then, the decoding unit 806 of the receiving device 800 decodes the access unit output by the fourth buffer (S2106).

[0704] This allows the receiving device 800 to perform a decoding operation that does not cause an underflow or overflow.

[0705] (Sixth embodiment) [overview] In the sixth embodiment, a specific transmission method and reception method using the reception buffer model explained in the fifth embodiment will be explained.

[0706] FIG. 78 is a diagram showing a protocol stack of the MMT / TLV method defined in ARIB STD-B60.

[0707] In the MMT method, packets contain data such as video and audio transmitted through multiple MPUs (Media Processing Units). It stores each specified data unit, such as a Presentation Unit (MFU) or Media Fragment Unit (MFU), and generates a specified MMTP packet by adding an MMTP packet header. MMTP packet headers are also added to control information, such as MMTP control messages, to generate a specified MMTP packet. The MMTP packet header has a field for storing a 32-bit short-format NTP (Network Time Protocol, specified in IETF RFC 5905), which can be used for QoS control of communication lines, etc.

[0708] In addition, the reference clock of the transmitting side (transmitting device) is synchronized with the 64-bit long format NTP defined in RFC 5905, and timestamps such as PTS (Presentation Time Stamp) and DTS (Decode Time Stamp) are added to the synchronous media based on the synchronized reference clock. Furthermore, the reference clock information of the transmitting side is transmitted to the receiving side, and the receiving device generates a system clock in the receiving device based on the reference clock information received from the transmitting side.

[0709] Specifically, the PTS and DTS are stored in the MPU timestamp descriptor and the MPU extended timestamp descriptor, which are MMTP control information, and stored in the MP table for each asset. They are then packetized as MMTP control messages and transmitted.

[0710] The MMTP packetized data is given a UDP header and an IP header and encapsulated in an IP packet. At this time, a set of packets with the same source IP address, destination IP address, source port number, destination port number, and protocol type in the IP header or UDP header is called an IP data flow. Note that IP packets of the same IP data flow have redundant headers, so some IP packets are header compressed.

[0711] In addition, a 64-bit NTP timestamp is stored in the NTP packet as reference clock information, and is then stored in the IP packet. At this time, the source IP address, destination IP address, source port number, destination port number, and protocol type of the IP packet that stores the NTP packet are fixed values, and the IP packet header is not compressed.

[0712] FIG. 79 is a diagram showing the structure of a TLV packet.

[0713] As shown in Figure 79, a TLV packet can contain transmission control information such as an IP packet, a compressed IP packet, an AMT (Address Map Table), or an NIT (Network Information Table) as data, and these data are identified using an 8-bit data type. In addition, a TLV packet indicates the data length (in bytes) using a 16-bit field, followed by the data value. In addition, a TLV packet has 1 byte of header information before the data type, and this header information is stored in a total of 4 bytes of header area. In addition, a TLV packet is mapped to a transmission slot in the advanced BS transmission system, and the mapping information is stored in the TMCC (Transmission and Multiplexing Configuration Control) control information.

[0714] FIG. 80 is a diagram illustrating an example of a block diagram of a receiving device.

[0715] In the receiving device 900, first, a demodulation means 902 decodes the channel-encoded data from the broadcast signal received by a tuner 901, performs error correction, and extracts TLV packets. Then, a TLV / IP DEMUX means 903 performs TLV DEMUX processing and IP DEMUX processing. The TLV DEMUX processing performs processing according to the data type of the TLV packet. For example, if the TLV packet contains a compressed IP packet, the compressed header of the compressed IP packet is restored. The IP DEMUX performs processing such as header analysis of IP packets and UDP packets to extract MMTP packets and NTP packets.

[0716] The NTP clock generation means 904 regenerates the NTP clock from the extracted NTP packets. The MMTP DEMUX means 905 performs filtering of components such as video and audio and control information based on the packet ID stored in the extracted MMTP packet header. The control information acquisition means 906 acquires the timestamp descriptor stored in the MP table, and the PTS / DTS calculation means 907 calculates the PTS and DTS for each access unit. The timestamp descriptor includes both the MPU timestamp descriptor and the MPU extended timestamp descriptor.

[0717] The access unit reproduction means 908 converts the video and audio filtered from the MMTP packets into data units to be presented. Specifically, the data units to be presented include NAL units and access units of the video signal, audio frames, and subtitle presentation units. The decoding and presentation means 909 decodes and presents the access unit at the time when the PTS / DTS of the access unit match, based on the reference time information of the NTP clock.

[0718] However, the configuration of the receiving device is not limited to this.

[0719] Figure 81 shows the configuration of a transmission slot. One frame consists of 120 transmission slots. Modulation methods are assigned to transmission slots in units of 5 slots. A maximum of 16 streams can be transmitted within one frame. The tlv_stream_id used to identify a TLV stream consisting of a sequence of TLV packets is stored in the TMCC, making it possible to specify the tlv_stream_id stored in the slot. The modulation method and coding rate can be changed for each slot, and once the number of slots to be used is determined, the transmission speed of the TLV stream is uniquely determined.

[0720] Next, a method for defining the buffer size of the TLV packet buffer in the receive buffer model described in Figure 68 of the fifth embodiment will be described. Note that the extraction rate Rxi of the TLV packet buffer uses the extraction rate described in the fifth embodiment. The receive buffer model is the same as that of the fifth embodiment, so the description will be omitted.

[0721] The buffer size of the TLV packet buffer is a value specified by calculating the maximum buffer occupancy expected when a TLV stream composed of TLV packets with an average packet length S per unit time T is input at a transmission rate Rt and extracted at an extraction rate Rxi. Note that the average packet length S is the average when packets are transmitted using the entire bandwidth of the TLV stream, not the average size of the TLV stream bandwidth excluding the bandwidth transmitted by TLV Null packets and TLV-SI. The packet with the largest amount of header compression and expansion is the smallest packet among variable-length packets, and the maximum occupancy in the TLV packet buffer occurs when the most consecutive packets have the minimum packet size. The minimum packet size may be specified as the header size of the TLV buffer, or as the minimum value of the TLV / IP / UDP / MMT headers when storing MMT packets.

[0722] In other words, the buffer size of the TLV packet buffer in the above receiver buffer model is equivalent to specifying the maximum number of TLV packets transmitted per unit time. Also, the buffer size specification can be said to be a specification that restricts the number of consecutive packets of the minimum packet size per unit time.

[0723] The transmission rate Rt of the input TLV stream used to define the buffer size is preferably based on the maximum transmission rate expected in the system. However, it does not necessarily have to be the maximum transmission rate. The unit time T can be set in any way, but it is desirable to set it in consideration of ease of control on the sending side.

[0724] The sending facility (transmitter) must packetize data so as not to overflow a buffer size determined, for example, using the input transmission rate Rt and extraction rate Rxi.

[0725] Note that the transmission rate Rt, unit time T, and average packet length S used in calculating the buffer size of the TLV packet buffer are values ​​used as calculation standards and do not necessarily specify that these values ​​be used for transmission. In reality, the transmission speed of a TLV stream varies depending on the transmission rate. For example, when transmitting using the advanced BS transmission method (ARIB STD-B44), the transmission rate varies depending on how many slots out of 120 are used for transmission, or on the combination of modulation method and coding rate used to transmit the transmitted slots. Therefore, regardless of the transmission rate, the transmission equipment only needs to transmit so as not to overflow the above-mentioned buffer. As long as transmission can be performed without causing the buffer to overflow, any control or algorithm can be used for transmission.

[0726] In other words, the maximum number of TLV packets transmitted per unit time may include a plurality of packet numbers that are different from each other and are set according to the above transmission rate, and the TLV packet buffer may be prevented from overflowing by transmitting a plurality of TLV packets per unit time that is equal to or less than the number of packets corresponding to the transmission rate of a predetermined data unit.

[0727] Regardless of the transmission rate, the receiving device can perform stable reception without buffer overflow by implementing a buffer based on the buffer size and extraction rate described above. In other words, the receiving device should implement a buffer with a buffer size larger than the maximum buffer occupancy calculated using the transmission rate Rt, average packet length S, and extraction rate Rxi.

[0728] Next, a method for defining a TLV buffer model size that can define a predetermined maximum number of TLV packets per transmission frame length unit (unit transmission period) will be described.

[0729] FIG. 82 is a diagram showing a transmission frame and one TLV stream (TLV packet sequence) having a specific tlv_stream_id stored in the transmission frame.

[0730] Also, (1) in Figure 82 shows a case where the unit time T is the transmission frame length when the transmission frame is transmitted at a predetermined transmission rate in the above-mentioned buffer size specification method. The transmission frame length is the average section of the average packet length S, or the section where the maximum number of TLV packets and the maximum number of consecutive packets of the minimum packet size are restricted. For example, in the advanced BS transmission system (ARIB STD-B44), the transmission frame length is 33.04647 ms. Here, it is assumed that the maximum number of consecutive packets of the minimum packet size that can be transmitted per transmission frame length is M.

[0731] In this case, the sending side (transmitting device) may transmit the TLV stream so as not to overflow the TLV packet buffer, regardless of the boundary of the transmission frame, in any of the sections A, B, and C shown in Figure 82 (1), so as to satisfy the buffer model for generating the TLV stream so that the number of TLV packets stored in the TLV stream is equal to or less than the maximum number of packets. Here, the section length of each of the sections A, B, and C is the transmission frame length. In other words, in this case, the sending equipment (transmitting device) may transmit (send) multiple TLV packets per unit time, the number of packets per unit time being equal to or less than a predetermined number (maximum number of TLV packets).

[0732] The predetermined number of packets may be a value set so that TLV packets of the minimum packet size are not consecutive. For example, as described above, the predetermined number of packets may be a value calculated when the maximum number of consecutive packets of the minimum packet size is set to M.

[0733] Furthermore, the maximum number of packets per transmission frame, the average packet length, or the number of consecutive packets of the minimum packet size may be defined in predetermined data units such as transmission frame units.

[0734] For example, the maximum number of packets, the average packet length, or the number of consecutive packets of the minimum packet size may be specified in transmission frame units, such as in sections D, E, and F shown in (2) in Figure 82. In other words, the sending facility (transmitting device) may transmit (send) a number of TLV packets equal to or less than a predetermined number of packets (maximum number of TLV packets) per unit time by transmitting transmission frames in each of a number of consecutive unit transmission periods (periods of a number of transmission frame lengths corresponding to a number of transmission frames) that do not overlap with each other in a unit time. In this case, a number of TLV packets equal to or less than the predetermined number of packets are stored in a transmission frame unit (i.e., one transmission frame) as a predetermined data unit.

[0735] Next, a specific example of a method for calculating the buffer size of a TLV packet buffer in a reception buffer model will be described.

[0736] Here, if the maximum number of consecutive packets of the minimum packet size that can be transmitted per transmission frame unit is set to M, the maximum number of consecutive packets of the minimum packet size in the TLV stream that is actually transmitted is 2 × M. Note that the maximum number of consecutive packets M is the same value as the maximum number of consecutive packets of the minimum packet size that can be transmitted per the transmission frame length. In the example of Figure 82, the maximum number of consecutive packets is the case where the TLV stream has M consecutive packets in the latter half of section D and M consecutive packets in the first half of section E, resulting in a total of (2 × M) consecutive packets. In order for the receiving device to have a buffer size that can withstand consecutive packets of the minimum packet size of (2 × M), a buffer size twice the buffer size calculated using the above method is required when the maximum number of consecutive packets of the minimum packet size that can be transmitted per transmission frame length is M.

[0737] In other words, by specifying the TLV buffer size so that it has a buffer size at least twice the buffer size calculated using the above method with unit time T as the transmission frame length, the transmission equipment can transmit data so as to satisfy the buffer model even if the maximum number of packets, average packet length, or number of consecutive packets of the minimum packet size calculated when unit time T is the transmission frame length is set to the maximum number of packets, average packet length, or number of consecutive packets of the minimum packet size per transmission frame.

[0738] As described above, by setting the buffer size and extraction rate of the TLV packet buffer, it is possible to specify the maximum number of packets per transmission frame, the average packet length, or the number of consecutive packets of the minimum packet size. Note that the maximum number of packets, the average packet length, and the number of consecutive packets of the minimum packet size can be calculated from the buffer size of the TLV packet buffer, the input rate (transmission rate), the extraction rate, and the minimum packet size.

[0739] Therefore, if the buffer size of the TLV packet buffer, extraction rate, and minimum packet size are fixed, the maximum number of packets, average packet length, or number of consecutive packets of the minimum packet size stored in a transmission frame can be specified for each transmission speed (transmission rate) of the TLV stream, making it easier to control transmission (packetization or mapping to a transmission frame).

[0740] Next, a transmission method when the buffer size of the TLV packet buffer and the maximum number of packets per transmission frame (or the maximum number of consecutive packets of the minimum packet size) are specified using the method explained in Figure 82 will be explained using Figure 83. Figure 83 is a flowchart of the transmission method when the buffer size of the TLV packet buffer and the maximum number of packets per transmission frame are specified.

[0741] Although the flowchart and explanation show an example of storing a TLV packet in a transmission frame, the same flow can be used not only when storing a TLV packet in a transmission frame, but also when considering storage.

[0742] First, we start considering the storage of TLV packets in which MMT packets are stored in transmission frame units.

[0743] It is determined whether the TLV packet stored in the transmission frame is a TLV packet (actual packet) that stores an MMTP packet (S2201).

[0744] If it is determined that the TLV packet is a TLV packet that stores an MMTP packet (Yes in S2201), the TLV packet is stored in the transmission frame, and the number of packets stored in the transmission frame is incremented by 1 (S2202).

[0745] On the other hand, if it is determined that the TLV packet is a TLV packet (NULL packet or TLV-SI) that stores a packet other than an MMT packet (No in S2201), the size of the TLV packets that store a packet other than an MMT packet is accumulated, and the number of packets is counted as one for every 1.5 kB (MTU size) (S2203).

[0746] That is, there are two types of TLV packets: first transmission packets that store actual packets (MMTP packets) that are different from invalid packets (NULL packets), and second transmission packets that store NULL packets. Also, it can be said that the first transmission packets have a specified maximum packet size, and the second transmission packets may have a packet size that exceeds the maximum packet size of the first transmission packets. When storing TLV packets in a transmission frame, if the first transmission packets are stored in the transmission frame, the number of first transmission packets is counted as the number of first packets. If the second transmission packets are stored in the transmission frame, the integer value obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet (1.5 kB) is counted as the number of second packets. A plurality of TLV packets are stored in the transmission frame so that the sum of the number of first packets and the number of second packets obtained by counting is equal to or less than the specified number of packets. More specifically, the number of second packets is an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet. For example, if the packet size of the second packets is always smaller than the maximum packet size of the first packets, it goes without saying that the number of second transmission packets will be equal to the second packet count.

[0747] After step S2202 or step S2203, the cumulative TLV packet size (total data size) in the transmission frame is calculated (S2204). That is, the sum of the packet sizes of the TLV packets stored in one transmission frame is calculated.

[0748] Next, it is determined whether the total data size calculated in step S2204 reaches the maximum data size in transmission frame units (S2205).

[0749] If the total data size has reached the maximum data size per transmission frame (Yes in step S2205), a null packet is placed without inserting a TLV packet, and the process ends.

[0750] On the other hand, if the total data size has not reached the maximum data size per transmission frame (No in step S2205), it is determined whether the sum of the first number of packets calculated in step S2202 and the second number of packets calculated in step S2203 has reached the maximum number of packets per transmission frame (i.e., a predetermined number of packets) (S2206).

[0751] Then, if it is determined that the sum of the first packet count and the second packet count has reached the maximum packet count (Yes in step S2206), the storage of the TLV packet (first transmission packet) containing the MMTP packet in the transmission frame is completed, and the TLV packet containing the next MMTP packet is stored in the next transmission frame.

[0752] In addition, the remaining transmission bandwidth of the transmission frame is used to store a TLV packet (second transmission packet) other than the TLV packet (first transmission packet) that stores the MMTP packet, such as a NULL packet or a TLV packet (second transmission packet) that stores a TLV-SI (S2207).

[0753] As described above, the control for satisfying the TLV packet buffer specifications in the receive buffer model can be easily controlled using only the number of packets stored in a transmission frame, the packet size, the number of bytes that can be stored in a transmission frame, and the maximum number of packets that can be stored in a transmission frame.

[0754] Furthermore, the ability to specify the maximum number of TLV packets stored in a transmission frame also facilitates the design of a receiving device. That is, the receiving device can simply design the buffer size of the buffer that functions as a TLV packet buffer to be larger than the maximum buffer occupancy calculated using the transmission rate, the average packet length of the transmission packets, and the extraction rate.

[0755] When specifying the maximum number of packets per TLV stream per frame, the average packet length, or the number of consecutive packets of the minimum packet size for each maximum TLV stream transmission speed (for each modulation method, coding rate, and number of transmission slots used), other constraints may be specified together. For example, the average packet length may be restricted to be equal to or less than the TS packet (average 188 bytes or average 187 bytes).

[0756] This restricts the number of variable-length packets to not be greater than when fixed-length TS packets are stored, making implementation and design of the receiving device easier.

[0757] Although the data input to the TLV packet buffer is one TLV stream with a specific tlv_stream_id, it may be replaced with a specification for each IP data flow by making it one IP data flow with a specific tlv_stream_id, a specific IP address, and a specific UDP port number.

[0758] For example, in the case of advanced BS / CS broadcasting, multiple IP data flows (= services) are stored in one TLV stream. In this case, it is possible to specify the receive buffer model and the maximum number of packets per transmission frame for each service.

[0759] [Supplementary information: Transmitting and receiving devices] As described above, a transmitting device that transmits a predetermined data unit in which a plurality of transmission packets are stored while satisfying the regulations of a predetermined receiving buffer model to guarantee the buffer operation of the receiving device can also be configured as shown in Figure 84. Also, a receiving device that receives the predetermined data unit can also be configured as shown in Figure 85. Figure 84 is a diagram showing an example of a specific configuration of a transmitting device. Figure 85 is a diagram showing an example of a specific configuration of a receiving device.

[0760] The transmitting device 1000 includes a transmitting unit 1001. The transmitting unit 1001 is realized by, for example, a microcomputer, a processor, or a dedicated circuit. That is, each processing unit constituting the transmitting device 1000 may be realized by software or hardware.

[0761] The receiving device 1100 includes a receiving unit 1101 and a buffer 1102. The receiving unit 1101 and the buffer 1102 are each realized by, for example, a microcomputer, a processor, or a dedicated circuit. In other words, each processing unit constituting the receiving device 1100 may be realized by software or hardware.

[0762] Detailed explanations of each component of the transmitting device 1000 and the receiving device 1100 will be given in the explanations of the transmitting method and the receiving method, respectively.

[0763] First, the transmission method will be described with reference to Fig. 86. Fig. 86 shows the operation flow (transmission method) of the transmission device.

[0764] The transmission method is a transmission method for transmitting a plurality of transmission packets in a state where the regulations of a predetermined reception buffer model are satisfied in order to guarantee the buffer operation of the reception device 1100.

[0765] First, the transmitting unit 1001 of the transmitting device 1000 transmits a plurality of transmission packets, the number of which is equal to or less than a predetermined number of packets per unit time (S2301).

[0766] The transmission packet is composed of a variable-length packet header and a variable-length payload. The reception buffer model includes a buffer that receives the transmission packet, converts a first packet composed of a variable-length packet header and a variable-length payload stored in the received transmission packet into a second packet having a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate.

[0767] This allows the transmitting device 1000 to guarantee the buffer operation of the receiving device 1100 when transmitting data using a method such as MMT.

[0768] Next, the receiving method will be described with reference to Fig. 87. Fig. 87 shows the operation flow (receiving method) of the receiving device.

[0769] First, the receiving unit 1101 of the receiving device 1100 receives a plurality of transmission packets, each of which is made up of a variable-length packet header and a variable-length payload (S2401).

[0770] Next, buffer 1102 of receiving device 1100 converts the first packets, each consisting of a variable-length packet header and a variable-length payload stored in the received multiple transmission packets, into second packets having header-expanded fixed-length packet headers, and outputs the second packets obtained by the conversion at a predetermined extraction rate (S2402).

[0771] The buffer size of the buffer is larger than the maximum buffer occupancy calculated using the transmission rate of a plurality of transmission packets, the average packet length of the transmission packets, and the extraction rate.

[0772] This allows the receiving device 1100 to perform processing without causing underflow or overflow in the buffer.

[0773] (Other embodiments) Although the transmitting device, receiving device, transmitting method and receiving method according to the embodiments have been described above, the present invention is not limited to these embodiments.

[0774] Furthermore, each processing unit included in the transmitting device and receiving device according to the above-described embodiments is typically realized as an LSI, which is an integrated circuit. These may be individually implemented as single chips, or some or all of them may be integrated into a single chip.

[0775] Furthermore, the integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI fabrication, or reconfigurable processors, which allow the connections and settings of circuit cells within LSIs to be reconfigured, may also be used.

[0776] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may also be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0777] In other words, the transmitting device and the receiving device include processing circuitry and storage electrically connected to (accessible from) the processing circuitry. The processing circuitry includes at least one of dedicated hardware and a program execution unit. If the processing circuitry includes a program execution unit, the storage unit stores a software program to be executed by the program execution unit. The processing circuitry uses the storage unit to execute the transmitting method or the receiving method according to the above embodiment.

[0778] Furthermore, the present invention may be the above-mentioned software program, or a non-transitory computer-readable recording medium on which the above-mentioned program is recorded. Needless to say, the above-mentioned program can be distributed via a transmission medium such as the Internet.

[0779] Furthermore, all the numbers used above are merely examples for the purpose of specifically explaining the present invention, and the present invention is not limited to the numbers used as examples.

[0780] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or time-shared by a single piece of hardware or software.

[0781] The order in which the steps included in the above-described transmitting method or receiving method are executed is merely an example for specifically explaining the present invention, and an order other than the above may be used. Also, some of the above steps may be executed simultaneously (in parallel) with other steps.

[0782] While the transmitting device, receiving device, transmitting method, and receiving method according to one or more aspects of the present invention have been described based on the embodiments, the present invention is not limited to these embodiments. As long as they do not deviate from the spirit of the present invention, various modifications conceivable by those skilled in the art to the present embodiments, and configurations constructed by combining components of different embodiments, may also be included within the scope of one or more aspects of the present invention. [Industrial Applicability]

[0783] The present invention can be applied to devices or equipment that transport media such as video data and audio data. [Explanation of symbols]

[0784] 15, 100, 300, 500, 700, 1000 transmitter 16, 101, 301 Encoding section 17, 102 Multiplexing section 18, 104 Transmitter 20, 200, 400, 600, 800, 1100 receiving device 21 Packet filtering section 22 Transmission order type discrimination section 23 Random Access Section 24, 212 Control information acquisition unit 25 Data Acquisition Section 26 Calculation section 27 Initialization information acquisition unit 28, 206 Decoding instruction section 29, 204A, 204B, 204C, 204D, 402 Decoding unit 30 Presentation section 201 Tuner 202 Demodulation section 203 Demultiplexer 205 Display section 211 Type Identification Unit 213 Slice information acquisition unit 214 Decoded data generation unit 302 Granting Department 303, 503, 702, 1001 Transmitter 401, 601, 801, 1101 Receiver 501 Split section 502, 603 Components 602 Judgment section 701 Generation part 802 First Buffer 803 Second Buffer 804 Third Buffer 805 Fourth Buffer 806 Decoding Unit 1102 buffer

Claims

1. 1. A transmission method for transmitting a plurality of transmission packets, comprising: the transmission packet is composed of a variable-length packet header and a variable-length payload; A receiving buffer model for receiving the plurality of transmitted transmission packets includes: a buffer that receives the transmission packet, converts a first packet that is configured with a variable-length packet header and a variable-length payload stored in the received transmission packet into a second packet that has a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate; storing the plurality of transmission packets, the number of which is equal to or less than a predetermined number of packets, in a predetermined data unit; transmitting the plurality of transmission packets; the plurality of transmission packets include a first transmission packet storing an actual packet different from the invalid packet, and a second transmission packet storing the invalid packet; In the storing, When the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; When storing the second transmission packets in the predetermined data unit, counting the integer value obtained by dividing the packet size of the second transmission packets by the maximum packet size of the first transmission packets and adding 1 to the quotient as the number of second packets; storing the plurality of transmission packets in the predetermined data unit so that the sum of the first number of packets and the second number of packets obtained by counting is equal to or less than the predetermined number of packets; The predetermined number of packets is a value set so that the transmission packets of the minimum packet size are not consecutive. Sending method.

2. A receiving method in a receiving device having a buffer, comprising: receiving a plurality of transmission packets, each of which comprises a variable-length packet header and a variable-length payload; converting, using the buffer, first packets each having a variable-length packet header and a variable-length payload stored in the received transmission packets into second packets each having a header-expanded fixed-length packet header, and outputting the second packets obtained by the conversion from the buffer at a predetermined extraction rate; a buffer size of the buffer is larger than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; The plurality of transmission packets include: a first transmission packet storing an actual packet different from the invalid packet, and a second transmission packet storing the invalid packet; The plurality of transmission packets stored in a predetermined data unit are If the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; If the second transmission packet is stored in the predetermined data unit, an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet is counted as the number of second packets; the packets are stored in the predetermined data unit so that the sum of the first number of packets and the second number of packets obtained by counting is equal to or less than the predetermined number of packets, The predetermined number of packets is a value set so that the transmission packets of the minimum packet size are not consecutive. Receiving method.

3. A transmitting device that transmits a predetermined data unit containing a plurality of transmission packets, the transmission packet is composed of a variable-length packet header and a variable-length payload; A receiving buffer model for receiving the plurality of transmitted transmission packets includes: a buffer that receives the transmission packet, converts a first packet that is configured with a variable-length packet header and a variable-length payload stored in the received transmission packet into a second packet that has a header-expanded fixed-length packet header, and outputs the second packet obtained by the conversion at a predetermined extraction rate; a transmitting unit that stores the plurality of transmission packets, the number of which is equal to or less than a predetermined number of packets, in a predetermined data unit and transmits the plurality of transmission packets; the plurality of transmission packets include a first transmission packet storing an actual packet different from the invalid packet, and a second transmission packet storing the invalid packet; The transmission unit When the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; When storing the second transmission packets in the predetermined data unit, counting the integer value obtained by dividing the packet size of the second transmission packets by the maximum packet size of the first transmission packets and adding 1 to the quotient as the number of second packets; storing the plurality of transmission packets in the predetermined data unit so that the sum of the first number of packets and the second number of packets obtained by counting is equal to or less than the predetermined number of packets; The predetermined number of packets is a value set so that the transmission packets of the minimum packet size are not consecutive. Transmitting device.

4. a receiver for receiving a plurality of transmission packets, each of which comprises a variable-length packet header and a variable-length payload; a buffer that converts first packets, each of which is composed of a variable-length packet header and a variable-length payload and is stored in the plurality of received transmission packets, into second packets having a header-expanded fixed-length packet header, and outputs the second packets obtained by the conversion at a predetermined extraction rate; a buffer size of the buffer is larger than a maximum buffer occupancy calculated using a transmission rate of the plurality of transmission packets, an average packet length of the transmission packets, and the extraction rate; The plurality of transmission packets include: a first transmission packet storing an actual packet different from the invalid packet, and a second transmission packet storing the invalid packet; The plurality of transmission packets stored in a predetermined data unit are If the first transmission packets are stored in the predetermined data unit, the number of the first transmission packets is counted as a first packet number; If the second transmission packet is stored in the predetermined data unit, an integer value obtained by adding 1 to the quotient obtained by dividing the packet size of the second transmission packet by the maximum packet size of the first transmission packet is counted as the number of second packets; the packets are stored in the predetermined data unit so that the sum of the first number of packets and the second number of packets obtained by counting is equal to or less than the predetermined number of packets, The predetermined number of packets is a value set so that the transmission packets of the minimum packet size are not consecutive. Receiving device.