Reception method, transmission method, reception device and transmission device

The method addresses real-time decoding challenges of ultra-high-definition video by generating and using presentation time information with leap second adjustments, ensuring accurate playback of encoded streams.

JP2025146912AActive Publication Date: 2025-10-03PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025124296
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2015-03-31
Filing Date
2025-07-24
Publication Date
2025-10-03
Estimated Expiration
2036-01-05

AI Technical Summary

Technical Problem

The introduction of ultra-high-definition video content like 8K and 4K poses challenges in real-time decoding due to high processing loads, and existing MMT transmission methods fail to accurately decode multiple access units at intended times during leap second adjustments.

Method used

A method involving the generation and transmission of presentation time information with identification information about leap second adjustments, allowing the receiving device to correctly reproduce encoded streams by adjusting presentation times based on reference time information.

Benefits of technology

Ensures accurate reproduction of multiple second data units at intended times, even during leap second adjustments, thereby enhancing the decoding efficiency of ultra-high-definition video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146912000001_ABST
    Figure 2025146912000001_ABST
Patent Text Reader

Abstract

To provide a reception method or the like capable of reproducing a plurality of second data units at an intended time.SOLUTION: Presentation time information showing a presentation time of a prescribed data unit is generated on the basis of reference time information, (i) the prescribed data unit, (ii) first control information storing the generated presentation time information, and (iii) second control information storing identification information showing whether to be generated on the basis of the reference time information before the presentation time information undergoes leap second adjustment is transmitted, the identification information shows whether the presentation time information is generated on the basis of the reference time information from a time before a predetermined period right before a time when the leap second adjustment is performed to the time right before, a presentation time of the prescribed data unit is generated by subtracting the predetermined time from the reference time information in generation, and the first control information is a MPU time stamp descriptor.SELECTED DRAWING: Figure 103
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a receiving method, a transmitting method, a receiving device, and a transmitting device. [Background technology]

[0002] As broadcasting and communication services become more sophisticated, the introduction of ultra-high definition video content such as 8K (7680 x 4320 pixels, hereinafter also referred to as 8K4K) and 4K (3840 x 2160 pixels, hereinafter also referred to as 4K2K) is being considered. A receiving device needs to decode and display the received encoded data of ultra-high definition video in real time. However, video with a resolution such as 8K imposes a large processing load during decoding, making it difficult to decode such video in real time using a single decoder. Therefore, methods are being considered for achieving real-time processing by using multiple decoders to parallelize the decoding process, thereby reducing the processing load per decoder.

[0003] Furthermore, the encoded data is multiplexed based on a multiplexing method such as MPEG-2 TS (Transport Stream) or MMT (MPEG Media Transport) and then transmitted. For example, Non-Patent Document 1 discloses a technique for transmitting encoded media data packet by packet in accordance with MMT. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Information technology - High efficiency coding and media delivery in heterogeneous environment - Part1:MPEG media transport(MMT), ISO / IEC DIS 23008-1 Summary of the Invention [Problem to be solved by the invention]

[0005] As broadcasting and communication services become more sophisticated, the introduction of ultra-high-definition video content such as 8K and 4K (3840x2160 pixels) is being considered. In the MMT / TLV method, the reference clock on the transmitting side is synchronized with the 64-bit long-format NTP defined in RFC 5905, and timestamps such as Presentation Time Stamp (PTS) and Decode Time Stamp (DTS) are added to synchronized media based on the reference clock. Furthermore, the reference clock information on the transmitting side is transmitted to the receiving side, and the receiving device generates its own system clock based on the reference clock information.

[0006] However, with such an MMT transmission method, when leap second adjustments are made to reference time information such as NTP, there is a problem in that the multiple access units, which are the multiple second data units stored in the MPU, which is the first data unit received by the receiving device, cannot be decoded or presented at the intended time even in accordance with the DTS or PTS associated with the multiple access units.

[0007] The present invention provides a receiving method that can reproduce multiple second data units stored in a first data unit at the intended time, even when leap second adjustments are made to the reference time information that serves as the basis for the reference clocks of the transmitting and receiving devices. [Means for solving the problem]

[0008] In order to achieve the above-mentioned object, a transmission method according to one embodiment of the present invention is a transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, which generates presentation time information indicating the presentation time of the predetermined data unit based on reference time information, and transmits (i) the predetermined data unit, (ii) first control information in which the generated presentation time information is stored, and (iii) second control information in which identification information indicating whether the presentation time information was generated based on the reference time information before a leap second adjustment is stored, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, and wherein the generation generates the presentation time of the predetermined data unit by subtracting a predetermined time from the reference time information, and the first control information is an MPU timestamp descriptor.

[0009] Furthermore, a receiving method according to one embodiment of the present disclosure is a receiving method for receiving a predetermined data unit containing data constituting an encoded stream, the method including receiving (i) the predetermined data unit, (ii) first control information containing presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information containing identification information indicating whether the presentation time information was generated based on reference time information before a leap second adjustment, and reproducing the received predetermined data unit based on the received first control information and second control information, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit being a time generated by subtracting a predetermined time from the reference time information, and the first control information being an MPU timestamp descriptor.

[0010] In order to achieve the above object, a transmission method according to one embodiment of the present invention is a transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, which generates presentation time information indicating the presentation time of the predetermined data unit based on reference time information, and transmits (i) the predetermined data unit, (ii) first control information in which the generated presentation time information is stored, and (iii) second control information in which identification information indicating whether the presentation time information was generated based on the reference time information before a leap second adjustment is stored, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, and wherein the generation generates the presentation time of the predetermined data unit by subtracting a predetermined time from the reference time information, wherein the presentation time of the predetermined data unit is the presentation time of the first access unit in presentation order, and the reference time information is NTP (Network Time Protocol).

[0011] Furthermore, a receiving method according to one embodiment of the present disclosure is a receiving method for receiving a predetermined data unit containing data constituting an encoded stream, the method including receiving (i) the predetermined data unit, (ii) first control information containing presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information containing identification information indicating whether the presentation time information was generated based on reference time information before a leap second adjustment, and reproducing the received predetermined data unit based on the received first control information and second control information, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit being a time generated by subtracting a predetermined time from the reference time information, the presentation time of the predetermined data unit being the presentation time of the first access unit in presentation order, and the reference time information being NTP (Network Time Protocol).

[0012] In order to achieve the above-mentioned object, a transmission method according to one embodiment of the present invention is a transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, which generates presentation time information indicating the presentation time of the predetermined data unit based on reference time information, and transmits (i) the predetermined data unit, (ii) first control information in which the generated presentation time information is stored, and (iii) second control information in which identification information indicating whether the presentation time information was generated based on the reference time information before a leap second adjustment is stored, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, and wherein the generation generates the presentation time of the predetermined data unit by subtracting a predetermined time from the reference time information, and the presentation time of the predetermined data unit is the presentation time of the first access unit in presentation order.

[0013] Furthermore, a receiving method according to one aspect of the present disclosure is a receiving method for receiving a predetermined data unit containing data constituting an encoded stream, the method including receiving (i) the predetermined data unit, (ii) first control information containing presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information containing identification information indicating whether the presentation time information was generated based on reference time information before a leap second adjustment, and reproducing the received predetermined data unit based on the received first control information and second control information, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit being a time generated by subtracting a predetermined time from the reference time information, and the presentation time of the predetermined data unit being the presentation time of the first access unit in presentation order.

[0014] In order to achieve the above object, a transmission method according to one embodiment of the present invention is a transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, which generates presentation time information indicating the presentation time of the predetermined data unit based on reference time information, and transmits (i) the predetermined data unit, (ii) first control information in which the generated presentation time information is stored, and (iii) second control information in which identification information indicating whether the presentation time information was generated based on the reference time information before a leap second adjustment is stored, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, and wherein the generation generates the presentation time of the predetermined data unit by subtracting a predetermined time from the reference time information, and the reference time information is NTP (Network Time Protocol).

[0015] In order to achieve the above-mentioned object, a receiving method according to one embodiment of the present invention is a receiving method for receiving a predetermined data unit containing data constituting an encoded stream, the method receiving (i) the predetermined data unit, (ii) first control information containing presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information containing identification information indicating whether the presentation time information was generated based on reference time information before a leap second adjustment, and reproducing the received predetermined data unit based on the received first control information and second control information, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit being a time generated by subtracting a predetermined time from the reference time information, and the reference time information being NTP (Network Time Protocol).

[0016] In order to achieve the above-mentioned object, a transmission method according to one embodiment of the present invention is a transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, which generates presentation time information indicating the presentation time of the predetermined data unit based on reference time information, and transmits (i) the predetermined data unit, (ii) first control information in which the generated presentation time information is stored, and (iii) second control information in which identification information indicating whether the presentation time information was generated based on the reference time information before a leap second adjustment is performed is stored, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment is performed to the time immediately before the leap second adjustment, and in the generation, the presentation time of the predetermined data unit is generated by subtracting a predetermined time from the reference time information.

[0017] In order to achieve the above-mentioned object, a receiving method according to one embodiment of the present invention is a receiving method for receiving a predetermined data unit containing data constituting an encoded stream, which receives (i) the predetermined data unit, (ii) first control information containing presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information containing identification information indicating whether the presentation time information was generated based on reference time information before a leap second adjustment, and plays back the received predetermined data unit based on the received first control information and second control information, wherein the identification information indicates whether the presentation time information was generated based on the reference time information from a time that is a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, and the presentation time of the predetermined data unit is a time generated by subtracting a predetermined time from the reference time information.

[0018] In order to achieve the above-mentioned object, a receiving method according to one embodiment of the present invention is a receiving method for receiving a first data unit containing data constituting an encoded stream, wherein the first data unit contains a plurality of second data units, and the receiving method includes receiving the first data unit, first time information indicating the presentation time of the first data unit, second time information indicating the presentation time or decoding time of each of the plurality of second data units using the first time information, and identification information, calculating the presentation time or decoding time of each of the received plurality of second data units using the received first time information and the second time information, and correcting the calculated presentation time or decoding time of each of the plurality of second data units based on the received identification information, wherein the identification information indicates whether the first time information was generated based on reference time information before leap second adjustment, and the first time information is time information indicating the presentation time of the second data unit that is first in presentation order among the plurality of second data units.

[0019] In order to achieve the above-mentioned object, a transmission method according to one embodiment of the present invention is a transmission method for transmitting a first data unit storing data constituting an encoded stream, wherein the first data unit stores a plurality of second data units, and the transmission method generates first time information indicating the presentation time of the first data unit based on reference time information received from an external source, and transmits the first data unit, the generated first time information, second time information indicating the presentation time or decoding time of each of the plurality of second data units using the first time information, and identification information, wherein the identification information is information indicating whether the first time information was generated based on reference time information before leap second adjustment, and the first time information is time information indicating the presentation time of the second data unit that is first in presentation order among the plurality of second data units.

[0020] In order to achieve the above-mentioned object, a receiving method according to one embodiment of the present invention is a receiving method for receiving a first data unit storing data constituting an encoded stream, wherein the first data unit stores a plurality of second data units, and the receiving method receives the first data unit, first time information indicating the presentation time of the first data unit, second time information indicating the presentation time or decoding time of each of the plurality of second data units using the first time information, and identification information, calculates the presentation time or decoding time of each of the plurality of received second data units using the received first time information and the second time information, and corrects the calculated presentation time or the decoding time of each of the plurality of second data units based on the received identification information.

[0021] These general or specific aspects may be realized as a system, device, integrated circuit, computer program, or computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, device, integrated circuit, computer program, and recording medium. [Effects of the Invention]

[0022] The present invention can reproduce multiple second data units stored in a first data unit at the intended time, even when leap second adjustments are made to the reference time information that serves as the basis for the reference clocks of the transmitting and receiving devices. [Brief explanation of the drawings]

[0023] [Figure 1] FIG. 1 is a diagram showing an example of dividing a picture into slice segments. [Figure 2] FIG. 2 is a diagram showing an example of a PES packet sequence in which picture data is stored. [Figure 3] FIG. 3 is a diagram showing an example of division of a picture according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of dividing a picture according to a comparative example of the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of data of an access unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a transmission device according to the first embodiment. [Figure 7] FIG. 7 is a block diagram of a receiving device according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of an MMT packet according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating another example of an MMT packet according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of data input to each decoding unit according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 12] FIG. 12 is a diagram showing another example of data input to each decoding unit according to the first embodiment. [Figure 13] FIG. 13 is a diagram showing an example of dividing a picture according to the first embodiment. [Figure 14] FIG. 14 is a flowchart of a transmission method according to the first embodiment. [Figure 15] FIG. 15 is a block diagram of a receiving device according to the first embodiment. [Figure 16] FIG. 16 is a flowchart of a receiving method according to the first embodiment. [Figure 17] FIG. 17 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 19] FIG. 19 is a diagram illustrating the configuration of the MPU. [Figure 20] FIG. 20 is a diagram showing the structure of MF metadata. [Figure 21] FIG. 21 is a diagram for explaining the data transmission order. [Figure 22]FIG. 22 is a diagram showing an example of a method for performing decoding without using header information. [Figure 23] FIG. 23 is a block diagram of a transmission device according to the second embodiment. [Figure 24] FIG. 24 is a flowchart of a transmission method according to the second embodiment. [Figure 25] FIG. 25 is a block diagram of a receiving device according to the second embodiment. [Figure 26] FIG. 26 is a flowchart of the operation for identifying the MPU start position and the NAL unit position. [Figure 27] FIG. 27 is a flowchart of an operation of obtaining initialization information based on a transmission order type and decoding media data based on the initialization information. [Figure 28] FIG. 28 is a flowchart of the operation of a receiving device when a low-delay presentation mode is provided. [Figure 29] FIG. 29 is a diagram illustrating an example of the transmission order of MMT packets when auxiliary data is transmitted. [Figure 30] FIG. 30 is a diagram illustrating an example in which a transmitting device generates auxiliary data based on the configuration of moof. [Figure 31] FIG. 31 is a diagram for explaining reception of auxiliary data. [Figure 32] FIG. 32 is a flowchart of a receiving operation using auxiliary data. [Figure 33] FIG. 33 shows the structure of an MPU that is made up of multiple movie fragments. [Figure 34] FIG. 34 is a diagram for explaining the transmission order of MMT packets when the MPU having the configuration of FIG. 33 is transmitted. [Figure 35] FIG. 35 is a first diagram illustrating an example of the operation of a receiving device when one MPU is made up of multiple movie fragments. [Figure 36] FIG. 36 is a second diagram illustrating an example of the operation of a receiving device when one MPU is made up of multiple movie fragments. [Figure 37] FIG. 37 is a flowchart of the operation of the receiving method described with reference to FIGS. [Figure 38] FIG. 38 is a diagram showing a case where non-VCL NAL units are aggregated as individual data units. [Figure 39] FIG. 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. [Figure 40] FIG. 40 is a flowchart showing the operation of the receiving device when a packet loss occurs. [Figure 41] FIG. 41 is a flowchart of the receiving operation when the MPU is divided into multiple movie fragments. [Figure 42] FIG. 42 is a diagram showing an example of a prediction structure of a picture for each TemporalId when temporal scalability is realized. [Figure 43] FIG. 43 shows the relationship between the decoding time (DTS) and the display time (PTS) for each picture in FIG. [Figure 44] FIG. 44 is a diagram showing an example of a prediction structure of a picture that requires picture delay processing and reorder processing. [Figure 45] Figure 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and an MMTP packet. [Figure 46] FIG. 46 is a diagram for explaining the calculation method and problems of the PTS and DTS. [Figure 47] FIG. 47 is a flowchart of the receiving operation when the DTS is calculated using information for DTS calculation. [Figure 48] FIG. 48 is a diagram for explaining a method of storing a data unit in a payload in MMT. [Figure 49] FIG. 49 is a flow chart showing the operation of the transmitting device according to the third embodiment. [Figure 50] FIG. 50 shows an operation flow of the receiving device according to the third embodiment. [Figure 51]FIG. 51 is a diagram illustrating an example of a specific configuration of a transmission device according to the third embodiment. [Figure 52] FIG. 52 is a diagram illustrating an example of a specific configuration of a receiving device according to the third embodiment. [Figure 53] Figure 53 shows how non-timed media is stored in the MPU and how it is transmitted in MMTP packets. [Figure 54] FIG. 54 shows an example in which a file is divided into a plurality of divided data pieces, each of which is packetized and transmitted. [Figure 55] FIG. 55 shows another example in which a plurality of divided data pieces obtained by dividing a file are packetized and transmitted. [Figure 56] FIG. 56 is a diagram showing the syntax of a loop for each file in the asset management table. [Figure 57] FIG. 57 shows the operational flow for identifying the divided data number in the receiving device. [Figure 58] FIG. 58 shows the operational flow for identifying the number of divided data pieces in the receiving device. [Figure 59] FIG. 59 shows an operational flow for determining whether to operate a fragment counter in a transmitting device. [Figure 60] FIG. 60 is a diagram for explaining a method for identifying the number of divided data pieces and divided data numbers (when a fragment counter is used). [Figure 61] Figure 61 shows the operational flow of a transmitting device when utilizing a fragment counter. [Figure 62] FIG. 62 shows the operational flow of a receiving device when utilizing a fragment counter. [Figure 63] FIG. 63 shows a service configuration in which the same program is transmitted using multiple IP data flows. [Figure 64] FIG. 64 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 65] FIG. 65 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 66] FIG. 66 shows the operational flow of the transmitting device. [Figure 67] FIG. 67 shows the operational flow of the receiving device. [Figure 68] FIG. 68 shows a reception buffer model based on the reception buffer model defined in ARIB STD-B60, particularly when only a broadcast transmission channel is used. [Figure 69] FIG. 69 is a diagram showing an example in which multiple data units are aggregated and stored in one payload. [Figure 70] FIG. 70 shows an example in which a plurality of data units are aggregated and stored in one payload, where a video signal in NAL size format is used as one data unit. [Figure 71] Figure 71 shows the structure of the payload of an MMTP packet in which the data unit length is not indicated. [Figure 72] FIG. 72 shows an example of the extend area assigned to each packet. [Figure 73] FIG. 73 shows the operation flow of the receiving device. [Figure 74] FIG. 74 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 75] FIG. 75 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 76] FIG. 76 shows the operational flow of the transmitting device. [Figure 77] FIG. 77 shows the operational flow of the receiving device. [Figure 78] FIG. 78 is a diagram showing a protocol stack of the MMT / TLV method defined in ARIB STD-B60. [Figure 79] FIG. 79 is a diagram showing the structure of a TLV packet. [Figure 80] FIG. 80 is a diagram illustrating an example of a block diagram of a receiving device. [Figure 81] FIG. 81 is a diagram illustrating a timestamp descriptor. [Figure 82]FIG. 82 is a diagram for explaining leap second adjustment. [Figure 83] FIG. 83 is a diagram showing the relationship between the NTP time, the MPU timestamp, and the MPU presentation timing. [Figure 84] FIG. 84 is a diagram for explaining a method for correcting a timestamp on the transmitting side. [Figure 85] FIG. 85 is a diagram for explaining a method for correcting a timestamp in a receiving device. [Figure 86] FIG. 86 shows the operational flow of the transmitting side (transmitting device) when correcting the MPU timestamp on the transmitting side (transmitting device). [Figure 87] FIG. 87 shows the operational flow of the receiving device when correcting the MPU timestamp on the transmitting side (transmitting device). [Figure 88] FIG. 88 shows the operational flow on the transmitting side (transmitting device) when correcting the MPU timestamp in the receiving device. [Figure 89] FIG. 89 shows the operational flow of the receiving device when correcting the MPU timestamp in the receiving device. [Figure 90] FIG. 90 is a diagram illustrating an example of a specific configuration of a transmitting device. [Figure 91] FIG. 91 is a diagram illustrating an example of a specific configuration of a receiving device. [Figure 92] FIG. 92 shows the operational flow of the transmitting device. [Figure 93] FIG. 93 shows the operational flow of the receiving device. [Figure 94] FIG. 94 is a diagram illustrating an example of an extension of the MPU extended timestamp descriptor. [Figure 95] FIG. 95 is a diagram for explaining a case where discontinuity occurs in the MPU sequence numbers due to adjustment of the MPU sequence numbers. [Figure 96] FIG. 96 is a diagram for explaining a case where packet sequence numbers become discontinuous at the timing of switching from normal equipment to redundant equipment. [Figure 97] FIG. 97 shows the operational flow of the receiving device when a discontinuity occurs in the MPU sequence number or packet sequence number. [Figure 98] FIG. 98 is a diagram for explaining a method for correcting a timestamp in a receiving device when a leap second is inserted. [Figure 99] FIG. 99 is a diagram for explaining a method for correcting a timestamp in a receiving device when leap seconds are deleted. [Figure 100] FIG. 100 shows the operational flow of the receiving device. [Figure 101] FIG. 101 is a diagram showing an example of a specific configuration of a transmission / reception system. [Figure 102] FIG. 102 is a diagram showing a specific configuration of a receiving device. [Figure 103] FIG. 103 shows the operational flow of the receiving device. DETAILED DESCRIPTION OF THE INVENTION

[0024] A receiving method according to one embodiment of the present invention is a receiving method for receiving a first data unit containing data constituting an encoded stream, wherein the first data unit contains a plurality of second data units, and the receiving method includes receiving the first data unit, first time information indicating the presentation time of the first data unit, second time information indicating the presentation time or decoding time of each of the plurality of second data units using the first time information, and identification information, calculating the presentation time or decoding time of each of the plurality of received second data units using the received first time information and second time information, and correcting the calculated presentation time or decoding time of each of the plurality of second data units based on the received identification information.

[0025] According to this, such a receiving method can reproduce the encoded stream consisting of the first data unit at the intended time even if leap second adjustment is made to the reference time information that serves as the basis for the reference clocks of the transmitting and receiving devices.

[0026] The first time information may be absolute time information indicating the presentation time of the first second data unit in presentation order among the plurality of second data units.

[0027] The first time information may be time information that is not corrected on the transmitting side of the first data unit during leap second adjustment.

[0028] The second time information may be relative time information for calculating the presentation time or the decoding time of each of the plurality of second data units in conjunction with the first time information.

[0029] In addition, the identification information is information indicating whether the first time information was generated based on the reference time information before the leap second adjustment, and the correction may involve determining, for each of the multiple second data units, whether the first time information of the first data unit that stores the second data unit was generated based on the reference time information before the leap second adjustment, and whether the calculated presentation time or decoding time of the second data unit satisfies a correction condition that is after a predetermined time, and correcting the presentation time or decoding time of the second data unit that is determined to satisfy the correction condition in the determination.

[0030] The reference time information may be NTP (Network Time Protocol).

[0031] The identification information may also be information indicating whether the first time information was generated based on the reference time information from a time a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment.

[0032] Furthermore, the receiving may include receiving a predetermined packet that stores control information in which the first time information corresponding to the first data unit is stored, and a plurality of the first data units.

[0033] Further, the receiving may include receiving an MMTP (MPEG Media Transport Protocol) packet as the specified packet, which stores a plurality of MPUs (Media Presentation Units) as the first data units, and the MMTP packet includes, as control information, an MPU timestamp descriptor including the first time information, an MPU extended timestamp descriptor including the second time information, and identification information, and the MPU may store a plurality of access units as the plurality of second data units.

[0034] These comprehensive or specific aspects may be realized in a device, a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or in any combination of a device, a system, an integrated circuit, a computer program, or a recording medium.

[0035] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0036] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.

[0037] (Findings that form the basis of the present invention) In recent years, the resolution of displays such as TVs, smartphones, and tablet devices has been increasing. In particular, 8K4K (8K x 4K resolution) services are scheduled for broadcasting in Japan in 2020. Since it is difficult to decode ultra-high resolution moving images such as 8K4K in real time using a single decoder, methods for decoding images in parallel using multiple decoders are being investigated.

[0038] Since the coded data is multiplexed and transmitted based on a multiplexing method such as MPEG-2 TS or MMT, the receiving device must separate the coded data of the video from the multiplexed data before decoding. Hereinafter, the process of separating the coded data from the multiplexed data will be referred to as demultiplexing.

[0039] When parallelizing the decoding process, it is necessary to allocate the coded data to be decoded to each decoder. When allocating the coded data, the coded data itself needs to be analyzed, and the processing load for analysis is large, especially for content such as 8K, where the bit rate is very high. Therefore, the demultiplexing part becomes a bottleneck, making real-time playback impossible.

[0040] In video coding standards such as H.264 and H.265 standardized by MPEG and ITU, a transmitting device divides a picture into multiple regions called slices or slice segments and encodes the picture so that each region can be decoded independently. Therefore, in the case of H.265, for example, a receiving device that receives a broadcast can separate data for each slice segment from received data and output the data for each slice segment to a separate decoder, thereby achieving parallel decoding.

[0041] 1 is a diagram showing an example in which one picture is divided into four slice segments in HEVC. For example, a receiving device includes four decoders, and each decoder decodes one of the four slice segments.

[0042] In conventional broadcasting, a transmitting device stores one picture (access unit in the MPEG system standard) in one PES packet and multiplexes the PES packet into a sequence of TS packets. Therefore, a receiving device must separate the payload of the PES packet, analyze the data of the access unit stored in the payload, separate each slice segment, and output the data of each separated slice segment to a decoder.

[0043] However, the present inventors have found that there is a problem in that it is difficult to perform this processing in real time because the amount of processing required to analyze the data of an access unit and separate the slice segments is large.

[0044] FIG. 2 is a diagram showing an example in which data of a picture divided into slice segments is stored in the payload of a PES packet.

[0045] 2, for example, data of a plurality of slice segments (slice segments 1 to 4) is stored in the payload of one PES packet, and the PES packets are multiplexed into a sequence of TS packets.

[0046] (Embodiment 1) In the following, an example will be described in which H.265 is used as the video encoding method, but this embodiment can also be applied to cases in which other encoding methods such as H.264 are used.

[0047] 3 is a diagram showing an example of dividing an access unit (picture) into division units in this embodiment. The access unit is divided into two equal parts horizontally and vertically, into a total of four tiles, by a function called tile introduced by H.265. Furthermore, there is a one-to-one correspondence between slice segments and tiles.

[0048] The reason for dividing the data into two equal parts horizontally and vertically will be explained below. First, when decoding, a line memory is generally required to store one horizontal line of data. However, when it comes to ultra-high resolutions such as 8K4K, the horizontal size increases, so the size of the line memory also increases. When implementing a receiving device, it is desirable to be able to reduce the size of the line memory. In order to reduce the size of the line memory, vertical division is necessary. Vertical division requires a data structure called a tile. For these reasons, tiles are used.

[0049] On the other hand, since images generally have high correlation in the horizontal direction, being able to refer to a wider range in the horizontal direction improves coding efficiency. Therefore, from the viewpoint of coding efficiency, it is desirable to divide the access unit horizontally.

[0050] By dividing the access unit into two equal parts horizontally and vertically, these two characteristics can be achieved simultaneously, and both implementation and coding efficiency can be taken into consideration.If a single decoder can decode 4K2K video in real time, by dividing an 8K4K image into four equal parts and dividing each slice segment into 4K2K, the receiving device can decode the 8K4K image in real time.

[0051] Next, we will explain why tiles obtained by dividing an access unit in the horizontal and vertical directions correspond one-to-one to slice segments. In H.265, an access unit is made up of multiple units called NAL (Network Adaptation Layer) units.

[0052] The payload of an NAL unit stores any of the following: an access unit delimiter indicating the start position of an access unit; an SPS (Sequence Parameter Set) which is initialization information used in common for each sequence at the time of decoding; a PPS (Picture Parameter Set) which is initialization information used in common for each picture at the time of decoding; SEI (Supplemental Enhancement Information) which is not required for the decoding process itself but is required for processing and displaying the decoding results; and coded data of a slice segment. The header of an NAL unit contains type information for identifying the data stored in the payload.

[0053] Here, the transmitting device transmits the coded data in MPEG-2 TS, MMT (MPEG Media Transport), MPEG DASH (Dynamic Adaptive Streaming) format. When multiplexing using a multiplexing format such as Streaming over HTTP (Streaming over HTTP) or RTP (Real-time Transport Protocol), the basic unit can be set to an NAL unit. In order to store one slice segment in one NAL unit, it is desirable to divide the access unit into regions by slice segment units. For this reason, the transmitting device associates tiles with slice segments in a one-to-one correspondence.

[0054] As shown in Fig. 4, the transmitting device can also set tiles 1 to 4 together in one slice segment. In this case, however, all tiles are stored in one NAL unit, making it difficult for the receiving device to separate the tiles in the multiplexing layer.

[0055] Note that there are two types of slice segments: independent slice segments that can be decoded independently, and reference slice segments that refer to independent slice segments. Here, a case where an independent slice segment is used will be described.

[0056] Fig. 5 is a diagram showing an example of data of an access unit divided so that the boundaries between tiles and slice segments coincide as shown in Fig. 3. The data of the access unit includes a NAL unit in which an access unit delimiter placed at the beginning is stored, followed by NAL units of SPS, PPS, and SEI, and data of slice segments in which data from tiles 1 to 4 is stored. Note that the data of the access unit does not have to include some or all of the NAL units of SPS, PPS, and SEI.

[0057] Next, a configuration of transmitting apparatus 100 according to this embodiment will be described. Fig. 6 is a block diagram showing an example configuration of transmitting apparatus 100 according to this embodiment. This transmitting apparatus 100 includes an encoding unit 101, a multiplexing unit 102, a modulating unit 103, and a transmitting unit 104.

[0058] The encoding unit 101 generates encoded data by encoding an input image according to, for example, H.265. Furthermore, the encoding unit 101 divides an access unit into four slice segments (tiles) and encodes each slice segment, for example, as shown in FIG. 3 .

[0059] The multiplexing unit 102 multiplexes the coded data generated by the coding unit 101. The modulation unit 103 modulates the data obtained by the multiplexing. The transmission unit 104 transmits the modulated data as a broadcast signal.

[0060] Next, a configuration of receiving apparatus 200 according to this embodiment will be described. Fig. 7 is a block diagram showing an example configuration of receiving apparatus 200 according to this embodiment. Receiving apparatus 200 includes tuner 201, demodulation unit 202, demultiplexing unit 203, a plurality of decoding units 204A to 204D, and display unit 205.

[0061] The tuner 201 receives a broadcast signal. The demodulation unit 202 demodulates the received broadcast signal. The demodulated data is input to the demultiplexing unit 203.

[0062] The demultiplexing unit 203 separates the demodulated data into division units and outputs the data for each division unit to the decoding units 204A to 204D. Here, a division unit is a divided area obtained by dividing an access unit, such as a slice segment in H.265. Also, here, an 8K4K image is divided into four 4K2K images. Therefore, there are four decoding units 204A to 204D.

[0063] The decoding units 204A to 204D operate in synchronization with one another based on a predetermined reference clock. Each decoding unit decodes the coded data in division units in accordance with the DTS (Decoding Time Stamp) of the access unit, and outputs the decoding result to the display unit 205.

[0064] The display unit 205 generates an 8K4K output image by integrating the multiple decoding results output from the multiple decoding units 204A to 204D. The display unit 205 displays the generated output image in accordance with the PTS (Presentation Time Stamp) of the access unit acquired separately. When integrating the decoding results, the display unit 205 may perform filtering such as a deblocking filter on boundary regions between adjacent division units, such as tile boundaries, so that the boundaries become less noticeable visually.

[0065] In the above description, the transmitting device 100 and the receiving device 200 that transmit or receive broadcasts are used as examples, but content may be transmitted and received via a communication network. When the receiving device 200 receives content via a communication network, the receiving device 200 separates the multiplexed data from IP packets received via a network such as Ethernet.

[0066] In broadcasting, the transmission path delay from when a broadcast signal is transmitted until it reaches the receiving device 200 is constant. On the other hand, in communication networks such as the Internet, due to the effects of congestion, the transmission path delay from when data transmitted from a server reaches the receiving device 200 is not constant. Therefore, the receiving device 200 often does not perform strictly synchronized playback based on a reference clock such as PCR in the broadcast MPEG-2 TS. Therefore, the receiving device 200 may display the 8K4K output image on the display device according to the PTS without strictly synchronizing each decoding unit.

[0067] Furthermore, due to congestion in the communication network, etc., the decoding process for all division units may not be completed by the time indicated by the PTS of the access unit. In this case, the receiving device 200 skips displaying the access unit or delays displaying it until decoding of at least four division units is completed and generation of the 8K4K image is completed.

[0068] The content may be transmitted and received by combining broadcasting and communication. The present method can also be applied to playing back multiplexed data stored on a recording medium such as a hard disk or memory.

[0069] Next, a method for multiplexing access units divided into slice segments when MMT is used as the multiplexing method will be described.

[0070] 8 is a diagram showing an example of packetizing data of an HEVC access unit into MMT. Although the SPS, PPS, SEI, and the like do not necessarily need to be included in the access unit, the case where they are present is illustrated here.

[0071] NAL units that are located before the first slice segment in an access unit, such as the access unit delimiter, SPS, PPS, and SEI, are stored together in MMT packet #1. Subsequent slice segments are stored in separate MMT packets for each slice segment.

[0072] As shown in FIG. 9, an NAL unit that is arranged before the first slice segment in an access unit may be stored in the same MMT packet as the first slice segment.

[0073] Furthermore, when NAL units such as End-of-Sequence or End-of-Bitstream, which indicate the end of a sequence or stream, are added after the last slice segment, they are stored in the same MMT packet as the last slice segment. However, since NAL units such as End-of-Sequence or End-of-Bitstream are inserted at the end point of the decoding process or the connection point of two streams, it may be desirable for receiving device 200 to easily obtain these NAL units in the multiplexing layer. In this case, these NAL units may be stored in an MMT packet separate from the slice segment. This allows receiving device 200 to easily separate these NAL units in the multiplexing layer.

[0074] Note that TS, DASH, RTP, or the like may be used as the multiplexing method. In these methods, the transmitting device 100 stores different slice segments in different packets. This ensures that the receiving device 200 can separate the slice segments in the multiplexing layer.

[0075] For example, when TS is used, coded data is packetized in slice segment units as PES packets. When RTP is used, coded data is packetized in slice segment units as RTP packets. Even in these cases, NAL units placed before slice segments and slice segments may be packetized separately, as in MMT packet #1 shown in FIG. 8.

[0076] When TS is used, the transmitting device 100 indicates the unit of data stored in a PES packet by using a data alignment descriptor, etc. Furthermore, since DASH is a method of downloading MP4 format data units called segments via HTTP, the transmitting device 100 does not packetize the encoded data for transmission. For this reason, the transmitting device 100 may create subsamples in slice segment units and store information indicating the storage locations of the subsamples in the MP4 header so that the receiving device 200 can detect slice segments in the multiplexing layer in MP4.

[0077] The MMT packetization of slice segments will be described in detail below.

[0078] As shown in Fig. 8, by packetizing the encoded data, data commonly referenced when decoding all slice segments in an access unit, such as SPS and PPS, is stored in MMT packet #1. In this case, receiving device 200 concatenates the payload data of MMT packet #1 with the data of each slice segment and outputs the obtained data to the decoding unit. In this way, receiving device 200 can easily generate input data for the decoding unit by concatenating the payloads of multiple MMT packets.

[0079] FIG. 10 is a diagram showing an example in which input data to decoding units 204A to 204D is generated from the MMT packets shown in FIG. 8. Demultiplexing unit 203 generates data necessary for decoding unit 204A to decode slice segment 1 by concatenating payload data of MMT packet #1 and MMT packet #2. Demultiplexing unit 203 similarly generates input data for decoding units 204B to 204D. That is, demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #3 to generate input data for decoding unit 204B. Demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #4 to generate input data for decoding unit 204C. Demultiplexing unit 203 concatenates payload data of MMT packet #1 and MMT packet #5 to generate input data for decoding unit 204D.

[0080] In addition, the demultiplexing unit 203 may remove NAL units that are not necessary for the decoding process, such as the access unit delimiter and SEI, from the payload data of MMT packet #1, and separate only the NAL units of SPS and PPS that are necessary for the decoding process and add them to the data of the slice segment.

[0081] 9, when encoded data is packetized, demultiplexing unit 203 outputs MMT packet #1 including the start data of an access unit in the multiplexing layer to the first decoding unit 204A. In addition, demultiplexing unit 203 analyzes the MMT packet including the start data of an access unit in the multiplexing layer, separates the SPS and PPS NAL units, and adds the separated SPS and PPS NAL units to each of the data of the second and subsequent slice segments to generate input data for each of the second and subsequent decoding units.

[0082] Furthermore, it is desirable that receiving device 200 can identify the type of data stored in the MMT payload and the index number of the slice segment within an access unit when a slice segment is stored in the payload, using information included in the header of the MMT packet. Here, the data type refers to either pre-slice segment data (a collective term for NAL units arranged before the first slice segment in an access unit) or slice segment data. When storing a unit obtained by fragmenting an MPU, such as a slice segment, in an MMT packet, a mode for storing MFUs (Media Fragment Units) is used. When using this mode, transmitting device 100 can, for example, set the Data Unit, which is the basic unit of data in an MFU, to a sample (a data unit in MMT, equivalent to an access unit) or a subsample (a unit obtained by dividing a sample).

[0083] At this time, the header of the MMT packet includes a field called a fragmentation indicator and a field called a fragment counter.

[0084] The Fragmentation indicator indicates whether the data stored in the payload of an MMT packet is a fragment of a Data unit, and if so, whether the fragment is the first or last fragment of the Data unit, or a fragment that is neither the first nor the last. In other words, the Fragmentation indicator included in the header of a certain packet is identification information that indicates whether (1) only the packet in question is included in the Data unit, which is the basic data unit, (2) the Data unit is divided into multiple packets and stored, and the packet in question is the first packet of the Data unit, (3) the Data unit is divided into multiple packets and stored, and the packet in question is a packet other than the first or last packet of the Data unit, or (4) the Data unit is divided into multiple packets and stored, and the packet in question is the last packet of the Data unit.

[0085] The fragment counter is an index number that indicates which fragment in the data unit the data stored in the MMT packet corresponds to.

[0086] Therefore, by transmitting device 100 setting samples in MMT to Data units and setting data before a slice segment and each slice segment to fragment units of Data units, receiving device 200 can identify the type of data stored in the payload using information included in the header of the MMT packet. That is, demultiplexing unit 203 can generate input data for each of decoding units 204A to 204D by referring to the header of the MMT packet.

[0087] FIG. 11 is a diagram showing an example in which a sample is set to a data unit, and data before a slice segment and a slice segment are packetized as fragments of the data unit.

[0088] The data before the slice segment and the slice segment are divided into five fragments, fragment #1 to fragment #5. Each fragment is stored in a separate MMT packet. At this time, the values ​​of the fragmentation indicator and fragment counter included in the header of the MMT packet are as shown in the figure.

[0089] For example, the Fragment indicator is a 2-bit binary value. The Fragment indicator of MMT packet #1, which is the head of the Data unit, the Fragment indicator of MMT packet #5, which is the last packet, and the Fragment indicators of MMT packet #2 to MMT packet #4, which are the packets in between, are all set to different values. Specifically, the Fragment indicator of MMT packet #1, which is the head of the Data unit, is set to 01, the Fragment indicator of MMT packet #5, which is the last packet, is set to 11, and the Fragment indicators of MMT packet #2 to MMT packet #4, which are the packets in between, are set to 10. Note that if a Data unit contains only one MMT packet, the Fragment indicator is set to 00.

[0090] In addition, the fragment counter is 4 in MMT packet #1, which is the total number of fragments, 5, minus 1, and is decremented by 1 in subsequent packets, reaching 0 in the final MMT packet #5.

[0091] Therefore, receiving apparatus 200 can identify an MMT packet that stores pre-slice segment data by using either the fragment indicator or the fragment counter. Also, receiving apparatus 200 can identify an MMT packet that stores the N-th slice segment by referring to the fragment counter.

[0092] The header of the MMT packet also includes the sequence number within the MPU of the Movie Fragment to which the Data Unit belongs, the sequence number of the MPU itself, and the sequence number within the Movie Fragment of the sample to which the Data Unit belongs. By referencing these, the demultiplexer 203 can uniquely determine the sample to which the Data Unit belongs.

[0093] Furthermore, since demultiplexing unit 203 can determine the index number of a fragment within a data unit from a fragment counter or the like, it can uniquely identify the slice segment stored in the fragment even if packet loss occurs. For example, even if fragment #4 shown in FIG. 11 cannot be acquired due to packet loss, demultiplexing unit 203 can determine that the fragment received next after fragment #3 is fragment #5, and can therefore correctly output slice segment 4 stored in fragment #5 to decoding unit 204D rather than decoding unit 204C.

[0094] Note that, when a transmission path that guarantees no packet loss is used, demultiplexing unit 203 can periodically process the arrived packets without referring to the header of the MMT packet to determine the type of data stored in the MMT packet or the index number of the slice segment. For example, when an access unit is transmitted using five MMT packets, including pre-slice data and four slice segments, receiving device 200 can sequentially acquire the pre-slice data and the data of the four slice segments by sequentially processing the received MMT packets after determining the pre-slice data of the access unit from which decoding is to be started.

[0095] A variation of packetization will now be described.

[0096] Slice segments do not necessarily have to be divided both horizontally and vertically within the plane of an access unit; as shown in Figure 1, they may be divided only horizontally or only vertically within the plane of an access unit.

[0097] Also, if the access unit is divided only horizontally, tiles do not need to be used.

[0098] Furthermore, the number of divisions within an access unit is arbitrary and is not limited to 4. However, the area sizes of slice segments and tiles must be equal to or greater than the lower limit of encoding standards such as H.265.

[0099] Transmitting device 100 may store identification information indicating the division method for an access unit in an MMT message, a TS descriptor, or the like. For example, information indicating the number of divisions in the horizontal and vertical directions within the plane may be stored. Alternatively, unique identification information may be assigned to the division method, such as dividing the access unit into two equal parts in each of the horizontal and vertical directions as shown in FIG. 3, or dividing the access unit into four equal parts in the horizontal direction as shown in FIG. 1. For example, if the access unit is divided as shown in FIG. 3, the identification information indicates mode 1, and if the access unit is divided as shown in FIG. 1, the identification information indicates mode 1.

[0100] Furthermore, information indicating constraints on coding conditions related to the intra-plane division method may be included in the multiplexing layer. For example, information indicating that one slice segment is composed of one tile may be used. Alternatively, information indicating that reference blocks when performing motion compensation during decoding of a slice segment or tile are limited to slice segments or tiles at the same position within the screen, or limited to blocks within a predetermined range in adjacent slice segments may be used.

[0101] Furthermore, the transmitting device 100 may switch whether to divide an access unit into a plurality of slice segments depending on the resolution of the video. For example, the transmitting device 100 may not perform intra-frame division when the video to be processed has a resolution of 4K2K, but may divide the access unit into four when the video to be processed has a resolution of 8K4K. By predefining the division method for 8K4K video, the receiving device 200 can determine whether to perform intra-frame division and the division method by acquiring the resolution of the video to be received, and switch the decoding operation.

[0102] Furthermore, receiving device 200 can detect whether or not a frame is fragmented by referring to the header of the MMT packet. For example, if an access unit is not fragmented, if the MMT Data unit is set to Sample, the Data unit is not fragmented. Therefore, receiving device 200 can determine that an access unit is not fragmented if the value of the Fragment counter included in the header of the MMT packet is always zero. Alternatively, receiving device 200 may detect whether the value of the Fragmentation indicator is always 01. Receiving device 200 can also determine that an access unit is not fragmented if the value of the Fragmentation indicator is always 01.

[0103] Furthermore, the receiving device 200 can also handle a case where the number of divisions in a plane of an access unit does not match the number of decoding units. For example, if the receiving device 200 includes two decoding units 204A and 204B that can decode 8K2K encoded data in real time, the demultiplexing unit 203 outputs two of the four slice segments that make up the 8K4K encoded data to the decoding unit 204A.

[0104] Fig. 12 is a diagram showing an example of operation when MMT packetized data as shown in Fig. 8 is input to two decoding units 204A and 204B. Here, it is desirable that receiving device 200 be able to integrate and output the decoding results of decoding units 204A and 204B as they are. Therefore, demultiplexing unit 203 selects slice segments to output to each of decoding units 204A and 204B so that the decoding results of each of decoding units 204A and 204B are spatially continuous.

[0105] Furthermore, the demultiplexing unit 203 may select a decoding unit to use depending on the resolution or frame rate of the encoded data of the video. For example, if the receiving device 200 is equipped with four 4K2K decoding units, and the resolution of the input image is 8K4K, the receiving device 200 performs the decoding process using all four decoding units. Furthermore, if the resolution of the input image is 4K2K, the receiving device 200 performs the decoding process using only one decoding unit. Alternatively, even if the plane is divided into four, if 8K4K can be decoded in real time by a single decoding unit, the demultiplexing unit 203 integrates all the division units and outputs them to a single decoding unit.

[0106] Furthermore, the receiving device 200 may determine the decoding unit to be used taking the frame rate into consideration. For example, if the receiving device 200 is equipped with two decoding units, each of which has an upper limit of 60 fps for the frame rate that can be decoded in real time when the resolution is 8K4K, there may be a case where 120 fps coded data in 8K4K is input. In this case, if the plane is configured with four division units, slice segment 1 and slice segment 2 are input to the decoding unit 204A, and slice segment 3 and slice segment 4 are input to the decoding unit 204B, as in the example of FIG. 12. Each of the decoding units 204A and 204B can decode up to 120 fps in real time at 8K2K (resolution half that of 8K4K), and therefore the decoding process is performed by these two decoding units 204A and 204B.

[0107] Furthermore, even if the resolution and frame rate are the same, the processing amount differs depending on the profile or level of the encoding method, or the encoding method itself, such as H.264 or H.265. Therefore, the receiving device 200 may select a decoder to use based on this information. Note that if the receiving device 200 is unable to decode all of the encoded data received via broadcasting or communication, or if it is unable to decode all of the slice segments or tiles constituting a region selected by the user, the receiving device 200 may automatically determine slice segments or tiles that can be decoded within the processing range of the decoding device. Alternatively, the receiving device 200 may provide a user interface that allows the user to select a region to decode. In this case, the receiving device 200 may display a warning message indicating that all regions cannot be decoded, or may display information indicating the number of regions, slice segments, or tiles that can be decoded.

[0108] The above method can also be applied to cases where MMT packets storing slice segments of the same coded data are transmitted and received using multiple transmission paths such as broadcasting and communication.

[0109] Furthermore, the transmitting device 100 may perform encoding so that the slice segments overlap each other to make the boundaries between division units less noticeable. In the example shown in FIG. 13, an 8K4K picture is divided into four slice segments 1 to 4. Each of slice segments 1 to 3 is, for example, 8K x 1.1K, and slice segment 4 is 8K x 1K. Adjacent slice segments overlap each other. This allows efficient motion compensation during encoding at the boundaries of the four divisions indicated by the dotted lines, improving the image quality of the boundary areas. In this way, image quality degradation at the boundary areas is reduced.

[0110] In this case, the display unit 205 extracts an 8K×1K region from the 8K×1.1K region and integrates the obtained regions. Note that the transmitting device 100 may include information indicating whether the slice segments are coded with overlapping and the extent of the overlap in the multiplexing layer or coded data and transmit the information separately.

[0111] Note that a similar approach can be applied when tiles are used.

[0112] The following describes the flow of operations of the transmission device 100. FIG.

[0113] First, the encoding unit 101 divides a picture (access unit) into a plurality of slice segments (tiles), which are a plurality of regions (S101). Next, the encoding unit 101 generates encoded data corresponding to each of the plurality of slice segments by encoding each of the plurality of slice segments so that each of the plurality of slice segments can be decoded independently (S102). Note that the encoding unit 101 may encode the plurality of slice segments using a single encoding unit, or may encode the plurality of slice segments in parallel using a plurality of encoding units.

[0114] Next, multiplexing unit 102 multiplexes the plurality of coded data generated by coding unit 101 by storing the plurality of coded data in a plurality of MMT packets (S103). Specifically, as shown in FIGS. 8 and 9, multiplexing unit 102 stores the plurality of coded data in a plurality of MMT packets so that coded data corresponding to different slice segments are not stored in one MMT packet. Furthermore, as shown in FIG. 8, multiplexing unit 102 stores control information commonly used for all decoding units in a picture in MMT packet #1, which is different from the plurality of MMT packets #2 to #5 in which the plurality of coded data are stored. Here, the control information includes at least one of an access unit delimiter, an SPS, a PPS, and an SEI.

[0115] Note that multiplexing unit 102 may store the control information in the same MMT packet as one of the plurality of MMT packets in which the plurality of encoded data are stored. For example, as shown in Fig. 9, multiplexing unit 102 may store the control information in the first MMT packet (MMT packet #1 in Fig. 9) of the plurality of MMT packets in which the plurality of encoded data are stored.

[0116] Finally, the transmitting device 100 transmits the multiple MMT packets. Specifically, the modulating unit 103 modulates the data obtained by multiplexing, and the transmitting unit 104 transmits the modulated data (S104).

[0117] Fig. 15 is a block diagram showing an example of the configuration of receiving device 200, and is a diagram showing in detail the configuration of demultiplexing unit 203 and the subsequent stages shown in Fig. 7. As shown in Fig. 15, receiving device 200 further includes a decoding command unit 206. In addition, demultiplexing unit 203 includes a type discrimination unit 211, a control information acquisition unit 212, a slice information acquisition unit 213, and a decoded data generation unit 214.

[0118] The following describes the flow of operation of receiving apparatus 200. Fig. 16 is a flowchart showing an example of operation of receiving apparatus 200. Here, the operation for one access unit is shown. When decoding processing for multiple access units is executed, the processing of this flowchart is repeated.

[0119] First, the receiving device 200 receives, for example, a plurality of packets (MMT packets) generated by the transmitting device 100 (S201).

[0120] Next, the type determination unit 211 analyzes the header of the received packet to obtain the type of encoded data stored in the received packet (S202).

[0121] Next, the type determination unit 211 determines whether the data stored in the received packet is pre-slice segment data or slice segment data, based on the type of the acquired coded data (S203).

[0122] If the data stored in the received packet is pre-slice segment data (Yes in S203), the control information acquisition unit 212 acquires the pre-slice segment data of the access unit to be processed from the payload of the received packet and stores the pre-slice segment data in memory (S204).

[0123] On the other hand, if the data stored in the received packet is data of a slice segment (No in S203), the receiving device 200 uses the header information of the received packet to determine which of the multiple areas the data stored in the received packet is encoded data for. Specifically, the slice information acquisition unit 213 analyzes the header of the received packet to acquire the index number Idx of the slice segment stored in the received packet (S205). Specifically, the index number Idx is the index number within the Movie Fragment of the access unit (sample in MMT).

[0124] The process of step S205 may be performed together with step S202.

[0125] Next, the decoding data generation unit 214 determines a decoding unit that will decode the slice segment (S206). Specifically, the index number Idx and a plurality of decoding units are associated in advance, and the decoding data generation unit 214 determines the decoding unit that will decode the slice segment as the decoding unit that corresponds to the index number Idx acquired in step S205.

[0126] 12, the decoding data generation unit 214 may determine the decoding unit that decodes the slice segment based on at least one of the resolution of the access unit (picture), the division method of the access unit into multiple slice segments (tiles), and the processing capabilities of the multiple decoding units included in the receiving device 200. For example, the decoding data generation unit 214 determines the division method of the access unit based on identification information in a descriptor such as an MMT message or a TS section.

[0127] Next, the decoded data generation unit 214 generates multiple pieces of input data (combined data) to be input to multiple decoding units by combining control information, which is included in one of the multiple packets and is used commonly for all decoding units in the picture, with each of the multiple pieces of encoded data for the multiple slice segments. Specifically, the decoded data generation unit 214 acquires slice segment data from the payload of the received packet. The decoded data generation unit 214 generates input data to the decoding unit determined in step S206 by combining the pre-slice segment data stored in memory in step S204 with the acquired slice segment data (S207).

[0128] After step S204 or S207, if the data of the received packet is not the final data of the access unit (No in S208), the processes from step S201 onward are performed again. That is, the above processes are repeated until input data for the multiple decoding units 204A to 204D corresponding to all slice segments included in the access unit are generated.

[0129] The timing at which the packets are received is not limited to the timing shown in FIG. 16, and a plurality of packets may be received in advance or sequentially and stored in a memory or the like.

[0130] On the other hand, if the data of the received packet is the final data of the access unit (Yes in S208), the decoding command unit 206 outputs the plurality of input data generated in step S207 to the corresponding decoding units 204A to 204D (S209).

[0131] Next, the multiple decoding units 204A to 204D decode the multiple pieces of input data in parallel in accordance with the DTS of the access unit, thereby generating multiple decoded images (S210).

[0132] Finally, the display unit 205 generates a display image by combining the decoded images generated by the decoding units 204A to 204D, and displays the display image in accordance with the PTS of the access unit (S211).

[0133] The receiving device 200 acquires the DTS and PTS of an access unit by analyzing the header information of an MPU or the payload data of an MMT packet that stores the header information of a Movie Fragment. Furthermore, when TS is used as the multiplexing method, the receiving device 200 acquires the DTS and PTS of an access unit from the header of a PES packet. When RTP is used as the multiplexing method, the receiving device 200 acquires the DTS and PTS of an access unit from the header of an RTP packet.

[0134] Furthermore, when integrating the decoding results of the multiple decoders, the display unit 205 may perform filtering, such as deblocking filtering, at the boundaries between adjacent division units. Note that, since filtering is not necessary when displaying the decoding results of a single decoder, the display unit 205 may switch the filtering process depending on whether or not to perform filtering at the boundaries between the decoding results of the multiple decoders. Whether filtering is necessary may be determined in advance depending on whether or not division is performed. Alternatively, information indicating whether filtering is necessary may be separately stored in the multiplexing layer. Information necessary for filtering, such as filter coefficients, may be stored in the SPS, PPS, SEI, or slice segment. The decoding units 204A to 204D or the demultiplexing unit 203 acquire this information by analyzing the SEI and output the acquired information to the display unit 205. The display unit 205 performs filtering using this information. Note that, if this information is stored in the slice segment, it is preferable that the decoding units 204A to 204D acquire this information.

[0135] In the above description, an example was shown in which the types of data stored in a fragment are two types: pre-slice segment data and slice segments. However, the types of data may be three or more. In this case, in step S203, the cases are classified according to the type.

[0136] Furthermore, when the data size of a slice segment is large, transmitting device 100 may fragment the slice segment and store the fragmented slice segment in an MMT packet. That is, transmitting device 100 may fragment the data before the slice segment and the slice segment. In this case, if the access unit and the data unit are set equal as in the packetization example shown in FIG. 11, the following problem occurs.

[0137] For example, if slice segment 1 is divided into three fragments, slice segment 1 is divided and transmitted as three packets with fragment counter values ​​of 1 to 3. Furthermore, from slice segment 2 onwards, the fragment counter value becomes 4 or greater, and it becomes impossible to associate the fragment counter value with the data stored in the payload. Therefore, receiving device 200 cannot identify the packet that stores the first data of the slice segment from the information in the header of the MMT packet.

[0138] In such a case, receiving apparatus 200 may analyze data in the payload of the MMT packet to identify the start position of the slice segment. Here, there are two types of formats for storing NAL units in the multiplexing layer in H.264 or H.265: a format called a byte stream format in which a start code consisting of a specific bit string is added immediately before the NAL unit header, and a format called an NAL size format in which a field indicating the size of the NAL unit is added.

[0139] The byte stream format is used in MPEG-2 systems and RTP, etc. The NAL size format is used in MP4, and DASH and MMT, which use MP4, etc.

[0140] When the byte stream format is used, the receiving device 200 analyzes whether the leading data of the packet matches the start code. If the leading data of the packet matches the start code, the receiving device 200 can detect whether the data included in the packet is data of a slice segment by obtaining the type of NAL unit from the NAL unit header that follows.

[0141] On the other hand, in the case of the NAL size format, the receiving device 200 cannot detect the start position of the NAL unit based on the bit string. Therefore, in order to obtain the start position of the NAL unit, the receiving device 200 needs to shift the pointer by reading data by the size of the NAL unit, starting from the first NAL unit of the access unit.

[0142] However, if the size of the subsample unit is indicated in the header of an MPU or Movie Fragment in MMT, and the subsample corresponds to pre-slice data or a slice segment, receiving device 200 can identify the start position of each NAL unit based on the size information of the subsample. Therefore, transmitting device 100 may include information indicating whether subsample unit information exists in an MPU or Movie Fragment in information that receiving device 200 acquires when starting to receive data, such as an MPT in MMT.

[0143] Note that MPU data is an extension of the MP4 format. MP4 has a mode in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and a mode in which they cannot be stored. Information for identifying this mode is indicated as the entry name of SampleEntry. When a mode in which parameter sets can be stored is used and a parameter set is included in a sample, receiving device 200 acquires the parameter set by the method described above.

[0144] On the other hand, when a mode that cannot store parameter sets is used, the parameter sets are stored as Decoder Specific Information in SampleEntry, or are stored using a stream for the parameter sets. Here, since streams for parameter sets are not generally used, it is desirable that the transmitting device 100 stores the parameter sets in Decoder Specific Information. In this case, the receiving device 200 analyzes the SampleEntry transmitted as metadata of the MPU or metadata of the Movie Fragment in the MMT packet, and acquires the parameter sets referenced by the access unit.

[0145] When a parameter set is stored as sample data, the receiving device 200 can acquire the parameter set required for decoding by referring only to the sample data without referring to the SampleEntry. In this case, the transmitting device 100 does not need to store the parameter set in the SampleEntry. This allows the transmitting device 100 to use the same SampleEntry in different MPUs, thereby reducing the processing load on the transmitting device 100 when generating an MPU. Another advantage is that the receiving device 200 does not need to refer to the parameter set in the SampleEntry.

[0146] Alternatively, transmitting device 100 may store one default parameter set in SampleEntry, and store the parameter set referenced by the access unit in the sample data. In conventional MP4, parameter sets were typically stored in SampleEntry, so there was a possibility that some receiving devices would stop playback if no parameter set existed in SampleEntry. This problem can be solved by using the above method.

[0147] Alternatively, transmitting device 100 may store a parameter set in sample data only when a parameter set different from the default parameter set is used.

[0148] In both modes, it is possible to store parameter sets in SampleEntry, so transmitting device 100 may always store parameter sets in VisualSampleEntry, and receiving device 200 may always acquire parameter sets from VisualSampleEntry.

[0149] In the MMT standard, MP4 header information such as Moov and Moof is transmitted as MPU metadata or movie fragment metadata, but the transmitting device 100 does not necessarily have to transmit MPU metadata and movie fragment metadata. Furthermore, the receiving device 200 can also determine whether an SPS and a PPS are stored in the sample data based on the service, asset type, or whether MPU meta is transmitted in the ARIB (Association of Radio Industries and Businesses) standard.

[0150] FIG. 17 is a diagram showing an example in which data before a slice segment and each slice segment are set to different data units.

[0151] 17, the data sizes of the data before the slice segment and slice segments 1 to 4 are Length #1 to Length #5, respectively. The field values ​​of the Fragmentation indicator, Fragment counter, and Offset included in the header of the MMT packet are as shown in the figure.

[0152] Here, Offset is offset information indicating the bit length (offset) from the beginning of the coded data of the sample (access unit or picture) to which the payload data belongs to to the first byte of the payload data (coded data) included in the MMT packet. Note that although the value of the Fragment counter is explained as starting from a value obtained by subtracting 1 from the total number of fragments, it may start from another value.

[0153] Fig. 18 is a diagram showing an example of when a Data unit is fragmented. In the example shown in Fig. 18, slice segment 1 is divided into three fragments, which are stored in MMT packets #2 to #4, respectively. In this case, if the data size of each fragment is Length #2_1 to Length #2_3, respectively, the values ​​of each field are as shown in the diagram.

[0154] In this way, when a data unit such as a slice segment is set to Data unit, the start of the access unit and the start of the slice segment can be determined as follows based on the field values ​​of the MMT packet header.

[0155] The start of the payload in a packet with an Offset value of 0 is the start of the access unit.

[0156] The start of the payload of a packet in which the Offset value is a value other than 0 and the Fragmentation indicator value is 00 or 01 is the start of the slice segment.

[0157] In addition, if no fragmentation of the data unit occurs and no packet loss occurs, the receiving device 200 can identify the index number of the slice segment to be stored in the MMT packet based on the number of slice segments obtained after detecting the beginning of the access unit.

[0158] Similarly, even when the data unit of the data before the slice segment is fragmented, the receiving device 200 can detect the beginning of the access unit and the slice segment.

[0159] Furthermore, even when packet loss occurs or when the SPS, PPS, and SEI included in the data before the slice segment are set in different Data units, receiving device 200 can identify the MMT packet that stores the start data of the slice segment based on the analysis result of the MMT header, and then analyze the header of the slice segment to identify the start position of the slice segment or tile within the picture (access unit). The amount of processing involved in analyzing the slice header is small, so the processing load is not a problem.

[0160] In this way, each of the plurality of coded data of the plurality of slice segments is in one-to-one correspondence with a basic data unit, which is a unit of data stored in one or more packets. Also, each of the plurality of coded data is stored in one or more MMT packets.

[0161] The header information of each MMT packet includes a fragmentation indicator (identification information) and an offset (offset information).

[0162] Receiving device 200 determines that the start of payload data included in a packet having header information including a Fragmentation indicator whose value is 00 or 01 is the start of the encoded data of each slice segment. Specifically, receiving device 200 determines that the start of payload data included in a packet having header information including an Offset whose value is not 0 and a Fragmentation indicator whose value is 00 or 01 is the start of the encoded data of each slice segment.

[0163] 17, the start of a Data Unit is either the start of an access unit or the start of a slice segment, and the value of the Fragmentation indicator is 00 or 01. Furthermore, receiving device 200 can also detect the start of an access unit or the start of a slice segment without referring to the Offset by referring to the type of NAL unit and determining whether the start of a Data Unit is an access unit delimiter or a slice segment.

[0164] In this way, transmitting device 100 performs packetization so that the beginning of the NAL unit always starts at the beginning of the payload of the MMT packet, and thus receiving device 200 can detect the beginning of an access unit or slice segment by analyzing the fragmentation indicator and the NAL unit header, even when the data before the slice segment is divided into multiple Data units. The type of the NAL unit is present in the first byte of the NAL unit header. Therefore, when analyzing the header portion of the MMT packet, receiving device 200 can acquire the type of the NAL unit by analyzing an additional byte of data.

[0165] In the case of audio, receiving apparatus 200 only needs to be able to detect the beginning of an access unit, and can make a determination based on whether the value of the fragmentation indicator is 00 or 01.

[0166] Furthermore, as described above, when storing coded data that has been coded so that it can be divided and decoded in PES packets of MPEG-2 TS, the transmitting device 100 can use a data alignment descriptor. An example of a method for storing coded data in PES packets will be described in detail below.

[0167] For example, in HEVC, the transmission device 100 can indicate whether the data stored in the PES packet is an access unit, a slice segment, or a tile by using a data alignment descriptor. The alignment types in HEVC are specified as follows:

[0168] Alignment type=8 indicates an HEVC slice segment. Alignment type=9 indicates an HEVC slice segment or access unit. Alignment type=12 indicates an HEVC slice segment or tile.

[0169] Therefore, the transmitting device 100 can indicate that the data of the PES packet is either a slice segment or pre-slice segment data by using, for example, type 9. Since a type indicating a slice rather than a slice segment is also separately defined, the transmitting device 100 may use a type indicating a slice rather than a slice segment.

[0170] Furthermore, the DTS and PTS included in the header of a PES packet are set only in the PES packet that contains the first data of an access unit. Therefore, if the type is 9 and the PES packet contains a DTS or PTS field, receiving device 200 can determine that the PES packet stores the entire access unit or the first division unit of the access unit.

[0171] Furthermore, the transmitting device 100 may enable the receiving device 200 to distinguish the data contained in the packet using a field such as transport_priority, which indicates the priority of a TS packet that stores a PES packet containing the first data of an access unit. The receiving device 200 may also determine the data contained in the packet by analyzing whether the payload of the PES packet is an access unit delimiter. Furthermore, the data_alignment_indicator in the PES packet header indicates whether data is stored in the PES packet according to these types. If this flag (data_alignment_indicator) is set to 1, it is guaranteed that the data stored in the PES packet complies with the type indicated in the data alignment descriptor.

[0172] Furthermore, the transmitting device 100 may use the data alignment descriptor only when PES packetizing is performed in units that can be divided and decoded, such as slice segments. As a result, if the data alignment descriptor is present, the receiving device 200 can determine that the coded data has been PES packetized in units that can be divided and decoded, and if the data alignment descriptor is not present, the receiving device 200 can determine that the coded data has been PES packetized in units of access units. Note that if the data_alignment_indicator is set to 1 and the data alignment descriptor is not present, the MPEG-2 TS standard specifies that the unit of PES packetization is the access unit.

[0173] If the PMT includes a data alignment descriptor, the receiving device 200 determines that the PES packetization is performed in units that can be divided and decoded, and can generate input data for each decoder based on the packetized units. If the PMT does not include a data alignment descriptor and the receiving device 200 determines that parallel decoding of the encoded data is necessary based on program information or other descriptor information, the receiving device 200 generates input data for each decoder by analyzing the slice header of the slice segment, etc. If the encoded data can be decoded by a single decoder, the receiving device 200 decodes the data of the entire access unit by that decoder. If information indicating whether the encoded data is composed of units that can be divided and decoded, such as slice segments or tiles, is separately indicated by a descriptor in the PMT, the receiving device 200 may determine whether the encoded data can be parallel decoded based on the analysis result of the descriptor.

[0174] Furthermore, since the DTS and PTS included in the header of a PES packet are set only in the PES packet containing the first data of an access unit, when an access unit is divided and packetized as PES packets, the second and subsequent PES packets do not contain information indicating the DTS and PTS of the access unit. Therefore, when decoding processes are performed in parallel, each of the decoding units 204A to 204D and the display unit 205 uses the DTS and PTS stored in the header of the PES packet containing the first data of an access unit.

[0175] (Embodiment 2) In the second embodiment, a method for storing data in the NAL size format in an MPU based on the MP4 format in MMT will be described. Note that, although a storage method in an MPU used in MMT will be described below as an example, such a storage method can also be applied to DASH, which is also based on the MP4 format.

[0176] [Storage method in MPU] In the MP4 format, multiple access units are stored together in a single MP4 file. The MPU used in MMT stores data for each media in a single MP4 file, and the data can contain any number of access units. Since an MPU is a unit that can be decoded independently, for example, an MPU stores access units in units of GOPs.

[0177] 19 is a diagram showing the structure of an MPU. At the beginning of an MPU are ftyp, mmpu, and moov, which are collectively defined as MPU metadata. moov stores initialization information common to the file and an MMT hint track.

[0178] Additionally, moof stores initialization information and size for each sample and subsample, information (sample_duration, sample_size, sample_composition_time_offset) that can identify the presentation time (PTS) and decoding time (DTS), and data_offset that indicates the position of the data.

[0179] Furthermore, each of the multiple access units is stored as a sample in mdat (mdat box). Data in moof and mdat excluding samples is defined as movie fragment metadata (hereinafter referred to as MF metadata), and sample data in mdat is defined as media data.

[0180] Fig. 20 is a diagram showing the structure of MF metadata. As shown in Fig. 20, the MF metadata is more specifically made up of the type, length, and data of a moof box (moof), and the type and length of an mdat box (mdat).

[0181] When storing access units in MP4 data, there are two modes: one in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and one in which they cannot be stored.

[0182] In the non-storable mode, the parameter set is stored in the Decoder Specific Information of the SampleEntry in the moov, and in the storable mode, the parameter set is included in the sample.

[0183] The MPU metadata, MF metadata, and media data are each stored in an MMT payload, and a fragment type (FT) is stored in the header of the MMT payload as an identifier for identifying these data. FT=0 indicates MPU metadata, FT=1 indicates MF metadata, and FT=2 indicates media data.

[0184] Note that, although Fig. 19 illustrates an example in which MPU metadata units and MF metadata units are stored as data units in the MMT payload, units such as ftyp, mmpu, moov, and moof may also be stored as data units in the MMT payload in data unit units. Similarly, Fig. 19 illustrates an example in which sample units are stored as data units in the MMT payload. However, data units may be configured in sample units or NAL unit units, and such data units may be stored in the MMT payload in data unit units. Such data units may also be further fragmented and stored in the MMT payload.

[0185] [Conventional transmission methods and issues] Conventionally, when multiple access units are encapsulated in the MP4 format, moov and moof are created when all samples to be stored in the MP4 are available.

[0186] When transmitting MP4 format in real time via broadcasting, for example, if the samples stored in one MP4 file are in GOP units, delays occur due to encapsulation because the moov and moof are created after the GOP-unit time samples are accumulated. This encapsulation on the transmitting side always increases the end-to-end delay by the GOP unit time. This makes it difficult to provide services in real time, and leads to degradation of the service for viewers, especially when transmitting live content.

[0187] Fig. 21 is a diagram for explaining the data transmission order. When MMT is applied to broadcasting, as shown in Fig. 21(a), if data is loaded onto MMT packets and transmitted in the order of the MPU configuration (transmitting MMT packets #1, #2, #3, #4, #5, and #6 in that order), a delay occurs in the transmission of the MMT packets due to encapsulation.

[0188] To prevent this delay due to encapsulation, a method has been proposed in which MPU header information such as MPU metadata and MF metadata is not sent (packets #1 and #2 are not sent, and packets #3 to #6 are sent in this order), as shown in (b) of Figure 21. Another possible method is to send media data first without waiting for the creation of MPU header information, and then send the MPU header information after the media data has been sent (sending packets #3 to #6, #1, and #2 in that order), as shown in (c) of Figure 20.

[0189] If the MPU header information is not transmitted, the receiving device decodes without using the MPU header information. If the MPU header information is delayed relative to the media data, the receiving device waits until it obtains the MPU header information before decoding.

[0190] However, conventional MP4-compliant receiving devices are not guaranteed to be able to decode without using MPU header information. Furthermore, if a receiving device performs special processing to decode without using the MPU header, using conventional transmission methods can complicate the decoding process, potentially making real-time decoding difficult. Furthermore, if a receiving device waits for MPU header information before decoding, media data must be buffered until the receiving device acquires the header information. However, because no buffer model is specified, decoding is not guaranteed.

[0191] Therefore, the transmitting device according to the second embodiment stores only common information in the MPU metadata, as shown in (d) of Fig. 20, so that the MPU metadata is transmitted before the media data.The transmitting device according to the second embodiment then transmits the MF metadata, which is generated with a delay, after the media data.This provides a transmitting method or receiving method that can guarantee decoding of the media data.

[0192] The following describes the reception method when using each of the transmission methods (a) to (d) in FIG.

[0193] In each transmission method shown in FIG. 21, first, MPU data is configured in the order of MPU metadata, MFU metadata, and media data.

[0194] After constructing the MPU data, if the transmitting device transmits data in the order of MPU metadata, MF metadata, and media data, as shown in (a) of Figure 21, the receiving device can perform decoding using either of the following methods (A-1) and (A-2).

[0195] (A-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data using the MPU header information.

[0196] (A-2) The receiving device decodes the media data without using the MPU header information.

[0197] Although these methods all incur delays due to encapsulation on the transmitting side, they have the advantage that the receiving device does not need to buffer the media data to obtain the MPU header. If buffering is not performed, there is no need to install memory for buffering, and buffering delays do not occur. Furthermore, method (A-1) is applicable to conventional receiving devices because it uses MPU header information for decoding.

[0198] When the transmitting device transmits only media data as shown in (b) of FIG. 21, the receiving device can perform decoding using the following method (B-1).

[0199] (B-1) The receiving device decodes the media data without using the MPU header information.

[0200] Although not shown, if MPU metadata is transmitted before the transmission of the media data in (b) of FIG. 21, decoding can be performed using the following method (B-2).

[0201] (B-2) The receiving device decodes the media data using the MPU metadata.

[0202] The advantages of both methods (B-1) and (B-2) above are that there is no delay due to encapsulation on the sending side, and there is no need to buffer media data to obtain the MPU header. However, since neither method (B-1) nor (B-2) performs decoding using MPU header information, special processing may be required for decoding.

[0203] When the transmitting device transmits data in the order of media data, MPU metadata, and MF metadata, as shown in (c) of Figure 21, the receiving device can perform decoding using either of the following methods (C-1) and (C-2).

[0204] (C-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data.

[0205] (C-2) The receiving device decodes the media data without using the MPU header information.

[0206] When the above method (C-1) is used, it is necessary to buffer the media data in order to obtain the MPU header information. On the other hand, when the above method (C-2) is used, it is not necessary to buffer the media data in order to obtain the MPU header information.

[0207] In addition, neither method (C-1) nor (C-2) causes delays due to encapsulation on the sending side. Furthermore, method (C-2) may require special processing because it does not use MPU header information.

[0208] When the transmitting device transmits data in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Figure 21, the receiving device can perform decoding using either of the following methods (D-1) and (D-2).

[0209] (D-1) After acquiring the MPU metadata, the receiving device further acquires the MF metadata, and then decodes the media data.

[0210] (D-2) After acquiring the MPU metadata, the receiving device decodes the media data without using the MF metadata.

[0211] When the above method (D-1) is used, it is necessary to buffer the media data in order to obtain the MF metadata, but when the above method (D-2) is used, it is not necessary to buffer the media data in order to obtain the MF metadata.

[0212] The above method (D-2) does not perform decoding using MF metadata, and therefore may require special processing.

[0213] As described above, when decoding is possible using MPU metadata and MF metadata, there is an advantage that decoding can also be performed by a conventional MP4 receiving device.

[0214] 21, the MPU data is structured in the order of MPU metadata, MFU metadata, and media data, and in moof, position information (offset) for each sample and subsample is determined based on this structure. Also, the MF metadata includes data other than the media data in the mdat box (box size and type).

[0215] Therefore, when a receiving device identifies media data based on MF metadata, the receiving device reconstructs the data in the order in which the MPU data was constructed, regardless of the order in which the data was transmitted, and then decodes it using the moov in the MPU metadata or the moof in the MF metadata.

[0216] In FIG. 21, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data, but the MPU data may be configured in an order different from that shown in FIG. 21, and the position information (offset) may be determined.

[0217] For example, MPU data may be configured in the order of MPU metadata, media data, and MF metadata, and negative position information (offset) may be indicated in the MF metadata. In this case, regardless of the order in which the data is transmitted, the receiving device reconstructs the data in the order in which the MPU data was configured on the transmitting side, and then performs decoding using moov or moof.

[0218] The transmitting device may signal information indicating the order in which MPU data is constructed, and the receiving device may reconstruct the data based on the signaled information.

[0219] As described above, the receiving device receives packetized MPU metadata, packetized media data (sample data), and packetized MF metadata in this order, as shown in (d) of Fig. 21. Here, the MPU metadata is an example of first metadata, and the MF metadata is an example of second metadata.

[0220] Next, the receiving device reconstructs MPU data (an MP4 format file) including the received MPU metadata, the received MF metadata, and the received sample data. Then, the receiving device decodes the sample data included in the reconstructed MPU data using the MPU metadata and MF metadata. The MF metadata is metadata including data that can be generated only after the sample data is generated on the transmitting side (for example, the length stored in the mbox).

[0221] The operation of the receiving device is more specifically performed by each component of the receiving device. For example, the receiving device includes a receiving unit that receives the data, a reconstructing unit that reconstructs the MPU data, and a decoding unit that decodes the MPU data. The receiving unit, generating unit, and decoding unit are each realized by a microcomputer, a processor, a dedicated circuit, etc.

[0222] [Method of decrypting without using header information] Next, a method for decoding without using header information will be described. Here, a method for decoding without using header information in a receiving device will be described, regardless of whether header information is sent on the transmitting side or not. That is, this method is applicable to any of the transmission methods described with reference to FIG. 21. However, some decoding methods are applicable only to specific transmission methods.

[0223] Figure 22 is a diagram showing an example of a method for decoding without using header information. Figure 22 shows only MMT payloads and MMT packets containing only media data, and does not show MMT payloads and MMT packets containing MPU metadata or MF metadata. In the following description of Figure 22, it is assumed that media data belonging to the same MPU are transmitted continuously. In addition, although an example will be described in which samples are stored in the payload as media data, in the following description of Figure 22, it goes without saying that NAL units or fragmented NAL units may be stored.

[0224] To decode media data, a receiving device must first obtain initialization information necessary for decoding. If the media is video, the receiving device must obtain initialization information for each sample, identify the start position of the MPU (a random access unit), and obtain the start positions of the samples and NAL units. The receiving device must also identify the decoding time (DTS) and presentation time (PTS) of each sample.

[0225] Therefore, the receiving device can perform decoding without using header information, for example, by using the following method: Note that when NAL unit units or units obtained by fragmenting NAL units are stored in the payload, "sample" in the following description can be read as "NAL unit in a sample."

[0226] <Random access (=identify the first sample of the MPU)> When header information is not transmitted, the receiving device can identify the first sample of the MPU using the following methods 1 and 2. Note that when header information is transmitted, method 3 can be used.

[0227] [Method 1] The receiving device acquires samples contained in MMT packets with 'RAP_flag=1' in the MMT packet header.

[0228] [Method 2] The receiving device acquires the sample with 'sample number=0' in the MMT payload header.

[0229] [Method 3] When at least one of MPU metadata and MF metadata is transmitted before or after the media data, the receiving device acquires samples contained in the MMT payload whose fragment type (FT) in the MMT payload header has been switched to media data.

[0230] In Methods 1 and 2, if a single payload contains a mixture of multiple samples belonging to different MPUs, it is impossible to determine which NAL unit is a random access point (RAP_flag = 1 or sample number = 0). For this reason, it is necessary to impose a constraint such as not mixing samples from different MPUs in a single payload, or, if a single payload contains a mixture of samples from different MPUs, to impose a constraint such as setting RAP_flag to 1 if the last (or first) sample is a random access point.

[0231] Furthermore, in order for the receiving device to obtain the start position of the NAL unit, it is necessary to shift the data read pointer by the size of the NAL unit, starting from the first NAL unit of the sample.

[0232] If the data is fragmented, the receiving device can identify the data unit by referring to the fragment_indicator and fragment_number.

[0233] <Determining the DTS of a sample> There are two methods for determining the DTS of a sample: Method 1 and Method 2 below.

[0234] [Method 1] The receiver determines the DTS of the first sample based on the prediction structure. However, this method requires analysis of the coded data, which may make real-time decoding difficult. Therefore, the following method 2 is preferable.

[0235] [Method 2] The receiving device separately transmits the DTS for the first sample and acquires the transmitted DTS for the first sample. Examples of methods for transmitting the DTS for the first sample include transmitting the DTS for the MPU's first sample using MMT-SI, or transmitting a DTS for each sample using the MMT packet header extension field. The DTS may be an absolute value or a relative value to the PTS. The transmitting side may also signal whether the DTS for the first sample is included.

[0236] In both Method 1 and Method 2, the DTS of subsequent samples is calculated assuming a fixed frame rate.

[0237] In addition to using an extension field, a method for storing the DTS for each sample in a packet header also includes storing the DTS of the sample contained in the MMT packet in the 32-bit NTP timestamp field in the MMT packet header. If the DTS cannot be expressed using the number of bits in one packet header (32 bits), the DTS may be expressed using multiple packet headers. Alternatively, the DTS may be expressed by combining the NTP timestamp field and extension field in the packet header. If DTS information is not included, a known value (e.g., ALL 0) is used.

[0238] <Determining the PTS of a sample> The receiving device obtains the PTS of the first sample from the MPU timestamp descriptor for each asset included in the MPU. The receiving device calculates the PTS of subsequent samples based on a fixed frame rate, using parameters such as POC that indicate the display order of the samples. In this way, transmission at a fixed frame rate is essential to calculate the DTS and PTS without using header information.

[0239] In addition, when MF metadata is transmitted, the receiving device can calculate the absolute values ​​of the DTS and PTS from the relative time information of the DTS and PTS from the first sample indicated in the MF metadata and the absolute value of the timestamp of the MPU first sample indicated in the MPU timestamp descriptor.

[0240] When analyzing the coded data and calculating the DTS and PTS, the receiving device may use SEI information included in the access unit.

[0241] <Initialization information (parameter set)> [For video] In the case of video, the parameter set is stored in the sample data. Furthermore, if the MPU metadata and MF metadata are not transmitted, it is guaranteed that the parameter set required for decoding can be obtained by referring only to the sample data.

[0242] Also, when MPU metadata is transmitted before media data, as in (a) and (d) of Figure 21, it may be specified that parameter sets are not stored in SampleEntry. In this case, the receiving device does not refer to the parameter sets in SampleEntry, but only refers to the parameter sets in the sample.

[0243] Furthermore, when MPU metadata is transmitted before media data, SampleEntry stores a parameter set common to the MPU or a default parameter set, and a receiving device may refer to the parameter set in SampleEntry and the parameter set in the sample. Storing a parameter set in SampleEntry enables decoding even on conventional receiving devices that cannot play back data unless a parameter set is present in SampleEntry.

[0244] [For audio] For audio, an LATM header is required for decoding, and in MP4, the LATM header must be included in the sample entry. However, if the header information is not transmitted, it is difficult for the receiving device to obtain the LATM header, so the LATM header is separately included in control information such as SI. The LATM header may also be included in a message, table, or descriptor. The LATM header may also be included in the sample.

[0245] The receiving device acquires the LATM header from the SI or the like before starting decoding, and starts decoding the audio. Alternatively, as shown in (a) and (d) of Figure 21, if the MPU metadata is transmitted before the media data, the receiving device can receive the LATM header before the media data. Therefore, if the MPU metadata is transmitted before the media data, decoding can be performed even using a conventional receiving device.

[0246] <Other> The transmission order and the type of transmission order may be notified as control information such as an MMT packet header, a payload header, or an MPT or other table, message, descriptor, etc. Note that the type of transmission order here refers to, for example, the four types of transmission order shown in (a) to (d) of Fig. 21, and an identifier for identifying each type may be stored in a location that can be obtained before decoding begins.

[0247] Furthermore, different transmission order types may be used for audio and video, or a common type may be used for audio and video. Specifically, for example, audio may be transmitted in the order of MPU metadata, MF metadata, and media data, as shown in (a) of Fig. 21, and video may be transmitted in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Fig. 21.

[0248] The above-described method allows a receiving device to decode without using header information. Also, if the MPU metadata is transmitted before the media data (see (a) and (d) in FIG. 21), decoding becomes possible even with a conventional receiving device.

[0249] In particular, by transmitting the MF metadata after the media data ((d) in FIG. 21), delay due to encapsulation does not occur, and decoding can be performed even by a conventional receiving device.

[0250] [Configuration and operation of transmitting device] Next, the configuration and operation of the transmission device will be described. Fig. 23 is a block diagram of the transmission device according to the second embodiment, and Fig. 24 is a flowchart of the transmission method according to the second embodiment.

[0251] As shown in FIG. 23, the transmitting device 15 includes an encoding unit 16, a multiplexing unit 17, and a transmitting unit .

[0252] The encoding unit 16 generates encoded data by encoding the video or audio to be encoded according to, for example, H.265 (S10).

[0253] The multiplexing unit 17 multiplexes (packetizes) the encoded data generated by the encoding unit 16 (S11). Specifically, the multiplexing unit 17 packetizes each of the sample data, MPU metadata, and MF metadata that make up an MP4 format file. The sample data is data obtained by encoding a video signal or an audio signal, the MPU metadata is an example of first metadata, and the MF metadata is an example of second metadata. Both the first metadata and the second metadata are metadata used to decode the sample data, but the difference between them is that the second metadata includes data that can be generated only after the sample data is generated.

[0254] Here, the data that can be generated only after the sample data is generated is, for example, data other than the sample data stored in mdat in MP4 format (data in the header of mdat, i.e., type and length shown in Fig. 20). Here, the second metadata only needs to include the length, which is at least a part of this data.

[0255] The transmitting unit 18 transmits the packetized MP4 format file (S12). The transmitting unit 18 transmits the MP4 format file, for example, by the method shown in (d) of Fig. 21. That is, the transmitting unit 18 transmits the packetized MPU metadata, packetized sample data, and packetized MF metadata in this order.

[0256] Each of the encoding unit 16, multiplexing unit 17, and transmitting unit 18 is realized by a microcomputer, a processor, a dedicated circuit, or the like.

[0257] [Configuration of receiving device] Next, the configuration and operation of the receiving device will be described. Fig. 25 is a block diagram of the receiving device according to the second embodiment.

[0258] As shown in FIG. 25, the receiving device 20 includes a packet filtering unit 21, a transmission order type discrimination unit 22, a random access unit 23, a control information acquisition unit 24, a data acquisition unit 25, a PTS, DTS calculation unit 26, an initialization information acquisition unit 27, a decoding command unit 28, a decoding unit 29, and a presentation unit 30.

[0259] [Receiver operation 1] First, an operation of the receiving device 20 for identifying the MPU start position and the NAL unit position when the media is video will be described. Fig. 26 is a flowchart of such an operation of the receiving device 20. Note that it is assumed here that the transmission order type of the MPU data is stored in the SI information by the transmitting device 15 (multiplexing unit 17).

[0260] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering and acquires the transmission order type of the MPU data (S21).

[0261] Next, the transmission order type determination unit 22 determines (discriminates) whether or not MPU header information (at least one of MPU metadata and MF metadata) is included in the data after packet filtering (S22). If the MPU header information is included (Yes in S22), the random access unit 23 detects that the fragment type of the MMT payload header has switched to media data, thereby identifying the MPU first sample (S23).

[0262] On the other hand, if the MPU header information is not included (No in S22), the random access unit 23 identifies the MPU first sample based on the RAP_flag in the MMT packet header or the sample number in the MMT payload header (S24).

[0263] Furthermore, the transmission order type determination unit 22 determines whether or not MF metadata is included in the packet-filtered data (S25). If it is determined that MF metadata is included (Yes in S25), the data acquisition unit 25 acquires NAL units by reading the NAL units based on the sample, subsample offset, and size information included in the MF metadata (S26). On the other hand, if it is determined that MF metadata is not included (No in S25), the data acquisition unit 25 acquires NAL units by reading data of the size of the NAL units in order from the first NAL unit of the sample (S27).

[0264] Note that even if it is determined in step S22 that MPU header information is included, receiving device 20 may identify the MPU first sample using the process of step S24 instead of step S23. Furthermore, if it is determined that MPU header information is included, the process of step S23 and the process of step S24 may be used together.

[0265] Furthermore, even if it is determined in step S25 that MF metadata is included, the receiving device 20 may acquire the NAL unit using the process of step S27 without using the process of step S26. Furthermore, if it is determined that MF metadata is included, the process of step S23 and the process of step S24 may be used in combination.

[0266] Furthermore, when it is determined in step S25 that MF metadata is included, it is assumed that the MF data is transmitted after the media data. In this case, the receiving device 20 may buffer the media data and wait until the MF metadata is acquired before performing the process of step S26, or the receiving device 20 may determine whether to perform the process of step S27 without waiting for the MF metadata to be acquired.

[0267] For example, receiving device 20 may determine whether to wait for acquisition of MF metadata based on whether it has a buffer with a buffer size capable of buffering media data. Also, receiving device 20 may determine whether to wait for acquisition of MF metadata based on whether the end-to-end delay will be small. Also, receiving device 20 may perform the decoding process mainly using the process of step S26, and use the process of step S27 in the case of a processing mode when packet loss or the like occurs.

[0268] In addition, if the transmission order type is predetermined, steps S22 and S26 may be omitted, and in this case, the receiving device 20 may determine the method for identifying the MPU first sample and the method for identifying the NAL unit, taking into account the buffer size and end-to-end delay.

[0269] If the transmission order type is known in advance, the transmission order type discriminator 22 in the receiving device 20 is not necessary.

[0270] 26, a decoding instruction unit 28 outputs the data acquired by the data acquisition unit to a decoding unit 29 based on the PTS and DTS calculated by the PTS / DTS calculation unit 26 and the initialization information acquired by the initialization information acquisition unit 27. The decoding unit 29 decodes the data, and a presentation unit 30 presents the decoded data.

[0271] [Receiver operation 2] Next, a description will be given of an operation of the receiving device 20 to obtain initialization information based on the transmission order type and decode media data based on the initialization information. Fig. 27 is a flowchart of such an operation.

[0272] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering and acquires the transmission order type (S301).

[0273] Next, the transmission order type determination unit 22 determines whether or not MPU metadata has been transmitted (S302). If it is determined that MPU metadata has been transmitted (Yes in S302), the transmission order type determination unit 22 determines, based on the analysis result of step S301, whether or not the MPU metadata has been transmitted before the media data (S303). If the MPU metadata has been transmitted before the media data (Yes in S303), the initialization information acquisition unit 27 decodes the media data based on the common initialization information included in the MPU metadata and the initialization information of the sample data (S304).

[0274] On the other hand, if it is determined that the MPU metadata was transmitted after the media data (No in S303), the data acquisition unit 25 buffers the media data until the MPU metadata is acquired (S305), and performs the processing of step S304 after the MPU metadata is acquired.

[0275] Furthermore, if it is determined in step S302 that the MPU metadata has not been transmitted (No in S302), the initialization information acquisition unit 27 decodes the media data based only on the initialization information of the sample data (S306).

[0276] If decoding of the media data is guaranteed only based on the initialization information of the sample data on the transmitting side, the processes based on the determinations of steps S302 and S303 are not performed, and the process of step S306 is used.

[0277] Furthermore, receiving device 20 may determine whether or not to buffer the media data before step S305. In this case, if receiving device 20 determines to buffer the media data, it proceeds to the process of step S305, and if receiving device 20 determines not to buffer the media data, it proceeds to the process of step S306. The determination of whether or not to buffer the media data may be made based on the buffer size and occupancy of receiving device 20, or may be made taking into account end-to-end delay, for example, by selecting the buffer with the smaller end-to-end delay.

[0278] [Receiver operation 3] Here, we will explain in detail the transmission method and reception method when MF metadata is transmitted after media data ((c) of FIG. 21 and (d) of FIG. 21). The following explains the case of (d) of FIG. 21 as an example. Note that in transmission, only the method of (d) of FIG. 21 is used, and signaling of the transmission order type is not performed.

[0279] As mentioned above, when data is transmitted in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Figure 21, (D-1) The receiving device 20 acquires the MPU metadata, and then acquires the MF metadata, and then decodes the media data. (D-2) After acquiring the MPU metadata, the receiving device 20 decodes the media data without using the MF metadata. There are two possible decoding methods:

[0280] Here, D-1 requires buffering of media data to acquire MF metadata, but since decoding can be performed using MPU header information, it can be decoded by a conventional MP4-compliant receiving device.D-2 does not require buffering of media data to acquire MF metadata, but since decoding cannot be performed using MF metadata, special processing is required for decoding.

[0281] Furthermore, the method of FIG. 21(d) has the advantage that the MF metadata is transmitted after the media data, so no delay occurs due to encapsulation, and the end-to-end delay can be reduced.

[0282] The receiving device 20 can select one of the two decoding methods described above depending on the capabilities of the receiving device 20 and the quality of service that the receiving device 20 provides.

[0283] The transmitting device 15 must ensure that decoding can be performed with reduced occurrence of buffer overflow and underflow during the decoding operation in the receiving device 20. For example, the following parameters can be used as elements for defining the decoder model when decoding using the D-1 method.

[0284] Buffer size for reconfiguring the MPU (MPU buffer) For example, buffer size = maximum rate × maximum MPU time × α, where the maximum rate is the upper limit rate of the profile and level of the encoded data + MPU header overhead. The maximum MPU time is the maximum time length of a GOP when 1 MPU = 1 GOP (video).

[0285] Here, audio may be in the GOP unit common to video, or in a different unit. α is a margin to prevent overflow, and may be multiplied or added to the maximum rate x maximum MPU time. When multiplied, α≧1, and when added, α≧0.

[0286] Upper limit of decoding delay time from when data is input to the MPU buffer until it is decoded (TSTD_delay in MPEG-TS STD) For example, at the time of transmission, the DTS is set so that the time when acquisition of the MPU data at the receiver is completed<=DTS, taking into consideration the maximum MPU time and the upper limit of the decoding delay time.

[0287] Furthermore, the transmitting device 15 may assign a DTS and a PTS according to a decoder model for decoding using the D-1 method, thereby ensuring that decoding is possible for a receiving device that performs decoding using the D-1 method, and may also transmit auxiliary information required for decoding using the D-2 method.

[0288] For example, the transmitting device 15 can guarantee the operation of a receiving device that decodes using the D-2 method by signaling the pre-buffering time in the decoder buffer when decoding using the D-2 method.

[0289] The pre-buffering time may be included in SI control information such as a message, table, or descriptor, or may be included in the header of an MMT packet or MMT payload. Alternatively, the SEI in the encoded data may be overwritten. The DTS and PTS for decoding using the D-1 method may be stored in an MPU timestamp descriptor or SampleEntry, and the DTS and PTS for decoding using the D-2 method or the pre-buffering time may be described in the SEI.

[0290] If the receiving device 20 only supports MP4-compliant decoding operations using an MPU header, it may select decoding method D-1, and if it supports both D-1 and D-2, it may select either one.

[0291] The transmitting device 15 may assign a DTS and a PTS to one of the streams (D-1 in this example) so as to ensure the decoding operation, and may also transmit auxiliary information to assist the decoding operation of the other stream.

[0292] Furthermore, when the D-2 method is used, compared to when the D-1 method is used, there is a high possibility that the end-to-end delay will be larger due to the delay caused by pre-buffering of the MF metadata. Therefore, the receiving device 20 may select the D-2 method for decoding when it is desired to reduce the end-to-end delay. For example, the receiving device 20 may always use the D-2 method when it is desired to always reduce the end-to-end delay. Furthermore, the receiving device 20 may use the D-2 method only when operating in a low-delay presentation mode in which it is desired to present live content, channel selection, zapping, etc. with low delay.

[0293] FIG. 28 is a flowchart of such a receiving method.

[0294] First, the receiving device 20 receives an MMT packet and acquires MPU data (S401). Then, the receiving device 20 (transmission order type determination unit 22) determines whether to present the program in low-delay presentation mode (S402).

[0295] If the program is not presented in low-latency presentation mode (No in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using the header information (S405). Also, the receiving device 20 (PTS, DTS calculation unit 26, decoding instruction unit 28, decoding unit 29, presentation unit 30) performs decoding and presentation processing based on the PTS and DTS assigned by the transmitting side (S406).

[0296] On the other hand, when the program is presented in low-latency presentation mode (Yes in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using a decoding method that does not use header information (S403). The receiving device 20 also performs decoding and presentation processing based on auxiliary information for decoding without using PTS, DTS, and header information that is assigned by the transmitting side (S404). Note that in steps S403 and S404, processing may be performed using MPU metadata.

[0297] [Transmission and reception method using auxiliary data] The above has described the transmission and reception operations in the case where MF metadata is transmitted after media data (cases (c) and (d) in Figure 21). Next, a method will be described in which the transmitting device 15 transmits auxiliary data having some of the functions of the MF metadata, thereby enabling decoding to begin earlier and reducing end-to-end delay. Here, an example will be described in which auxiliary data is further transmitted based on the transmission method shown in (d) in Figure 21, but the method using auxiliary data is also applicable to the transmission methods shown in (a) to (c) in Figure 21.

[0298] Figure 29(a) is a diagram showing an MMT packet transmitted using the method shown in Figure 21(d). That is, data is transmitted in the order of MPU metadata, media data, and MF metadata.

[0299] Here, sample #1, sample #2, sample #3, and sample #4 are samples included in the media data. Note that, although an example in which media data is stored in MMT packets in sample units is described here, the media data may be stored in MMT packets in NAL unit units, or in units obtained by dividing an NAL unit. Note that there are also cases in which multiple NAL units are aggregated and stored in an MMT packet.

[0300] As explained in D-1 above, in the case of the method shown in (d) of Fig. 21, that is, when data is transmitted in the order of MPU metadata, media data, and MF metadata, there is a method in which the MPU metadata is acquired, then the MF metadata is acquired, and then the media data is decoded. This method of D-1 requires buffering of the media data to acquire the MF metadata, but has the advantage that the method of D-1 can be applied to conventional MP4-compliant receiving devices because decoding is performed using MPU header information. On the other hand, it has the disadvantage that the receiving device 20 must wait to start decoding until the MF metadata is acquired.

[0301] In contrast to this, as shown in (b) of FIG. 29, in the method using auxiliary data, the auxiliary data is transmitted before the MF metadata.

[0302] MF metadata contains information indicating the DTS, PTS, offsets, and sizes of all samples contained in a movie fragment, while ancillary data contains information indicating the DTS, PTS, offsets, and sizes of some of the samples contained in a movie fragment.

[0303] For example, the MF metadata includes information on all samples (samples #1-#4), whereas the auxiliary data includes information on some samples (samples #1-#2).

[0304] In the case shown in (b) of Figure 29, the auxiliary data enables decoding of sample #1 and sample #2, so the end-to-end delay is smaller than in the transmission method D-1. Note that the auxiliary data may include any combination of sample information, and the auxiliary data may be transmitted repeatedly.

[0305] 29(c), when transmitting auxiliary information at timing A, transmitting device 15 includes information on sample #1 in the auxiliary information, and when transmitting auxiliary information at timing B, transmitting device 15 includes information on sample #1 and sample #2 in the auxiliary information. When transmitting auxiliary information at timing C, transmitting device 15 includes information on sample #1, sample #2, and sample #3 in the auxiliary information.

[0306] The MF metadata includes information on sample #1, sample #2, sample #3, and sample #4 (information on all samples in the movie fragment).

[0307] The auxiliary data does not necessarily have to be transmitted immediately after it is generated.

[0308] In addition, in the header of an MMT packet or an MMT payload, a type indicating that auxiliary data is stored is specified.

[0309] For example, when auxiliary data is stored in the MMT payload using the MPU mode, a data type indicating that it is auxiliary data is specified as the fragment_type field value (e.g., FT=3). The auxiliary data may be data based on the moof structure or may have other structures.

[0310] When auxiliary data is stored as a control signal (descriptor, table, message) in the MMT payload, a descriptor tag, table ID, message ID, etc. that indicate that it is auxiliary data are specified.

[0311] In addition, the PTS or DTS may be stored in the header of the MMT packet or MMT payload.

[0312] [Example of generating auxiliary data] An example in which a transmission device generates auxiliary data based on the configuration of moof will be described below. Fig. 30 is a diagram for explaining an example in which a transmission device generates auxiliary data based on the configuration of moof.

[0313] In a normal MP4, a moof is created for a movie fragment, as shown in Fig. 20. The moof contains information indicating the DTS, PTS, offset, and size of the samples included in the movie fragment.

[0314] Here, the transmitting device 15 composes an MP4 file using only a portion of the sample data that composes the MPU, and generates auxiliary data.

[0315] For example, as shown in (a) of Figure 30, the transmitting device 15 generates an MP4 using only sample #1 of samples #1-#4 that make up the MPU, and the header of moof+mdat is used as auxiliary data.

[0316] Next, as shown in (b) of Figure 30, the transmitting device 15 generates an MP4 using samples #1 and #2 of samples #1-#4 that make up the MPU, and the header of moof+mdat is used as the next auxiliary data.

[0317] Next, as shown in (c) of Figure 30, the transmitting device 15 generates an MP4 using samples #1, #2, and #3 of samples #1-#4 that make up the MPU, and the header of moof+mdat is set as the next auxiliary data.

[0318] Next, as shown in (d) of Figure 30, the transmission device 15 generates all MP4s from samples #1-#4 that make up the MPU, and the header of moof+mdat among them becomes movie fragment metadata.

[0319] Although transmitting device 15 generates auxiliary data for each sample here, it may generate auxiliary data for every N samples. The value of N is an arbitrary number, and for example, if auxiliary data is transmitted M times when transmitting one MPU, N may be set to total samples / M.

[0320] The information indicating the offset of a sample in moof may be an offset value after the sample entry area for the subsequent number of samples is secured as a NULL area.

[0321] The auxiliary data may be generated so as to have a configuration in which the MF metadata is fragmented.

[0322] [Example of reception using auxiliary data] The following describes reception of auxiliary data generated as described in Fig. 30. Fig. 31 is a diagram for explaining reception of auxiliary data. In Fig. 31(a), the number of samples constituting the MPU is 30, and auxiliary data is generated and transmitted every 10 samples.

[0323] In FIG. 30(a), auxiliary data #1 includes sample information for samples #1-#10, auxiliary data #2 includes sample information for samples #1-#20, and MF metadata includes sample information for samples #1-#30.

[0324] Although samples #1-#10, samples #11-#20, and samples #21-#30 are stored in one MMT payload, they may also be stored in sample units or NAL units, or in fragment or aggregate units.

[0325] The receiving device 20 receives the packets of MPU meta, samples, MF meta, and auxiliary data, respectively.

[0326] The receiving device 20 concatenates the sample data in the order in which they are received (to the end), and after receiving the latest auxiliary data, updates the previous auxiliary data.Furthermore, the receiving device 20 can configure a complete MPU by finally replacing the auxiliary data with MF metadata.

[0327] Upon receiving auxiliary data #1, receiving device 20 concatenates the data to form an MP4, as shown in the upper part of (b) of Fig. 31. This allows receiving device 20 to parse samples #1-#10 using the MPU metadata and information in auxiliary data #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0328] Furthermore, upon receiving auxiliary data #2, receiving device 20 concatenates the data to form an MP4, as shown in the middle part of (b) of Fig. 31. This allows receiving device 20 to parse samples #1-#20 using the MPU metadata and information in auxiliary data #2, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0329] Furthermore, upon receiving the MF metadata, the receiving device 20 concatenates the data as shown in the lower part of (b) of Fig. 31 to construct an MP4 file. This enables the receiving device 20 to parse samples #1-#30 using the MPU metadata and MF metadata, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF metadata.

[0330] In the absence of auxiliary data, receiving device 20 could only acquire sample information after receiving MF metadata, and therefore had to start decoding after receiving MF metadata. However, by transmitting device 15 generating and transmitting auxiliary data, receiving device 20 can acquire sample information using the auxiliary data without waiting for reception of MF metadata, thereby speeding up the decoding start time. Furthermore, by transmitting device 15 generating auxiliary data based on the moof described with reference to Figure 30, receiving device 20 can parse using a conventional MP4 parser as is.

[0331] Furthermore, newly generated auxiliary data and MF metadata contain sample information that overlaps with previously transmitted auxiliary data. Therefore, even if previous auxiliary data cannot be obtained due to packet loss, etc., it is possible to reconstruct the MP4 and obtain sample information (PTS, DTS, size, and offset) by using the newly obtained auxiliary data and MF metadata.

[0332] It should be noted that the auxiliary data does not necessarily have to include information on past sample data. For example, auxiliary data #1 may correspond to sample data #1-#10, and auxiliary data #2 may correspond to sample data #11-#20. For example, as shown in (c) of Figure 31, transmitting device 15 may use complete MF metadata as a data unit and sequentially transmit fragments of the data unit as auxiliary data.

[0333] Furthermore, the transmitting device 15 may repeatedly transmit the auxiliary data or the MF metadata in order to deal with packet loss.

[0334] The MMT packet and MMT payload in which auxiliary data is stored contain an MPU sequence number and an asset ID, as well as MPU metadata, MF metadata, and sample data.

[0335] The above-described receiving operation using auxiliary data will be described with reference to the flowchart in Fig. 32. Fig. 32 is a flowchart of the receiving operation using auxiliary data.

[0336] First, receiving device 20 receives an MMT packet and analyzes the packet header and payload header (S501). Next, receiving device 20 analyzes whether the fragment type is auxiliary data or MF metadata (S502). If the fragment type is auxiliary data, receiving device 20 overwrites and updates the previous auxiliary data (S503). At this time, if there is no previous auxiliary data for the same MPU, receiving device 20 uses the received auxiliary data as new auxiliary data. Then, receiving device 20 acquires samples based on the MPU metadata, auxiliary data, and sample data, and performs decoding (S507).

[0337] On the other hand, if the fragment type is MF metadata, the receiving device 20 overwrites the previous auxiliary data with the MF metadata in step S505 (S505).Then, the receiving device 20 obtains the sample in the form of a complete MPU based on the MPU metadata, MF metadata, and sample data, and performs decoding (S506).

[0338] Although not shown in Figure 32, in step S502, if the fragment type is MPU metadata, the receiving device 20 stores the data in a buffer, and if the fragment type is sample data, it stores the data concatenated at the end for each sample in a buffer.

[0339] If the auxiliary data cannot be obtained due to packet loss, the receiving device 20 can either overwrite the sample with the latest auxiliary data or decode the sample using the previous auxiliary data.

[0340] The transmission cycle and the number of transmissions of the auxiliary data may be predetermined values. Information on the transmission cycle and the number of times (count, countdown) may be transmitted together with the data. For example, the transmission cycle, the number of times of transmission, and a timestamp such as initial_cpb_removal_delay may be stored in the data unit header.

[0341] By transmitting ancillary data including information on the first sample of the MPU at least once before the initial_cpb_removal_delay, it is possible to comply with the CPB buffer model. In this case, the MPU timestamp descriptor is set to a value based on the picture timing SEI.

[0342] The transmission method for receiving operations using such auxiliary data is not limited to the MMT method, but can also be applied to streaming transmission of packets configured in ISOBMFF file format, such as MPEG-DASH.

[0343] [Transmission method when one MPU consists of multiple movie fragments] In the explanation from Fig. 19 onwards, one MPU is composed of one movie fragment, but here we will explain the case where one MPU is composed of multiple movie fragments. Fig. 33 shows the configuration of an MPU composed of multiple movie fragments.

[0344] In Figure 33, samples (#1-#6) stored in one MPU are divided into two movie fragments. The first movie fragment is generated based on samples #1-#3, and a corresponding moof box is generated. The second movie fragment is generated based on samples #4-#6, and a corresponding moof box is generated.

[0345] The headers of the moof box and mdat box in the first movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #1. Meanwhile, the headers of the moof box and mdat box in the second movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #2. Note that in Figure 33, the MMT payload in which movie fragment metadata is stored is hatched.

[0346] The number of samples constituting an MPU and the number of samples constituting a movie fragment are arbitrary. For example, the number of samples constituting an MPU may be the number of samples in a GOP unit, and two movie fragments may be composed by using half the number of samples in a GOP unit as movie fragments.

[0347] Note that, although an example is shown here in which one MPU contains two movie fragments (a moof box and an mdat box), one MPU may contain three or more movie fragments instead of two. Also, the samples stored in a movie fragment do not have to be divided equally, but may be divided into any number of samples.

[0348] 33, the MPU metadata unit and the MF metadata unit are each stored as a data unit in the MMT payload. However, transmitting device 15 may store units such as ftyp, mmpu, moov, and moof as data units in the MMT payload in data unit units, or may store data units in the MMT payload in fragmented units. Furthermore, transmitting device 15 may store data units in the MMT payload in aggregated units.

[0349] 33, samples are stored in the MMT payload in sample units. However, transmitting device 15 may configure data units in NAL unit units or units aggregating multiple NAL units instead of sample units, and store the data units in the MMT payload. Also, transmitting device 15 may store data units in fragmented units or aggregated units in the MMT payload.

[0350] In Fig. 33, the MPU is configured in the order of moof#1, mdat#1, moof#2, mdat#2, and an offset is assigned to moof#1, assuming that the corresponding mdat#1 is attached after it. However, an offset may also be assigned to mdat#1, assuming that it is attached before moof#1. In this case, however, movie fragment metadata cannot be generated in the form of moof+mdat, and the headers of moof and mdat are transmitted separately.

[0351] Next, a description will be given of the transmission order of MMT packets when transmitting an MPU having the configuration described in Fig. 33. Fig. 34 is a diagram for explaining the transmission order of MMT packets.

[0352] Figure 34(a) shows the transmission order when MMT packets are transmitted in the configuration order of the MPUs shown in Figure 33. Figure 34(a) specifically shows an example in which MPU meta, MF meta #1, media data #1 (samples #1-#3), MF meta #2, and media data #2 (samples #4-#6) are transmitted in this order.

[0353] FIG. 34(b) shows an example in which MPU meta, media data #1 (samples #1-#3), MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in this order.

[0354] FIG. 34(c) shows an example in which media data #1 (samples #1-#3), MPU meta, MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in this order.

[0355] MF meta #1 is generated using samples #1-#3, and MF meta #2 is generated using samples #4-#6. Therefore, when the transmission method of Figure 34(a) is used, a delay occurs in the transmission of sample data due to encapsulation.

[0356] In contrast, when the transmission methods of Figures 34(b) and 34(c) are used, samples can be transmitted without waiting for the MF meta to be generated, so no delay due to encapsulation occurs and end-to-end delay can be reduced.

[0357] Also, in the transmission order (a) of Figure 34, one MPU is divided into multiple movie fragments, and the number of samples stored in the MF meta is smaller than in the case of Figure 19, so the amount of delay due to encapsulation can be reduced compared to the case of Figure 19.

[0358] In addition to the method shown here, for example, transmitting device 15 may concatenate MF meta #1 and MF meta #2 and transmit them together at the end of the MPU. In this case, MF meta of different movie fragments may be aggregated and stored in one MMT payload. Also, MF meta of different MPUs may be aggregated and stored in an MMT payload.

[0359] [How to receive when one MPU consists of multiple movie fragments] Here, a description will be given of an example of operation of receiving device 20 that receives and decodes MMT packets transmitted in the transmission order described in (b) of Fig. 34. Figs. 35 and 36 are diagrams for explaining such an example of operation.

[0360] The receiving device 20 receives each of the MMT packets including the MPU meta, samples, and MF meta transmitted in the transmission order shown in Fig. 35. The sample data is concatenated in the order in which it is received.

[0361] At T1, which is the time when MF meta #1 is received, receiving device 20 concatenates the data as shown in (1) of Fig. 36 to construct an MP4. This allows receiving device 20 to obtain samples #1-#3 based on the MPU metadata and information on MF meta #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF meta.

[0362] Furthermore, receiving device 20 concatenates the data as shown in (2) of Fig. 36 at T2, which is the time when MF meta #2 is received, to construct an MP4. This allows receiving device 20 to acquire samples #4-#6 based on the MPU metadata and information in MF meta #2, and to perform decoding based on the PTS, DTS, offset, and size information in the MF meta. Receiving device 20 may also acquire samples #1-#6 based on the information in MF meta #1 and MF meta #2 by concatenating the data as shown in (3) of Fig. 36 to construct an MP4.

[0363] By dividing a single MPU into multiple movie fragments, the time it takes for the MPU to acquire the initial MF meta is shortened, which allows for earlier decoding start time and reduces the buffer size for storing samples before decoding.

[0364] The transmitting device 15 may set the division unit of the movie fragment so that the time from transmitting (or receiving) the first sample in the movie fragment to transmitting (or receiving) the MF meta corresponding to the movie fragment is shorter than the initial_cpb_removal_delay specified by the encoder. By setting it in this way, the receiving buffer can follow the cpb buffer, realizing low-delay decoding. In this case, absolute times based on the initial_cpb_removal_delay can be used for the PTS and DTS.

[0365] Alternatively, the transmitting device 15 may divide the movie fragments at equal intervals, or divide subsequent movie fragments at shorter intervals than the previous movie fragments, which allows the receiving device 20 to always receive MF meta containing information about a sample before decoding that sample, enabling continuous decoding.

[0366] The absolute time of the PTS and DTS can be calculated using the following two methods.

[0367] (1) The absolute times of the PTS and DTS are determined based on the reception time (T1 or T2) of the MF meta #1 or MF meta #2 and the relative times of the PTS and DTS included in the MF meta.

[0368] (2) The absolute time of the PTS and DTS is determined based on the absolute time signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative time of the PTS and DTS included in the MF meta.

[0369] Also, (2-A) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on the initial_cpb_removal_delay specified by the encoder.

[0370] Also, (2-B) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on a predicted value of the reception time of the MF meta.

[0371] Note that MF meta #1 and MF meta #2 may be transmitted repeatedly. By repeatedly transmitting MF meta #1 and MF meta #2, the receiving device 20 can acquire the MF meta again even if it was unable to acquire it due to packet loss or the like.

[0372] The payload header of an MFU containing samples that constitute a movie fragment can store an identifier indicating the order of the movie fragment. On the other hand, an identifier indicating the order of the MF meta that constitutes a movie fragment is not included in the MMT payload. Therefore, the receiving device 20 identifies the order of the MF meta by the packet_sequence_number. Alternatively, the transmitting device 15 may store and signal an identifier indicating the ordinal number of the movie fragment to which the MF meta belongs in control information (message, table, descriptor), the MMT header, the MMT payload header, or the data unit header.

[0373] The transmitting device 15 may transmit the MPU meta, MF meta, and samples in a predetermined transmission order, and the receiving device 20 may perform the receiving process based on the predetermined transmission order. Alternatively, the transmitting device 15 may signal the transmission order, and the receiving device 20 may select (determine) the receiving process based on the signaling information.

[0374] The above-described receiving method will be explained using Fig. 37. Fig. 37 is a flowchart of the operation of the receiving method explained in Figs.

[0375] First, the receiving device 20 determines (identifies) whether the data included in the payload is MPU metadata, MF metadata, or sample data (MFU) based on the fragment type indicated in the MMT payload (S601, S602). If the data is sample data, the receiving device 20 buffers the sample and waits for reception of MF metadata corresponding to the sample and for the start of decoding (S603).

[0376] On the other hand, in step S602, if the data is MF metadata, the receiving device 20 obtains sample information (PTS, DTS, position information, and size) from the MF metadata, obtains a sample based on the obtained sample information, and decodes and presents the sample based on the PTS and DTS (S604).

[0377] Although not shown, if the data is MPU metadata, the MPU metadata contains initialization information necessary for decoding, which the receiving device 20 stores and uses to decode the sample data in step S604.

[0378] When the receiving device 20 stores the received MPU data (MPU metadata, MF metadata, and sample data) in a storage device, it stores the data after rearranging it into the MPU configuration described in Figure 19 or Figure 33.

[0379] On the transmitting side, packet sequence numbers are assigned to MMT packets with the same packet ID. At this time, packet sequence numbers may be assigned after MMT packets containing MPU metadata, MF metadata, and sample data are rearranged in transmission order, or packet sequence numbers may be assigned in the order before rearrangement.

[0380] If packet sequence numbers are assigned in the order before rearrangement, reception device 20 can rearrange the data in the order configured in the MPU based on the packet sequence numbers, facilitating storage.

[0381] [Method for detecting the beginning of an access unit and the beginning of a slice segment] A method for detecting the beginning of an access unit or a slice segment based on information in the MMT packet header and the MMT payload header will be described.

[0382] Here, two examples are shown: one where non-VCL NAL units (such as access unit delimiters, VPS, SPS, PPS, and SEI) are collectively stored as data units in an MMT payload, and one where each non-VCL NAL unit is treated as a data unit, and the data units are aggregated and stored in a single MMT payload.

[0383] FIG. 38 is a diagram showing a case where non-VCL NAL units are aggregated as individual data units.

[0384] 38, the start of the access unit is an MMT packet whose fragment_type value is MFU, and is the start data of an MMT payload that includes a data unit whose aggregation_flag value is 1 and whose offset value is 0. In this case, the Fragmentation_indicator value is 0.

[0385] Also, in the case of Figure 38, the beginning of the slice segment is an MMT packet whose fragment_type value is MFU, and is the beginning data of an MMT payload whose aggregation_flag value is 0 and whose fragmentation_indicator value is 00 or 01.

[0386] 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. Note that the field values ​​of the packet header are as shown in FIG. 17 (or FIG. 18).

[0387] In the case of FIG. 39, the head of the access unit is the head data of the payload in the packet with an Offset value of 0.

[0388] In addition, in the case of FIG. 39, the start of a slice segment is the first data of the payload of a packet whose Offset value is a value other than 0 and whose fragmentation indicator value is 00 or 01.

[0389] [Reception process when packet loss occurs] Generally, when transmitting MP4 format data in an environment where packet loss occurs, the receiving device 20 restores packets using ALFEC (Application Layer FEC), packet retransmission control, or the like.

[0390] However, if packet loss occurs in streaming such as broadcasting when AL-FEC cannot be used, the packets cannot be restored.

[0391] After data is lost due to packet loss, the receiving device 20 needs to resume decoding of video and audio. To do this, the receiving device 20 needs to detect the beginning of an access unit or NAL unit and start decoding from the beginning of the access unit or NAL unit.

[0392] However, since there is no start code at the beginning of an NAL unit in MP4 format, the receiving device 20 cannot detect the beginning of an access unit or an NAL unit even if it analyzes the stream.

[0393] FIG. 40 is a flowchart of the operation of the receiving device 20 when a packet loss occurs.

[0394] The receiving device 20 detects packet loss using the packet sequence number, packet counter, fragment counter, etc. in the header of the MMT packet or MMT payload (S701), and determines which packet has been lost based on the context (S702).

[0395] If it is determined that no packet loss has occurred (No in S702), the receiving device 20 constructs an MP4 file and decodes the access units or NAL units (S703).

[0396] If it is determined that a packet loss has occurred (Yes in S702), the receiving device 20 generates a NAL unit corresponding to the NAL unit that has experienced the packet loss using dummy data, and constructs an MP4 file (S704). When inserting dummy data into a NAL unit, the receiving device 20 indicates that the NAL unit type is dummy data.

[0397] In addition, the receiving device 20 can resume decoding by detecting the beginning of the next access unit or NAL unit and inputting the beginning data into the decoder based on the methods described in Figures 17, 18, 38, and 39 (S705).

[0398] In addition, if packet loss occurs, the receiving device 20 may resume decoding from the beginning of the access unit and NAL unit based on information detected based on the packet header, or may resume decoding from the beginning of the access unit and NAL unit based on header information of the reconstructed MP4 file, which includes a dummy data NAL unit.

[0399] When storing an MP4 file (MPU), the receiving device 20 may separately acquire packet data (NAL units, etc.) lost due to packet loss and store (replace) it from broadcasting or communication.

[0400] At this time, when receiving device 20 acquires the lost packet from the communication, it notifies the server of information about the lost packet (packet ID, MPU sequence number, packet sequence number, IP data flow number, IP address, etc.) and acquires the packet. The receiving device 20 is not limited to acquiring only the lost packet, but may also acquire a group of packets before and after the lost packet at the same time.

[0401] [How to compose a movie fragment] Here we will explain in detail how to configure movie fragments.

[0402] As described in Fig. 33, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are arbitrary. For example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU may be a fixed predetermined number, or may be dynamically determined.

[0403] Here, by configuring the movie fragments on the transmitting side (transmitting device 15) so as to satisfy the following conditions, low-delay decoding in the receiving device 20 can be guaranteed.

[0404] The conditions are as follows:

[0405] The transmitting device 15 generates and transmits MF meta as movie fragments, which are units obtained by dividing sample data, so that the receiving device 20 can always receive MF meta containing information about any sample (Sample(i)) before the decoding time (DTS(i)) of that sample.

[0406] Specifically, the transmitting device 15 constructs a movie fragment using samples (including the i-th sample) that have been coded before DTS(i).

[0407] To ensure low-latency decoding, the following method, for example, is used to dynamically determine the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU.

[0408] (1) At the start of decoding, the decoding time DTS(0) of the first sample Sample(0) of the GOP is based on initial_cpb_removal_delay. The transmitting device constructs a first movie fragment using samples that have already been coded at a time before DTS(0). The transmitting device 15 also generates MF metadata corresponding to the first movie fragment and transmits it at a time before DTS(0).

[0409] (2) The transmitting device 15 constructs movie fragments so that the above conditions are met for subsequent samples as well.

[0410] For example, if the first sample of a movie fragment is the kth sample, the MF meta of the movie fragment including the kth sample is transmitted by the decoding time DTS(k) of the kth sample. If the encoding completion time of the lth sample is before DTS(k) and the encoding completion time of the (l+1)th sample is after DTS(k), the transmitting device 15 constructs a movie fragment using the kth sample to the lth sample.

[0411] In addition, the transmitting device 15 may construct a movie fragment using the kth sample to a sample less than the lth sample.

[0412] (3) After completing the encoding of the last sample in the MPU, the transmitting device 15 constructs a movie fragment using the remaining samples, generates MF metadata corresponding to the movie fragment, and transmits it.

[0413] It should be noted that the transmitting device 15 may construct a movie fragment using only a portion of the samples that have been coded, rather than using all of the samples that have been coded.

[0414] In the above example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are dynamically determined based on the above conditions to ensure low-latency decoding. However, the method for determining the number of samples and the number of movie fragments is not limited to this method. For example, the number of movie fragments constituting one MPU may be fixed to a predetermined value, and the number of samples may be determined to satisfy the above conditions. Furthermore, the number of movie fragments constituting one MPU and the time at which the movie fragments are divided (or the code amount of the movie fragments) may be fixed to predetermined values, and the number of samples may be determined to satisfy the above conditions.

[0415] In addition, if the MPU is divided into multiple movie fragments, information indicating whether the MPU is divided into multiple movie fragments, attributes of the divided movie fragments, or MF meta attributes for the divided movie fragments may be transmitted.

[0416] Here, the attribute of a movie fragment is information indicating whether the movie fragment is the first movie fragment of an MPU, the last movie fragment of an MPU, or some other movie fragment.

[0417] In addition, the attributes of the MF meta are information that indicates whether the MF meta corresponds to the first movie fragment of the MPU, the last movie fragment of the MPU, or any other movie fragment.

[0418] The transmitting device 15 may store and transmit the number of samples that make up a movie fragment and the number of movie fragments that make up one MPU as control information.

[0419] [Operation of receiving device] The operation of the receiving device 20 based on the movie fragment configured as above will now be described.

[0420] The receiving device 20 determines the absolute times of the PTS and DTS based on the absolute times signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative times of the PTS and DTS included in the MF meta.

[0421] Based on information on whether the MPU is divided into multiple movie fragments, the receiving device 20 performs the following processing based on the attributes of the divided movie fragments if the MPU is divided.

[0422] (1) When the movie fragment is the first movie fragment of the MPU, the receiving device 20 generates the absolute time of the PTS and DTS using the absolute time of the PTS of the first sample included in the MPU timestamp descriptor and the relative times of the PTS and DTS included in the MF meta.

[0423] (2) If the movie fragment is not the first movie fragment of the MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the relative times of the PTS and DTS included in the MF meta, without using the information in the MPU timestamp descriptor.

[0424] (3) If the movie fragment is the last movie fragment in the MPU, the receiving device 20 calculates the absolute times of the PTS and DTS of all samples and then resets the PTS and DTS calculation process (addition process of relative times). Note that the reset process may also be performed on the first movie fragment in the MPU.

[0425] The receiving device 20 may determine whether a movie fragment is divided as follows: The receiving device 20 may also acquire attribute information of the movie fragment as follows.

[0426] For example, the receiving device 20 may determine whether the movie fragment has been divided based on the identifier movie_fragment_sequence_number field value that indicates the order of the movie fragments indicated in the MMTP (MMT Protocol) payload header.

[0427] Specifically, the receiving device 20 may determine that an MPU is divided into multiple movie fragments if the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 1, and there is a value of 2 or greater in the field value.

[0428] In addition, the receiving device 20 may determine that an MPU is divided into multiple movie fragments if the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 0, and there is a value other than 0 in the field value.

[0429] The attribute information of the movie fragment may also be determined based on the movie_fragment_sequence_number.

[0430] It should be noted that, without using the movie_fragment_sequence_number, it is also possible to determine whether a movie fragment is divided and the attribute information of the movie fragment by counting the transmission of movie fragments and MF meta contained in one MPU.

[0431] With the above-described configurations of the transmitting device 15 and the receiving device 20, the receiving device 20 can receive movie fragment metadata at intervals shorter than that of an MPU, enabling decoding to start with low latency. Also, decoding with low latency can be performed using a decoding process based on the MP4 parsing method.

[0432] The reception operation when the MPU is divided into multiple movie fragments as described above will be explained using a flowchart. Figure 41 is a flowchart of the reception operation when the MPU is divided into multiple movie fragments. Note that this flowchart illustrates the operation of step S604 in Figure 37 in more detail.

[0433] First, based on the data type indicated in the MMTP payload header, if the data type is MF meta, the receiving device 20 acquires the MF meta data (S801).

[0434] Next, the receiving device 20 determines whether the MPU is divided into multiple movie fragments (S802), and if the MPU is divided into multiple movie fragments (Yes in S802), determines whether the received MF metadata is the first metadata in the MPU (S803). If the received MF metadata is the first MF metadata in the MPU (Yes in S803), the receiving device 20 calculates the absolute times of the PTS and DTS from the absolute time of the PTS indicated in the MPU timestamp descriptor and the relative times of the PTS and DTS indicated in the MF metadata (S804), and determines whether the metadata is the last metadata in the MPU (S805).

[0435] On the other hand, if the received MF metadata is not the MF metadata at the beginning of the MPU (No in S803), the receiving device 20 calculates the absolute times of the PTS and DTS using the relative times of the PTS and DTS indicated in the MF metadata without using the information in the MPU timestamp descriptor (S808), and proceeds to processing in step S805.

[0436] If it is determined in step S805 that this is the last MF metadata of the MPU (Yes in S805), the receiving device 20 calculates the absolute times of the PTS and DTS of all samples and then resets the PTS and DTS calculation process.If it is determined in step S805 that this is not the last MF metadata of the MPU (No in S805), the receiving device 20 ends the process.

[0437] Also, if it is determined in step S802 that the MPU is not divided into multiple movie fragments (No in S802), the receiving device 20 acquires sample data based on the MF metadata transmitted after the MPU and determines the PTS and DTS (S807).

[0438] Finally, although not shown, the receiving device 20 performs decoding and presentation processes based on the determined PTS and DTS.

[0439] [Issues that arise when splitting movie fragments and their solutions] So far, we have explained how to reduce end-to-end delay by dividing movie fragments. From here, we will explain the new issues that arise when dividing movie fragments and how to solve them.

[0440] First, as background, the picture structure in coded data will be described. Figure 42 is a diagram showing an example of a prediction structure of a picture for each TemporalId when implementing temporal scalability.

[0441] In coding formats such as MPEG-4 AVC and HEVC (High Efficiency Video Coding), temporal scalability can be achieved by using B-pictures (bidirectional reference predictive pictures) that can be referenced from other pictures.

[0442] TemporalId shown in (a) of Figure 42 is an identifier for a layer in the coding structure, with larger TemporalId values ​​indicating deeper layers. Square blocks indicate pictures, with Ix in each block indicating an I-picture (intra-picture predicted picture), Px indicating a P-picture (forward reference predicted picture), and Bx and bx indicating B-pictures (bidirectional reference predicted pictures). The x in Ix / Px / Bx indicates the display order, indicating the order in which the pictures are displayed. Arrows between pictures indicate reference relationships; for example, picture B4 indicates that a predicted image is generated using I0 and B8 as reference images. It is prohibited for a picture to use as a reference image another picture with a TemporalId higher than its own. The layers are specified to allow for temporal scalability. For example, in Figure 42, decoding all pictures results in 120 fps (frames per second) video, but decoding only layers with TemporalId from 0 to 3 results in 60 fps video.

[0443] Figure 43 shows the relationship between the decoding time (DTS) and display time (PTS) for each picture in Figure 42. For example, picture I0 shown in Figure 43 is displayed after the decoding of B4 is completed so that no gap occurs in the decoding and display.

[0444] As shown in Figure 43, when the prediction structure includes a B picture, the decoding order and the display order are different, so after decoding the picture in the receiving device 20, picture delay processing and picture reordering processing are required.

[0445] Although examples of picture prediction structures in temporal scalability have been described above, even when temporal scalability is not used, picture delay processing and reordering processing may be required depending on the prediction structure. Figure 44 is a diagram showing an example of a prediction structure of a picture that requires picture delay processing and reordering processing. Note that the numbers in Figure 44 indicate the decoding order.

[0446] As shown in Figure 44, depending on the prediction structure, the first sample in decoding order may differ from the first sample in presentation order, and in Figure 44, the first sample in presentation order is the fourth sample in decoding order. Note that Figure 44 shows an example of a prediction structure, and the prediction structure is not limited to this structure. In other prediction structures, the first sample in decoding order may differ from the first sample in presentation order.

[0447] Like FIG. 33, FIG. 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and MMTP packets. Note that the number of samples constituting an MPU and the number of samples constituting a movie fragment are arbitrary. For example, the number of samples constituting an MPU may be the number of samples per GOP, and two movie fragments may be constituted by using half the number of samples per GOP. One sample may be treated as one movie fragment, or the samples constituting an MPU may not be divided.

[0448] While FIG. 45 shows an example in which one MPU contains two movie fragments (a moof box and an mdat box), the number of movie fragments contained in one MPU does not have to be two. The number of movie fragments contained in one MPU may be three or more, or may be the number of samples contained in the MPU. Furthermore, the samples stored in a movie fragment do not have to be divided equally, but may be divided into any number of samples.

[0449] Movie fragment metadata (MF metadata) contains information on the PTS, DTS, offset, and size of the samples contained in the movie fragment, and when the receiving device 20 decodes a sample, it extracts the PTS and DTS from the MF meta containing information about the sample and determines the decoding timing and presentation timing.

[0450] Hereinafter, for the sake of detailed explanation, the absolute value of the decoding time of the i sample will be referred to as DTS(i), and the absolute value of the presentation time will be referred to as PTS(i).

[0451] The information of the i-th sample among the timestamp information stored in the moof in the MF meta is specifically the relative value of the decoding time of the i-th sample and the (i+1)-th sample, and the relative value of the decoding time of the i-th sample and the presentation time, which will be referred to as DT(i) and CT(i) hereafter.

[0452] Movie fragment metadata #1 contains DT(i) and CT(i) for samples #1-#3, and movie fragment metadata #2 contains DT(i) and CT(i) for samples #4-#6.

[0453] The absolute PTS value of the access unit at the beginning of the MPU is stored in the MPU timestamp descriptor or the like, and the receiving device 20 calculates the PTS and DTS based on the PTS_MPU of the access unit at the beginning of the MPU, the CT, and the DT.

[0454] FIG. 46 is a diagram for explaining a method of calculating PTS and DTS and problems involved when an MPU is configured using samples #1 to #10.

[0455] (a) of Figure 46 shows an example where the MPU is not divided into movie fragments, (b) of Figure 46 shows an example where the MPU is divided into two movie fragments of 5 sample units, and (c) of Figure 46 shows an example where the MPU is divided into 10 movie fragments of sample units.

[0456] As explained in Figure 45, when the PTS and DTS are calculated using the MPU timestamp descriptor and the timestamp information (CT and DT) in the MP4, the first sample in the presentation order in Figure 44 is the fourth sample in decoding order. Therefore, the PTS stored in the MPU timestamp descriptor is the PTS (absolute value) of the fourth sample in decoding order. Note that, hereinafter, this sample will be referred to as sample A. Furthermore, the first sample in decoding order will be referred to as sample B.

[0457] Because the only absolute time information related to the timestamp is the information in the MPU timestamp descriptor, the receiving device 20 cannot calculate the PTS (absolute time) and DTS (absolute time) of other samples until the arrival of sample A. The receiving device 20 also cannot calculate the PTS and DTS of sample B.

[0458] 46(a), sample A is included in the same movie fragment as sample B and is stored in one MF meta, so receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.

[0459] 46(b), sample A is included in the same movie fragment as sample B and is stored in one MF meta. Therefore, receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.

[0460] In the example of (c) in Figure 46, sample A is included in a different movie fragment from sample B. Therefore, receiving device 20 cannot determine the DTS of sample B until it receives MF meta including the CT and DT of the movie fragment that includes sample A.

[0461] Therefore, in the example of FIG. 46(c), the receiving device 20 cannot start decoding immediately after the arrival of the B sample.

[0462] In this way, if a movie fragment containing a B sample does not contain an A sample, the receiving device 20 cannot start decoding the B sample until it has received the MF meta for the movie fragment containing the A sample.

[0463] This issue occurs when the first sample in presentation order does not match the first sample in decoding order, and the movie fragment is split to the point where sample A and sample B are no longer stored in the same movie fragment. This issue also occurs regardless of whether the MF meta is forward or backward.

[0464] In this way, if the first sample in presentation order does not match the first sample in decoding order, and if sample A and sample B are not stored in the same movie fragment, the DTS cannot be determined immediately after receiving sample B. Therefore, transmitting device 15 separately transmits the DTS (absolute value) of sample B or information that allows the receiving side to calculate the DTS (absolute value) of sample B. Such information may be transmitted using control information, a packet header, etc.

[0465] Using this information, the receiving device 20 calculates the DTS (absolute value) of sample B. Fig. 47 is a flowchart of the receiving operation when the DTS is calculated using this information.

[0466] The receiving device 20 receives the movie fragment at the beginning of the MPU (S901), and determines whether the A sample and the B sample are stored in the same movie fragment (S902). If they are stored in the same movie fragment (Yes in S902), the receiving device 20 calculates the DTS using only the MF meta information, without using the DTS (absolute time) of the B sample, and starts decoding (S904). Note that in step S904, the receiving device 20 may determine the DTS using the DTS of the B sample.

[0467] On the other hand, if sample A and sample B are not stored in the same movie fragment in step S902 (No in S902), receiving device 20 obtains the DTS (absolute time stamp) of sample B, determines the DTS, and starts decoding (S903).

[0468] In the above description, an example has been described in which the absolute value of the decoding time and the absolute value of the presentation time of each sample are calculated using MF meta (timestamp information stored in the moof in MP4 format) in the MMT standard, but it goes without saying that the MF meta may be replaced with any control information that can be used to calculate the absolute value of the decoding time and the absolute value of the presentation time of each sample. Examples of such control information include control information in which the above-mentioned relative value CT(i) of the decoding time between the i-th sample and the (i+1)-th sample is replaced with the relative value of the presentation time between the i-th sample and the (i+1)-th sample, and control information that includes both the relative value CT(i) of the decoding time between the i-th sample and the (i+1)-th sample and the relative value of the presentation time between the i-th sample and the (i+1)-th sample.

[0469] (Embodiment 3) [overview] In the third embodiment, a content transmission method and data structure when transmitting content such as video, audio, subtitles, and data broadcasting via broadcasting will be described. That is, a content transmission method and data structure specialized for playing back broadcast streams will be described.

[0470] In the third embodiment, an example will be described in which the MMT method (hereinafter also simply referred to as MMT) is used as the multiplexing method, but other multiplexing methods such as MPEG-DASH or RTP may also be used.

[0471] First, a method for storing a data unit (DU) in a payload in MMT will be described in detail. Fig. 48 is a diagram for explaining a method for storing a data unit in a payload in MMT.

[0472] In MMT, a transmitting device stores part of the data that constitutes an MPU in an MMTP payload as a data unit, and transmits it with a header. The header includes an MMTP payload header and an MMTP packet header. The data unit may be in units of NAL units or samples.

[0473] (a) of Fig. 48 shows an example in which a transmitting device aggregates multiple data units and stores them in one payload. In the example of (a) of Fig. 48, a data unit header (DUH) and a data unit length (DUL) are added to the beginning of each of the multiple data units, and multiple data units with the data unit header and data unit length added are stored together in the payload.

[0474] Figure 48(b) shows an example in which one data unit is stored in one payload. In the example of Figure 48(b), a data unit header is added to the beginning of the data unit and stored in the payload. Figure 48(c) shows an example in which one data unit is divided, and the divided data units are added with data unit headers and stored in the payload.

[0475] There are various types of data units, such as timed-MFU, which is media including information related to synchronization of video, audio, or subtitles, non-timed-MFU, which is media including no information related to synchronization such as files, MPU metadata, and MF metadata, and a data unit header is defined depending on the type of data unit. Note that MPU metadata and MF metadata do not have a data unit header.

[0476] Furthermore, although a transmitting device cannot in principle aggregate different types of data units, it may be specified to be able to aggregate different types of data units. For example, when the size of MF metadata is small, such as when it is divided into movie fragments for each sample, aggregating the MF metadata and media data can reduce the number of packets and also the transmission capacity.

[0477] If the data unit is an MFU, some information about the MPU, such as information for configuring the MPU (MP4), is stored as a header.

[0478] For example, the header of a timed-MFU includes movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter, while the header of a non-timed-MFU includes item_iD. The meaning of each field is specified in standards such as ISO / IEC23008-1 or ARIB STD-B60. The meaning of each field specified in such standards will be explained below.

[0479] The movie_fragment_sequence_number indicates the sequence number of the movie fragment to which the MFU belongs, and is also specified in ISO / IEC14496-12.

[0480] The sample_number indicates the sample number to which the MFU belongs, and is also specified in ISO / IEC14496-12.

[0481] The offset indicates the offset amount of the MFU in the sample to which the MFU belongs, in bytes.

[0482] The priority indicates the relative importance of the MFU in the MPU to which the MFU belongs, and an MFU with a larger priority number is more important than an MFU with a smaller priority number.

[0483] The dependency_counter indicates the number of MFUs whose decoding process depends on the MFU (i.e., the number of MFUs whose decoding process cannot be performed unless the MFU is decoded). For example, when the MFU is HEVC and a B picture or a P picture refers to an I picture, the B picture or the P picture cannot be decoded unless the I picture is decoded.

[0484] Therefore, when the MFU is in sample units, the dependency_counter in the MFU of an I-picture indicates the number of pictures that reference the I-picture. When the MFU is in NAL unit units, the dependency_counter in the MFU belonging to the I-picture indicates the number of NAL units that belong to the picture that references the I-picture. Furthermore, in the case of a video signal that has been temporally hierarchically coded, the MFU of the enhancement layer depends on the MFU of the base layer, so the dependency_counter in the MFU of the base layer indicates the number of MFUs of the enhancement layer. This field can only be generated after the number of dependent MFUs has been determined.

[0485] The item_iD indicates an identifier that uniquely identifies the item.

[0486] [MP4 non-support mode] As explained in Figures 19 and 21, the transmitting device can transmit the MPU in MMT by transmitting MPU metadata or MF metadata before or after the media data, or by transmitting only the media data. In addition, the receiving device can decode using a receiving device or method that complies with MP4, or by decoding without using a header.

[0487] As a method of transmitting data specialized for broadcast stream playback, there is, for example, a transmission method that does not support MP4 reconstruction in a receiving device.

[0488] An example of a transmission method that does not support MP4 reconstruction in a receiving device is a method that does not transmit metadata (MPU metadata and MF metadata), as shown in (b) of Figure 21. In this case, the field value of the fragment type (information indicating the type of data unit) included in the MMTP packet is fixed to 2 (=MFU).

[0489] If metadata is not transmitted, as explained above, an MP4-compliant receiving device cannot decode the received data as MP4, but it can decode it without using the metadata (header).

[0490] Therefore, metadata is not necessarily essential information for decoding and playing back a broadcast stream. Similarly, the information in the data unit header in timed-MFU, as explained in Fig. 48, is information for reconstructing MP4 in a receiving device. Since there is no need to reconstruct MP4 for broadcast stream playback, the information in the data unit header in timed-MFU (hereinafter also referred to as timed-MFU header) is not necessarily information required for broadcast stream playback.

[0491] A receiving device can easily reconstruct an MP4 file by using the metadata and the information for reconstructing an MP4 file in the data unit header (hereinafter also referred to as MP4 configuration information). However, a receiving device cannot reconstruct an MP4 file even if only one of the metadata and the MP4 configuration information in the data unit header is transmitted. There is little benefit to transmitting only one of the metadata and the information for reconstructing an MP4 file, and generating and transmitting unnecessary information increases processing and reduces transmission efficiency.

[0492] Therefore, the transmitting device controls the data structure and transmission of the MP4 configuration information using the following method. The transmitting device determines whether to indicate the MP4 configuration information in the data unit header based on whether metadata is transmitted. Specifically, if metadata is transmitted, the transmitting device indicates the MP4 configuration information in the data unit header, and if metadata is not transmitted, the transmitting device does not indicate the MP4 configuration information in the data unit header.

[0493] As a method for not indicating MP4 configuration information in the data unit header, for example, the following method can be used.

[0494] 1. The transmitting device sets the MP4 configuration information as reserved and does not use it. This reduces the amount of processing on the sending side (the amount of processing on the transmitting device) that generates the MP4 configuration information.

[0495] 2. The transmitting device deletes the MP4 configuration information and compresses the header, which reduces the amount of processing on the sending side that generates the MP4 configuration information and also reduces transmission capacity.

[0496] When the transmitting device deletes the MP4 configuration information and compresses the header, the transmitting device may indicate a flag indicating that the MP4 configuration information has been deleted (compressed). The flag is indicated in the header (MMTP packet header, MMTP payload header, data unit header) or control information.

[0497] Furthermore, information as to whether metadata is transmitted may be determined in advance, or may be separately signaled in the header or control information and transmitted to the receiving device.

[0498] For example, the MFU header may store information indicating whether metadata corresponding to the MFU has been transmitted.

[0499] On the other hand, the receiving device can determine whether MP4 configuration information is indicated based on whether metadata is transmitted.

[0500] Here, if the data transmission order (for example, an order such as MPU metadata, MF metadata, and media data) is fixed, the receiving device may make a determination based on whether the metadata is received before the media data.

[0501] If MP4 configuration information is indicated, the receiving device can use the MP4 configuration information to reconstruct the MP4, or the receiving device can use the MP4 configuration information to detect the beginning of other access units or NAL units.

[0502] The MP4 configuration information may be the entire timed-MFU header or a part of it.

[0503] Similarly, the transmitting device determines whether metadata is transmitted in the non-timed-MFU header. You may decide whether to show the id.

[0504] The transmitting device may indicate MP4 configuration information in only one of timed-MFU and non-timed-MFU. If the transmitting device indicates MP4 configuration information in only one of timed-MFU and non-timed-MFU, the transmitting device determines whether to indicate MP4 configuration information based on whether metadata is transmitted and whether the MFU is timed or non-timed. The receiving device can determine whether to indicate MP4 configuration information based on whether metadata is transmitted and the timed / non-timed flag.

[0505] In the above description, the transmitting device determines whether to indicate the MP4 configuration information based on whether the metadata (both the MPU metadata and the MF metadata) is transmitted. However, the transmitting device may not indicate the MP4 configuration information if some of the metadata (either the MPU metadata or the MF metadata) is not transmitted.

[0506] The transmitting device may also determine whether to indicate MP4 configuration information based on information other than metadata.

[0507] For example, modes such as an MP4 support mode / an MP4 non-support mode may be defined, and the transmitting device may indicate MP4 configuration information in a data unit header in the MP4 support mode, and not indicate MP4 configuration information in the data unit header in the MP4 non-support mode. Also, the transmitting device may transmit metadata and indicate MP4 configuration information in a data unit header in the MP4 support mode, and not transmit metadata and not indicate MP4 configuration information in the data unit header in the MP4 non-support mode.

[0508] [Transmitter operation flow] Next, the operation flow of the transmitting device will be described with reference to Figure 49.

[0509] The transmitting device first determines whether to transmit metadata (S1001). If the transmitting device determines to transmit metadata (Yes in S1002), the transmitting device proceeds to step S1003, generates MP4 configuration information, stores it in a header, and transmits it (S1003). In this case, the transmitting device also generates and transmits metadata.

[0510] On the other hand, if the transmitting device determines not to transmit metadata (No in S1002), it transmits the MP4 configuration information without generating it and storing it in the header (S1004). In this case, the transmitting device does not generate or transmit metadata.

[0511] Note that whether or not to transmit metadata in step S1001 may be determined in advance, or may be determined based on whether metadata has been generated within the transmitting device or whether metadata is being transmitted within the transmitting device.

[0512] [Operation flow of receiving device] Next, the operation flow of the receiving device will be explained, as shown in Figure 50.

[0513] The receiving device first determines whether metadata is being transmitted (S1101). Whether metadata is being transmitted can be determined by monitoring the fragment type in the MMTP packet payload. Alternatively, whether metadata is being transmitted may be determined in advance.

[0514] If the receiving device determines that metadata has been transmitted (Yes in S1102), it reconstructs the MP4 and executes a decoding process using the MP4 configuration information (S1103). On the other hand, if the receiving device determines that metadata has not been transmitted (No in S1102), it does not reconstruct the MP4 and executes a decoding process without using the MP4 configuration information (S1104).

[0515] In addition, using the methods described above, the receiving device can detect random access points, the beginning of access units, the beginning of NAL units, etc. without using MP4 configuration information, and can perform decoding processes, packet loss detection, and recovery from packet loss.

[0516] For example, the beginning of an access unit is the beginning data of an MMT payload in which the aggregation_flag value is 1. In this case, the fragmentation_indicator value is 0.

[0517] In addition, the beginning of a slice segment is the beginning data of an MMT payload in which the aggregation_flag value is 0 and the fragmentation_indicator value is 00 or 01.

[0518] Based on the above information, the receiving device can detect the beginning of an access unit and a slice segment.

[0519] In addition, the receiving device may analyze the NAL unit header in a packet including the beginning of a data unit whose fragmentation_indicator value is 00 or 01, and detect that the type of NAL unit is an AU delimiter and that the type of NAL unit is a slice segment.

[0520] [Broadcast Simple Mode] So far, we have described a method for transmitting data specialized for broadcast stream playback that does not support MP4 configuration information in the receiving device, but the method for transmitting data specialized for broadcast stream playback is not limited to this.

[0521] As a method for transmitting data specialized for broadcast stream playback, for example, the following method may be used.

[0522] In a fixed broadcast reception environment, the transmitting device does not need to use AL-FEC. If AL-FEC is not used, the FEC_type in the MMTP packet header is always fixed to 0.

[0523] The transmitting device may always use AL-FEC in a mobile broadcast reception environment and in the communication UDP transmission mode. When AL-FEC is used, the FEC_type in the MMTP packet header is always 0 or 1.

[0524] The transmitting device may not transmit assets in bulk. If assets are not transmitted in bulk, location_infolocation, which indicates the number of asset transmission locations within the MPT, may be fixed to 1.

[0525] · The transmitting device does not have to perform hybrid transmission of assets, programs, and messages.

[0526] Also, for example, if a broadcast simple mode is defined, the transmitting device may set the MP4 non-support mode when in the broadcast simple mode, or may use the data transmission method specialized for broadcast stream playback described above. Whether the mode is the broadcast simple mode may be determined in advance, or the transmitting device may store a flag indicating the broadcast simple mode as control information and transmit it to the receiving device.

[0527] In addition, the transmitting device may, based on whether the MP4 non-support mode is in effect (whether metadata is transmitted) as described in Figure 49, use the data transmission method specialized for broadcast stream playback shown above as broadcast simple mode if the MP4 non-support mode is in effect.

[0528] When the receiving device is in broadcast simple mode, it is considered to be in MP4 non-support mode and can perform decoding processing without reconstructing the MP4.

[0529] Furthermore, when the receiving device is in the broadcast simple mode, it determines that the function is specialized for broadcasting, and can perform reception processing specialized for broadcasting.

[0530] As a result, when in broadcast simple mode, by using only functions specialized for broadcasting, not only can unnecessary processing be reduced for the transmitting device and receiving device, but transmission overhead can also be reduced by not compressing and transmitting unnecessary information.

[0531] When the MP4 non-support mode is used, hint information supporting a storage method other than the MP4 format may be indicated.

[0532] Storage methods other than MP4 configuration include, for example, directly storing MMT packets or IP packets, or converting MMT packets into MPEG-2 TS packets.

[0533] In the case of an MP4 non-support mode, a format that does not conform to the MP4 structure may be used.

[0534] For example, in the case of an MP4 non-support mode, the data stored in the MFU may be in a format with a byte start code at the beginning of the NAL unit, rather than in the MP4 format where the size of the NAL unit is added to the beginning of the NAL unit.

[0535] In MMT, the asset type indicating the type of asset is described in 4CC registered in MP4REG (http: / / www.mp4ra.org), and when HEVC is used as the video signal, 'HEV1' or 'HVC1' is used. 'HVC1' is a format that may include parameter sets in samples, while 'HEV1' is a format that does not include parameter sets in samples but includes parameter sets in the sample entries in the MPU metadata.

[0536] In the case of broadcast simple mode or MP4 non-support mode, if MPU metadata and MF metadata are not transmitted, it may be specified that a parameter set must be included in the sample. Also, whether 'HEV1' or 'HVC1' is indicated in the asset type, it may be specified that the format must be 'HVC1'.

[0537] [Supplement 1: Transmitter] As described above, when metadata is not transmitted, the MP4 configuration information is set to reserved, and a transmitting device that is not in operation can also be configured as shown in Fig. 51. Fig. 51 is a diagram showing an example of a specific configuration of a transmitting device.

[0538] The transmitting device 300 includes an encoding unit 301, an assigning unit 302, and a transmitting unit 303. Each of the encoding unit 301, the assigning unit 302, and the transmitting unit 303 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0539] The encoding unit 301 encodes a video signal or an audio signal to generate sample data, which is specifically a data unit.

[0540] The adding unit 302 adds header information including MP4 configuration information to sample data, which is data obtained by encoding a video signal or an audio signal. The MP4 configuration information is information for reconstructing the sample data as an MP4 format file on the receiving side, and the content of the information varies depending on whether the presentation time of the sample data is specified.

[0541] As described above, the attachment unit 302 includes MP4 configuration information such as movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter in the header (header information) of timed-MFU, which is an example of sample data with a defined presentation time (sample data including information about synchronization).

[0542] On the other hand, the attachment unit 302 includes MP4 configuration information such as item_id in the header (header information) of timed-MFU, which is an example of sample data for which the presentation time is not specified (sample data that does not include information about synchronization).

[0543] Then, when the transmitting unit 303 does not transmit metadata corresponding to the sample data (for example, in the case of (b) in Figure 21), the attaching unit 302 attaches header information that does not include MP4 configuration information to the sample data, depending on whether the presentation time of the sample data is specified or not.

[0544] Specifically, when the presentation time of the sample data is determined, the attachment unit 302 attaches header information that does not include the first MP4 configuration information to the sample data, and when the presentation time of the sample data is not determined, the attachment unit 302 attaches header information that includes the second MP4 configuration information to the sample data.

[0545] 49, when the transmitting unit 303 does not transmit metadata corresponding to the sample data, the adding unit 302 sets the MP4 configuration information to reserved (a fixed value), thereby not substantially generating the MP4 configuration information and not substantially storing the MP4 configuration information in the header (header information). Note that the metadata includes MPU metadata and movie fragment metadata.

[0546] The transmitting unit 303 transmits the sample data to which the header information has been added. More specifically, the transmitting unit 303 packetizes the sample data to which the header information has been added in accordance with the MMT method and transmits the packetized data.

[0547] As described above, in the transmission method and reception method specialized for playing back a broadcast stream, the receiving device does not need to reconstruct data units into MP4. If the receiving device does not need to reconstruct data into MP4, the processing load of the transmitting device is reduced by not generating unnecessary information such as MP4 configuration information.

[0548] On the other hand, the transmitting device must transmit the necessary information, but must maintain compatibility with the standard so that it does not have to transmit any additional information separately.

[0549] With a configuration such as that of the transmitting device 300, by setting the area in which the MP4 configuration information is stored to a fixed value, the MP4 configuration information is not transmitted, and only necessary information is transmitted based on the standard, thereby eliminating the need to transmit unnecessary additional information. In other words, the configuration of the transmitting device and the amount of processing performed by the transmitting device can be reduced. Furthermore, since unnecessary data is not transmitted, transmission efficiency can be improved.

[0550] [Supplement 2: Receiving device] Furthermore, a receiving device corresponding to transmitting device 300 may be configured, for example, as shown in Fig. 52. Fig. 52 is a diagram showing another example of the configuration of a receiving device.

[0551] The receiving device 400 includes a receiving unit 401 and a decoding unit 402. The receiving unit 401 and the decoding unit 402 are realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0552] The receiving unit 401 receives sample data, which is data in which a video signal or an audio signal has been encoded, and which is provided with header information including MP4 configuration information for reconstructing the sample data as a file in MP4 format.

[0553] If the receiving unit does not receive metadata corresponding to the sample data and the presentation time of the sample data is determined, the decoding unit 402 decodes the sample data without using the MP4 configuration information.

[0554] For example, as shown in step S1104 in FIG. 50, if the receiving unit 401 does not receive metadata corresponding to the sample data, the decoding unit 402 performs the decoding process without using the MP4 configuration information.

[0555] This allows the configuration of the receiving device 400 and the amount of processing in the receiving device 400 to be reduced.

[0556] (Fourth embodiment) [overview] In the fourth embodiment, a method for storing asynchronous (non-timed) media that does not contain information related to synchronization, such as files, in an MPU and a method for transmitting it in an MMTP packet will be described. Note that in the fourth embodiment, an MPU in MMT will be used as an example, but the method can also be applied to DASH, which is also MP4-based.

[0557] First, the details of how non-timed media (hereinafter also referred to as "asynchronous media data") is stored in the MPU will be explained using Figure 53. Figure 53 shows how non-timed media is stored in the MPU and how it is transmitted in MMTP packets.

[0558] An MPU that stores non-timed media consists of boxes such as ftyp, mmpu, moov, and meta, and stores information about the files stored in the MPU. Multiple idat boxes can be stored in a meta box, and an idat box stores one file as an item.

[0559] A part of the ftyp, mmpu, moov, and meta boxes constitute one data unit as MPU metadata, and the item or idat box constitutes a data unit as MFU.

[0560] After the data units are aggregated or fragmented, they are given a data unit header, an MMTP payload header, and an MMTP packet header, and then transmitted as an MMTP packet.

[0561] Note that Figure 53 shows an example in which File #1 and File #2 are stored in one MPU. The MPU metadata is not divided, and the MFU is divided and stored in an MMTP packet, but this is not limited to this, and data units may be aggregated or fragmented depending on the size of the data unit. Also, the MPU metadata does not have to be transmitted, in which case only the MFU is transmitted.

[0562] Header information such as the data unit header indicates an itemID (an identifier that uniquely identifies an item), and the MMTP payload header and MMTP packet header include a packet sequence number (a sequence number for each packet) and an MPU sequence number (a sequence number for the MPU, a number unique within an asset).

[0563] Note that the data structures of the MMTP payload header and MMTP packet header other than the data unit header are the same as those of the timed media (hereinafter also referred to as "synchronized media data") described so far, and include an aggregation_flag, fragmentation_indicator, fragment_counter, etc.

[0564] Next, a specific example of the header information when a file (= Item = MFU) is divided and packetized will be described using FIGS. 54 and 55.

[0565] FIGS. 54 and 55 are diagrams showing examples of packetizing and transmitting for each of a plurality of divided data obtained by dividing a file. FIGS. 54 and 55 specifically show information (packet sequence number, fragment counter, fragmentation indicator, MPU sequence number, item ID) included in any of the data unit header, MMTP payload header, and MMTP packet header, which is the header information for each divided MMTP packet. Note that FIG. 54 shows an example in which File #1 is divided into M (M <= 256) parts, and FIG. 55 shows an example in which File #2 is divided into N (256 < N) parts.

[0566] The divided data number indicates the index of the divided data from the beginning of the file, and this information is not transmitted. That is, the divided data number is not included in the header information. Also, the divided data number is a number assigned to each packet corresponding to each of the plurality of divided data obtained by dividing the file, and is a number assigned by incrementing by 1 in ascending order from the first packet.

[0567] The packet sequence number is the sequence number of packets having the same packet ID. In FIGS. 54 and 55, assuming the divided data at the beginning of the file is A, consecutive numbers are assigned up to the divided data at the end of the file. The packet sequence number is a number assigned by incrementing by 1 in ascending order from the divided data at the beginning of the file, and is a number corresponding to the divided data number.

[0568] The fragment counter indicates the number of split data pieces that come after the split data piece in question among the multiple split data pieces obtained by splitting a single file. Furthermore, if the number of split data pieces, which is the number of split data pieces obtained by splitting a single file, exceeds 256, the fragment counter indicates the remainder when the number of split data pieces is divided by 256. In the example of Figure 54, the number of split data pieces is 256 or less, so the field value of the fragment counter is (M-split data number). On the other hand, in the example of Figure 55, the number of split data pieces exceeds 256, so the value obtained by dividing (N-split data number) by 256 is ((N-split data number)%256).

[0569] The fragmentation indicator indicates the fragmentation state of the data stored in the MMTP packet, and is a value indicating whether the fragment is the first fragment of a fragmented data unit, the last fragment, any other fragment, or one or more unfragmented data units. Specifically, the fragmentation indicator is "01" for the first fragment, "11" for the last fragment, "10" for the remaining fragments, and "00" for unfragmented data units.

[0570] In this embodiment, when the number of divided data pieces exceeds 256, it is explained as indicating the remainder when the number of divided data pieces is divided by 256, but the number of divided data pieces is not limited to 256 and may be another number (predetermined number).

[0571] 54 and 55, when a file is divided and conventional header information is added to each of the multiple data segments obtained by dividing the file and transmitted, the receiving device does not have information that can be used to determine the position of the data segment in the original file (data segment number) that the data stored in the received MMTP packet is, the number of data segments in the file, or the data segment numbers and number of data segments. For this reason, with conventional transmission methods, even if an MMTP packet is received, it is not possible to uniquely detect the data segment number or number of data segments for the data stored in the received MMTP packet.

[0572] For example, if the number of divided data pieces is 256 or less, as shown in Figure 54, and it is known in advance that the number of divided data pieces is 256 or less, it is possible to identify the divided data piece numbers and the number of divided data pieces by referencing the fragment counter. However, if the number of divided data pieces is 256 or more, it is not possible to identify the divided data piece numbers and the number of divided data pieces.

[0573] Furthermore, if the number of divided data segments of a file is limited to 256 or less, and the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed, and in this case the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes]. Therefore, if you want to transmit a file that is larger than 1 [Mbytes], you cannot limit the number of divided data segments of the file to 256 or less.

[0574] Furthermore, for example, since the first and last fragments of a file can be detected by referencing the fragmentation indicator, it is possible to count the number of MMTP packets until the MMTP packet containing the last fragment of the file is received, or to calculate the fragment number and number of fragments by combining it with the packet sequence number after receiving the MMTP packet containing the last fragment of the file. Therefore, the fragment number and number of fragments may be signaled by combining the fragmentation indicator and the packet sequence number. However, if reception starts from an MMTP packet containing fragment data in the middle of a file (i.e., fragment data that is neither the first nor the last fragment of the file), the fragment number and number of fragments of that fragment cannot be identified. The fragment number and number of fragments of that fragment can only be identified after receiving the MMTP packet containing the last fragment of the file.

[0575] To address the issue described in Figures 54 and 55, that is, to uniquely determine the data segment number and number of data segments of a file when a packet containing the file's data segment is received midway, the following method is used.

[0576] First, the divided data number will be explained.

[0577] For the divided data number, the packet sequence number in the divided data at the beginning of the file (item) is signaled.

[0578] As a signaling method, it is stored in the control information that manages the file. Specifically, in Figures 54 and 55, the packet sequence number A of the divided data at the beginning of the file is stored in the control information. The receiving device obtains the value of A from the control information and calculates the divided data number from the packet sequence number indicated in the packet header.

[0579] The divided data number of the divided data is obtained by subtracting the packet sequence number A of the first divided data from the packet sequence number of the divided data.

[0580] An example of control information for managing files is the asset management table specified in ARIB STD-B60. The asset management table indicates the file size, version information, etc. for each file, and is stored in a data transmission message for transmission. Figure 56 shows the syntax of a loop for each file in the asset management table.

[0581] If the area of ​​the existing asset management table cannot be expanded, signaling may be performed using a 32-bit area in part of the item_info_byte field that indicates item information. A flag indicating whether the packet sequence number in the first divided data of the file (item) is indicated in part of the item_info_byte area may be indicated in, for example, a reserved_future_use field of the control information.

[0582] When a file is repeatedly transmitted, such as in a data carousel, multiple packet sequence numbers may be indicated, or the packet sequence number of the first packet of the file to be transmitted immediately after may be indicated.

[0583] The packet sequence number is not limited to the packet sequence number of the divided data at the beginning of the file, and may be any information that links the divided data number of the file with the packet sequence number.

[0584] Next, the number of divided data items will be described.

[0585] The order of loops for each file included in the asset management table may be defined as the transmission order of the files. This allows the first packet sequence numbers of two consecutive files in transmission order to be known, and the number of divided data pieces of the previously transmitted file can be determined by subtracting the first packet sequence number of the previously transmitted file from the first packet sequence number of the later transmitted file. That is, for example, if File #1 shown in Figure 54 and File #2 shown in Figure 55 are consecutive files in this order, the last packet sequence number of File #1 and the first packet sequence number of File #2 are assigned consecutive numbers.

[0586] The number of divided data pieces of a file may also be specified by specifying a file division method. For example, if the number of divided data pieces is N, the size of each of the 1st to (N-1)th divided data pieces is set to L, and the size of the Nth divided data piece is specified as a fraction (item_size-L*(N-1)), so that the number of divided data pieces can be calculated backward from the item_size shown in the asset management table. In this case, the integer value obtained by rounding up (item_size / L) becomes the number of divided data pieces. However, the file division method is not limited to this.

[0587] The number of divided data items may also be stored directly in the asset management table.

[0588] By using the above method, the receiving device receives the control information and calculates the number of divided data pieces based on the control information. It can also calculate a packet sequence number corresponding to the divided data piece number of the file based on the control information. If the timing of receiving the divided data packets is earlier than the timing of receiving the control information, the divided data piece number and the number of divided data pieces may be calculated at the timing of receiving the control information.

[0589] When the fragment data number or the number of fragment data is signaled using the above method, the fragment data number or the number of fragment data is not identified based on the fragment counter, and the fragment counter becomes unnecessary data. Therefore, in the transmission of asynchronous media, when information that can identify the fragment data number and the number of fragment data is signaled using the above method, the fragment counter may not be used, or header compression may be performed. This reduces the processing load of the transmitting device and receiving device, and also improves transmission efficiency. In other words, when transmitting asynchronous media, the fragment counter may be reserved (disabled). Specifically, the value of the fragment counter may be set to a fixed value, for example, "0." Furthermore, when receiving asynchronous media, the fragment counter may be ignored.

[0590] When storing synchronous media such as video and audio, the order in which MMTP packets are sent at the sending device matches the order in which they arrive at the receiving device, and packets are not retransmitted. In such a case, if there is no need to detect packet loss and reconstruct packets, the fragment counter may not be used. In other words, in this case, the fragment counter may be reserved (disabled).

[0591] In addition, it is possible to detect random access points, the beginning of access units, the beginning of NAL units, etc. without using a fragment counter, and it is possible to perform decoding processes, detect packet loss, and recover from packet loss.

[0592] Furthermore, transmission of real-time content such as live broadcasts requires even lower latency, and requires that data be packetized and transmitted sequentially starting from the data that has been completely encoded. However, in the transmission of real-time content, conventional fragment counters cannot determine the number of divided data pieces when transmitting the first divided data piece, so the first divided data piece is transmitted after all the encoding of the data unit has been completed and the number of divided data pieces has been determined, resulting in a delay. Even in such cases, this delay can be reduced by using the above method and not operating a fragment counter.

[0593] FIG. 57 shows the operational flow for identifying the divided data number in the receiving device.

[0594] The receiving device acquires control information that describes file information (S1201). The receiving device determines whether the control information indicates the packet sequence number of the beginning of the file (S1202), and if the control information indicates the packet sequence number of the beginning of the file (Yes in S1202), calculates the packet sequence number that corresponds to the divided data number of the divided data of the file (S1203). Then, after acquiring MMTP packets that store the divided data, the receiving device identifies the divided data number of the file from the packet sequence number stored in the packet header of the acquired MMTP packet (S1204).

[0595] On the other hand, if the control information does not indicate the packet sequence number at the beginning of the file (No in S1202), the receiving device acquires the MMTP packet containing the last divided data of the file, and then identifies the divided data number using the fragment indicator stored in the packet header of the acquired MMTP packet and the packet sequence number (S1205).

[0596] FIG. 58 shows the operational flow for identifying the number of divided data pieces in the receiving device.

[0597] The receiving device acquires control information that describes file information (S1301). The receiving device determines whether the control information includes information that allows the number of divided data pieces of the file to be calculated (S1302). If it determines that the control information includes information that allows the number of divided data pieces to be calculated (Yes in S1302), it calculates the number of divided data pieces based on the information included in the control information (S1303). On the other hand, if it determines that the number of divided data pieces cannot be calculated (No in S1302), the receiving device acquires an MMTP packet that includes the last divided data piece of the file, and then identifies the number of divided data pieces using the fragment indicator and packet sequence number stored in the packet header of the acquired MMTP packet (S1304).

[0598] FIG. 59 shows an operational flow for determining whether to operate a fragment counter in a transmitting device.

[0599] First, the transmitting device determines whether the media to be transmitted (hereinafter also referred to as "media data") is synchronous media or asynchronous media (S1401).

[0600] If the result of the determination in step S1401 is synchronous media (Synchronous Media in S1402), the transmitting device determines whether the order of MMTP packets sent and received matches in the environment in which the synchronous media is transmitted, and whether packet reassembly is unnecessary in the event of packet loss (S1403). If the transmitting device determines that it is unnecessary (Yes in S1403), it does not operate a fragment counter (S1404). On the other hand, if the transmitting device determines that it is not unnecessary (No in S1403), it operates a fragment counter (S1405).

[0601] If the result of the determination in step S1401 is asynchronous media (Asynchronous Media in S1402), the transmitting device determines whether or not to operate a fragment counter based on whether the fragment data number and the number of fragment data are signaled using the method described above. Specifically, if the fragment data number and the number of fragment data are signaled (Yes in S1406), the transmitting device does not operate a fragment counter (S1404). On the other hand, if the fragment data number and the number of fragment data are not signaled (No in S1406), the transmitting device operates a fragment counter (S1405).

[0602] In addition, if the transmitting device does not operate a fragment counter, it may set the value of the fragment counter to reserved or may perform header compression.

[0603] In addition, the transmitting device may determine whether to signal the above-mentioned fragment data number and number of fragment data based on whether or not a fragment counter is operated.

[0604] If the synchronous media does not operate a fragment counter, the transmitting device may signal the divided data number and the number of divided data using the method described above for the asynchronous media. Conversely, the operation of the synchronous media may be determined based on whether the asynchronous media operates a fragment counter. In this case, the synchronous media and the asynchronous media can be operated in the same manner regarding whether or not to operate fragments.

[0605] Next, a method for identifying the number of divided data pieces and the divided data numbers (when a fragment counter is used) will be described. Figure 60 is a diagram for explaining a method for identifying the number of divided data pieces and the divided data numbers (when a fragment counter is used).

[0606] As explained using Figure 54, if the number of divided data pieces is 256 or less and it is known in advance that the number of divided data pieces is 256 or less, it is possible to identify the divided data piece number and the number of divided data pieces by referring to the fragment counter.

[0607] If the number of divided data segments in a file is limited to 256 or less, and the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed, and in this case the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes].

[0608] If the file size exceeds the maximum size of a transmittable file, the file is split in advance so that each split file is no larger than x*256 bytes. Each of the multiple split files obtained by splitting the file is treated as a single file (item) and is further split into no more than 256 files. Each of the split data obtained by further splitting is stored in an MMTP packet and transmitted.

[0609] Note that information indicating that the item is a split file, the number of split files, and the sequence numbers of the split files may be stored in the control information and transmitted to the receiving device. This information may also be stored in the asset management table, or may be indicated using part of the existing field item_info_byte.

[0610] When an item is one of multiple split files obtained by splitting a single file, the receiving device can identify the other split files and reconstruct the original file. Furthermore, the receiving device can uniquely identify the number of split data pieces and the split data numbers by using the number of split files, the split file index, and the fragment counter in the control information. Furthermore, the number of split data pieces and the split data numbers can be uniquely identified without using packet sequence numbers, etc.

[0611] Here, it is desirable that the item_id of each of the split files obtained by splitting one file is the same. If a different item_id is assigned, the item_id of the first split file may be indicated in order to uniquely refer to the file from other control information, etc.

[0612] Alternatively, multiple split files may always belong to the same MPU. When multiple files are stored in an MPU, files of different types may not be stored, and files that are split from a single file may always be stored. A receiving device can detect file updates by checking the version information for each MPU, without checking the version information for each item.

[0613] Figure 61 shows the operational flow of a transmitting device when utilizing a fragment counter.

[0614] First, the transmitting device checks the size of the file to be transmitted (S1501). Next, the transmitting device determines whether the file size exceeds x*256 [bytes] (x is the data size that can be transmitted in one packet, for example, the MTU size) (S1502). If the file size exceeds x*256 [bytes] (Yes in S1502), the transmitting device divides the file so that the size of each divided file is less than x*256 [bytes] (S1503). Then, the transmitting device transmits the divided files as items, and transmits information about the divided files (for example, the fact that they are divided files, the sequence numbers in the divided files, etc.) in control information (S1504). On the other hand, if the file size is less than x*256 [bytes] (No in S1502), the transmitting device transmits the file as an item as usual (S1505).

[0615] FIG. 62 shows the operational flow of a receiving device when utilizing a fragment counter.

[0616] First, the receiving device acquires and analyzes control information related to file transmission, such as an asset management table (S1601). Next, the receiving device determines whether the desired item is a split file (S1602). If the receiving device determines that the desired file is a split file (Yes in S1602), it acquires information for reconstructing the file, such as the split file and its index, from the control information (S1603). Then, the receiving device acquires the items that make up the split file and reconstructs the original file (S1604). On the other hand, if the receiving device determines that the desired file is not a split file (No in S1602), it acquires the file as usual (S1605).

[0617] In short, the transmitting device signals the packet sequence number of the divided data at the beginning of the file. The transmitting device also signals information that can identify the number of divided data. Alternatively, the transmitting device defines a fragmentation rule that can identify the number of divided data. The transmitting device also performs reserved or header compression without using a fragment counter.

[0618] When the packet sequence number of the data at the beginning of the file is signaled, the receiving device determines the divided data number and the number of divided data from the packet sequence number of the divided data at the beginning of the file and the packet sequence number of the MMTP packet.

[0619] From another perspective, the transmitting device divides a file, divides data into individual divided files, and transmits the divided files by signaling information linking the divided files (such as a sequence number and the number of divisions).

[0620] The receiving device identifies the divided data number and the number of divided data pieces based on the fragment counter and the sequence number of the divided file.

[0621] This allows the divided data number and divided data to be uniquely identified. In addition, since the divided data number of a divided data item can be identified when the divided data item is received, waiting time and memory usage can be reduced.

[0622] Furthermore, by not using a fragment counter, the configuration of the transmitting / receiving device can reduce the amount of processing and improve transmission efficiency.

[0623] Figure 63 shows a service configuration in which the same program is transmitted over multiple IP data flows. This example shows a case in which part of the data (video and audio) of a program with service ID = 2 is transmitted over an IP data flow using the MMT method, and data with the same service ID but different from the part of the data is transmitted over an IP data flow using the advanced BS data transmission method (in this example, the file transmission protocols are different, but may be the same protocol).

[0624] The transmitting device multiplexes the IP data so as to ensure that data consisting of multiple IP data flows is ready by the time of decoding at the receiving device.

[0625] The receiving device can realize guaranteed receiver operation by processing data consisting of multiple IP data flows based on the decoding time.

[0626] [Supplementary information: Transmitting and receiving devices] As described above, a transmitting device that transmits data without operating a fragment counter can also be configured as shown in Figure 64. Also, a receiving device that receives data without operating a fragment counter can also be configured as shown in Figure 65. Figure 64 is a diagram showing an example of a specific configuration of a transmitting device. Figure 65 is a diagram showing an example of a specific configuration of a receiving device.

[0627] The transmitting device 500 includes a dividing unit 501, a composing unit 502, and a transmitting unit 503. Each of the dividing unit 501, the composing unit 502, and the transmitting unit 503 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0628] The receiving device 600 includes a receiving unit 601, a determining unit 602, and a configuration unit 603. The receiving unit 601, the determining unit 602, and the configuration unit 603 are each realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0629] Detailed explanations of each component of the transmitting device 500 and the receiving device 600 will be given in the explanations of the transmitting method and the receiving method, respectively.

[0630] First, the transmission method will be described with reference to Fig. 66. Fig. 66 shows the operation flow (transmission method) performed by the transmission device.

[0631] First, the division unit 501 of the transmission device 500 divides data into a plurality of divided data (S1701).

[0632] Next, the configuration unit 502 of the transmitting device 500 configures a plurality of packets by adding header information to each of the plurality of divided data and packetizing the data (S1702).

[0633] Then, the transmitting unit 503 of the transmitting device 500 transmits the configured plurality of packets (S1703). The transmitting unit 503 transmits the divided data information and the value of the invalidated fragment counter. The divided data information is information for identifying the divided data number and the number of divided data. The divided data number is a number indicating the ordinal number of the divided data among the plurality of divided data. The number of divided data is the number of the plurality of divided data.

[0634] This allows the amount of processing by the transmission device 500 to be reduced.

[0635] Next, the receiving method will be described with reference to Fig. 67. Fig. 67 shows the operation flow (receiving method) of the receiving device.

[0636] First, the receiving unit 601 of the receiving device 600 receives a plurality of packets (S1801).

[0637] Next, the determination unit 602 of the receiving device 600 determines whether or not divided data information has been acquired from the received plurality of packets (S1802).

[0638] Then, if the determination unit 602 determines that the divided data information has been acquired (Yes in S1802), the construction unit 603 of the receiving device 600 constructs data from the multiple received packets without using the value of the fragment counter included in the header information (S1803).

[0639] On the other hand, if the judgment unit 602 determines that the divided data information has not been acquired (No in S1802), the construction unit 603 may construct data from the multiple received packets using the value of the fragment counter included in the header information (S1804).

[0640] This allows the amount of processing by the receiving device 600 to be reduced.

[0641] (Embodiment 5) [overview] In the fifth embodiment, a method for transmitting transport packets (TLV packets) when NAL units are stored in the multiplexing layer in the NAL size format will be described.

[0642] As described in the first embodiment, there are two types of formats for storing H.264 or H.265 NAL units in the multiplexing layer. One is a format called byte stream format, in which a start code consisting of a specific bit string is added immediately before the NAL unit header. The other is a format called NAL size format, in which a field indicating the size of the NAL unit is added. The byte stream format is used in MPEG-2 systems and RTP, and the NAL size format is used in MP4, or DASH and MMT that use MP4.

[0643] In the byte stream format, the start code consists of three bytes, and an optional byte (a byte whose value is 0) can be added.

[0644] On the other hand, in the general NAL size format in MP4, size information is indicated as one byte, two bytes, or four bytes. This size information is indicated by the lengthSizeMinusOne field in the HEVC sample entry. A value of "0" in this field indicates one byte, "1" indicates two bytes, and "3" indicates four bytes.

[0645] In ARIB STD-B60 "Media Transport Method Using MMT in Digital Broadcasting," standardized in July 2014, when storing NAL units in the multiplexing layer, if the output of the HEVC encoder is a byte stream, the byte start code is removed and the size of the NAL unit in bytes, expressed as a 32-bit (unsigned integer), is added immediately before the NAL unit as length information. Note that MPU metadata including HEVC sample entries is not transmitted, and the size information is fixed at 32 bits (4 bytes).

[0646] In addition, ARIB STD-B60 "Media Transport Method Using MMT in Digital Broadcasting" specifies that in the reception buffer model that a transmitting device takes into account during transmission to ensure buffer operation in a receiving device, the pre-decoding buffer for video signals is CPB.

[0647] However, there are the following issues: CPB in the MPEG-2 system and HRD in HEVC are specified on the assumption that the video signal is in byte stream format. Therefore, for example, if transmission packet rate control is performed on the assumption that the video signal is in byte stream format with a 3-byte start code, a receiving device that receives transmission packets in the NAL size format with a 4-byte size field may not be able to satisfy the receiving buffer model in ARIB STD-B60. Furthermore, the receiving buffer model in ARIB STD-B60 does not specify a specific buffer size or extraction rate, making it difficult to guarantee buffer operation in the receiving device.

[0648] Therefore, in order to solve the above problem, a receiving buffer model for guaranteeing buffer operation in the receiver is defined as follows.

[0649] FIG. 68 shows a reception buffer model based on the reception buffer model defined in ARIB STD-B60, particularly when only a broadcast transmission channel is used.

[0650] The receiving buffer model includes a TLV packet buffer (first buffer), an IP packet buffer (second buffer), an MMTP buffer (third buffer), and a pre-decoding buffer (fourth buffer). Note that a de-jitter buffer and a buffer for FEC are not required for broadcast transmission paths, so they are omitted.

[0651] The TLV packet buffer receives TLV packets (transmission packets) from the broadcast transmission path, converts the IP packets, which consist of the variable-length packet headers (IP packet headers, full headers when the IP packets are compressed, and compressed headers when the IP packets are compressed) stored in the received TLV packets and the variable-length payloads, into IP packets (first packets) with header-expanded fixed-length IP packet headers, and outputs the IP packets obtained by the conversion at a constant bit rate.

[0652] The IP packet buffer converts IP packets into MMTP packets (second packets) with a packet header and a variable-length payload, and outputs the MMTP packets obtained by the conversion at a constant bit rate. Note that the IP packet buffer may be merged with the MMTP buffer.

[0653] The MMTP buffer converts the output MMTP packets into NAL units, and outputs the NAL units obtained by the conversion at a constant bit rate.

[0654] The pre-decoding buffer sequentially stores the output NAL units, generates access units from the stored NAL units, and outputs the generated access units to the decoder at the decoding time corresponding to the access units.

[0655] The receive buffer model shown in Figure 68 is characterized in that the MMTP buffer and pre-decoding buffer, which are buffers other than the upstream TLV packet buffer and IP packet buffer, follow the receive buffer model in MPEG-2 TS.

[0656] For example, the MMTP buffer for video (MMTP B1) is composed of buffers equivalent to the transport buffer (TB) and multiplexing buffer (MB) in MPEG-2 TS, and the MMTP buffer for audio (MMTP Bn) is composed of buffers equivalent to the transport buffer (TB) in MPEG-2 TS.

[0657] The buffer size of the transport buffer is the same as that of MPEG-2 TS and is a fixed value, for example, n times the MTU size (n can be a decimal or an integer, and is 1 or greater).

[0658] The MMTP packet size is also specified so that the overhead rate of the MMTP packet header is smaller than the overhead rate of the PES packet header, which allows the transport buffer extraction rates RX1, RXn, and RXs in MPEG-2 TS to be applied as is to the transport buffer extraction rates.

[0659] The size of the multiplexing buffer and the extraction rate are the same as those of MPEG-2 The MB size in TS and RBX1.

[0660] In addition to the above receive buffer model, the following constraints are set to solve the problem.

[0661] The HEVC HRD specification assumes a byte stream format, and MMT uses a NAL size format that adds a 4-byte size field to the beginning of the NAL unit. Therefore, during encoding, rate control is performed in the NAL size format to satisfy the HRD.

[0662] That is, the transmitting device controls the rate of transmission packets based on the above-mentioned receiving buffer model and constraints.

[0663] In the receiving device, by performing receiving processing using the above signals, decoding operations can be performed without underflow or overflow.

[0664] Even if the size field at the beginning of the NAL unit is not 4 bytes, rate control is performed to satisfy the HRD, taking into account the size field at the beginning of the NAL unit.

[0665] The extraction rate of the TLV packet buffer (the bit rate at which the TLV packet buffer outputs IP packets) is set taking into consideration the transmission rate after the IP header is expanded.

[0666] That is, the transmission rate of the output IP packet is taken into account after inputting a TLV packet with a variable data size, removing the TLV header, and expanding (restoring) the IP header. In other words, the amount of header increase or decrease is taken into account relative to the input transmission rate.

[0667] Specifically, the transmission rate of output IP packets is not unique because the data size is variable, packets with compressed IP headers are mixed with packets without compressed IP headers, and the IP header size differs depending on the packet type, such as IPv4 or IPv6. For this reason, the average packet length of variable-length data sizes is determined, and the transmission rate of IP packets output from TLV packets is determined.

[0668] Here, in order to define the maximum transmission speed after the IP header is expanded, the transmission rate is determined assuming that the IP header is always compressed.

[0669] In addition, when IPv4 and IPv6 packet types are mixed, or when specifying without distinguishing between packet types, the transmission rate is determined assuming IPv6 packets, which have a large header size and a large growth rate after header expansion.

[0670] For example, if the average packet length of TLV packets input to the TLV packet buffer is S, and all IP packets stored in the TLV packets are IPv6 packets and are header-compressed, the maximum output transmission rate after removing the TLV header and expanding the IP header is: Input rate × {S / (S + IPv6 header compression amount)} This becomes:

[0671] More specifically, the average packet length S of a TLV packet is defined as S = 0.75 x 1500 (1500 is the maximum MTU size assumed) is set as the standard, Amount of IPv6 header compression = TLV header length - IPv6 header length - UDP header length =3-40-8 In this case, the maximum output transmission rate after removing the TLV header and expanding the IP header is Input rate × 1.0417 ≒ Input rate × 1.05 This becomes:

[0672] FIG. 69 is a diagram showing an example in which multiple data units are aggregated and stored in one payload.

[0673] In the MMT method, when data units are aggregated, a data unit length and a data unit header are added before the data units, as shown in FIG.

[0674] However, for example, when a video signal in the NAL size format is stored as one data unit, as shown in Figure 70, there are two fields indicating the size for one data unit, and the information is redundant. Figure 70 shows an example of aggregating and storing multiple data units into one payload, where a video signal in the NAL size format is treated as one data unit. Specifically, the first size field in the NAL size format (hereinafter referred to as the "size field") and the data unit length field located before the data unit header in the MMTP payload header are both fields indicating the size, and the information is redundant. For example, if the length of the NAL unit is L bytes, the size field indicates L bytes, and the data unit length field indicates L bytes + "length of the size field" (bytes). Although the values ​​indicated in the size field and the data unit length field do not exactly match, they can be said to be redundant because one value can be easily calculated from the other.

[0675] In this way, when data containing data size information is stored as a data unit and multiple such data units are aggregated and stored in a single payload, there is a problem that the size information is duplicated, resulting in large overhead and poor transmission efficiency.

[0676] Therefore, in a transmitting device, when data containing data size information is stored as a data unit and multiple such data units are aggregated and stored in a single payload, it is possible to store them as shown in Figures 71 and 72.

[0677] As shown in Figure 71, it is conceivable to store a NAL unit including a size field as a data unit, and not indicate the data unit length that is conventionally included in the MMTP payload header. Figure 71 shows the structure of the payload of an MMTP packet in which the data unit length is not indicated.

[0678] Also, as shown in Figure 72, a flag indicating whether the data unit length is indicated and information indicating the length of the size field may be newly stored in the header. The location where the flag and information indicating the length of the size field are stored may be indicated on a data unit basis, such as in a data unit header, or may be indicated on a unit where multiple data units are aggregated (packet basis). Figure 72 shows an example of the extend field assigned on a packet basis. Note that the storage location of the above newly indicated information is not limited to this, and may also be the MMTP payload header, MMTP packet header, or control information.

[0679] On the receiving side, if the flag indicating whether the data unit length is compressed indicates that the data unit length is compressed, the length information of the size area inside the data unit is obtained, and the size area is obtained based on the length information of the size area, and the data unit length can be calculated using the length information of the obtained size area and the size area.

[0680] By using the above method, the amount of data can be reduced on the sending side, and transmission efficiency can be improved.

[0681] Note that overhead may be reduced by reducing the size field instead of reducing the data unit length. When reducing the size field, information indicating whether the size field has been reduced or the length of the data unit length field may be stored.

[0682] The MMTP payload header also contains length information.

[0683] When a NAL unit containing a size field is stored as a data unit, the payload size field in the MMTP payload header may be reduced regardless of whether aggregation is performed or not.

[0684] Also, even when data that does not include a size field is stored as a data unit, if it is aggregated and the data unit length is indicated, the payload size field in the MMTP payload header may be reduced.

[0685] When reducing the payload size area, a flag indicating whether reduction has been performed, length information of the reduced size field, or length information of the non-reduced size field may be indicated, as described above.

[0686] FIG. 73 shows the operation flow of the receiving device.

[0687] As described above, the transmitting device stores NAL units containing size fields as data units, and the data unit length contained in the MMTP payload header is not indicated in the MMTP packet.

[0688] In the following, we will explain an example in which whether the data unit length is indicated is indicated by a flag or the length information in the size field in the MMTP packet.

[0689] The receiving device determines whether the data unit includes a size field and whether the data unit length has been reduced based on information transmitted from the transmitting side (S1901).

[0690] If it is determined that the data unit length has been reduced (Yes in S1902), the length information of the size field inside the data unit is obtained, and then the size field inside the data unit is analyzed and the data unit length is calculated (S1903).

[0691] On the other hand, if it is determined that the data unit length has not been reduced (No in S1902), the data unit length is calculated as usual from either the data unit length or the size field inside the data unit (S1904).

[0692] Note that the flag indicating whether the data unit length has been reduced or the length information in the size field need not be transmitted if the receiving device knows this in advance. In this case, the receiving device performs the processing shown in Figure 73 based on predetermined information.

[0693] [Supplementary information: Transmitting and receiving devices] As described above, a transmitting device that performs rate control so as to satisfy the requirements of the receiving buffer model during encoding can also be configured as shown in Fig. 74. Also, a receiving device that receives and decodes transmission packets transmitted from the transmitting device can also be configured as shown in Fig. 75. Fig. 74 is a diagram showing an example of a specific configuration of a transmitting device. Fig. 75 is a diagram showing an example of a specific configuration of a receiving device.

[0694] The transmission device 700 includes a generation unit 701 and a transmission unit 702. Each of the generation unit 701 and the transmission unit 702 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0695] The receiving device 800 includes a receiving unit 801, a first buffer 802, a second buffer 803, a third buffer 804, a fourth buffer 805, and a decoding unit 806. Each of the receiving unit 801, the first buffer 802, the second buffer 803, the third buffer 804, the fourth buffer 805, and the decoding unit (decoder) 806 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0696] Detailed explanations of each component of the transmitting device 700 and the receiving device 800 will be given in the explanations of the transmitting method and the receiving method, respectively.

[0697] First, the transmission method will be described with reference to Fig. 76. Fig. 76 shows the operation flow (transmission method) performed by the transmission device.

[0698] First, the generating unit 701 of the transmitting device 700 generates a coded stream by performing rate control so as to satisfy the regulations of a predetermined receiving buffer model in order to guarantee the buffer operation of the receiving device (S2001).

[0699] Next, the transmitting unit 702 of the transmitting device 700 packetizes the generated coded stream and transmits the transmission packets obtained by the packetization (S2002).

[0700] The receiving buffer model used in the transmitting device 700 has the same configuration as the receiving device 800, including the first to fourth buffers 802 to 805, and therefore a description thereof will be omitted.

[0701] This allows the transmitting device 700 to guarantee the buffer operation of the receiving device 800 when transmitting data using a method such as MMT.

[0702] Next, the receiving method will be described with reference to Fig. 77. Fig. 77 shows the operation flow (receiving method) of the receiving device.

[0703] First, the receiving unit 801 of the receiving device 800 receives a transmission packet made up of a fixed-length packet header and a variable-length payload (S2101).

[0704] Next, the first buffer 802 of the receiving device 800 converts the packet consisting of a variable-length packet header and a variable-length payload stored in the received transmission packet into a first packet having a header-expanded fixed-length packet header, and outputs the first packet obtained by the conversion at a constant bit rate (S2102).

[0705] Next, the second buffer 803 of the receiving device 800 converts the first packet obtained by the conversion into a second packet consisting of a packet header and a variable-length payload, and outputs the second packet obtained by the conversion at a constant bit rate (S2103).

[0706] Next, the third buffer 804 of the receiving device 800 converts the output second packets into NAL units, and outputs the NAL units obtained by the conversion at a constant bit rate (S2104).

[0707] Next, the fourth buffer 805 of the receiving device 800 sequentially accumulates the output NAL units, generates an access unit from the accumulated multiple NAL units, and outputs the generated access unit to the decoder at the decoding time corresponding to the access unit (S2105).

[0708] Then, the decoding unit 806 of the receiving device 800 decodes the access unit output by the fourth buffer (S2106).

[0709] This allows the receiving device 800 to perform a decoding operation that does not cause an underflow or overflow.

[0710] (Sixth embodiment) [overview] In the sixth embodiment, a transmission method and a reception method will be described when leap second adjustment is performed on the time information that is the reference of the reference clock in the MMT / TLV transmission method.

[0711] FIG. 78 is a diagram showing a protocol stack of the MMT / TLV method defined in ARIB STD-B60.

[0712] In the MMT method, packets contain data such as video and audio transmitted through multiple MPUs (Media Processing Units). It stores each primary data unit, such as a Presentation Unit (MFU) or Media Fragment Unit (MFU), and generates a specified MMTP packet by adding an MMTP packet header. Also, it generates a specified MMTP packet by adding an MMTP packet header to control information, such as MMTP control messages. The MMTP packet header has a field for storing a 32-bit short-format NTP (Network Time Protocol, as defined in IETF RFC 5905), which can be used for QoS control of communication lines.

[0713] In addition, the reference clock of the transmitting side (transmitting device) is synchronized with the 64-bit long format NTP defined in RFC 5905, and timestamps such as PTS (Presentation Time Stamp) and DTS (Decode Time Stamp) are added to the synchronous media based on the synchronized reference clock. Furthermore, the reference clock information of the transmitting side is transmitted to the receiving side, and the receiving device generates a system clock in the receiving device based on the reference clock information received from the transmitting side.

[0714] Specifically, the PTS and DTS are stored in the MPU timestamp descriptor and the MPU extended timestamp descriptor, which are MMTP control information, and stored in the MP table for each asset. They are then packetized as MMTP control messages and transmitted.

[0715] The MMTP packetized data is given a UDP header and an IP header and encapsulated in an IP packet. At this time, a set of packets with the same source IP address, destination IP address, source port number, destination port number, and protocol type in the IP header or UDP header is called an IP data flow. Note that IP packets of the same IP data flow have redundant headers, so some IP packets are header compressed.

[0716] In addition, a 64-bit NTP timestamp is stored in the NTP packet as reference clock information, and is then stored in the IP packet. At this time, the source IP address, destination IP address, source port number, destination port number, and protocol type of the IP packet that stores the NTP packet are fixed values, and the IP packet header is not compressed.

[0717] FIG. 79 is a diagram showing the structure of a TLV packet.

[0718] As shown in Figure 79, a TLV packet can contain transmission control information such as an IP packet, a compressed IP packet, an AMT (Address Map Table), or an NIT (Network Information Table) as data, and these data are identified using an 8-bit data type. In addition, a TLV packet indicates the data length (in bytes) using a 16-bit field, followed by the data value. In addition, a TLV packet has 1 byte of header information before the data type, and this header information is stored in a total of 4 bytes of header area. In addition, a TLV packet is mapped to a transmission slot in the advanced BS transmission system, and the mapping information is stored in the TMCC (Transmission and Multiplexing Configuration Control) control information.

[0719] FIG. 80 is a diagram illustrating an example of a block diagram of a receiving device.

[0720] In the receiving device, first, the demodulation means decodes the channel-encoded data from the broadcast signal received by the tuner, performs error correction, etc., and extracts TLV packets. Then, the TLV / IP DEMUX means performs TLV DEMUX processing and IP DEMUX processing. The TLV DEMUX processing performs processing according to the data type of the TLV packet. For example, if the TLV packet contains a compressed IP packet, the compressed header of the compressed IP packet is restored. The IP DEMUX performs processing such as header analysis of IP packets and UDP packets, and extracts MMTP packets and NTP packets.

[0721] The NTP clock generation means regenerates the NTP clock from the extracted NTP packets. The MMTP DEMUX performs filtering of components such as video and audio and control information based on the packet ID stored in the extracted MMTP packet header. The control information acquisition means acquires the timestamp descriptor stored in the MP table, and the PTS / DTS calculation means calculates the PTS and DTS for each access unit. The timestamp descriptor includes both the MPU timestamp descriptor and the MPU extended timestamp descriptor.

[0722] The access unit playback means converts the video and audio filtered from the MMTP packets into data units to be presented. Specifically, the data units to be presented include NAL units and access units of the video signal, audio frames, and subtitle presentation units. The decoding and presentation means decodes and presents the access unit at the time when the PTS / DTS of the access unit match, based on the reference time information of the NTP clock.

[0723] However, the configuration of the receiving device is not limited to this.

[0724] Next, the timestamp descriptor will be described.

[0725] FIG. 81 is a diagram illustrating a timestamp descriptor.

[0726] The PTS and DTS are stored in the MMT control information, which is the MPU timestamp descriptor as the first control information and the MPU extended timestamp descriptor as the second control information, and are stored in the MP table for each asset, packetized as an MMTP control message, and then transmitted.

[0727] 81(a) is a diagram showing the structure of an MPU timestamp descriptor specified in ARIB STD-B60. The MPU timestamp descriptor stores, for each of a plurality of MPUs, presentation time information (first time information) indicating the PTS (absolute value represented by 64-bit NTP) of the leading (first) AU (hereinafter referred to as the "leading AU") in presentation order among a plurality of access units (AUs) serving as second data units stored in the MPU. In other words, the presentation time information of the MPU assigned to the MPU is stored in the control information of the MMTP packet and transmitted.

[0728] (b) of Figure 81 shows the configuration of an MPU extended timestamp descriptor. The MPU extended timestamp descriptor stores information for calculating the PTS and DTS of the AUs included in each of multiple MPUs. The MPU extended timestamp descriptor includes relative information (second time information) from the PTS of the first AU of the MPU stored in the MPU timestamp descriptor, and the PTS and DTS of each of the multiple AUs included in the MPU can be calculated based on both the MPU timestamp descriptor and the MPU extended timestamp descriptor. In other words, the PTS and DTS of AUs other than the first AU included in the MPU can be calculated based on the PTS of the first AU stored in the MPU timestamp descriptor and the relative information stored in the MPU extended timestamp descriptor.

[0729] In other words, the second time information is relative time information for calculating the PTS or DTS of each of the multiple AUs in conjunction with the first time information. That is, the second time information is information that indicates the PTS or DTS of each of the multiple AUs in conjunction with the first time information.

[0730] NTP is a standard time information system based on Coordinated Universal Time (UTC). UTC adjusts for leap seconds (hereafter referred to as "leap second adjustment") to adjust for the difference with astronomical time, which is based on the Earth's rotational speed. Leap second adjustments are specifically carried out at 9:00 a.m. Japan time, and involve the insertion or deletion of one second.

[0731] FIG. 82 is a diagram for explaining leap second adjustment.

[0732] Figure 82(a) shows an example of the insertion of a leap second in Japan time. As shown in Figure 82(a), when a leap second is inserted, after 8:59:59 Japan time, it becomes 8:59:59 when it would normally be 9:00:00, and the 8:59:59 range is repeated twice.

[0733] Figure 82(b) shows an example of leap second deletion in Japan time. As shown in Figure 82(b), leap second deletion causes the time after 8:59:58 Japan time to become 9:00:00 when it would normally be 8:59:59, and one second around 8:59:59 is deleted.

[0734] In addition to a 64-bit timestamp, an NTP packet also contains a 2-bit leap_indicator. The leap_indicator is a flag that notifies in advance that a leap second adjustment will be performed; leap_indicator=1 indicates a leap second insertion, and leap_indicator=2 indicates a leap second deletion. Advance notification can be given from the beginning of the month in which the leap second adjustment will be performed, 24 hours in advance, or at any other time. The leap_indicator will become 0 when the leap second adjustment is completed (9:00:00). For example, if advance notification is given 24 hours in advance, the leap_indicator will be set to "1" or "2" between 9:00 Japan time on the day before the leap second adjustment and the time immediately before the leap second adjustment on the day of the leap second adjustment (that is, the first time around 8:59:59 in the case of a leap second insertion, or the time around 8:59:58 in the case of a leap second deletion).

[0735] Next, we will explain the issues involved in adjusting leap seconds.

[0736] Figure 83 is a diagram showing the relationship between NTP time, MPU timestamp, and MPU presentation timing. Note that NTP time is the time indicated by NTP. Furthermore, MPU timestamp is a timestamp indicating the PTS of the first AU in the MPU. Furthermore, MPU presentation timing is the timing at which a receiving device should present an MPU in accordance with the MPU timestamp. Specifically, (a) to (c) of Figure 83 are diagrams showing the relationship between NTP time, MPU timestamp, and MPU presentation time when no leap second adjustment occurs, when a leap second is inserted, and when a leap second is deleted, respectively.

[0737] Here, we will explain an example in which the NTP time (reference clock) on the sending side is synchronized with the NTP server, and the NTP time (system clock) on the receiving side is synchronized with the NTP time on the sending side. In this case, the receiving device reproduces the time based on the timestamp stored in the NTP packet transmitted from the sending side. In this case, since both the NTP time on the sending side and the NTP time on the receiving side are synchronized with the NTP server, an adjustment of ±1 second is made when adjusting for leap seconds. Also, the NTP time in Figure 83 is assumed to be the same for the NTP time on the sending side and the NTP time on the receiving side. Note that this explanation assumes that there is no transmission delay.

[0738] The MPU timestamp in Figure 83 indicates the timestamp of the first AU in presentation order among multiple AUs included in each of multiple MPUs, and is generated (set) based on the NTP time indicated by the base of the arrow. Specifically, the presentation time information of the MPU is generated by adding a predetermined time (for example, 1.5 seconds in Figure 83) to the NTP time, which serves as reference time information at the timing of generating the presentation time information of the MPU. The generated MPU timestamp is stored in the MPU timestamp descriptor.

[0739] The receiving device presents the MPU at the MPU presentation time based on the timestamp stored in the MPU timestamp descriptor.

[0740] In Figure 83, the playback time of one MPU is described as being 1 second, but the playback time of one MPU may be other playback times, for example, 0.5 seconds or 0.1 seconds.

[0741] In the example of (a) of Figure 83, the receiving device can present MPUs #1-#5 in order based on the timestamps stored in the MPU timestamp descriptor.

[0742] However, in Figure 83 (b), the presentation times of MPU #2 and MPU #3 overlap due to the insertion of leap seconds. Therefore, if the receiving device presents an MPU based on the timestamp stored in the MPU timestamp descriptor, there will be two MPUs to present in the same time slot around 9:00:00, and the receiving device will not be able to determine which of the two MPUs to present. Furthermore, the MPU presentation time (8:59:59) indicated in the timestamp of MPU #1 will exist twice, so the receiving device will not be able to determine which of the two MPU presentation times to present.

[0743] Also, in (c) of Figure 83, the receiving device cannot present MPU#3 because the MPU presentation time (8:59:59) indicated in the MPU timestamp of MPU#3 does not exist in NTP time due to the deletion of leap seconds.

[0744] While a receiving device can solve the above problems by performing MPU decoding and presentation processing that is not based on timestamps, a receiving device that performs processing based on timestamps has difficulty performing different processing (processing that is not based on timestamps) only when a leap second occurs.

[0745] Next, we will explain how to solve the problem of leap second adjustment by correcting the timestamp on the transmitting side.

[0746] Figure 84 is a diagram for explaining a method for correcting a timestamp on the transmitting side. Specifically, (a) of Figure 84 shows an example of inserting a leap second, and (b) of Figure 84 shows an example of deleting a leap second.

[0747] First, the case of leap second insertion will be described.

[0748] As shown in Figure 84(a), when a leap second is inserted, the time up to just before the leap second is inserted (i.e., up to the first 8:59:59 in NTP time) is designated as area A, and the time after the leap second is inserted (i.e., after the second 8:59:59 in NTP time) is designated as area B. Note that areas A and B are temporal areas, and are time periods or periods. The MPU timestamp in Figure 84(a) is the same as the timestamp described in Figure 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is added.

[0749] A specific method for correcting timestamps on the transmitting side when leap seconds are inserted will now be described.

[0750] The transmitting side (transmitting device) performs the following processing.

[0751] 1. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (a) of Figure 84) is included in area A and the MPU timestamp (the value of the MPU timestamp before correction) indicates 9:00:00 or later, the MPU timestamp is corrected by subtracting 1 second, and the corrected MPU timestamp is stored in the MPU timestamp descriptor. In other words, if the MPU timestamp was generated based on the NTP time included in area A and indicates 9:00:00 or later, the MPU timestamp is corrected by -1 second. Note that "9:00:00" here refers to the time that corresponds to Japan time as the reference time for leap second adjustment (i.e., the time calculated by adding 9 hours to UTC time). In addition, correction information indicating the correction is separately sent to the receiving device.

[0752] 2. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (a) of Figure 84) is in area B, the MPU timestamp is not corrected. In other words, if the MPU timestamp is generated based on the NTP time included in area B, the MPU timestamp is not corrected.

[0753] The receiving device presents the MPU based on the MPU timestamp and correction information indicating whether the MPU timestamp has been corrected (that is, whether information indicating that the MPU timestamp has been corrected is included).

[0754] If the receiving device determines that the MPU timestamp has not been corrected (i.e., if it determines that it does not contain information indicating that it has been corrected), it presents the MPU at the time when the timestamp stored in the MPU timestamp descriptor matches the NTP time of the receiving device (both before and after correction).In other words, if the MPU timestamp stored in the MPU timestamp descriptor of an MPU transmitted before the MPU whose MPU timestamp is to be corrected is an MPU, the receiving device presents the MPU at the time when the MPU timestamp matches the NTP time before the leap second insertion (i.e., before the first 8:59:59 timestamp).Also, if the MPU timestamp stored in the MPU timestamp descriptor of the received MPU is a timestamp transmitted after the MPU timestamp to be corrected, the receiving device presents the MPU at the time when the MPU timestamp matches the NTP time after the leap second insertion (i.e., after the second 8:59:59 timestamp).

[0755] In addition, if the MPU timestamp stored in the MPU timestamp descriptor of the received MPU has been corrected, the receiving device presents the MPU based on the NTP time after the leap second insertion (i.e., after the second 8:59:59 timeframe) in the receiving device.

[0756] Information indicating that the MPU timestamp value has been corrected is stored in a control message, descriptor, table, MPU metadata, MF metadata, MMTP packet header, etc. and transmitted.

[0757] Next, the case of deleting leap seconds will be described.

[0758] As shown in Figure 84(b), when a leap second is deleted, the time immediately before the leap second is deleted (i.e., immediately before 9:00:00 in NTP time) is designated as area C, and the time after the leap second is deleted (i.e., after 9:00:00 in NTP time) is designated as area D. Note that areas C and D are temporal areas, and represent time periods or periods. The MPU timestamp in Figure 84(b) is the same as the timestamp described in Figure 83, and is a timestamp generated (set) based on the NTP time at the time the MPU timestamp is assigned.

[0759] A specific method for correcting timestamps on the transmitting side when leap seconds are deleted will now be described.

[0760] The transmitting side (transmitting device) performs the following processing.

[0761] 1. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (b) of Figure 84) is included in Area C and the MPU timestamp (the value of the MPU timestamp before correction) indicates 8:59:59 or later, the MPU timestamp is corrected by adding 1 second, and the corrected MPU timestamp is stored in the MPU timestamp descriptor. In other words, if the MPU timestamp was generated based on the NTP time included in Area C and indicates 8:59:59 or later, the MPU timestamp is corrected by +1 second. Note that "8:59:59" here is the time obtained by subtracting -1 second from the time that corresponds the reference time for leap second adjustment to Japan time (i.e., the time calculated by adding 9 hours to UTC time). In addition, correction information indicating the correction is separately transmitted to the receiving device. Note that in this case, the correction information does not necessarily have to be transmitted.

[0762] 2. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (b) of Figure 84) is in area D, the MPU timestamp is not corrected. In other words, if the MPU timestamp is generated based on the NTP time included in area D, the MPU timestamp is not corrected.

[0763] The receiving device presents the MPU based on the MPU timestamp. If correction information indicating whether the MPU timestamp has been corrected is available, the receiving device may present the MPU based on the MPU timestamp and the correction information.

[0764] By performing the above processing, even if a leap second adjustment is made to the NTP time, the receiving device can present a normal MPU using the MPU timestamp stored in the MPU timestamp descriptor.

[0765] Note that whether the MPU timestamp was assigned in area A, area B, area C, or area D may be signaled and notified to the receiving side. That is, whether the MPU timestamp was generated based on the NTP time included in area A, area B, area C, or area D may be signaled and notified to the receiving side. In other words, identification information indicating whether the MPU timestamp (presented time) was generated based on the reference time information (NTP time) before the leap second adjustment may be transmitted. This identification information is assigned based on the leap_indicator included in the NTP packet, and therefore indicates whether the MPU timestamp was set based on the time between 9:00 Japan time on the day before the leap second adjustment and the time immediately before the leap second adjustment on the day of the leap second adjustment (that is, the first time around 8:59:59 in the case of leap second insertion, or the time around 8:59:58 in the case of leap second deletion). In other words, the identification information indicates whether the MPU timestamp was generated based on the NTP time from a time a predetermined period (e.g., 24 hours) before the time immediately before the leap second adjustment to the time immediately before that time.

[0766] Next, a method for solving the problem of leap second adjustment by correcting the timestamp in the receiving device will be described.

[0767] Fig. 85 is a diagram for explaining a method for correcting a timestamp in a receiving device. Specifically, Fig. 85(a) shows an example of inserting a leap second, and Fig. 85(b) shows an example of deleting a leap second.

[0768] First, the case of leap second insertion will be described.

[0769] As shown in Figure 85(a), when a leap second is inserted, as in Figure 84(a), the time up to just before the leap second insertion (i.e., up to the first 8:59:59 in NTP time) is designated as area A, and the time after the leap second insertion (i.e., from the second 8:59:59 in NTP time) is designated as area B. Note that areas A and B are temporal areas, and are time periods or periods. The MPU timestamp in Figure 85(a) is the same as the timestamp described in Figure 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is added.

[0770] A specific method for correcting timestamps in a receiving device when leap seconds are inserted will now be described.

[0771] The transmitting side (transmitting device) performs the following processing.

[0772] The generated MPU timestamp is not corrected, but is stored in an MPU timestamp descriptor and transmitted to the receiving device.

[0773] - Information indicating whether the timing of adding the MPU timestamp is in area A or area B is transmitted as identification information to the receiving device. In other words, identification information indicating whether the MPU timestamp was generated based on the NTP time included in area A or the NTP time included in area B is transmitted to the receiving device.

[0774] The receiving device performs the following processing.

[0775] The receiving device corrects the MPU timestamp based on the MPU timestamp and the identification information indicating whether the timing at which the MPU timestamp was added is in area A or area B.

[0776] Specifically, the following processing is performed.

[0777] 1. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (a) of Figure 85) is included in area A and the MPU timestamp (the value of the MPU timestamp before correction) indicates 9:00:00 or later, the MPU timestamp is corrected by subtracting 1 second. In other words, if the MPU timestamp was generated based on the NTP time included in area A and indicates 9:00:00 or later, the MPU timestamp is corrected by -1 second. Note that "9:00:00" here refers to the time that corresponds to Japan time as the reference time for leap second adjustments (in other words, the time calculated by adding 9 hours to UTC time).

[0778] 2. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (a) of Figure 85) is in area B, the MPU timestamp is not corrected. In other words, if the MPU timestamp is generated based on the NTP time included in area B, the MPU timestamp is not corrected.

[0779] If the MPU timestamp is not corrected, the MPU is presented at the time when the MPU timestamp stored in the MPU timestamp descriptor matches the NTP time of the receiving device (including before and after correction).

[0780] In other words, for MPU timestamps received before the MPU timestamp to be corrected, the MPU will present the time that matches the NTP time before the leap second insertion (i.e., before the first 8:59:59 timestamp). Also, for MPU timestamps received after the MPU timestamp to be corrected, the MPU will present the time that matches the NTP time after the leap second insertion (i.e., after the second 8:59:59 timestamp).

[0781] When correcting the MPU timestamp, the corrected MPU timestamp is presented to the MPU based on the NTP time after the leap second insertion (that is, after the second 8:59:59 time).

[0782] The transmitting side stores identification information indicating whether the timing at which the MPU timestamp was added is in area A or area B in a control message, descriptor, table, MPU metadata, MF metadata, MMTP packet header, etc., and transmits it.

[0783] Next, the case of deleting leap seconds will be described.

[0784] As shown in (b) of Figure 85, when a leap second is deleted, as in (a) of Figure 84, the time immediately before the leap second is deleted (i.e., until immediately before 9:00:00 in NTP time) is designated as area C, and the time after the leap second is deleted (i.e., after 9:00:00 in NTP time) is designated as area D. Note that areas C and D are temporal areas, and represent time periods or periods. The MPU timestamp in (b) of Figure 85 is the same as the timestamp described in Figure 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is assigned.

[0785] A specific method for correcting timestamps in a receiving device when leap seconds are deleted will now be described.

[0786] The transmitting side (transmitting device) performs the following processing.

[0787] The generated MPU timestamp is not corrected, but is stored in an MPU timestamp descriptor and transmitted to the receiving device.

[0788] - Information indicating whether the timing of adding the MPU timestamp is in area C or area D is transmitted as identification information to the receiving device. In other words, identification information indicating whether the MPU timestamp was generated based on the NTP time included in area C or the NTP time included in area B is transmitted to the receiving device.

[0789] The receiving device performs the following processing.

[0790] The receiving device corrects the MPU timestamp based on the MPU timestamp and the identification information indicating whether the timing at which the MPU timestamp was added is in the C region or the D region.

[0791] Specifically, the following processing is performed.

[0792] 1. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in Figure 85(b)) is included in area C and the MPU timestamp (the value of the MPU timestamp before correction) indicates 8:59:59 or later, then the MPU timestamp value is corrected by +1 second, adding 1 second. In other words, if the MPU timestamp was generated based on the NTP time included in area C and indicates 8:59:59 or later, then the MPU timestamp is corrected by +1 second. Note that "8:59:59" here is the time obtained by subtracting -1 second from the time that corresponds the reference time for leap second adjustment to Japan time (in other words, the time calculated by adding 9 hours to UTC time).

[0793] 2. If the timing at which the MPU timestamp is assigned (the timing indicated by the origin of the arrow in (b) of Figure 85) is in area D, the MPU timestamp is not corrected. In other words, if the MPU timestamp is generated based on the NTP time included in area D, the MPU timestamp is not corrected.

[0794] The receiving device presents the MPU based on the MPU timestamp and the corrected MPU timestamp.

[0795] By performing the above processing, even if a leap second adjustment is made to the NTP time, the receiving device can present a normal MPU using the MPU timestamp stored in the MPU timestamp descriptor.

[0796] Even in this case, similar to the case where the MPU timestamp is corrected on the transmitting side as explained in Fig. 84, the receiving side may be notified by signaling whether the timing of adding the MPU timestamp is in area A, area B, area C, or area D. Details of the notification are the same, so explanation will be omitted.

[0797] 84 and 85, additional information (identification information) such as information indicating whether the MPU timestamp has been corrected or information indicating whether the timing at which the MPU timestamp was added is in area A, area B, area C, or area D may be enabled when the leap_indicator in the NTP packet indicates the deletion (leap_indicator=2) or insertion (leap_indicator=1) of a leap second. This information may be enabled at any predetermined time (for example, 3 seconds before the leap second adjustment), or may be enabled dynamically.

[0798] In addition, the timing at which the validity period of the additional information ends and the timing at which the signaling of the additional information ends may be set to coincide with the leap_indicator in the NTP packet, or may become valid from any predetermined time (for example, 3 seconds before the leap second adjustment), or may be dynamically enabled.

[0799] Note that the 32-bit timestamp stored in the MMTP packet and the timestamp information stored in the TMCC are also generated and assigned based on the NTP time, which causes the same problem. Therefore, in the case of the 32-bit timestamp and the timestamp information stored in the TMCC, the timestamp can be corrected using the same method as in Figures 84 and 85, and processing based on the timestamp can be performed in the receiving device. For example, when storing the above additional information for the 32-bit timestamp stored in the MMTP packet, it may be indicated using the extension field of the MMTP packet header. In this case, the additional information is indicated in the extension type of the multi-header type. Furthermore, the time when the leap_indicator is set may indicate the additional information using some bits of the 32-bit or 64-bit timestamp.

[0800] In the example of FIG. 84, the corrected MPU timestamp is stored in the MPU timestamp descriptor. However, both the pre-correction and post-correction MPU timestamps (i.e., the uncorrected MPU timestamp and the corrected MPU timestamp) may be transmitted to the receiving device. For example, the pre-correction MPU timestamp descriptor and the post-correction MPU timestamp descriptor may be stored in the same MPU timestamp descriptor, or may be stored in two MPU timestamp descriptors. In this case, whether the MPU timestamp is pre-correction or post-correction may be identified based on the arrangement order of the two MPU timestamp descriptors or the order of descriptions within the MPU timestamp descriptor. Alternatively, the pre-correction MPU timestamp may always be stored in the MPU timestamp descriptor, and when a correction is performed, the post-correction timestamp may be stored in the MPU extended timestamp descriptor.

[0801] In this embodiment, the NTP time has been described as being Japan time (9:00 AM) as an example, but it is not limited to Japan time. Leap second adjustment is based on UTC time and is corrected simultaneously around the world. Japan time is 9 hours ahead of UTC time and is expressed as a value of (+9) relative to UTC time. In this way, it is also possible to use a time adjusted to a different time depending on the time difference depending on the location.

[0802] As described above, correcting the timestamp on the transmitting side or the receiving device based on the information indicating the timing at which the timestamp was added enables normal reception processing using the timestamp.

[0803] Note that a receiving device that continues decoding processing from well before the leap second adjustment time may be able to perform decoding processing and presentation processing without using a timestamp, but a receiving device that tunes in just before the leap second adjustment time may not be able to determine a timestamp and may not be able to present it until after the leap second adjustment is complete. Even in such cases, the correction method of this embodiment makes it possible to perform reception processing using a timestamp, making it possible to tune in even just before the leap second adjustment time.

[0804] FIG. 86 shows the operational flow of the transmitting side (transmitting device) when correcting the MPU timestamp on the transmitting side (transmitting device) as explained in FIG. 84, and FIG. 87 shows the operational flow of the receiving device.

[0805] First, the operational flow on the transmitting side (transmitting device) will be described with reference to FIG.

[0806] When a leap second is inserted or deleted, it is determined whether the timing of adding the MPU timestamp is in area A, area B, area C, or area D (S2201). Note that cases when a leap second is not inserted or deleted are not shown.

[0807] Here, the areas A to D are defined as follows.

[0808] Area A: Time up to just before the leap second insertion (up to the first 8:59:59 mark) Area B: Time after the leap second insertion (after the second 8:59:59) Area C: The time immediately before the leap second deletion (up to 9:00:00) D area: Time after leap second deletion (after 9:00:00)

[0809] In step S2201, if it is determined that the area is A and the MPU timestamp indicates 9:00:00 or later, the MPU timestamp is corrected by -1 second and the corrected MPU timestamp is stored in the MPU timestamp descriptor (S2202).

[0810] Then, correction information indicating that the MPU timestamp has been corrected is signaled and transmitted to the receiving device (S2203).

[0811] Also, if it is determined in step S2201 that the area is C and the MPU timestamp indicates 8:59:59 or later, the MPU timestamp is corrected by +1 second and the corrected MPU timestamp is stored in the MPU timestamp descriptor (S2205).

[0812] If it is determined in step S2201 that the data is in area B or area D, the MPU timestamp is stored in the MPU timestamp descriptor without being corrected (S2204).

[0813] Next, the operation flow of the receiving device will be explained using FIG.

[0814] Based on the information signaled from the transmitting side, it is determined whether the MPU timestamp has been corrected (S2301).

[0815] If it is determined that the MPU timestamp has been corrected (Yes in S2301), the receiving device presents the MPU based on the MPU timestamp at the NTP time after the leap second adjustment was performed (S2302).

[0816] If it is determined that the MPU timestamp has not been corrected (No in S2301), the MPU is presented based on the MPU timestamp at the NTP time.

[0817] When correcting an MPU timestamp, the MPU corresponding to that MPU timestamp is presented in the interval where leap seconds are inserted.

[0818] Conversely, if the MPU timestamp is not corrected, the MPU corresponding to that MPU timestamp is not presented in an interval where leap seconds are inserted, but is presented in an interval where leap seconds are not inserted.

[0819] FIG. 88 shows the operational flow on the transmitting side when correcting the MPU timestamp in the receiving device, as explained in FIG. 85, and FIG. 89 shows the operational flow of the receiving device.

[0820] First, the operational flow on the transmitting side (transmitting device) will be explained using FIG.

[0821] Based on the leap_indicator in the NTP packet, it is determined whether a leap second adjustment (insertion or deletion) has been made (S2401).

[0822] If it is determined that a leap second adjustment is to be performed (Yes in S2401), the timing for adding an MPU timestamp is determined, and the identification information is signaled and transmitted to the receiving device (S2402).

[0823] On the other hand, if it is determined that leap second adjustment is not to be performed (No in S2041), the process ends with normal operation.

[0824] Next, the operation flow of the receiving device will be explained using FIG.

[0825] Based on the identification information signaled by the transmitting side (transmitting device), it is determined whether the timing of adding the MPU timestamp is in area A, area B, area C, or area D (S2501). Areas A to D are the same as those defined above, so their explanation will be omitted. As with Figure 87, cases where leap seconds are not inserted or deleted are not shown.

[0826] In step S2501, if it is determined that the area is area A and the MPU timestamp indicates 9:00:00 or later, the MPU timestamp is corrected by -1 second (S2502).

[0827] In step S2501, if it is determined that the area is C and the MPU timestamp indicates 8:59:59 or later, the MPU timestamp is corrected by +1 second (S2504).

[0828] If it is determined in step S2501 that the area is area B or area D, the MPU timestamp is not corrected (S2503).

[0829] The receiving device then presents the MPU based on the corrected MPU timestamp in a process not shown.

[0830] When correcting the MPU timestamp, the MPU corresponding to the MPU timestamp is presented based on the MPU timestamp at the NTP time after the leap second adjustment has been performed.

[0831] Conversely, if the MPU timestamp is not corrected, the MPU corresponding to that MPU timestamp is not presented in an interval where leap seconds are inserted, but is presented in an interval where leap seconds are not inserted.

[0832] In short, the transmitting side (transmitting device) determines, for each of multiple MPUs, the timing for assigning an MPU timestamp corresponding to that MPU. If the determination results in the timing being up to the time immediately prior to the leap second insertion and the MPU timestamp indicating 9:00:00 or later, the MPU timestamp is corrected by -1 second. Correction information indicating that the MPU timestamp has been corrected is signaled and transmitted to the receiving device. If the determination results in the timing being up to the time immediately prior to the leap second deletion and the MPU timestamp indicating 8:59:59 or later, the MPU timestamp is corrected by +1 second.

[0833] Furthermore, the receiving device determines whether the MPU timestamp indicated by the correction information signaled by the sending side (transmitting device) has been corrected, and if it indicates that the MPU timestamp has been corrected, it presents an MPU based on the MPU timestamp at the NTP time after the leap second adjustment.If the MPU timestamp has not been corrected, it presents an MPU based on the MPU timestamp at the NTP time before ...

Claims

1. A transmission method for storing data constituting an encoded stream in a predetermined data unit and transmitting the data, comprising: generating presentation time information indicating a presentation time of the predetermined data unit based on the reference time information; (i) transmitting the predetermined data unit, (ii) first control information storing the generated presentation time information, and (iii) second control information storing identification information indicating whether the presentation time information was generated based on the reference time information before leap second adjustment; the identification information indicates whether the presented time information was generated based on the reference time information from a time a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, In the generation, a presentation time of the predetermined data unit is generated by subtracting a predetermined time from the reference time information; The first control information is an MPU timestamp descriptor. Sending method.

2. A receiving method for receiving a predetermined data unit in which data constituting an encoded stream is stored, comprising: (i) receiving the predetermined data unit; (ii) first control information storing presentation time information indicating the presentation time of the predetermined data unit; and (iii) second control information storing identification information indicating whether the presentation time information was generated based on reference time information before leap second adjustment; Regenerating the received predetermined data unit based on the received first control information and the received second control information; the identification information indicates whether the presented time information was generated based on the reference time information from a time a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit is a time generated by subtracting a predetermined time from the reference time information, The first control information is an MPU timestamp descriptor. Receiving method.

3. A transmitting device that stores data constituting an encoded stream in a predetermined data unit and transmits the data, a generating unit that generates presentation time information indicating a presentation time of the predetermined data unit based on the reference time information; a transmitter that transmits (i) the predetermined data unit, (ii) first control information that stores the presentation time information generated by the generator, and (iii) second control information that stores identification information that indicates whether the presentation time information was generated based on the reference time information before leap second adjustment, the identification information indicates whether the presented time information was generated based on the reference time information from a time a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the generation unit generates a presentation time of the predetermined data unit by subtracting a predetermined time from the reference time information; The first control information is an MPU timestamp descriptor. Transmitting device.

4. A receiving device for receiving a predetermined data unit in which data constituting an encoded stream is stored, a receiving unit that receives (i) the predetermined data unit, (ii) first control information that stores presentation time information indicating the presentation time of the predetermined data unit, and (iii) second control information that stores identification information that indicates whether the presentation time information was generated based on reference time information before leap second adjustment; a reproduction unit that reproduces the predetermined data unit received by the receiving unit based on the first control information and the second control information received by the receiving unit, the identification information indicates whether the presented time information was generated based on the reference time information from a time a predetermined period before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment, the presentation time of the predetermined data unit is a time generated by subtracting a predetermined time from the reference time information, The first control information is an MPU timestamp descriptor. Receiving device.

Citation Information

Patent Citations

  • Terminal device

    JP2010272071A

  • Transmission method and reception method

    JP2015005977A

  • Transmission system, multiplexer, and leap second correction handling method

    JP2016171385A

  • Leap second support in content timestamps

    US20140282791A1

  • Transmission apparatus, transmission method, reception apparatus, and reception method

    WO2016129413A1