Transmission method, reception method, transmission device, and reception device

The transmission method addresses the challenge of maintaining real-time decoding and playback of ultra-high-definition video content by generating and transmitting time specification information that accounts for leap second adjustments, ensuring accurate synchronization and continuous playback.

JP7693907B2Active Publication Date: 2025-06-17PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024091487
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-02-18
Filing Date
2024-06-05
Publication Date
2025-06-17
Estimated Expiration
2035-10-28

AI Technical Summary

Technical Problem

Existing transmission methods for ultra-high-definition video content, such as 8K, face challenges in maintaining real-time decoding and playback due to high processing loads, especially when leap second adjustments occur, which can disrupt the synchronization of reference clocks between transmission and receiving devices.

Method used

A transmission method that generates and transmits time specification information, including UTC and NPT times, along with control information indicating whether the time information is before or after a leap second adjustment, allowing receiving devices to accurately synchronize and reproduce data units at the intended time.

Benefits of technology

Enables the receiving device to execute applications composed of predetermined data units at the intended time even during leap second adjustments, ensuring continuous and synchronized playback of ultra-high-definition video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693907000001
    Figure 0007693907000001
  • Figure 0007693907000002
    Figure 0007693907000002
  • Figure 0007693907000003
    Figure 0007693907000003
Patent Text Reader

Abstract

To provide a transmitting method which enables a receiving device to execute, at an intended time, an application including predetermined data units.SOLUTION: A transmitting method for storing data of an application in a predetermined data unit and transmitting the data includes: generating time specification information indicating an operation time of the application, based on reference time information received from an external source; and transmitting (i) the predetermined data unit and (ii) control information indicating the generated time specification information. The control information includes the generated time specification information and identification information indicating whether or not the time specification information is time information indicating a time that is before a leap second adjustment. The time specification information of the predetermined data unit is generated by adding a predetermined period of time to the time information. The time specification information is a UTC time and an NPT. The control information is a UTC-NPT reference descriptor indicating a relationship between the UTC time and the NPT.SELECTED DRAWING: Figure 92
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a transmission method, a reception method, a transmission apparatus, and a reception apparatus.

Background Art

[0002] With the advancement of broadcast and communication services, the introduction of ultra-high-definition moving image contents such as 8K (7680×4320 pixels: hereinafter also referred to as 8K4K) and 4K (3840×2160 pixels: hereinafter also referred to as 4K2K) has been under consideration. A receiving apparatus needs to decode and display the encoded data of the received ultra-high-definition moving image in real time. However, in particular, moving images with a resolution such as 8K have a large processing load during decoding, and it is difficult to decode such a moving image in real time with a single decoder. Therefore, a method of reducing the processing load per decoder and achieving real-time processing by parallelizing the decoding process using a plurality of decoders has been under consideration.

[0003] Also, the encoded data is multiplexed based on a multiplexing method such as MPEG-2 TS (Transport Stream) or MMT (MPEG Media Transport) and then transmitted. For example, Non-Patent Document 1 discloses a technique of transmitting encoded media data for each packet according to MMT.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, with the advancement of broadcasting and communication services, the introduction of ultra-high-definition video content such as 8K and 4K (3840x2160 pixels) is being considered. In the MMT / TLV method, the reference clock on the transmission side is synchronized with the 64-bit long format NTP defined in RFC 5905. Based on the reference clock, time stamps such as PTS (Presentation Time Stamp) and DTS (Decode Time Stamp) are added to the synchronized media. Further, the reference clock information on the transmission side is transmitted to the reception side, and in the receiving device, a system clock in the receiving device is generated based on the reference clock information.

[0006] However, in such a transmission method of MMT, when leap second adjustment is performed on the reference time information serving as the reference for the reference clocks of the transmission side and the receiving device, there is a problem that an MPU, which is a predetermined data unit included in the packet received by the receiving device, cannot be decoded or presented at the intended time according to the DTS or PTS associated with the MPU.

[0007] The present invention provides a transmission method and the like that enable a receiving device to reproduce a predetermined data unit at the intended time even when leap second adjustment is performed on the reference time information serving as the reference for the reference clocks of the transmission side and the receiving device.

Means for Solving the Problem

[0008] In order to achieve the above object, a transmission method according to an aspect of the present invention is a transmission method for storing and transmitting data constituting an application in a predetermined data unit, and based on reference time information, generating time specification information indicating an operation time of the application, and transmitting (i) the predetermined data unit and (ii) control information indicating the generated time specification information. The control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment. The time specification information of the predetermined data unit is time information obtained by adding a predetermined time to the time information. The time specification information is Universal Time Coordinated (UTC) time and Normal Play Time (NPT), and the control information is a UTC-NPT reference descriptor indicating the relationship between UTC time and NPT time.

[0009] Also, a reception method according to an aspect of the present invention is a reception method for receiving a predetermined data unit in which data constituting an application is stored, and receiving (i) the predetermined data unit, (ii) time specification information indicating an operation time of the application, and control information storing identification information indicating whether the time specification information is time information before leap second adjustment. Based on the received control information, executing the application stored in the received predetermined data unit. The time specification information of the predetermined data unit is time information obtained by adding a predetermined time to the time information. The time specification information is Universal Time Coordinated (UTC) time and Normal Play Time (NPT), and the control information is a UTC-NPT reference descriptor indicating the relationship between UTC time and NPT time.

[0010] In order to achieve the above object, a transmission method according to an aspect of the present invention is a transmission method for storing and transmitting data constituting an application in a predetermined data unit, which generates time specification information indicating an operation time of the application based on reference time information, and transmits (i) the predetermined data unit and (ii) control information indicating the generated time specification information. The control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment. The time specification information is Universal Time Coordinated (UTC) time and Normal Play Time (NPT), and the control information is a UTC-NPT reference descriptor indicating the relationship between UTC time and NPT time.

[0011] Further, a reception method according to an aspect of the present invention is a reception method for receiving a predetermined data unit in which data constituting an application is stored, which receives (i) the predetermined data unit, (ii) time specification information indicating an operation time of the application, and control information storing identification information indicating whether the time specification information is time information before leap second adjustment, and executes the application stored in the received predetermined data unit based on the received control information. The time specification information is Universal Time Coordinated (UTC) time and Normal Play Time (NPT), and the control information is a UTC-NPT reference descriptor indicating the relationship between UTC time and NPT time.

[0012] In order to achieve the above object, a transmission method according to an aspect of the present invention is a transmission method for storing and transmitting data constituting an application in a predetermined data unit, generating time specification information indicating an operation time of the application based on reference time information, and transmitting (i) the predetermined data unit and (ii) control information indicating the generated time specification information, wherein the control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment, the identification information indicates whether the time specification information is generated based on the time information from a time predetermined before the time immediately before leap second adjustment to the time immediately before the adjustment, and the reference time information is NTP (Network Time Protocol).

[0013] Also, a reception method according to an aspect of the present invention is a reception method for receiving a predetermined data unit in which data constituting an application is stored, receiving (i) the predetermined data unit, (ii) time specification information indicating an operation time of the application, and control information storing identification information indicating whether the time specification information is time information before leap second adjustment, and executing the application stored in the received predetermined data unit based on the received control information, wherein the identification information indicates whether the time specification information is generated based on the time information from a time predetermined before the time immediately before leap second adjustment to the time immediately before the adjustment, and the time specification information is generated based on NTP (Network Time Protocol).

[0014] In order to achieve the above object, a transmission method according to an aspect of the present invention is a transmission method for storing data constituting an application in a predetermined data unit and transmitting the data. Based on reference time information, time specification information indicating the operation time of the application is generated, and (i) the predetermined data unit and (ii) control information indicating the generated time specification information are transmitted. The control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment. The identification information indicates whether the time specification information is generated based on the time information from a time predetermined period before the time immediately before leap second adjustment to the time immediately before leap second adjustment. The time specification information of the predetermined data unit is time information obtained by adding a predetermined time to the time information.

[0015] Further, a reception method according to an aspect of the present invention is a reception method for receiving a predetermined data unit in which data constituting an application is stored. (i) The predetermined data unit, (ii) time specification information indicating the operation time of the application, and control information storing identification information indicating whether the time specification information is time information before leap second adjustment are received. Based on the received control information, the application stored in the received predetermined data unit is executed. The identification information indicates whether the time specification information is generated based on the time information from a time predetermined period before the time immediately before leap second adjustment to the time immediately before leap second adjustment. The time specification information of the predetermined data unit is time information generated by adding a predetermined time to the time information.

[0016] In order to achieve the above object, a transmission method according to an aspect of the present invention is a transmission method for storing and transmitting data constituting an application in a predetermined data unit, wherein time specification information indicating an operation time of the application is generated based on reference time information received from the outside, and (i) the predetermined data unit and (ii) control information indicating the generated time specification information are transmitted, and the control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment.

[0017] Note that these general or specific aspects may be realized by a system, a device, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a device, an integrated circuit, a computer program, and a recording medium.

Effects of the Invention

[0018] Even when leap second adjustment is performed on the reference time information that is the reference of the reference clocks of the transmission side and the receiving device, the present invention enables the receiving device to execute an application composed of predetermined data units at the intended time.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66

Figure 67

Figure 68

Figure 69

Figure 70

Figure 71

Figure 72

Figure 73

Figure 74

Figure 75

Figure 76

Figure 77

Figure 78

Figure 79

Figure 80

Figure 81

Figure 82

Figure 83

Figure 84

Figure 85

Figure 86

Figure 87

Figure 88

Figure 89

Figure 90

Figure 91

Figure 92

Figure 93

Figure 94

Figure 95

Figure 96

Figure 97

DETAILED DESCRIPTION OF THE INVENTION

[0020] A transmission method according to an aspect of the present invention is a transmission method for storing and transmitting data constituting an application in a predetermined data unit, and generating time specification information indicating an operation time of the application based on reference time information received from the outside, and transmitting (i) the predetermined data unit and (ii) control information indicating the generated time specification information, wherein the control information stores the generated time specification information and identification information indicating whether the time specification information is time information before leap second adjustment.

[0021] Even when leap second adjustment is performed on the reference time information that is the reference of the reference clocks of the transmission side and the receiving device, such a transmission method can execute an application composed of predetermined data units at an intended time.

[0022] In the generation, the presentation time information of the MPU may be generated by adding a predetermined time to the time information.

[0023] The identification information may indicate whether the presentation time information is generated based on the time information from a time predetermined period before the time immediately before leap second adjustment to the time immediately before the leap second adjustment.

[0024] The time specification information may be UTC (Universal Time Coordinated) time and NPT (Normal Play Time), and the control information may be a UTC-NPT reference descriptor indicating the relationship between UTC time and NPT time.

[0025] The reference time information may be NTP (Network Time Protocol).

[0026] Also, a receiving method according to one aspect of the present invention is a receiving method for receiving a predetermined data unit in which data constituting an application is stored, the method comprising: receiving (i) the predetermined data unit, (ii) time specifying information indicating an operation time of the application, and control information storing identification information indicating whether the time specifying information is time information before leap second adjustment; and executing the application stored in the received predetermined data unit based on the received control information.

[0027] Such a receiving method can execute an application composed of a predetermined data unit at an intended time even when a leap second adjustment is made to the reference time information serving as a reference for the reference clock.

[0028] These general or specific aspects may be implemented by a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, an apparatus, an integrated circuit, a computer program, or a recording medium.

[0029] Hereinafter, embodiments will be specifically described with reference to the drawings.

[0030] Note that each of the embodiments described below shows general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present invention. In addition, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.

[0031] (Findings on which the present invention is based) In recent years, the resolution of displays such as TVs, smartphones, or tablet terminals has been increasing. In particular, in the broadcasting in Japan, 8K4K (resolution of 8K×4K) services are planned for 2020. In ultra-high resolution moving images such as 8K4K, it is difficult to decode in real time with a single decoder, so a method of performing decoding processing in parallel using multiple decoders has been studied.

[0032] Since the encoded data is multiplexed and transmitted based on a multiplexing method such as MPEG-2 TS or MMT, the receiving device needs to separate the video encoded data from the multiplexed data prior to decoding. Hereinafter, the process of separating the encoded data from the multiplexed data is called inverse multiplexing.

[0033] When parallelizing the decoding process, it is necessary to distribute the encoded data to be decoded to each of the decoders. When distributing the encoded data, it is necessary to analyze the encoded data itself. Especially in the case of content such as 8K, the bit rate is very high, so the processing load related to the analysis is large. Therefore, there has been a problem that the inverse multiplexing part becomes a bottleneck and real-time playback cannot be performed.

[0034] By the way, in moving image encoding methods such as H.264 and H.265 standardized by MPEG and ITU, the transmitting device can divide a picture into a plurality of regions called slices or slice segments and encode them so that each of the divided regions can be decoded independently. Therefore, for example, in the case of H.265, the receiving device that receives the broadcast can separate the data for each slice segment from the received data and output the data for each slice segment to a separate decoder, thereby realizing parallelization of the decoding process.

[0035] FIG. 1 is a diagram showing an example of dividing one picture into four slice segments in HEVC. For example, the receiving device includes four decoders, and each decoder decodes one of the four slice segments.

[0036] In conventional broadcasting, a transmission device stores one picture (access unit in the MPEG system standard) in one PES packet and multiplexes the PES packets into a TS packet stream. Therefore, a receiving device has to separate the payload of the PES packet, analyze the data of the access unit stored in the payload to separate each slice segment, and output the data of each separated slice segment to a decoder.

[0037] However, the inventor has found that there is a problem that it is difficult to perform this process in real time because the processing amount when analyzing the data of the access unit to separate the slice segments is large.

[0038] FIG. 2 is a diagram showing an example in which data of a picture divided into slice segments is stored in the payload of a PES packet.

[0039] As shown in FIG. 2, for example, data of a plurality of slice segments (slice segments 1 to 4) is stored in the payload of one PES packet. Also, the PES packet is multiplexed into a TS packet stream.

[0040] (Embodiment 1) Hereinafter, the case where H.265 is used as a moving image encoding method will be described as an example, but this embodiment can also be applied to the case where other encoding methods such as H.264 are used.

[0041] FIG. 3 is a diagram showing an example in which an access unit (picture) in this embodiment is divided into division units. The access unit is divided into a total of four tiles by being bisected in the horizontal and vertical directions by a function called a tile introduced by H.265. Also, a slice segment and a tile are associated one-to-one.

[0042] The reasons for bisecting in the horizontal and vertical directions as described above will be explained. First, at the time of decoding, generally a line memory for storing data of one horizontal line is required. However, in the case of ultra-high resolutions such as 8K4K, since the size in the horizontal direction becomes large, the size of the line memory increases. In the implementation of the receiving apparatus, it is desirable to be able to reduce the size of the line memory. In order to reduce the size of the line memory, division in the vertical direction is necessary. For division in the vertical direction, a data structure called a tile is required. For these reasons, tiles are used.

[0043] On the other hand, since an image generally has a high correlation in the horizontal direction, the encoding efficiency is improved if a wider range can be referenced in the horizontal direction. Therefore, from the viewpoint of encoding efficiency, it is desirable that the access unit be divided in the horizontal direction.

[0044] By bisecting the access unit in the horizontal and vertical directions, these two characteristics can be made compatible, and both the implementation aspect and the encoding efficiency can be considered. When a single decoder can decode a 4K2K moving image in real time, the 8K4K image is divided into four equal parts, and each slice segment is divided so as to be 4K2K, whereby the receiving apparatus can decode the 8K4K image in real time.

[0045] Next, the reason for associating the tiles and slice segments obtained by dividing the access unit in the horizontal and vertical directions on a one-to-one basis will be explained. In H.265, an access unit is composed of units called a plurality of NAL (Network Adaptation Layer) units.

[0046] The payload of the NAL unit stores either an access unit delimiter indicating the start position of the access unit, an SPS (Sequence Parameter Set) which is initialization information commonly used at the sequence level during decoding, a PPS (Picture Parameter Set) which is initialization information commonly used within a picture during decoding, an SEI (Supplemental Enhancement Information) which is not necessary for the decoding process itself but is required for processing and displaying the decoding result, or encoded data of a slice segment. The header of the NAL unit contains type information for identifying the data stored in the payload.

[0047] Here, when the transmitting device multiplexes the encoded data using a multiplexing format such as MPEG-2 TS, MMT (MPEG Media Transport), MPEG DASH (Dynamic Adaptive Streaming over HTTP), or RTP (Real-time Transport Protocol), the basic unit can be set to the NAL unit. In order to store one slice segment in one NAL unit, it is desirable to divide the access unit into units of slice segments when dividing the access unit into regions. For this reason, the transmitting device associates tiles and slice segments one-to-one.

[0048] As shown in FIG. 4, the transmitting device can also set tiles 1 to 4 together in one slice segment. However, in this case, all the tiles will be stored in one NAL unit, making it difficult for the receiving device to separate the tiles at the multiplexing layer.

[0049] Note that there are independent slice segments that can be decoded independently and reference slice segments that reference independent slice segments in the slice segment. Here, the case where independent slice segments are used will be described.

[0050] FIG. 5 is a diagram showing an example of data of an access unit divided so that the boundaries between tiles and slice segments coincide as shown in FIG. 3. The data of the access unit includes an NAL unit storing an access unit delimiter arranged at the head, NAL units of SPS, PPS, and SEI arranged thereafter, and data of slice segments storing data from tile 1 to tile 4 arranged thereafter. Note that the data of the access unit may not include some or all of the NAL units of SPS, PPS, and SEI.

[0051] Next, the configuration of the transmission device 100 according to the present embodiment will be described. FIG. 6 is a block diagram showing a configuration example of the transmission device 100 according to the present embodiment. This transmission device 100 includes an encoding unit 101, a multiplexing unit 102, a modulation unit 103, and a transmission unit 104.

[0052] The encoding unit 101 generates encoded data by encoding an input image according to, for example, H.265. Further, the encoding unit 101 divides an access unit into four slice segments (tiles) as shown in FIG. 3, for example, and encodes each slice segment.

[0053] The multiplexing unit 102 multiplexes the encoded data generated by the encoding unit 101. The modulation unit 103 modulates the data obtained by multiplexing. The transmission unit 104 transmits the modulated data as a broadcast signal.

[0054] Next, the configuration of the reception device 200 according to the present embodiment will be described. FIG. 7 is a block diagram showing a configuration example of the reception device 200 according to the present embodiment. This reception device 200 includes a tuner 201, a demodulation unit 202, a demultiplexing unit 203, a plurality of decoding units 204A to 204D, and a display unit 205.

[0055] The tuner 201 receives a broadcast signal. The demodulation unit 202 demodulates the received broadcast signal. The demodulated data is input to the demultiplexing unit 203.

[0056] The inverse multiplexing unit 203 separates the demodulated data into divided units and outputs the data for each divided unit to the decoding units 204A to 204D. Here, the divided unit is a divided area obtained by dividing the access unit, for example, a slice segment in H.265. Also, here, an 8K×4K image is divided into four 4K×2K images. Therefore, there are four decoding units 204A to 204D.

[0057] The plurality of decoding units 204A to 204D operate synchronously with each other based on a predetermined reference clock. Each decoding unit decodes the encoded data of the divided unit according to the DTS (Decoding Time Stamp) of the access unit and outputs the decoding result to the display unit 205.

[0058] The display unit 205 generates an 8K×4K output image by integrating the plurality of decoding results output from the plurality of decoding units 204A to 204D. The display unit 205 displays the generated output image according to the PTS (Presentation Time Stamp) of the access unit obtained separately. Note that when integrating the decoding results, the display unit 205 may perform filter processing such as a deblocking filter so that the boundary is not visually prominent in the boundary region of adjacent divided units such as the tile boundary.

[0059] In the above description, the transmission device 100 and the reception device 200 that perform broadcast transmission or reception are taken as examples, but the content may be transmitted and received via a communication network. When the reception device 200 receives the content via the communication network, the reception device 200 separates the multiplexed data from the IP packet received by a network such as Ethernet.

[0060] In broadcasting, the transmission path delay from when the broadcast signal is transmitted until it reaches the receiving device 200 is constant. On the other hand, in a communication network such as the Internet, due to the influence of congestion, the transmission path delay from when the data transmitted from the server reaches the receiving device 200 is not constant. Therefore, the receiving device 200 often does not perform precise synchronous playback based on a reference clock such as the PCR in the broadcast MPEG-2 TS. For this reason, the receiving device 200 may display the 8K4K output image on the display unit according to the PTS without strictly synchronizing each decoding unit.

[0061] Also, due to congestion in the communication network or the like, there may be a case where the decoding process for all divided units is not completed at the time indicated by the PTS of the access unit. In this case, the receiving device 200 skips the display of the access unit, or delays the display until at least four divided units have been decoded and the generation of the 8K4K image is completed.

[0062] Note that content may be transmitted and received using both broadcasting and communication in combination. Also, this method can be applied when playing back multiplexed data stored in a recording medium such as a hard disk or memory.

[0063] Next, a multiplexing method for access units divided into slice segments when MMT is used as the multiplexing method will be described.

[0064] FIG. 8 is a diagram showing an example when packetizing the data of the access unit of HEVC into MMT. SPS, PPS, SEI, etc. are not necessarily included in the access unit, but here the case where they exist is illustrated.

[0065] NAL units such as the access unit delimiter, SPS, PPS, and SEI, which are arranged before the first slice segment in the access unit, are grouped together and stored in the MMT packet #1. Subsequent slice segments are stored in separate MMT packets for each slice segment.

[0066] Note that, as shown in FIG. 9, a NAL unit arranged in the access unit before the leading slice segment may be stored in the same MMT packet as the leading slice segment.

[0067] Also, when NAL units such as End-of-Sequence or End-of-Bitstream indicating the end of a sequence or stream are added after the final slice segment, these are stored in the same MMT packet as the final slice segment. However, since NAL units such as End-of-Sequence or End-of-Bitstream are inserted at the end point of the decoding process or the connection point of two streams, etc., there may be cases where it is desirable for the receiving device 200 to easily acquire these NAL units in the multiplexing layer. In this case, these NAL units may be stored in an MMT packet different from the slice segment. Thereby, the receiving device 200 can easily separate these NAL units in the multiplexing layer.

[0068] Note that, as a multiplexing method, TS, DASH, RTP, etc. may be used. Also in these methods, the transmitting device 100 stores different slice segments in different packets. Thereby, it can be guaranteed that the receiving device 200 can separate the slice segments in the multiplexing layer.

[0069] For example, when TS is used, the encoded data is packetized into PES packets in units of slice segments. When RTP is used, the encoded data is packetized into RTP packets in units of slice segments. Also in these cases, as in the MMT packet #1 shown in FIG. 8, the NAL unit arranged before the slice segment and the slice segment may be packetized separately.

[0070] When TS is used, the transmitting device 100 indicates the unit of data stored in the PES packet, such as by using a data alignment descriptor. Also, since DASH is a method of downloading data units in the MP4 format called segments via HTTP or the like, the transmitting device 100 does not perform packetization of the encoded data during transmission. For this reason, the transmitting device 100 may create subsamples in units of slice segments and store information indicating the storage position of the subsamples in the MP4 header so that the receiving device 200 can detect slice segments in the multiplexing layer in MP4.

[0071] Hereinafter, the MMT packetization of slice segments will be described in detail.

[0072] As shown in FIG. 8, when the encoded data is packetized, data that is commonly referred to during decoding of all slice segments in an access unit such as SPS and PPS is stored in MMT packet #1. In this case, the receiving device 200 concatenates the payload data of MMT packet #1 and the data of each slice segment, and outputs the obtained data to the decoding unit. In this way, the receiving device 200 can easily generate input data for the decoding unit by concatenating the payloads of a plurality of MMT packets.

[0073] FIG. 10 is a diagram showing an example in which input data to the decoding units 204A to 204D shown in FIG. 8 is generated. The demultiplexing unit 203 concatenates the payload data of MMT packet #1 and MMT packet #2, so that the decoding unit 204A generates data necessary for decoding slice segment 1. The demultiplexing unit 203 similarly generates input data for the decoding units 204B to 204D. That is, the demultiplexing unit 203 generates the input data for the decoding unit 204B by concatenating the payload data of MMT packet #1 and MMT packet #3. The demultiplexing unit 203 generates the input data for the decoding unit 204C by concatenating the payload data of MMT packet #1 and MMT packet #4. The demultiplexing unit 203 generates the input data for the decoding unit 204D by concatenating the payload data of MMT packet #1 and MMT packet #5.

[0074] Note that the demultiplexing unit 203 may remove NAL units that are not necessary for the decoding process, such as access unit delimiters and SEI, from the payload data of MMT packet #1, and separate only the NAL units of SPS and PPS that are necessary for the decoding process and add them to the data of the slice segment.

[0075] When the encoded data is packetized as shown in FIG. 9, the demultiplexing unit 203 outputs MMT packet #1 including the leading data of the access unit in the multiplexing layer to the first decoding unit 204A. Further, the demultiplexing unit 203 analyzes the MMT packet including the leading data of the access unit in the multiplexing layer, separates the NAL units of SPS and PPS, and generates input data for each of the second and subsequent decoding units by adding the separated NAL units of SPS and PPS to each of the data of the second and subsequent slice segments.

[0076] Furthermore, it is desirable that the receiving apparatus 200 can identify the type of data stored in the MMT payload and the index number of the slice segment in the access unit when the slice segment is stored in the payload, using the information included in the header of the MMT packet. Here, the type of data refers to either the data before the slice segment (collectively referred to as such, which is the NAL unit arranged before the first slice segment in the access unit) or the data of the slice segment. When storing a unit obtained by fragmenting an MPU such as a slice segment in the MMT packet, a mode for storing an MFU (Media Fragment Unit) is used. When using this mode, the transmitting apparatus 100 can set, for example, a Data Unit, which is a basic unit of data in the MFU, to a sample (a data unit in MMT, corresponding to an access unit) or a subsample (a unit obtained by dividing a sample).

[0077] At this time, the header of the MMT packet includes a field called Fragmentation indicator and a field called Fragment counter.

[0078] The Fragmentation indicator indicates whether the data stored in the payload of an MMT packet is a fragmentation of a Data unit, and if it is a fragmentation, whether the fragment is the first or last fragment in the Data unit, or a fragment that is neither the first nor the last. In other words, the Fragmentation indicator included in the header of a certain packet is identification information indicating which of the following cases it is: (1) the packet is the only one included in the Data unit which is the basic data unit, (2) the Data unit is divided and stored in multiple packets, and the packet is the first packet of the Data unit, (3) the Data unit is divided and stored in multiple packets, and the packet is a packet other than the first and last packets of the Data unit, and (4) the Data unit is divided and stored in multiple packets, and the packet is the last packet of the Data unit.

[0079] The Fragment counter is an index number indicating which fragment in the Data unit the data stored in the MMT packet corresponds to.

[0080] Therefore, when the transmitting device 100 sets the samples in the MMT as a Data unit and sets the pre-slice segment data and each slice segment in units of fragments of the Data unit respectively, the receiving device 200 can identify the type of data stored in the payload by using the information included in the header of the MMT packet. That is, the demultiplexing unit 203 can generate the input data to each decoding unit 204A to 204D by referring to the header of the MMT packet.

[0081] FIG. 11 is a diagram showing an example in which samples are set as a Data unit and the pre-slice segment data and the slice segments are packetized as fragments of the Data unit.

[0082] The pre-slice segment data and the slice segment are divided into five fragments from fragment #1 to fragment #5. Each fragment is stored in an individual MMT packet. At this time, the values of the Fragmentation indicator and the Fragment counter included in the header of the MMT packet are as shown in the figure.

[0083] For example, the Fragment indicator is a 2-bit binary value. The Fragment indicator of MMT packet #1, which is the start of the Data unit, the Fragment indicator of MMT packet #5, which is the last, and the Fragment indicators of the packets in between, i.e., MMT packets #2 to #4, are set to different values. Specifically, the Fragment indicator of MMT packet #1, which is the start of the Data unit, is set to 01, the Fragment indicator of MMT packet #5, which is the last, is set to 11, and the Fragment indicators of the packets in between, i.e., MMT packets #2 to #4, are set to 10. Note that when the Data unit contains only one MMT packet, the Fragment indicator is set to 00.

[0084] Also, the Fragment counter is 4, which is the value obtained by subtracting 1 from the total number of fragments, i.e., 5, in MMT packet #1, and it decreases by 1 in each subsequent packet, and is 0 in the last MMT packet #5.

[0085] Therefore, the receiving device 200 can identify the MMT packet storing the pre-slice segment data using either the Fragment indicator or the Fragment counter. Also, the receiving device 200 can identify the MMT packet storing the Nth slice segment by referring to the Fragment counter.

[0086] The header of the MMT packet separately includes the sequence number within the MPU of the Movie Fragment to which the Data unit belongs, the sequence number of the MPU itself, and the sequence number within the Movie Fragment of the sample to which the Data unit belongs. By referring to these, the demultiplexing unit 203 can uniquely determine the sample to which the Data unit belongs.

[0087] Furthermore, since the demultiplexing unit 203 can determine the index number of the fragment within the Data unit from the Fragment counter or the like, even when packet loss occurs, the slice segment stored in the fragment can be uniquely identified. For example, even if the fragment #4 shown in FIG. 11 cannot be obtained due to packet loss, since it can be known that the next received fragment after fragment #3 is fragment #5, the slice segment 4 stored in fragment #5 can be correctly output to the decoding unit 204D instead of the decoding unit 204C.

[0088] Note that when a transmission path that guarantees no packet loss is used, the demultiplexing unit 203 does not need to determine the type of data stored in the MMT packet or the index number of the slice segment by referring to the header of the MMT packet, and can simply process the arrived packets periodically. For example, when an access unit is transmitted by a total of 5 MMT packets including pre-slice data and 4 slice segments, after the receiving device 200 determines the pre-slice data of the access unit for which decoding is to be started, it can sequentially obtain the pre-slice data and the data of the 4 slice segments by processing the received MMT packets in order.

[0089] Hereinafter, a modification example of packetization will be described.

[0090] The slice segment does not necessarily have to be divided both horizontally and vertically within the plane of the access unit. As shown in FIG. 1, the access unit may be divided only horizontally or only vertically.

[0091] Also, when the access unit is divided only horizontally, it is not necessary to use tiles.

[0092] Also, the number of in-plane divisions in the access unit is arbitrary and is not limited to four. However, the area sizes of the slice segment and the tile need to be equal to or greater than the lower limit of the encoding standard such as H.265.

[0093] The transmission device 100 may store identification information indicating the in-plane division method in the access unit in an MMT message, a TS descriptor, or the like. For example, information indicating the number of horizontal and vertical divisions in the plane may be stored. Alternatively, as shown in FIG. 3, unique identification information may be assigned to the division method, such as being divided into two equal parts horizontally and vertically, or, as shown in FIG. 1, being divided into four equal parts horizontally. For example, when the access unit is divided as shown in FIG. 3, the identification information indicates mode 1, and when the access unit is divided as shown in FIG. 1, the identification information indicates mode 1.

[0094] Also, information indicating the constraints of the encoding conditions related to the in-plane division method may be included in the multiplexing layer. For example, information indicating that one slice segment is composed of one tile may be used. Alternatively, information indicating that the reference block for motion compensation during decoding of the slice segment or the tile is limited to the slice segment or tile at the same position in the screen, or is limited to blocks within a predetermined range in adjacent slice segments may be used.

[0095] Further, the transmission device 100 may switch whether to divide an access unit into a plurality of slice segments according to the resolution of the moving image. For example, when the moving image to be processed has a resolution of 4K2K, the transmission device 100 may not perform in-plane division, and when the moving image to be processed has a resolution of 8K4K, the access unit may be divided into four. By prescribing in advance the division method in the case of an 8K4K moving image, the reception device 200 can determine the presence or absence of in-plane division and the division method by acquiring the resolution of the received moving image, and can switch the decoding operation.

[0096] Further, the reception device 200 can detect the presence or absence of in-plane division by referring to the header of the MMT packet. For example, when the access unit is not divided, if the Data unit of MMT is set to sample, the fragmentation of the Data unit is not performed. Therefore, when the value of the Fragment counter included in the header of the MMT packet is always zero, the reception device 200 can determine that the access unit is not divided. Alternatively, the reception device 200 may detect whether the value of the Fragmentation indicator is always 01. The reception device 200 can also determine that the access unit is not divided when the value of the Fragmentation indicator is always 01.

[0097] Further, the reception device 200 can also handle the case where the number of in-plane divisions in the access unit does not match the number of decoding units. For example, when the reception device 200 includes two decoding units 204A and 204B that can decode 8K2K encoded data in real time, the demultiplexing unit 203 outputs two of the four slice segments that make up the 8K4K encoded data to the decoding unit 204A.

[0098] FIG. 12 is a diagram showing an operation example when the data packetized by MMT as shown in FIG. 8 is input to two decoding units 204A and 204B. Here, it is desirable that the receiving apparatus 200 can output by integrating the decoding results in the decoding units 204A and 204B as they are. Therefore, the de-multiplexing unit 203 selects slice segments to be output to each of the decoding units 204A and 204B so that the decoding results of each of the decoding units 204A and 204B are spatially continuous.

[0099] Further, the de-multiplexing unit 203 may select a decoding unit to be used according to the resolution or frame rate of the encoded data of the moving image. For example, when the receiving apparatus 200 includes four 4K2K decoding units, if the resolution of the input image is 8K4K, the receiving apparatus 200 performs decoding processing using all four decoding units. Also, if the resolution of the input image is 4K2K, the receiving apparatus 200 performs decoding processing using only one decoding unit. Alternatively, even if the in-plane is divided into four, when the 8K4K can be decoded in real time by a single decoding unit, the de-multiplexing unit 203 integrates all the divided units and outputs them to one decoding unit.

[0100] Furthermore, the receiving apparatus 200 may determine a decoding unit to be used in consideration of the frame rate. For example, when the receiving apparatus 200 includes two decoding units whose upper limit of the frame rate that can be decoded in real time is 60 fps when the resolution is 8K4K, there is a case where encoded data of 8K4K at 120 fps is input. At this time, assuming that the in-plane is composed of four divided units, similar to the example of FIG. 12, slice segment 1 and slice segment 2 are input to the decoding unit 204A, and slice segment 3 and slice segment 4 are input to the decoding unit 204B. Since each of the decoding units 204A and 204B can decode in real time up to 120 fps if it is 8K2K (half of the resolution of 8K4K), decoding processing is performed by these two decoding units 204A and 204B.

[0101] Also, even if the resolution and frame rate are the same, the processing amount will be different if the profile or level in the encoding method, or the encoding method itself such as H.264 or H.265, is different. Therefore, the receiving device 200 may select a decoding unit to be used based on this information. Note that when the receiving device 200 cannot decode all of the encoded data received by broadcasting or communication, or when all of the slice segments or tiles constituting the area selected by the user cannot be decoded, it may automatically determine slice segments or tiles that can be decoded within the processing range of the decoding unit. Alternatively, the receiving device 200 may provide a user interface for the user to select an area to be decoded. At this time, the receiving device 200 may display a warning message indicating that it cannot decode all areas, or may display information indicating the number of decodable areas, slice segments, or tiles.

[0102] Also, the above method can also be applied when MMT packets storing slice segments of the same encoded data are transmitted and received using a plurality of transmission paths such as broadcasting and communication.

[0103] Also, the transmitting device 100 may perform encoding so that the areas of the respective slice segments overlap in order to make the boundary of the division unit inconspicuous. In the example shown in FIG. 13, an 8K4K picture is divided into four slice segments 1 to 4. Each of the slice segments 1 to 3 is, for example, 8K×1.1K, and the slice segment 4 is 8K×1K. Also, adjacent slice segments overlap each other. By doing so, at the boundary in the case of four-way division indicated by the dotted line, motion compensation at the time of encoding can be efficiently executed, so that the image quality at the boundary portion is improved. In this way, the image quality deterioration at the boundary portion is reduced.

[0104] In this case, the display unit 205 cuts out an 8K×1K area from an 8K×1.1K area and integrates the obtained area. Note that the transmission device 100 may separately transmit information indicating whether the slice segments are encoded with overlap and the range of the overlap, included in the multiplexing layer or the encoded data.

[0105] Note that the same method can also be applied when tiles are used.

[0106] Hereinafter, the operation flow of the transmission device 100 will be described. FIG. 14 is a flowchart showing an operation example of the transmission device 100.

[0107] First, the encoding unit 101 divides a picture (access unit) into a plurality of slice segments (tiles) which are a plurality of areas (S101). Next, the encoding unit 101 generates encoded data corresponding to each of the plurality of slice segments by encoding each of the plurality of slice segments so that it can be independently decoded (S102). Note that the encoding unit 101 may encode the plurality of slice segments with a single encoding unit, or may perform parallel processing with a plurality of encoding units.

[0108] Next, the multiplexing unit 102 multiplexes the plurality of encoded data generated by the encoding unit 101 by storing the plurality of encoded data in a plurality of MMT packets (S103). Specifically, as shown in FIGS. 8 and 9, the multiplexing unit 102 stores the plurality of encoded data in a plurality of MMT packets so that the encoded data corresponding to different slice segments is not stored in one MMT packet. Further, as shown in FIG. 8, the multiplexing unit 102 stores control information commonly used for all decoding units in the picture in an MMT packet #1 different from the plurality of MMT packets #2 to #5 in which the plurality of encoded data is stored. Here, the control information includes at least one of an access unit delimiter, SPS, PPS, and SEI.

[0109] Note that the multiplexing unit 102 may store the control information in the same MMT packet as any one of the plurality of MMT packets in which the plurality of encoded data are stored. For example, as shown in FIG. 9, the multiplexing unit 102 may store the control information in the first MMT packet (MMT packet #1 in FIG. 9) among the plurality of MMT packets in which the plurality of encoded data are stored.

[0110] Finally, the transmission device 100 transmits a plurality of MMT packets. Specifically, the modulation unit 103 modulates the data obtained by multiplexing, and the transmission unit 104 transmits the modulated data (S104).

[0111] FIG. 15 is a block diagram showing a configuration example of the reception device 200, and is a diagram showing in detail the configuration of the demultiplexing unit 203 shown in FIG. 7 and the subsequent stage thereof. As shown in FIG. 15, the reception device 200 further includes a decoding command unit 206. The demultiplexing unit 203 includes a type discrimination unit 211, a control information acquisition unit 212, a slice information acquisition unit 213, and a decoded data generation unit 214.

[0112] Hereinafter, the operation flow of the reception device 200 will be described. FIG. 16 is a flowchart showing an operation example of the reception device 200. Here, the operation for one access unit is shown. When the decoding process for a plurality of access units is executed, the process of this flowchart is repeated.

[0113] First, the reception device 200 receives, for example, a plurality of packets (MMT packets) generated by the transmission device 100 (S201).

[0114] Next, the type discrimination unit 211 acquires the type of the encoded data stored in the received packet by analyzing the header of the received packet (S202).

[0115] Next, the type discrimination unit 211 determines whether the data stored in the received packet is pre-slice segment data or slice segment data based on the type of the acquired encoded data (S203).

[0116] When the data stored in the received packet is pre-slice segment data (Yes in S203), the control information acquisition unit 212 acquires the pre-slice segment data of the access unit to be processed from the payload of the received packet, and stores the pre-slice segment data in the memory (S204).

[0117] On the other hand, when the data stored in the received packet is slice segment data (No in S203), the receiving device 200 uses the header information of the received packet to determine which region among a plurality of regions the data stored in the received packet is encoded data of. Specifically, the slice information acquisition unit 213 acquires the index number Idx of the slice segment stored in the received packet by analyzing the header of the received packet (S205). Specifically, the index number Idx is the index number in the Movie Fragment of the access unit (sample in MMT).

[0118] Note that the process of this step S205 may be performed collectively in step S202.

[0119] Next, the decoded data generation unit 214 determines a decoding unit for decoding the slice segment (S206). Specifically, the index number Idx and a plurality of decoding units are associated in advance, and the decoded data generation unit 214 determines the decoding unit corresponding to the index number Idx acquired in step S205 as the decoding unit for decoding the slice segment.

[0120] Note that, as described in the example of FIG. 12, the decoded data generation unit 214 may determine a decoding unit for decoding the slice segment based on at least one of the resolution of the access unit (picture), the method of dividing the access unit into a plurality of slice segments (tiles), and the processing capabilities of the plurality of decoding units included in the receiving apparatus 200. For example, the decoded data generation unit 214 discriminates the method of dividing the access unit based on identification information in a descriptor such as an MMT message or a TS section.

[0121] Next, the decoded data generation unit 214 generates a plurality of input data (combined data) to be input to the plurality of decoding units by combining control information commonly used for all decoding units in the picture, which is included in any of the plurality of packets, and each of the plurality of encoded data of the plurality of slice segments. Specifically, the decoded data generation unit 214 acquires the data of the slice segment from the payload of the received packet. The decoded data generation unit 214 generates the input data to the decoding unit determined in step S206 by combining the pre-slice segment data stored in the memory in step S204 and the acquired data of the slice segment (S207).

[0122] After step S204 or S207, if the data of the received packet is not the final data of the access unit (No in S208), the processing after step S201 is performed again. That is, the above processing is repeated until the input data to the plurality of decoding units 204A to 204D corresponding to all the slice segments included in the access unit is generated.

[0123] Note that the timing at which the packet is received is not limited to the timing shown in FIG. 16, and a plurality of packets may be received in advance or sequentially and stored in a memory or the like.

[0124] On the other hand, when the data of the received packet is the final data of the access unit (Yes in S208), the decoding instruction unit 206 outputs the plurality of input data generated in step S207 to the corresponding decoding units 204A to 204D (S209).

[0125] Next, the plurality of decoding units 204A to 204D generate a plurality of decoded images by decoding the plurality of input data in parallel according to the DTS of the access unit (S210).

[0126] Finally, the display unit 205 generates a display image by combining the plurality of decoded images generated by the plurality of decoding units 204A to 204D, and displays the display image according to the PTS of the access unit (S211).

[0127] Note that the receiving device 200 obtains the DTS and PTS of the access unit by analyzing the header information of the MPU or the payload data of the MMT packet storing the header information of the Movie Fragment. Also, when TS is used as the multiplexing method, the receiving device 200 obtains the DTS and PTS of the access unit from the header of the PES packet. When RTP is used as the multiplexing method, the receiving device 200 obtains the DTS and PTS of the access unit from the header of the RTP packet.

[0128] Also, when integrating the decoding results of a plurality of decoding units, the display unit 205 may perform filter processing such as a deblocking filter at the boundary of adjacent divided units. Note that when displaying the decoding result of a single decoding unit, filter processing is not necessary. Therefore, the display unit 205 may switch the processing according to whether to perform filter processing at the boundary of the decoding results of the plurality of decoding units. Whether filter processing is necessary may be defined in advance according to the presence or absence of division. Alternatively, information indicating whether filter processing is necessary may be separately stored in the multiplexing layer. Also, information necessary for filter processing such as filter coefficients may be stored in the SPS, PPS, SEI, or within a slice segment. The decoding units 204A to 204D, or the inverse multiplexing unit 203, acquire this information by analyzing the SEI, and output the acquired information to the display unit 205. The display unit 205 performs filter processing using this information. Note that when this information is stored within a slice segment, it is desirable for the decoding units 204A to 204D to acquire this information.

[0129] Note that in the above description, an example in which the types of data stored in the fragment are two types, i.e., pre-slice segment data and slice segment, is shown. However, the number of data types may be three or more. In this case, branching according to the type is performed in step S203.

[0130] Also, when the data size of the slice segment is large, the transmission device 100 may fragment the slice segment and store it in the MMT packet. That is, the transmission device 100 may fragment the pre-slice segment data and the slice segment. In this case, when the access unit and the data unit are set to be equal as in the packetization example shown in FIG. 11, the following problems occur.

[0131] For example, when slice segment 1 is divided into three fragments, slice segment 1 is divided into three packets with fragment counter values from 1 to 3 and transmitted. Also, for slice segments 2 and later, the fragment counter value becomes 4 or more, and the association between the fragment counter value and the data stored in the payload cannot be established. Therefore, the receiving device 200 cannot identify the packet storing the start data of the slice segment from the information in the header of the MMT packet.

[0132] In such a case, the receiving device 200 may analyze the data in the payload of the MMT packet to identify the start position of the slice segment. Here, as a format for storing NAL units in multiplexing layers in H.264 or H.265, there are two types: a byte stream format called where a start code consisting of a specific bit sequence is added immediately before the NAL unit header, and a NAL size format called where a field indicating the size of the NAL unit is added.

[0133] The byte stream format is used in MPEG-2 systems, RTP, etc. The NAL size format is used in MP4, as well as DASH and MMT that use MP4.

[0134] When the byte stream format is used, the receiving device 200 analyzes whether the start data of the packet matches the start code. If the start data of the packet matches the start code, the receiving device 200 can detect whether the data contained in the packet is the data of the slice segment by obtaining the type of the NAL unit from the subsequent NAL unit header.

[0135] On the other hand, in the case of the NAL size format, the receiving device 200 cannot detect the start position of the NAL unit based on the bit string. Therefore, in order for the receiving device 200 to obtain the start position of the NAL unit, it is necessary to shift the pointer by reading data by the size of the NAL unit in order from the first NAL unit of the access unit.

[0136] However, in the header of the MPU or Movie Fragment in MMT, when the size in sub-sample units is indicated and the sub-sample corresponds to pre-slice data or a slice segment, the receiving device 200 can specify the start position of each NAL unit based on the size information of the sub-sample. Therefore, the transmitting device 100 may include information indicating whether information in sub-sample units exists in the MPU or Movie Fragment in information that the receiving device 200 obtains at the start of data reception, such as MPT in MMT.

[0137] Note that the data of the MPU is an extension based on the MP4 format. In MP4, there are a mode in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and a mode in which they cannot be stored. Also, information for specifying this mode is shown as the entry name of SampleEntry. When the mode in which storage is possible is used and the parameter set is included in the sample, the receiving device 200 obtains the parameter set by the method described above.

[0138] On the other hand, when an un-storable mode is used, the parameter set is stored as Decoder Specific Information in the SampleEntry, or is stored using a stream for the parameter set. Here, since the stream for the parameter set is not generally used, it is desirable for the transmitting device 100 to store the parameter set in the Decoder Specific Information. In this case, the receiving device 200 analyzes the SampleEntry transmitted as the metadata of the MPU or the metadata of the Movie Fragment in the MMT packet, and acquires the parameter set referred to by the access unit.

[0139] When the parameter set is stored as sample data, the receiving device 200 can acquire the parameter set necessary for decoding by referring only to the sample data without referring to the SampleEntry. At this time, the transmitting device 100 does not have to store the parameter set in the SampleEntry. By doing so, the transmitting device 100 can use the same SampleEntry in different MPUs, so that the processing load of the transmitting device 100 at the time of MPU generation can be reduced. Furthermore, there is an advantage that the receiving device 200 does not need to refer to the parameter set in the SampleEntry.

[0140] Alternatively, the transmitting device 100 may store one default parameter set in the SampleEntry and store the parameter set referred to by the access unit in the sample data. In conventional MP4, since it was common to store the parameter set in the SampleEntry, there may be a receiving device that stops playback when the parameter set does not exist in the SampleEntry. This problem can be solved by using the above method.

[0141] Alternatively, the transmitting device 100 may store the parameter set in the sample data only when a parameter set different from the default parameter set is used.

[0142] Note that since it is possible to store the parameter set in SampleEntry in both modes, the transmission device 100 may always store the parameter set in VisualSampleEntry, and the receiving device 200 may always obtain the parameter set from VisualSampleEntry.

[0143] Note that in the MMT standard, header information of MP4 such as Moov and Moof is transmitted as MPU metadata or movie fragment metadata, but the transmission device 100 does not necessarily have to transmit MPU metadata and movie fragment metadata. Further, the receiving device 200 can also determine whether SPS and PPS are stored in the sample data based on services of the ARIB (Association of Radio Industries and Businesses) standard, the type of asset, or the presence or absence of transmission of MPU meta.

[0144] FIG. 17 is a diagram showing an example in which pre-slice segment data and each slice segment are set in different Data units.

[0145] In the example shown in FIG. 17, the sizes of the pre-slice segment data and the data from slice segment 1 to slice segment 4 are Length#1 to Length#5, respectively. The field values of the Fragmentation indicator, Fragment counter, and Offset included in the header of the MMT packet are as shown in the figure.

[0146] Here, Offset is offset information indicating the bit length (offset) from the start of the encoded data of the sample (access unit or picture) to which the payload data belongs to the start byte of the payload data (encoded data) included in the MMT packet. Note that although the description is given assuming that the value of the Fragment counter starts from the value obtained by subtracting 1 from the total number of fragments, it may start from other values.

[0147] FIG. 18 is a diagram showing an example when a Data unit is fragmented. In the example shown in FIG. 18, slice segment 1 is divided into three fragments and stored in MMT packets #2 to #4 respectively. Also in this case, if the data sizes of the respective fragments are denoted as Length#2_1 to Length#2_3 respectively, the values of the respective fields are as shown in the figure.

[0148] Thus, when a data unit such as a slice segment is set as the Data unit, the start of the access unit and the start of the slice segment can be determined as follows based on the field values of the MMT packet header.

[0149] The start of the payload in a packet where the value of Offset is 0 is the start of the access unit.

[0150] The start of the payload of a packet where the value of Offset is different from 0 and the value of the Fragmentation indcatorno is 00 or 01 is the start of the slice segment.

[0151] Also, when fragmentation of the Data unit does not occur and packet loss does not occur, the receiving device 200 can specify the index number of the slice segment stored in the MMT packet based on the number of slice segments acquired after detecting the start of the access unit.

[0152] Also, even when the Data unit of the pre-slice segment data is fragmented, similarly, the receiving device 200 can detect the access unit and the start of the slice segment.

[0153] Also, when packet loss occurs, or when the SPS, PPS, and SEI included in the pre-slice segment data are set in separate Data units, the receiving device 200 identifies the MMT packet storing the start data of the slice segment based on the analysis result of the MMT header, and then analyzes the header of the slice segment to identify the start position of the slice segment or tile within the picture (access unit). The processing amount related to the analysis of the slice header is small, and the processing load is not a problem.

[0154] Thus, each of the encoded data of a plurality of slice segments is associated one-to-one with a basic data unit (Data unit) which is a unit of data stored in one or more packets. Also, each of the plurality of encoded data is stored in one or more MMT packets.

[0155] The header information of each MMT packet includes a Fragmentation indicator (identification information) and an Offset (offset information).

[0156] The receiving device 200 determines that the start of the payload data included in the packet having header information including a Fragmentation indicator with a value of 00 or 01 is the start of the encoded data of each slice segment. Specifically, it determines that the start of the payload data included in the packet having header information including an Offset whose value is not 0 and a Fragmentation indicator whose value is 00 or 01 is the start of the encoded data of each slice segment.

[0157] Also, in the example of FIG. 17, the start of the Data unit is either the start of the access unit or the start of the slice segment, and the value of the Fragmentation indicator is 00 or 01. Further, the receiving device 200 can detect the start of the access unit or the start of the slice segment without referring to the Offset by determining whether the start of the Data Unit is the access unit delimiter or the slice segment with reference to the type of the NAL unit.

[0158] In this way, by the transmitting device 100 performing packetization so that the start of the NAL unit always starts from the start of the payload of the MMT packet, including the case where the pre-slice segment data is divided into a plurality of Data units, the receiving device 200 can detect the start of the access unit or the slice segment by analyzing the Fragmentation indicator and the NAL unit header. The type of the NAL unit exists in the first byte of the NAL unit header. Therefore, when analyzing the header part of the MMT packet, the receiving device 200 can obtain the type of the NAL unit by additionally analyzing one byte of data.

[0159] In the case of audio, it is sufficient for the receiving device 200 to be able to detect the start of the access unit, and it may be determined based on whether the value of the Fragmentation indicator is 00 or 01.

[0160] Also, as described above, when storing the encoded data encoded so as to enable split decoding in the PES packet of MPEG-2 TS, the transmitting device 100 can use the data alignment descriptor. Hereinafter, an example of the method of storing the encoded data in the PES packet will be described in detail.

[0161] For example, in HEVC, by using a data alignment descriptor, the transmission device 100 can indicate whether the data stored in the PES packet is an access unit, a slice segment, or a tile. The types of alignment in HEVC are defined as follows.

[0162] Alignment type = 8 indicates a slice segment of HEVC. Alignment type = 9 indicates a slice segment or an access unit of HEVC. Alignment type = 12 indicates a slice segment or a tile of HEVC.

[0163] Therefore, the transmission device 100 can indicate, for example, by using type 9, that the data of the PES packet is either a slice segment or pre-slice-segment data. Since a type indicating a slice rather than a slice segment is also defined separately, the transmission device 100 may use a type indicating a slice rather than a slice segment.

[0164] Also, DTS and PTS included in the header of the PES packet are set only in the PES packet containing the leading data of the access unit. Therefore, the receiving device 200 can determine that if the type is 9 and there is a DTS or PTS field in the PES packet, the entire access unit or the leading division unit in the access unit is stored in the PES packet.

[0165] Further, the transmission device 100 may use a field such as transport_priority that indicates the priority of a TS packet storing a PES packet including the head data of an access unit, so that the receiving device 200 can distinguish the data included in the packet. Also, the receiving device 200 may determine the data included in the packet by analyzing whether the payload of the PES packet is an access unit delimiter. Further, the data_alignment_indicator in the PES packet header indicates whether data is stored in the PES packet according to these types. If this flag (data_alignment_indicator) is set to 1, it is guaranteed that the data stored in the PES packet follows the type indicated by the data alignment descriptor.

[0166] Also, the transmission device 100 may use the data alignment descriptor only when packetizing the PES packet in units that can be decoded separately, such as slice segments. Thereby, the receiving device 200 can determine that the encoded data is packetized in units that can be decoded separately when the data alignment descriptor exists, and can determine that the encoded data is packetized in access unit units when the data alignment descriptor does not exist. Note that in the case where the data_alignment_indicator is set to 1 and the data alignment descriptor does not exist, it is defined in the MPEG-2 TS standard that the unit of packetization is an access unit.

[0167] If the data alignment descriptor is included in the PMT, the receiving device 200 determines that the PES packets are packetized in units that can be split and decoded, and can generate input data for each decoding unit based on the packetized units. Further, if the data alignment descriptor is not included in the PMT and it is determined that parallel decoding of the encoded data is necessary based on the program information or other descriptor information, the receiving device 200 generates input data for each decoding unit by analyzing the slice header of the slice segment or the like. Also, when the encoded data can be decoded by a single decoding unit, the receiving device 200 decodes the data of the entire access unit by the corresponding decoding unit. Note that when information indicating whether the encoded data is composed of units that can be split and decoded, such as slice segments or tiles, is separately indicated by a descriptor in the PMT or the like, the receiving device 200 may determine whether the encoded data can be decoded in parallel based on the analysis result of the descriptor.

[0168] Also, since the DTS and PTS included in the header of the PES packet are set only in the PES packet containing the leading data of the access unit, when the access unit is split and packetized into PES packets, the second and subsequent PES packets do not contain information indicating the DTS and PTS of the access unit. Therefore, when performing the decoding process in parallel, each decoding unit 204A to 204D and the display unit 205 use the DTS and PTS stored in the header of the PES packet containing the leading data of the access unit.

[0169] (Embodiment 2) In Embodiment 2, in MMT, a method of storing data in the NAL size format in an MPU based on the MP4 format will be described. Note that hereinafter, as an example, a method of storing in the MPU used in MMT will be described, but such a storage method is also applicable to DASH, which is also based on the same MP4 format.

[0170] [Method of Storing in MPU] In the MP4 format, multiple access units are grouped together and stored in a single MP4 file. The MPU used in MMT has data for each media stored in a single MP4 file, and the data can contain any number of access units. Since the MPU is a unit that can be decoded alone, for example, access units in GOP units are stored in the MPU.

[0171] Figure 19 is a diagram showing the configuration of the MPU. The beginning of the MPU is ftyp, mmpu, and moov, which are collectively defined as MPU metadata. The moov stores initialization information common to the file and the MMT hint track.

[0172] Also, the moof stores initialization information and sizes for each sample and subsample, information that can identify the presentation time (PTS) and decoding time (DTS) (sample_duration, sample_size, sample_composition_time_offset), and a data_offset indicating the position of the data, etc.

[0173] Also, multiple access units are each stored as samples in mdat (mdat box). The data excluding samples in moof and mdat is defined as movie fragment metadata (hereinafter referred to as MF metadata), and the sample data in mdat is defined as media data.

[0174] Figure 20 is a diagram showing the configuration of the MF metadata. As shown in Figure 20, the MF metadata consists more specifically of the type, length, and data of the moof box (moof) and the type and length of the mdat box (mdat).

[0175] When storing access units in MP4 data, there are a mode where parameter sets such as the SPS and PPS of H.264 and H.265 can be stored as sample data, and a mode where they cannot be stored.

[0176] Here, in the non-storable mode, the parameter set is stored in the Decoder Specific Information of the SampleEntry in moov. Also, in the storable mode, the parameter set is included in the sample.

[0177] MPU metadata, MF metadata, and media data are each stored in the MMT payload, and as an identifier for identifying these data, a Fragment Type (FT) is stored in the header of the MMT payload. FT = 0 indicates MPU metadata, FT = 1 indicates MF metadata, and FT = 2 indicates media data.

[0178] Note that in FIG. 19, an example where MPU metadata units and MF metadata units are stored in the MMT payload as data units is illustrated. However, units such as ftyp, mmpu, moov, and moof may be stored in the MMT payload in data unit units as data units. Similarly, in FIG. 19, an example where sample units are stored in the MMT payload as data units is illustrated. However, data units may be composed of sample units or NAL unit units, and such data units may be stored in the MMT payload in data unit units. Such data units may be further stored in the MMT payload in fragmented units.

[0179] [Conventional Transmission Method and Problems] Conventionally, when encapsulating a plurality of access units in the MP4 format, moov and moof were created when all the samples to be stored in MP4 were complete.

[0180] When transmitting the MP4 format in real time using broadcasting or the like, for example, assuming that the samples stored in one MP4 file are in GOP units, since moov and moof are created after the time samples in GOP units are accumulated, a delay due to encapsulation occurs. Due to such encapsulation on the transmitting side, the End-to-End delay always becomes longer by the GOP unit time. This makes it difficult to provide services in real time, and in particular, when live content is transmitted, it leads to deterioration of the service for viewers.

[0181] Figure 21 is a diagram for explaining the transmission order of data. When applying MMT to broadcasting, as shown in Fig. 21(a), when placing MMT packets in the order of MPU composition and transmitting them (transmitting in the order of MMT packets #1, #2, #3, #4, #5, #6), a delay due to encapsulation occurs in the transmission of MMT packets.

[0182] In order to prevent this delay due to encapsulation, as shown in Fig. 21(b), a method has been proposed in which MPU header information such as MPU metadata and MF metadata is not sent (packets #1 and #2 are not transmitted, and packets #3 - #6 are transmitted in this order). Also, as shown in Fig. 20(c), a method can be considered in which media data is transmitted first without waiting for the creation of MPU header information, and MPU header information is transmitted after the transmission of media data (transmitting in the order of #3 - #6, #1, #2).

[0183] When the receiving device does not receive the MPU header information, it decodes without using the MPU header information. Also, when the MPU header information is sent after the media data in the receiving device, it waits to obtain the MPU header information and then decodes.

[0184] However, in a conventional MP4-compliant receiving device, it is not guaranteed that decoding can be performed without using MPU header information. Also, when the receiving device performs decoding without using the MPU header through special processing and uses the conventional transmission method, the decoding process becomes complicated, and real-time decoding is likely to be difficult. Further, when the receiving device performs decoding after waiting to acquire the MPU header information, it is necessary to buffer the media data until the receiving device acquires the header information. However, the buffer model is not defined, and decoding is not guaranteed.

[0185] Therefore, as shown in FIG. 20(d), the transmitting device according to Embodiment 2 transmits the MPU metadata before the media data by storing only the information common to the MPU metadata. Then, the transmitting device according to Embodiment 2 transmits the MF metadata that causes a delay in generation after the media data. Thereby, a transmission method or a reception method that can guarantee decoding of media data is provided.

[0186] Hereinafter, the reception method when each of the transmission methods shown in FIGS. 21(a) to 21(d) is used will be described.

[0187] In each of the transmission methods shown in FIG. 21, first, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data.

[0188] After configuring the MPU data, when the transmitting device transmits the data in the order of MPU metadata, MF metadata, and media data as shown in FIG. 21(a), the receiving device can perform decoding by any of the following methods (A-1) and (A-2).

[0189] (A-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data using the MPU header information.

[0190] (A-2) The receiving device decodes the media data without using the MPU header information.

[0191] All of these methods cause a delay due to encapsulation on the transmitting side, but on the receiving device, there is an advantage that it is not necessary to buffer the media data to obtain the MPU header. When not buffering, there is no need to mount memory for buffering, and furthermore, no buffering delay occurs. Also, the method of (A-1) can be applied to a conventional receiving device because decoding is performed using the MPU header information.

[0192] When the transmitting device transmits only the media data as shown in Fig. 21(b), the receiving device can perform decoding by the following method (B-1).

[0193] (B-1) The receiving device decodes the media data without using the MPU header information.

[0194] Also, although not shown, when the MPU metadata is transmitted before the transmission of the media data in Fig. 21(b), decoding can be performed by the following method (B-2).

[0195] (B-2) The receiving device decodes the media data using the MPU metadata.

[0196] Both of the above methods (B-1) and (B-2) have the advantage that no delay due to encapsulation occurs on the transmitting side and it is not necessary to buffer the media data to obtain the MPU header. However, both of the methods (B-1) and (B-2) do not perform decoding using the MPU header information, so special processing may be required for decoding.

[0197] When the transmitting device transmits data in the order of media data, MPU metadata, and MF metadata as shown in Fig. 21(c), the receiving device can perform decoding by either of the following methods (C-1) and (C-2).

[0198] After the receiving device acquires the MPU header information (MPU metadata and MF metadata), it decrypts the media data.

[0199] The receiving device decrypts the media data without using the MPU header information.

[0200] When the method of (C-1) above is used, it is necessary to buffer the media data for acquiring the MPU header information. On the contrary, when the method of (C-2) above is used, it is not necessary to perform buffering for acquiring the MPU header information.

[0201] Also, in either of the methods of (C-1) and (C-2) above, no delay due to encapsulation occurs on the transmission side. Also, since the method of (C-2) does not use the MPU header information, special processing may be required.

[0202] When the transmitting device transmits data in the order of MPU metadata, media data, and MF metadata as shown in (d) of FIG. 21, the receiving device can perform decryption by either of the following methods (D-1) and (D-2).

[0203] The receiving device acquires the MPU metadata, then further acquires the MF metadata, and then decrypts the media data.

[0204] The receiving device acquires the MPU metadata and decrypts the media data without using the MF metadata.

[0205] When the method of (D-1) above is used, it is necessary to buffer the media data for acquiring the MF metadata, but in the case of the method of (D-2) above, it is not necessary to perform buffering for acquiring the MF metadata.

[0206] Since the method of (D-2) above does not perform decryption using the MF metadata, special processing may be required.

[0207] As described above, when it is possible to perform decoding using MPU metadata and MF metadata, there is an advantage that even a conventional MP4 receiving device can perform decoding.

[0208] In FIG. 21, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data. In moof, position information (offset) for each sample and subsample is determined based on this configuration. Also, the MF metadata includes data other than the media data in the mdat box (the size and type of the box).

[0209] Therefore, when the receiving device identifies media data based on MF metadata, the receiving device reconfigures the data in the order in which the MPU data was configured regardless of the order in which the data was transmitted, and then performs decoding using moov of the MPU metadata or moof of the MF metadata.

[0210] In FIG. 21, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data. However, the MPU data may be configured in an order different from that in FIG. 21, and position information (offset) may be determined.

[0211] For example, the MPU data may be configured in the order of MPU metadata, media data, and MF metadata, and negative position information (offset) may be indicated in the MF metadata. Also in this case, regardless of the order in which the data is transmitted, the receiving device reconfigures the data in the order in which the MPU data was configured on the transmission side, and then performs decoding using moov or moof.

[0212] Note that the transmitting device may signal information indicating the order in which the MPU data is configured, and the receiving device may reconfigure the data based on the signaled information.

[0213] As described above, as shown in Fig. 21(d), the receiving device receives the packetized MPU metadata, the packetized media data (sample data), and the packetized MF metadata in this order. Here, the MPU metadata is an example of the first metadata, and the MF metadata is an example of the second metadata.

[0214] Next, the receiving device reconstructs the MPU data (a file in MP4 format) including the received MPU metadata, the received MF metadata, and the received sample data. Then, the sample data included in the reconstructed MPU data is decoded using the MPU metadata and the MF metadata. The MF metadata is metadata including data (for example, length stored in mbox) that can be generated only after the generation of the sample data on the transmission side.

[0215] Note that the operation of the receiving device is performed in more detail by each component constituting the receiving device. For example, the receiving device includes a receiving unit that receives the above data, a reconstructing unit that reconstructs the above MPU data, and a decoding unit that decodes the above MPU data. Each of the receiving unit, the generating unit, and the decoding unit is realized by a microcomputer, a processor, a dedicated circuit, or the like.

[0216] [Method of Decoding Without Using Header Information] Next, a method of decoding without using header information will be described. Here, a method of decoding without using header information in the receiving device will be described regardless of whether the header information is sent or not on the transmission side. That is, this method is applicable regardless of which transmission method described with reference to Fig. 21 is used. However, some decoding methods are decoding methods applicable only to specific transmission methods.

[0217] FIG. 22 is a diagram showing an example of a method of performing decoding without using header information. In FIG. 22, only the MMT payload including only media data and the MMT packet are illustrated, and the MMT payload and the MMT packet including the MPU metadata and the MF metadata are not illustrated. Further, in the following description of FIG. 22, it is assumed that the media data belonging to the same MPU is transmitted continuously. Further, a case where samples are stored in the payload as media data will be described as an example, but in the following description of FIG. 22, naturally, NAL units may be stored, or fragmented NAL units may be stored.

[0218] In order to decode media data, the receiving device must first obtain initialization information necessary for decoding. Further, if the media is video, the receiving device must obtain initialization information for each sample, identify the start position of the MPU which is a random access unit, and obtain the start positions of the sample and the NAL unit. Further, the receiving device needs to identify the decoding time (DTS) and the presentation time (PTS) of each sample.

[0219] Therefore, the receiving device can perform decoding without using header information by using, for example, the following method. When a NAL unit unit or a unit obtained by fragmenting a NAL unit is stored in the payload, in the following description, "sample" may be read as "NAL unit in the sample".

[0220] <Random access (= identifying the first sample of the MPU)> When header information is not transmitted, there are the following Method 1 and Method 2 for the receiving device to identify the first sample of the MPU. When header information is transmitted, Method 3 can be used.

[0221] [Method 1] The receiving device obtains the samples included in the MMT packet in which 'RAP_flag = 1' in the MMT packet header.

[0222] [Method 2] The receiving device acquires a sample with'sample number = 0' in the MMT payload header.

[0223] [Method 3] When at least one of the MPU metadata and the MF metadata is transmitted before or after the media data, the receiving device acquires a sample included in the MMT payload in which the fragment type (FT) in the MMT payload header has switched to the media data.

[0224] Note that in Method 1 and Method 2, when multiple samples belonging to different MPUs are mixed in one payload, it is impossible to determine which NAL unit is a random access point (RAP_flag = 1 or sample number = 0). Therefore, restrictions such as not mixing samples of different MPUs in one payload, or when samples of different MPUs are mixed in one payload, setting the RAP_flag to 1 when the last (or first) sample is a random access point are required.

[0225] Also, in order for the receiving device to acquire the start position of the NAL unit, it is necessary to shift the data read pointer by the size of the NAL unit in order from the first NAL unit of the sample.

[0226] When the data is fragmented, the receiving device can identify the data unit by referring to the fragment_indicator and the fragment_number.

[0227] <Determination of the DTS of the Sample> There are the following Method 1 and Method 2 for determining the DTS of the sample.

[0228] [Method 1] The receiving device determines the DTS of the first sample based on the predicted structure. However, since this method requires analysis of the encoded data and may be difficult to decode in real time, the following Method 2 is desirable.

[0229] [Method 2] The receiving device separately transmits the DTS of the first sample and acquires the transmitted DTS of the first sample. Examples of the method for transmitting the DTS of the first sample include a method of transmitting the DTS of the MPU first sample using MMT-SI, a method of transmitting the DTS for each sample using the MMT packet header extension area, etc. Note that the DTS may be an absolute value or a relative value with respect to the PTS. Also, it may be signaled whether the DTS of the first sample is included on the transmitting side.

[0230] Note that in both Method 1 and Method 2, the DTS of the subsequent samples is calculated assuming a fixed frame rate.

[0231] As a method of storing the DTS for each sample in the packet header, in addition to using the extension area, there is a method of storing the DTS of the sample included in the MMT packet in the 32-bit NTP timestamp field in the MMT packet header. When the DTS cannot be represented by the number of bits (32 bits) of one packet header, the DTS may be represented using a plurality of packet headers. Also, the DTS may be represented by combining the NTP timestamp field of the packet header and the extension area. When the DTS information is not included, it is set to a known value (for example, ALL0).

[0232] <Determination of the PTS of the Sample> The receiving device acquires the PTS of the first sample from the MPU timestamp descriptor for each asset included in the MPU. For the subsequent sample PTS, the receiving device calculates it from parameters indicating the display order of the samples, such as POC, assuming a fixed frame rate. Thus, in order to calculate the DTS and PTS without using the header information, transmission at a fixed frame rate is essential.

[0233] Also, when MF metadata is being transmitted, the receiving device can calculate the absolute values of DTS and PTS from the relative time information of DTS and PTS from the start sample indicated in the MF metadata and the absolute value of the timestamp of the MPU start sample indicated in the MPU timestamp descriptor.

[0234] Note that when calculating DTS and PTS by analyzing the encoded data, the receiving device may calculate using the SEI information included in the access unit.

[0235] <Initialization information (parameter set)> [In the case of video] In the case of video, the parameter set is stored in the sample data. Also, when MPU metadata and MF metadata are not transmitted, it is guaranteed that the parameter set necessary for decoding can be obtained by referring only to the sample data.

[0236] Also, as shown in FIGS. 21(a) and (d), when MPU metadata is transmitted before the media data, it may be stipulated that the parameter set is not stored in the SampleEntry. In this case, the receiving device refers only to the parameter set within the sample without referring to the parameter set of the SampleEntry.

[0237] Also, when MPU metadata is transmitted before the media data, a parameter set common to the MPU or a default parameter set is stored in the SampleEntry, and the receiving device may refer to the parameter set of the SampleEntry and the parameter set within the sample. By storing the parameter set in the SampleEntry, it becomes possible to perform decoding even with a conventional receiving device that cannot play back if there is no parameter set in the SampleEntry.

[0238] [In the case of audio] In the case of audio, an LATM header is required for decoding, and in MP4, it is essential that the LATM header be included in the sample entry. However, if the header information is not transmitted, it is difficult for the receiving device to obtain the LATM header. Therefore, the LATM header is included in separate control information such as SI. Note that the LATM header may be included in a message, table, or descriptor. Note also that the LATM header may be included within a sample.

[0239] Before starting decoding, the receiving device obtains the LATM header from SI or the like and starts decoding the audio. Alternatively, as shown in FIGS. 21(a) and 21(d), when the MPU metadata is transmitted before the media data, the receiving device can receive the LATM header before the media data. Therefore, when the MPU metadata is transmitted before the media data, decoding can be performed even using a conventional receiving device.

[0240] <Others> The transmission order and the type of transmission order may be notified as control information such as an MMT packet header, a payload header, or an MPT or other table, message, or descriptor. Here, the type of transmission order refers to, for example, the four types of transmission orders shown in FIGS. 21(a) to 21(d), and an identifier for identifying each type may be stored at a location where it can be obtained before the start of decoding.

[0241] Also, different types of transmission orders may be used for audio and video, or the same type of transmission order may be used for both audio and video. Specifically, for example, as shown in FIG. 21(a), audio may be transmitted in the order of MPU metadata, MF metadata, and media data, and as shown in FIG. 21(d), video may be transmitted in the order of MPU metadata, media data, and MF metadata.

[0242] By the method described above, the receiving device can perform decoding without using the header information. Also, when the MPU metadata is transmitted before the media data (Figs. 21(a) and 21(d)), even a conventional receiving device can perform decoding.

[0243] In particular, since the MF metadata is transmitted after the media data (Fig. 21(d)), it is possible to avoid causing a delay due to encapsulation and to perform decoding even with a conventional receiving device.

[0244] [Configuration and Operation of Transmitting Device] Next, the configuration and operation of the transmitting device will be described. Fig. 23 is a block diagram of the transmitting device according to Embodiment 2, and Fig. 24 is a flowchart of the transmitting method according to Embodiment 2.

[0245] As shown in Fig. 23, the transmitting device 15 includes an encoding unit 16, a multiplexing unit 17, and a transmitting unit 18.

[0246] The encoding unit 16 generates encoded data by encoding video or audio to be encoded, for example, according to H.265 (S10).

[0247] The multiplexing unit 17 multiplexes (packetizes) the encoded data generated by the encoding unit 16 (S11). Specifically, the multiplexing unit 17 packetizes each of sample data, MPU metadata, and MF metadata that constitute a file in the MP4 format. The sample data is data obtained by encoding a video signal or an audio signal, the MPU metadata is an example of the first metadata, and the MF metadata is an example of the second metadata. Both the first metadata and the second metadata are metadata used for decoding the sample data, but the difference between them is that the second metadata includes data that can be generated only after the generation of the sample data.

[0248] Here, data that can be generated only after the generation of sample data is, for example, data other than the sample data stored in mdat in the MP4 format (data in the header of mdat. That is, type and length illustrated in FIG. 20). Here, the second metadata may include at least a part of this data, which is length.

[0249] The transmission unit 18 transmits the packetized MP4 format file (S12). The transmission unit 18 transmits the MP4 format file by, for example, the method shown in FIG. 21(d). That is, the packetized MPU metadata, the packetized sample data, and the packetized MF metadata are transmitted in this order.

[0250] Note that each of the encoding unit 16, the multiplexing unit 17, and the transmission unit 18 is realized by a microcomputer, a processor, a dedicated circuit, or the like.

[0251] [Configuration of Receiver] Next, the configuration and operation of the receiver will be described. FIG. 25 is a block diagram of the receiver according to Embodiment 2.

[0252] As shown in FIG. 25, the receiver 20 includes a packet filtering unit 21, a transmission order type determination unit 22, a random access unit 23, a control information acquisition unit 24, a data acquisition unit 25, a PTS and DTS calculation unit 26, an initialization information acquisition unit 27, a decoding command unit 28, a decoding unit 29, and a presentation unit 30.

[0253] [Operation 1 of Receiver] First, when the media is video, the operation of the receiver 20 for specifying the MPU start position and the NAL unit position will be described. FIG. 26 is a flowchart of such an operation of the receiver 20. Here, it is assumed that the transmission order type of the MPU data is stored in the SI information by the transmission device 15 (multiplexing unit 17).

[0254] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering to acquire the transmission order type of the MPU data (S21).

[0255] Next, the transmission order type determination unit 22 determines (discriminates) whether the data after packet filtering contains MPU header information (at least one of MPU metadata or MF metadata) (S22). When the MPU header information is included (Yes in S22), the random access unit 23 specifies the MPU start sample by detecting that the fragment type of the MMT payload header switches to media data (S23).

[0256] On the other hand, when the MPU header information is not included (No in S22), the random access unit 23 specifies the MPU start sample based on the RAP_flag of the MMT packet header or the sample number of the MMT payload header (S24).

[0257] Also, the transmission order type determination unit 22 determines whether the data after packet filtering contains MF metadata (S25). When it is determined that the MF metadata is included (Yes in S25), the data acquisition unit 25 acquires the NAL unit by reading the NAL unit based on the samples, sub-sample offsets, and size information included in the MF metadata (S26). On the other hand, when it is determined that the MF metadata is not included (No in S25), the data acquisition unit 25 acquires the NAL unit by reading data of the size of the NAL unit in order from the start NAL unit of the sample (S27).

[0258] Note that even when it is determined in step S22 that the MPU header information is included, the receiving device 20 may specify the MPU start sample using the process of step S24 instead of step S23. Also, when it is determined that the MPU header information is included, the processes of step S23 and step S24 may be used in combination.

[0259] Also, even when it is determined in step S25 that the MF metadata is included, the receiving device 20 may acquire the NAL unit using the process of step S27 without using the process of step S26. Also, when it is determined that the MF metadata is included, the processes of step S23 and step S24 may be used in combination.

[0260] Also, it is assumed that in the case where it is determined in step S25 that the MF metadata is included and the MF data is transmitted after the media data. In this case, the receiving device 20 may buffer the media data, wait until the MF metadata is acquired, and then perform the process of step S26, or the receiving device 20 may determine whether to perform the process of step S27 without waiting for the acquisition of the MF metadata.

[0261] For example, the receiving device 20 may determine whether to wait for the acquisition of the MF metadata based on whether it has a buffer with a buffer size capable of buffering the media data. Also, the receiving device 20 may determine whether to wait for the acquisition of the MF metadata based on whether the end-to-end delay becomes smaller. Also, the receiving device 20 may mainly perform the decoding process using the process of step S26, and may use the process of step S27 in the case of the processing mode when packet loss or the like occurs.

[0262] In addition, if the transmission order type is predefined, steps S22 and S26 may be omitted. In this case, the receiving device 20 may determine the method for specifying the MPU start sample and the method for specifying the NAL unit in consideration of the buffer size and the End-to-End delay.

[0263] In addition, if the transmission order type is known in advance, the transmission order type determination unit 22 in the receiving device 20 is unnecessary.

[0264] Also, although not described in FIG. 26 above, the decoding command unit 28 outputs the data acquired by the data acquisition unit to the decoding unit 29 based on the PTS and DTS calculated by the PTS, DTS calculation unit 26 and the initialization information acquired by the initialization information acquisition unit 27. The decoding unit 29 decodes the data, and the presentation unit 30 presents the decoded data.

[0265] [Operation of Receiving Device 2] Next, the operation in which the receiving device 20 acquires initialization information based on the transmission order type and decodes media data based on the initialization information will be described. FIG. 27 is a flowchart of such an operation.

[0266] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type determination unit 22 analyzes the SI information obtained by the packet filtering and acquires the transmission order type (S301).

[0267] Next, the transmission order type determination unit 22 determines whether MPU metadata has been transmitted (S302). If it is determined that the MPU metadata has been transmitted (Yes in S302), the transmission order type determination unit 22 determines whether the MPU metadata has been transmitted before the media data as a result of the analysis in step S301 (S303). If the MPU metadata has been transmitted before the media data (Yes in S303), the initialization information acquisition unit 27 decrypts the media data based on the common initialization information included in the MPU metadata and the initialization information of the sample data (S304).

[0268] On the other hand, if it is determined that the MPU metadata has been transmitted after the media data (No in S303), the data acquisition unit 25 buffers the media data until the MPU metadata is acquired (S305), and performs the process of step S304 after the MPU metadata is acquired.

[0269] Also, in step S302, if it is determined that the MPU metadata has not been transmitted (No in S302), the initialization information acquisition unit 27 decrypts the media data based only on the initialization information of the sample data (S306).

[0270] Note that if decryption of the media data is guaranteed only when based on the initialization information of the sample data on the transmission side, the processes based on the determinations in step S302 and step S303 are not performed, and the process of step S306 is used.

[0271] Also, before step S305, the receiving device 20 may determine whether to buffer the media data. In this case, if the receiving device 20 determines to buffer the media data, it proceeds to the process of step S305, and if it determines not to buffer the media data, it proceeds to the process of step S306. The determination of whether to buffer the media data may be made based on the buffer size and occupancy of the receiving device 20, or for example, the determination may be made considering the End-to-End delay, such as selecting the one with a smaller End-to-End delay.

[0272] [Operation 3 of the receiving device] Here, the details of the transmission method and reception method when the MF metadata is transmitted after the media data (in (c) and (d) of FIG. 21) will be described. Hereinafter, the case of (d) in FIG. 21 will be described as an example. In transmission, it is assumed that only the method of (d) in FIG. 21 is used and no signaling of the transmission order type is performed.

[0273] As described above, when transmitting data in the order of MPU metadata, media data, and MF metadata as shown in (d) of FIG. 21, (D-1) After the receiving device 20 acquires the MPU metadata, and further after acquiring the MF metadata, it decrypts the media data. (D-2) After the receiving device 20 acquires the MPU metadata, it decrypts the media data without using the MF metadata. These two decryption methods are possible.

[0274] Here, D-1 requires buffering of the media data for acquiring the MF metadata, but since decryption can be performed using the MPU header information, it can be decrypted by a conventional MP4-compliant receiving device. Also, D-2 does not require buffering of the media data for acquiring the MF metadata, but since it cannot be decrypted using the MF metadata, special processing is required for decryption.

[0275] Also, in the method of (d) in FIG. 21, since the MF metadata is transmitted after the media data, there is no delay due to encapsulation, and it has the advantage of being able to reduce the End-to-End delay.

[0276] The receiving device 20 can select the above two decoding methods according to the capabilities of the receiving device 20 and the quality of service provided by the receiving device 20.

[0277] The transmitting device 15 must ensure that it can guarantee decoding while reducing the occurrence of buffer overflow and underflow in the decoding operation of the receiving device 20. As elements for defining the decoder model when decoding using the D-1 method, for example, the following parameters can be used.

[0278] · Buffer size for reconstructing the MPU (MPU buffer) For example, the buffer size = maximum rate × maximum MPU time × α, where the maximum rate is the upper limit rate of the profile and level of the encoded data + the overhead of the MPU header. Also, the maximum MPU time is the maximum time length of the GOP when 1 MPU = 1 GOP (video).

[0279] Here, the audio may be in the same GOP unit as the above video or in another unit. α is a margin for not causing an overflow, and it may be multiplied or added to the maximum rate × maximum MPU time. When multiplied, α ≧ 1, and when added, α ≧ 0.

[0280] · Upper limit of the decoding delay time from when data is input to the MPU buffer until decoding (TSTD_delay in the STD of MPEG-TS) For example, at the time of transmission, considering the maximum MPU time and the upper limit value of the decoding delay time, the DTS is set so that the acquisition completion time of the MPU data at the receiver <= DTS.

[0281] Further, when decoding using the method of D-1, the transmission device 15 may assign DTS and PTS according to the decoder model. Thereby, while ensuring the decoding of the receiving device that decodes using the method of D-1, the transmission device 15 may also transmit auxiliary information necessary when decoding is performed using the method of D-2.

[0282] For example, the transmission device 15 can ensure the operation of the receiving device that decodes using the method of D-2 by signaling the pre-buffering time in the decoder buffer when decoding using the method of D-2.

[0283] The pre-buffering time may be included in SI control information such as messages, tables, descriptors, etc., or may be included in the headers of MMT packets and MMT payloads. Also, the SEI in the encoded data may be overwritten. The DTS and PTS for decoding using the method of D-1 are stored in the MPU time stamp descriptor and SampleEntry, and the DTS and PTS for decoding using the method of D-2, or the pre-buffering time, may be described in the SEI.

[0284] When the receiving device 20 only supports the decoding operation compliant with MP4 using the MPU header, the receiving device 20 may select the decoding method D-1, and when it supports both D-1 and D-2, either one may be selected.

[0285] The transmission device 15 may assign DTS and PTS so as to ensure the decoding operation of one (in this description, D-1), and may also transmit auxiliary information for assisting the decoding operation of the other.

[0286] Also, when the method of D-2 is used, compared with the case where the method of D-1 is used, the End-to-End delay is likely to increase due to the delay caused by the pre-buffering of the MF metadata. Therefore, when the receiving device 20 wants to reduce the End-to-End delay, it may select and decode using the method of D-2. For example, when the receiving device 20 always wants to reduce the End-to-End delay, it may always use the method of D-2. Also, the receiving device 20 may use the method of D-2 only when operating in a low-latency presentation mode where it wants to present live content, channel selection, zapping operations, etc. with low latency.

[0287] FIG. 28 is a flowchart of such a receiving method.

[0288] First, the receiving device 20 receives an MMT packet and acquires MPU data (S401). Then, the receiving device 20 (transmission order type determination unit 22) determines whether to present the program in the low-latency presentation mode (S402).

[0289] When the program is not presented in the low-latency presentation mode (No in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) performs random access using the header information and acquires initialization information (S405). Also, the receiving device 20 (PTS, DTS calculation unit 26, decoding command unit 28, decoding unit 29, presentation unit 30) performs decoding and presentation processing based on the PTS and DTS given on the transmission side (S406).

[0290] On the other hand, when the program is presented in the low-latency presentation mode (Yes in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) performs random access and acquires initialization information using a decoding method that does not use the header information (S403). Also, the receiving device 20 performs decoding and presentation processing based on the auxiliary information for decoding without using the PTS, DTS, and header information given on the transmission side (S404). Note that in steps S403 and S404, processing may be performed using the MPU metadata.

[0291] [Transmission and Reception Method Using Auxiliary Data] As described above, the transmission and reception operations in the case where MF metadata is transmitted after media data (the cases of (c) and (d) in FIG. 21) have been explained. Next, a method will be described in which the transmission device 15 transmits auxiliary data having some functions of MF metadata, enabling earlier start of decoding and reduction of the End-to-End delay. Here, an example in which auxiliary data is further transmitted based on the transmission method shown in (d) of FIG. 21 will be described, but the method using auxiliary data is also applicable to the transmission methods shown in (a) to (c) of FIG. 21.

[0292] FIG. 29(a) is a diagram showing an MMT packet transmitted using the method shown in (d) of FIG. 21. That is, the data is transmitted in the order of MPU metadata, media data, and MF metadata.

[0293] Here, Sample #1, Sample #2, Sample #3, and Sample #4 are samples included in the media data. Here, an example in which the media data is stored in MMT packets in units of samples is described, but the media data may be stored in MMT packets in units of NAL units, or may be stored in units obtained by dividing NAL units. Note that there may be a case where a plurality of NAL units are aggregated and stored in an MMT packet.

[0294] As described in D-1 above, in the case of the method shown in (d) of FIG. 21, that is, when the data is transmitted in the order of MPU metadata, media data, and MF metadata, after obtaining the MPU metadata, the MF metadata is further obtained, and then there is a method of decoding the media data. In such a D-1 method, buffering of the media data for obtaining the MF metadata is required, but since decoding is performed using the MPU header information, the D-1 method has the advantage of being applicable to conventional MP4-compliant receiving devices. On the other hand, the receiving device 20 has the drawback that it must wait to start decoding until the MF metadata is obtained.

[0295] On the other hand, as shown in FIG. 29(b), in the method using auxiliary data, the auxiliary data is transmitted before the MF metadata.

[0296] The MF metadata includes information indicating the DTS, PTS, offset, and size of all samples included in the movie fragment. On the other hand, the auxiliary data includes information indicating the DTS, PTS, offset, and size of some of the samples included in the movie fragment.

[0297] For example, the MF metadata includes information of all samples (sample #1 - sample #4), while the auxiliary data includes information of some samples (sample #1 - #2).

[0298] In the case shown in FIG. 29(b), since the use of the auxiliary data enables the decoding of sample #1 and sample #2, the End-to-End delay is reduced compared to the transmission method of D-1. Note that the auxiliary data may include the sample information combined in any way, and the auxiliary data may be transmitted repeatedly.

[0299] For example, in FIG. 29(c), when transmitting the auxiliary information at the timing of A, the transmitting device 15 includes the information of sample #1 in the auxiliary information, and when transmitting the auxiliary information at the timing of B, the transmitting device 15 includes the information of sample #1 and sample #2 in the auxiliary information. When the transmitting device 15 transmits the auxiliary information at the timing of C, the auxiliary information includes the information of sample #1, sample #2, and sample #3.

[0300] Note that the MF metadata includes the information of sample #1, sample #2, sample #3, and sample #4 (the information of all samples in the movie fragment).

[0301] The auxiliary data does not necessarily need to be transmitted immediately after generation.

[0302] In the headers of MMT packets and MMT payloads, a type indicating that auxiliary data is stored is specified.

[0303] For example, when the auxiliary data is stored in the MMT payload using the MPU mode, a data type indicating that it is auxiliary data is specified as the fragment_type field value (e.g., FT = 3). The auxiliary data may be data based on the moof configuration or other configurations.

[0304] When the auxiliary data is stored in the MMT payload as control signals (descriptors, tables, messages), descriptor tags, table IDs, message IDs, etc. indicating that it is auxiliary data are specified.

[0305] Also, PTS or DTS may be stored in the headers of MMT packets and MMT payloads.

[0306] [Example of Generating Auxiliary Data] Hereinafter, an example in which a transmission device generates auxiliary data based on the moof configuration will be described. FIG. 30 is a diagram for explaining an example in which a transmission device generates auxiliary data based on the moof configuration.

[0307] In a normal MP4, as shown in FIG. 20, a moof is created for a movie fragment. The moof contains information indicating the DTS, PTS, offset, and size of the samples included in the movie fragment.

[0308] Here, the transmission device 15 configures an MP4 (MP4 file) using only some of the sample data among the sample data constituting the MPU and generates auxiliary data.

[0309] For example, as shown in Fig. 30(a), the transmitting device 15 generates an MP4 using only sample #1 out of samples #1 - #4 that make up the MPU, and among them, the moof + mdat header is used as auxiliary data.

[0310] Next, as shown in Fig. 30(b), the transmitting device 15 generates an MP4 using samples #1 and #2 out of samples #1 - #4 that make up the MPU, and among them, the moof + mdat header is used as the next auxiliary data.

[0311] Next, as shown in Fig. 30(c), the transmitting device 15 generates an MP4 using samples #1, #2, and #3 out of samples #1 - #4 that make up the MPU, and among them, the moof + mdat header is used as the next auxiliary data.

[0312] Next, as shown in Fig. 30(d), the transmitting device 15 generates all MP4s out of samples #1 - #4 that make up the MPU, and among them, the moof + mdat header becomes movie fragment metadata.

[0313] Here, although the transmitting device 15 generates auxiliary data for each sample, it may also generate auxiliary data for every N samples. The value of N is an arbitrary number. For example, when transmitting auxiliary data M times when transmitting one MPU, N = total samples / M may be used.

[0314] Note that the information indicating the offset of the sample in moof may also be the offset value after ensuring that the sample entry area for the subsequent number of samples is a NULL area.

[0315] Note that the auxiliary data may be generated so as to be configured to fragment the MF metadata.

[0316] [Example of reception operation using auxiliary data] The reception of the auxiliary data generated as described with reference to FIG. 30 will be described. FIG. 31 is a diagram for explaining the reception of the auxiliary data. In FIG. 31(a), it is assumed that the number of samples constituting the MPU is 30, and the auxiliary data is generated and transmitted every 10 samples.

[0317] In FIG. 30(a), the auxiliary data #1 includes the samples #1 - #10, the auxiliary data #2 includes the samples #1 - #20, and the MF metadata includes the sample information of the samples #1 - #30.

[0318] Although the samples #1 - #10, the samples #11 - #20, and the samples #21 - #30 are stored in one MMT payload, they may be stored in units of samples or NAL units, or may be stored in units of fragments or aggregations.

[0319] The receiving device 20 receives the packets of the MPU meta, samples, MF meta, and auxiliary data respectively.

[0320] The receiving device 20 concatenates the sample data in the order of reception (to the rear), and after receiving the latest auxiliary data, updates the auxiliary data so far. Further, the receiving device 20 can configure a complete MPU by finally replacing the auxiliary data with the MF metadata.

[0321] When the receiving device 20 receives the auxiliary data #1, it concatenates the data as shown in the upper part of FIG. 31(b) to configure an MP4. Thereby, the receiving device 20 can parse the samples #1 - #10 using the MPU metadata and the information of the auxiliary data #1, and can perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0322] Also, when the receiving device 20 receives the auxiliary data #2, it concatenates the data as shown in the middle row of Fig. 31(b) to form an MP4. Thereby, the receiving device 20 can parse samples #1 - #20 using the MPU metadata and the information of the auxiliary data #2, and can perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.

[0323] Also, when the receiving device 20 receives the MF metadata, it concatenates the data as shown in the lower row of Fig. 31(b) to form an MP4. Thereby, the receiving device 20 can parse samples #1 - #30 using the MPU metadata and the MF metadata, and can perform decoding based on the PTS, DTS, offset, and size information included in the MF metadata.

[0324] In the case where there is no auxiliary data, since the receiving device 20 can obtain sample information only after receiving the MF metadata, it was necessary to start decoding after receiving the MF metadata. However, by the transmitting device 15 generating and transmitting the auxiliary data, the receiving device 20 can obtain sample information using the auxiliary data without waiting for the reception of the MF metadata, so the decoding start time can be advanced. Furthermore, by the transmitting device 15 generating the auxiliary data based on moof described with reference to Fig. 30, the receiving device 20 can use the conventional MP4 parser as it is and perform parsing.

[0325] Also, the newly generated auxiliary data and MF metadata include sample information that overlaps with the auxiliary data transmitted in the past. Therefore, even when the past auxiliary data could not be obtained due to packet loss or the like, by using the newly obtained auxiliary data and MF metadata, it is possible to reconstruct the MP4 and obtain the sample information (PTS, DTS, size, and offset).

[0326] Note that the auxiliary data does not necessarily have to include information on past sample data. For example, auxiliary data #1 may correspond to sample data #1 - #10, and auxiliary data #2 may correspond to sample data #11 - #20. For example, as shown in (c) of FIG. 31, the transmission device 15 may sequentially transmit complete MF metadata as data units and units obtained by fragmenting the data units as auxiliary data.

[0327] Also, the transmission device 15 may repeatedly transmit the auxiliary data or repeatedly transmit the MF metadata for packet loss countermeasures.

[0328] Note that the MMT packets and MMT payloads in which the auxiliary data is stored include an MPU sequence number and an asset ID, similar to the MPU metadata, the MF metadata, and the sample data.

[0329] The reception operation using the auxiliary data as described above will be described with reference to the flowchart of FIG. 32. FIG. 32 is a flowchart of the reception operation using the auxiliary data.

[0330] First, the receiving device 20 receives an MMT packet and analyzes the packet header and the payload header (S501). Next, the receiving device 20 analyzes whether the fragment type is auxiliary data or MF metadata (S502). If the fragment type is auxiliary data, the receiving device 20 overwrites and updates the past auxiliary data (S503). At this time, if there is no past auxiliary data for the same MPU, the receiving device 20 uses the received auxiliary data as new auxiliary data as it is. Then, the receiving device 20 acquires samples and performs decoding based on the MPU metadata, the auxiliary data, and the sample data (S507).

[0331] On the other hand, when the fragment type is MF metadata, in step S505, the receiving device 20 overwrites the past auxiliary data with the MF metadata (S505). Then, the receiving device 20 obtains samples in the form of a complete MPU based on the MPU metadata, MF metadata, and sample data, and performs decoding (S506).

[0332] Although not shown in FIG. 32, in step S502, when the fragment type is MPU metadata, the receiving device 20 stores the data in a buffer, and when it is sample data, it stores the data concatenated at the back for each sample in the buffer.

[0333] If the auxiliary data cannot be obtained due to packet loss, the receiving device 20 can overwrite it with the latest auxiliary data or decode the samples using the past auxiliary data.

[0334] Note that the transmission period and the number of transmissions of the auxiliary data may be predetermined values. Information on the transmission period and the number of transmissions (count, countdown) may be transmitted together with the data. For example, a time stamp such as a transmission period, the number of transmissions, and initial_cpb_removal_delay may be stored in the data unit header.

[0335] By transmitting the auxiliary data including the information of the first sample of the MPU one or more times before initial_cpb_removal_delay, it becomes possible to comply with the CPB buffer model. At this time, a value based on the picture timing SEI is stored in the MPU time stamp descriptor.

[0336] Note that the transmission method in the receiving operation in which such auxiliary data is used is not limited to the MMT method, and is applicable to cases such as streaming transmission of packets configured in the ISOBMFF file format such as MPEG-DASH.

[0337] [Transmission method when one MPU is composed of multiple movie fragments] In the description after FIG. 19 above, one MPU was composed of one movie fragment. Here, however, the case where one MPU is composed of multiple movie fragments will be described. FIG. 33 is a diagram showing the configuration of an MPU composed of multiple movie fragments.

[0338] In FIG. 33, the samples (#1 - #6) stored in one MPU are divided and stored in two movie fragments. The first movie fragment is generated based on samples #1 - #3, and a corresponding moof box is generated. The second movie fragment is generated based on samples #4 - #6, and a corresponding moof box is generated.

[0339] The headers of the moof box and mdat box in the first movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #1. On the other hand, the headers of the moof box and mdat box in the second movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #2. In FIG. 33, the MMT payload storing the movie fragment metadata is hatched.

[0340] Note that the number of samples constituting the MPU and the number of samples constituting the movie fragment are arbitrary. For example, the number of samples constituting the MPU may be set as the number of samples per GOP unit, and two movie fragments may be formed with the number of samples being half of the number of samples per GOP unit.

[0341] Here, an example where one MPU contains two movie fragments (moof box and mdat box) is shown. However, the number of movie fragments contained in one MPU may not be two, but may be three or more. Also, the samples stored in the movie fragment may not be divided into equal numbers of samples, but may be divided into arbitrary numbers of samples.

[0342] Note that in FIG. 33, the MPU metadata unit and the MF metadata unit are each stored in the MMT payload as data units. However, the transmission device 15 may store units such as ftyp, mmpu, moov, and moof in the MMT payload in data unit units, or may store the data units in the MMT payload in units obtained by fragmenting the data units. Further, the transmission device 15 may store the data units in the MMT payload in units obtained by aggregating the data units.

[0343] Also, in FIG. 33, samples are stored in the MMT payload in sample units. However, the transmission device 15 may configure data units in NAL unit units or units obtained by grouping a plurality of NAL units instead of sample units, and store the data units in the MMT payload in data unit units. Further, the transmission device 15 may store the data units in the MMT payload in units obtained by fragmenting the data units, or may store the data units in the MMT payload in units obtained by aggregating the data units.

[0344] Note that in FIG. 33, the MPU is configured in the order of moof#1, mdat#1, moof#2, mdat#2, and an offset is given to moof#1 assuming that the corresponding mdat#1 follows. However, an offset may be given assuming that mdat#1 precedes moof#1. However, in this case, movie fragment metadata cannot be generated in the form of moof+mdat, and the headers of moof and mdat are transmitted separately.

[0345] Next, the transmission order of MMT packets when the MPU having the configuration described with reference to FIG. 33 is transmitted will be described. FIG. 34 is a diagram for explaining the transmission order of MMT packets.

[0346] Fig. 34(a) shows the transmission order when transmitting the MMT packet in the configuration order of the MPU shown in Fig. 33. Specifically, Fig. 34(a) shows an example of transmitting in the order of MPU meta, MF meta #1, media data #1 (samples #1 - #3), MF meta #2, and media data #2 (samples #4 - #6).

[0347] Fig. 34(b) shows an example of transmitting in the order of MPU meta, media data #1 (samples #1 - #3), MF meta #1, media data #2 (samples #4 - #6), and MF meta #2.

[0348] Fig. 34(c) shows an example of transmitting in the order of media data #1 (samples #1 - #3), MPU meta, MF meta #1, media data #2 (samples #4 - #6), and MF meta #2.

[0349] MF meta #1 is generated using samples #1 - #3, and MF meta #2 is generated using samples #4 - #6. Therefore, when the transmission method in Fig. 34(a) is used, a delay due to encapsulation occurs in the transmission of sample data.

[0350] On the other hand, when the transmission methods in Fig. 34(b) and Fig. 34(c) are used, samples can be transmitted without waiting to generate the MF meta. Therefore, no delay due to encapsulation occurs, and the End-to-End delay can be reduced.

[0351] Also, even in the transmission order of Fig. 34(a), since one MPU is divided into multiple movie fragments and the number of samples stored in the MF meta is smaller than that in the case of Fig. 19, the amount of delay due to encapsulation can be made smaller than in the case of Fig. 19.

[0352] In addition to the method shown here, for example, the transmission device 15 may concatenate MF meta #1 and MF meta #2 and transmit them together at the end of the MPU. In this case, the MF metas of different movie fragments may be aggregated and stored in one MMT payload. Also, the MF metas of different MPUs may be aggregated together and stored in the MMT payload.

[0353] [Receiving Method When One MPU Is Composed of Multiple Movie Fragments] Here, an operation example of the receiving device 20 that receives and decodes the MMT packet transmitted in the transmission order described in FIG. 34(b) will be described. FIGS. 35 and 36 are diagrams for explaining such an operation example.

[0354] The receiving device 20 receives each MMT packet including MPU meta, sample, and MF meta transmitted in the transmission order as shown in FIG. 35. The sample data is concatenated in the order of reception.

[0355] At time T1 when the receiving device 20 receives MF meta #1, the data is concatenated as shown in FIG. 36(1) to form an MP4. Thereby, the receiving device 20 can acquire samples #1-#3 based on the MPU meta data and the information of MF meta #1, and can perform decoding based on the PTS, DTS, offset, and size information included in the MF meta.

[0356] Also, at time T2 when the receiving device 20 receives MF meta #2, the data is concatenated as shown in FIG. 36(2) to form an MP4. Thereby, the receiving device 20 can acquire samples #4-#6 based on the MPU meta data and the information of MF meta #2, and can perform decoding based on the PTS, DTS, offset, and size information of the MF meta. Also, the receiving device 20 may concatenate the data as shown in FIG. 36(3) to form an MP4, and acquire samples #1-#6 based on the information of MF meta #1 and MF meta #2.

[0357] By dividing one MPU into multiple movie fragments, the time until the first MF meta in the MPU is obtained is shortened, so that the decoding start time can be advanced. Also, the buffer size for accumulating samples before decoding can be reduced.

[0358] Note that the transmission device 15 may set the division unit of the movie fragment so that the time from transmitting (or receiving) the first sample in the movie fragment to transmitting (or receiving) the MF meta corresponding to the movie fragment is shorter than the initial_cpb_removal_delay specified by the encoder. By setting it in this way, the reception buffer can follow the cpb buffer, and low-latency decoding can be realized. In this case, absolute time based on the initial_cpb_removal_delay can be used for PTS and DTS.

[0359] Also, the transmission device 15 may divide the movie fragment at equal intervals, or divide subsequent movie fragments at intervals shorter than the previous movie fragment. Thereby, the reception device 20 can always receive the MF meta including the information of the sample before decoding the sample, and continuous decoding becomes possible.

[0360] The following two methods can be used to calculate the absolute time of PTS and DTS.

[0361] (1) The absolute time of PTS and DTS is determined based on the reception time (T1 or T2) of MF meta #1 or MF meta #2, and the relative time of PTS and DTS included in the MF meta.

[0362] (2) The absolute time of PTS and DTS is determined based on the absolute time signaled from the transmission side, such as the MPU timestamp descriptor, and the relative time of PTS and DTS included in the MF meta.

[0363] Also, the absolute time signaled by the transmission device 15 may be the absolute time calculated based on the initial_cpb_removal_delay specified by the encoder.

[0364] Also, the absolute time signaled by the transmission device 15 may be the absolute time calculated based on the predicted value of the reception time of the MF meta.

[0365] Note that MF meta #1 and MF meta #2 may be repeatedly transmitted. By repeatedly transmitting MF meta #1 and MF meta #2, the receiving device 20 can acquire them again even if it fails to acquire the MF meta due to packet loss or the like.

[0366] An identifier indicating the order of the movie fragments can be stored in the payload header of the MFU including the samples constituting the movie fragments. On the other hand, the identifier indicating the order of the MF meta constituting the movie fragments is not included in the MMT payload. Therefore, the receiving device 20 identifies the order of the MF meta using the packet_sequence_number. Alternatively, the transmission device 15 may store and signal an identifier indicating which movie fragment the MF meta belongs to in the control information (message, table, descriptor), MMT header, MMT payload header, or data unit header.

[0367] Note that the transmission device 15 may transmit the MPU meta, MF meta, and samples in a predetermined transmission order determined in advance, and the receiving device 20 may perform reception processing based on the predetermined transmission order determined in advance. Also, the transmission device 15 may signal the transmission order, and the receiving device 20 may select (judge) the reception processing based on the signaling information.

[0368] The above-described reception method will be described with reference to FIG. 37. FIG. 37 is a flowchart of the operation of the reception method described in FIGS. 35 and 36.

[0369] First, the receiving device 20 determines (identifies) whether the data included in the payload is MPU metadata, MF metadata, or sample data (MFU) based on the fragment type indicated in the MMT payload (S601, S602). If the data is sample data, the receiving device 20 buffers the sample and waits for the reception of the MF metadata corresponding to the sample and the start of decoding (S603).

[0370] On the other hand, in step S602, if the data is MF metadata, the receiving device 20 acquires sample information (PTS, DTS, position information, and size) from the MF metadata, acquires samples based on the acquired sample information, and decodes and presents the samples based on PTS and DTS (S604).

[0371] Although not shown in the figure, when the data is MPU metadata, the MPU metadata includes initialization information necessary for decoding. Therefore, the receiving device 20 accumulates this and uses it for decoding the sample data in step S604.

[0372] When the receiving device 20 accumulates the received MPU data (MPU metadata, MF metadata, and sample data) in the storage device, it accumulates the data after rearranging it in the MPU configuration as described in FIG. 19 or FIG. 33.

[0373] On the transmission side, a packet sequence number is assigned to packets having the same packet ID in the MMT packet. At this time, the packet sequence number may be assigned after the MMT packets including MPU metadata, MF metadata, and sample data are rearranged in the transmission order, or may be assigned in the order before rearrangement.

[0374] When the packet sequence number is assigned in the order before rearrangement, in the receiving device 20, the data can be rearranged in the MPU configuration order based on the packet sequence number, facilitating accumulation.

[0375] [Method for Detecting the Start of an Access Unit and the Start of a Slice Segment] A method for detecting the start of an access unit and the start of a slice segment based on the information in the MMT packet header and the MMT payload header will be described.

[0376] Here, two examples are shown: the case where non-VCL NAL units (such as access unit delimiters, VPS, SPS, PPS, and SEI) are collectively stored in the MMT payload as data units, and the case where non-VCL NAL units are each treated as data units and the data units are aggregated and stored in one MMT payload.

[0377] FIG. 38 is a diagram showing the case where non-VCL NAL units are individually treated as data units and aggregated.

[0378] In the case of FIG. 38, the start of the access unit is the start data of the MMT payload including a data unit where the fragment_type value is MFU in the MMT packet, the aggregation_flag value is 1, and the offset value is 0. At this time, the Fragmentation_indicator value is 0.

[0379] Also, in the case of FIG. 38, the start of the slice segment is the start data of the MMT payload where the fragment_type value is MFU in the MMT packet, and the aggregation_flag value is 0, and the fragmentation_indicator value is 00 or 01.

[0380] FIG. 39 is a diagram showing the case where non-VCL NAL units are collectively treated as data units. Note that the field values of the packet header are as shown in FIG. 17 (or FIG. 18).

[0381] In the case of FIG. 39, the start of the access unit is the start data of the payload in the packet with an Offset value of 0, which becomes the start of the access unit.

[0382] Also, in the case of FIG. 39, the start of the slice segment is the start data of the payload of the packet with an Offset value different from 0 and a fragmentation indicator value of 00 or 01, which becomes the start of the slice segment.

[0383] [Receiving Process in Case of Packet Loss] Normally, when transmitting MP4 - formatted data in an environment where packet loss occurs, the receiving device 20 restores the packets by means of ALFEC (Application Layer FEC), packet re - transmission control, etc.

[0384] However, when packet loss occurs in streaming such as broadcasting where AL - FEC is not used, the packets cannot be restored.

[0385] After data is lost due to packet loss, the receiving device 20 needs to resume decoding video and audio again. For this purpose, the receiving device 20 needs to detect the start of the access unit or NAL unit and start decoding from the start of the access unit or NAL unit.

[0386] However, since there is no start code at the start of the MP4 - formatted NAL unit, the receiving device 20 cannot detect the start of the access unit or NAL unit even by analyzing the stream.

[0387] FIG. 40 is a flowchart of the operation of the receiving device 20 when packet loss occurs.

[0388] The receiving device 20 detects packet loss based on the Packet sequence number, packet counter, fragment counter, etc. in the header of the MMT packet or MMT payload (S701), and determines which packet has disappeared from the context before and after (S702).

[0389] When it is determined that no packet loss has occurred (No in S702), the receiving device 20 constructs an MP4 file and decodes the access unit or NAL unit (S703).

[0390] When it is determined that packet loss has occurred (Yes in S702), the receiving device 20 generates a NAL unit corresponding to the NAL unit with packet loss using dummy data and constructs an MP4 file. When the receiving device 20 inserts dummy data into the NAL unit, it indicates that the NAL unit is dummy data in terms of the NAL unit type.

[0391] Also, the receiving device 20 can resume decoding by detecting the start of the next access unit or NAL unit based on the methods described in FIGS. 17, 18, 38, and 39, and inputting the start data to the decoder from the start data (S705).

[0392] When packet loss occurs, the receiving device 20 may resume decoding from the start of the access unit and NAL unit based on the information detected based on the packet header, or may resume decoding from the start of the access unit and NAL unit based on the header information of the reconstructed MP4 file including the NAL unit of dummy data.

[0393] When accumulating the MP4 file (MPU), the receiving device 20 may separately acquire and accumulate (replace) the packet data (such as NAL units) that have disappeared due to packet loss from broadcasting or communication.

[0394] At this time, when the receiving device 20 acquires a lost packet from communication, it notifies the server of information on the lost packet (such as the packet ID, MPU sequence number, packet sequence number, IP data flow number, and IP address), and acquires the packet. The receiving device 20 may acquire not only the lost packet but also a packet group before and after the lost packet at the same time.

[0395] [Method for Configuring Movie Fragments] Here, the method for configuring movie fragments will be described in detail.

[0396] As described with reference to FIG. 33, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are arbitrary. For example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU may be a fixed predetermined number or may be determined dynamically.

[0397] Here, by configuring movie fragments so that the following conditions are satisfied on the transmission side (transmission device 15), low-latency decoding in the receiving device 20 can be guaranteed.

[0398] The conditions are as follows.

[0399] The transmission device 15 generates and transmits MF metadata with a unit obtained by dividing sample data as a movie fragment so that the receiving device 20 can surely receive MF metadata including information on the sample before the decoding time (DTS(i)) of any sample (Sample(i)).

[0400] Specifically, the transmission device 15 configures a movie fragment using the encoded sample (including the i-th sample) before DTS(i).

[0401] As a method for dynamically determining the number of samples that make up a movie fragment and the number of movie fragments that make up one MPU so as to ensure low-latency decoding, for example, the following method is used.

[0402] (1) At the start of decoding, the decoding time DTS(0) of the sample Sample(0) at the head of the GOP is a time based on the initial_cpb_removal_delay. The transmitting device configures the first movie fragment using the samples that have been encoded and completed at a time before DTS(0). Also, the transmitting device 15 generates MF metadata corresponding to the first movie fragment and transmits it at a time before DTS(0).

[0403] (2) The transmitting device 15 also configures movie fragments so as to satisfy the above conditions for subsequent samples.

[0404] For example, when the sample at the head of the movie fragment is the k-th sample, the MF meta of the movie fragment including the k-th sample is transmitted by the time of the decoding time DTS(k) of the k-th sample. When the encoding completion time of the l-th sample is before DTS(k) and the encoding completion time of the (l + 1)-th sample is after DTS(k), the transmitting device 15 configures a movie fragment using the k-th sample to the l-th sample.

[0405] Note that the transmitting device 15 may configure a movie fragment using the k-th sample to less than the l-th sample.

[0406] (3) After the encoding of the last sample of the MPU is completed, the transmitting device 15 configures a movie fragment using the remaining samples, generates MF metadata corresponding to the movie fragment, and transmits it.

[0407] Note that the transmission device 15 may configure a movie fragment using some of the samples that have been encoded, instead of using all the samples that have been encoded to configure the movie fragment.

[0408] Note that in the above, an example was shown in which the number of samples that configure a movie fragment and the number of movie fragments that configure one MPU are dynamically determined based on the above conditions so as to guarantee low-latency decoding. However, the method for determining the number of samples and the number of movie fragments is not limited to such a method. For example, the number of movie fragments that configure one MPU may be fixed to a predetermined value, and the number of samples may be determined so as to satisfy the above conditions. Also, the number of movie fragments that configure one MPU and the time at which the movie fragment is divided (or the amount of code of the movie fragment) may be fixed to predetermined values, and the number of samples may be determined so as to satisfy the above conditions.

[0409] Also, when the MPU is divided into a plurality of movie fragments, information indicating whether the MPU is divided into a plurality of movie fragments, the attributes of the divided movie fragments, or the attributes of the MF meta for the divided movie fragments may be transmitted.

[0410] Here, the attributes of the movie fragment are information indicating whether the movie fragment is the first movie fragment of the MPU, the last movie fragment of the MPU, or another movie fragment, etc.

[0411] Also, the attributes of the MF meta are information indicating whether the MF meta is the MF meta corresponding to the first movie fragment of the MPU, the MF meta corresponding to the last movie fragment of the MPU, or the MF meta corresponding to another movie fragment, etc.

[0412] Note that the transmission device 15 may store and transmit, as control information, the number of samples constituting the movie fragment and the number of movie fragments constituting one MPU.

[0413] [Operation of the receiving device] The operation of the receiving device 20 based on the movie fragment configured as described above will be described.

[0414] The receiving device 20 determines the absolute times of each of the PTS and DTS based on the absolute times signaled from the transmission side, such as the MPU timestamp descriptor, and the relative times of the PTS and DTS included in the MF meta.

[0415] The receiving device 20 performs the following processing based on the information on whether the MPU is divided into a plurality of movie fragments. When the MPU is divided, the processing is based on the attributes of the divided movie fragments.

[0416] (1) When the movie fragment is the first movie fragment of the MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the absolute time of the PTS of the first sample included in the MPU timestamp descriptor and the relative times of the PTS and DTS included in the MF meta.

[0417] (2) When the movie fragment is not the first movie fragment of the MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the relative times of the PTS and DTS included in the MF meta without using the information of the MPU timestamp descriptor.

[0418] (3) When the movie fragment is the last movie fragment of the MPU, after calculating the absolute times of the PTS and DTS of all samples, the receiving device 20 resets the calculation process (relative time addition process) of the PTS and DTS. Note that the reset process may be performed in the first movie fragment of the MPU.

[0419] The receiving device 20 may determine whether the movie fragment is split as follows. Also, the receiving device 20 may acquire the attribute information of the movie fragment as follows.

[0420] For example, the receiving device 20 may determine whether it is split based on the identifier movie_fragment_sequence_number field value indicating the order of the movie fragments shown in the MMTP (MMT Protocol) payload header.

[0421] Specifically, when the number of movie fragments included in one MPU is 1, and the movie_fragment_sequence_number field value is 1, and there are values of 2 or more for the said field value, the receiving device 20 may determine that the said MPU is split into a plurality of movie fragments.

[0422] Also, when the number of movie fragments included in one MPU is 1, and the movie_fragment_sequence_number field value is 0, and there are values other than 0 for the said field value, the receiving device 20 may determine that the said MPU is split into a plurality of movie fragments.

[0423] Similarly, the attribute information of the movie fragment may be determined based on movie_fragment_sequence_number.

[0424] Note that even without using movie_fragment_sequence_number, by counting the transmission of movie fragments and MF meta included in one MPU, it may be determined whether the movie fragment is split and the attribute information of the movie fragment.

[0425] With the configurations of the transmission device 15 and the reception device 20 as described above, the reception device 20 can receive movie fragment metadata at intervals shorter than those of the MPU, enabling decoding start with low latency. Also, it is possible to perform decoding with low latency by using a decoding process based on the method of MP4 parsing.

[0426] The reception operation in the case where the MPU is divided into a plurality of movie fragments as described above will be described using a flowchart. FIG. 41 is a flowchart of the reception operation in the case where the MPU is divided into a plurality of movie fragments. Note that this flowchart illustrates the operation of step S604 in FIG. 37 in more detail.

[0427] First, based on the data type indicated in the MMTP payload header, when the data type is MF meta, the reception device 20 acquires MF metadata (S801).

[0428] Next, the reception device 20 determines whether the MPU is divided into a plurality of movie fragments (S802). If the MPU is divided into a plurality of movie fragments (Yes in S802), it determines whether the received MF metadata is the metadata at the head of the MPU (S803). When the received MF metadata is the MF metadata at the head of the MPU (Yes in S803), the reception device 20 calculates the absolute times of PTS and DTS from the absolute time of PTS indicated in the MPU timestamp descriptor and the relative times of PTS and DTS indicated in the MF metadata (S804), and determines whether it is the last metadata of the MPU (S805).

[0429] On the other hand, when the received MF metadata is not the MF metadata at the head of the MPU (No in S803), the reception device 20 calculates the absolute times of PTS and DTS using the relative times of PTS and DTS indicated in the MF metadata without using the information in the MPU timestamp descriptor (S808), and proceeds to the process of step S805.

[0430] In step S805, if it is determined that the MPU is the last MF metadata (Yes in S805), the receiving device 20 calculates the absolute times of the PTS and DTS of all samples and then resets the calculation process of the PTS and DTS. If it is determined in step S805 that the MPU is not the last MF metadata (No in S805), the receiving device 20 ends the process.

[0431] Also, in step S802, if it is determined that the MPU is not divided into a plurality of movie fragments (No in S802), the receiving device 20 acquires sample data and determines the PTS and DTS based on the MF metadata transmitted after the MPU (S807).

[0432] Then, although not shown in the figure, the receiving device 20 finally performs decoding processing and presentation processing based on the determined PTS and DTS.

[0433] [Problems that occur when splitting movie fragments and solutions thereto] So far, a method for shortening the End-to-End delay by splitting movie fragments has been described. From here, problems newly occurring when splitting movie fragments and solutions thereto will be described.

[0434] First, as background, the picture structure in the encoded data will be described. FIG. 42 is a diagram showing an example of the prediction structure of pictures in each TemporalId when realizing temporal scalability.

[0435] In encoding methods such as MPEG-4 AVC and HEVC (High Efficiency Video Coding), temporal scalability (temporal scalability) can be realized by using B pictures (bidirectional reference prediction pictures) that can be referenced from other pictures.

[0436] The TemporalId shown in Fig. 42(a) is an identifier for the hierarchy of the coding structure. The larger the value of the TemporalId, the deeper the hierarchy it indicates. The square blocks represent pictures. In the block, Ix represents an I picture (intra-predicted picture within the frame), Px represents a P picture (forward reference predicted picture), and Bx and bx represent B pictures (bi-directional reference predicted pictures). The x in Ix / Px / Bx indicates the display order, representing the order in which the pictures are to be displayed. The arrows between pictures indicate the reference relationship. For example, the picture B4 indicates that it generates a predicted picture using I0 and B8 as reference pictures. Here, it is prohibited for one picture to use another picture with a TemporalId larger than its own as a reference picture. The hierarchy is defined to provide temporal scalability. For example, if all the pictures in Fig. 42 are decoded, a video at 120 fps (frames per second) can be obtained, but if only the hierarchy from TemporalId 0 to 3 is decoded, a video at 60 fps can be obtained.

[0437] Fig. 43 is a diagram showing the relationship between the decoding time (DTS) and the presentation time (PTS) for each picture in Fig. 42. For example, the picture I0 shown in Fig. 43 is presented after the decoding of B4 is completed so that no gap occurs in decoding and presentation.

[0438] As shown in Fig. 43, when the prediction structure includes B pictures, etc., since the decoding order and the presentation order are different, in the receiving device 20, after decoding a picture, delay processing of the picture and reordering (reorderer) processing of the picture are required.

[0439] Above, an example of the prediction structure of pictures in temporal scalability has been described. However, even when temporal scalability is not used, depending on the prediction structure, delay processing of pictures and reorderer processing may be required. Fig. 44 is a diagram showing an example of the prediction structure of pictures for which delay processing of pictures and reorderer processing are required. Note that the numbers in Fig. 44 indicate the decoding order.

[0440] As shown in FIG. 44, depending on the prediction structure, the sample that is the first in the decoding order and the sample that is the first in the presentation order may be different. In FIG. 44, the sample that is the first in the presentation order is the fourth sample in the decoding order. Note that FIG. 44 shows an example of the prediction structure, and the prediction structure is not limited to such a structure. In other prediction structures, the sample that is the first in the decoding order and the sample that is the first in the presentation order may also be different.

[0441] Similar to FIG. 33, FIG. 45 is a diagram showing an example in which an MPU configured in the MP4 format is divided into a plurality of movie fragments and stored in an MMTP payload and an MMTP packet. Note that the number of samples constituting the MPU and the number of samples constituting the movie fragment are arbitrary. For example, the number of samples constituting the MPU may be set to the number of samples in a GOP unit, and two movie fragments may be configured with the number of samples that is half of the GOP unit as the movie fragment. One sample may be set as one movie fragment, or the samples constituting the MPU may not be divided.

[0442] In FIG. 45, an example in which one MPU includes two movie fragments (moof box and mdat box) is shown, but the number of movie fragments included in one MPU does not have to be two. The number of movie fragments included in one MPU may be three or more, or may be the number of samples included in the MPU. Also, the samples stored in the movie fragment do not have to be divided into equal numbers of samples, and may be divided into arbitrary numbers of samples.

[0443] The movie fragment metadata (MF metadata) includes information on the PTS, DTS, offset, and size of the samples included in the movie fragment. When decoding the samples, the receiving device 20 extracts the PTS and DTS from the MF metadata including the information on the samples, and determines the decoding timing and the presentation timing.

[0444] Hereinafter, for the sake of detailed explanation, the absolute value of the decoding time of the i-th sample is denoted as DTS(i), and the absolute value of the presentation time is denoted as PTS(i).

[0445] Specifically, the information of the i-th sample among the timestamp information stored in moof in MF meta is the relative value of the decoding times of the i-th sample and the (i + 1)-th sample, and the relative value of the decoding time and the presentation time of the i-th sample, which are hereinafter denoted as DT(i) and CT(i).

[0446] Movie fragment metadata #1 includes DT(i) and CT(i) of samples #1 - #3, and movie fragment metadata #2 includes DT(i) and CT(i) of samples #4 - #6.

[0447] Also, the absolute value of PTS of the access unit at the head of the MPU is stored in the MPU timestamp descriptor or the like, and the receiving device 20 calculates PTS and DTS based on PTS_MPU of the access unit at the head of the MPU, CT, and DT.

[0448] FIG. 46 is a diagram for explaining the calculation method and problems of PTS and DTS when the MPU is composed of samples #1 - #10.

[0449] FIG. 46(a) shows an example where the MPU is not divided into movie fragments, FIG. 46(b) shows an example where the MPU is divided into two movie fragments in units of 5 samples, and FIG. 46(c) shows an example where the MPU is divided into 10 movie fragments in units of samples.

[0450] As described with reference to FIG. 45, when PTS and DTS are calculated using the MPU timestamp descriptor and the timestamp information (CT and DT) in the MP4, the sample that is the first in the presentation order in FIG. 44 is the fourth in the decoding order. Therefore, the PTS stored in the MPU timestamp descriptor is the PTS (absolute value) of the fourth sample in the decoding order. Hereinafter, this sample will be referred to as the A sample. Also, the sample that is the first in the decoding order will be referred to as the B sample.

[0451] Since the absolute time information related to the timestamp is only the information in the MPU timestamp descriptor, the receiving device 20 cannot calculate the PTS (absolute time) and DTS (absolute time) of other samples until the A sample arrives. The receiving device 20 also cannot calculate the PTS and DTS of the B sample.

[0452] In the example of FIG. 46(a), the A sample is included in the same movie fragment as the B sample and is stored in one MF meta. Therefore, after receiving the MF meta, the receiving device 20 can immediately determine the DTS of the B sample.

[0453] In the example of FIG. 46(b), the A sample is included in the same movie fragment as the B sample and is stored in one MF meta. Therefore, after receiving the MF meta, the receiving device 20 can immediately determine the DTS of the B sample.

[0454] In the example of FIG. 46(c), the A sample is included in a different movie fragment from the B sample. Therefore, the receiving device 20 cannot determine the DTS of the B sample until it receives the MF meta including the CT and DT of the movie fragment including the A sample.

[0455] Therefore, in the case of the example of FIG. 46(c), the receiving device 20 cannot start decoding immediately after the B sample arrives.

[0456] Thus, when the movie fragment containing the B sample does not contain the A sample, the receiving device 20 cannot start decrypting the B sample until it has received the MF meta related to the movie fragment containing the A sample.

[0457] This problem occurs when the first sample in the presentation order does not match the first sample in the decoding order and the movie fragment is split before the A sample and the B sample are no longer stored in the same movie fragment. This problem occurs regardless of whether the MF meta is post-forwarded or pre-forwarded.

[0458] Thus, when the first sample in the presentation order does not match the first sample in the decoding order and the A sample and the B sample are not stored in the same movie fragment, the DTS cannot be determined immediately after the B sample is received. Therefore, the transmitting device 15 separately transmits the DTS (absolute value) of the B sample or information that can be calculated on the receiving side for the DTS (absolute value) of the B sample. Such information may be transmitted using control information, packet headers, or the like.

[0459] The receiving device 20 calculates the DTS (absolute value) of the B sample using such information. FIG. 47 is a flowchart of the receiving operation when the DTS is calculated using such information.

[0460] The receiving device 20 receives the movie fragment at the head of the MPU (S901) and determines whether the A sample and the B sample are stored in the same movie fragment (S902). If they are stored in the same movie fragment (Yes in S902), the receiving device 20 calculates the DTS using only the information of the MF meta without using the DTS (absolute time) of the B sample and starts decoding (S904). Note that in step S904, the receiving device 20 may determine the DTS using the DTS of the B sample.

[0461] On the other hand, if in step S902, the A sample and the B sample are not stored in the same movie fragment (No in S902), the receiving apparatus 20 acquires the DTS (absolute time) of the B sample, determines the DTS, and starts decoding (S903).

[0462] In the above description, an example of calculating the absolute value of the decoding time and the absolute value of the presentation time of each sample using MF meta (timestamp information stored in moof in the MP4 format) in the MMT standard has been described. However, it goes without saying that MF meta can be replaced with any control information that can be used to calculate the absolute value of the decoding time and the absolute value of the presentation time of each sample and implemented. Examples of such control information include control information in which the relative value CT(i) of the decoding times of the i-th sample and the (i + 1)-th sample is replaced with the relative value of the presentation times of the i-th sample and the (i + 1)-th sample, and control information including both the relative value CT(i) of the decoding times of the i-th sample and the (i + 1)-th sample and the relative value of the presentation times of the i-th sample and the (i + 1)-th sample.

[0463] (Embodiment 3) [Overview] In Embodiment 3, a content transmission method and a data structure in the case of transmitting contents such as video, audio, subtitles, and data broadcasting by broadcasting will be described. That is, a content transmission method and a data structure specialized for the playback of a broadcast stream will be described.

[0464] In Embodiment 3, an example in which the MMT method (hereinafter also simply referred to as MMT) is used as the multiplexing method will be described. However, other multiplexing methods such as MPEG-DASH or RTP may be used.

[0465] First, details of the method of storing in the payload of a data unit (DU: Data Unit) in MMT will be described. FIG. 48 is a diagram for explaining the method of storing in the payload of a data unit in MMT.

[0466] In MMT, the transmitting device stores a part of the data that constitutes the MPU as a data unit in the MMTP payload, attaches a header, and transmits it. The header includes an MMTP payload header and an MMTP packet header. Note that the unit of the data unit may be a NAL unit or a sample unit.

[0467] Fig. 48(a) shows an example in which the transmitting device aggregates a plurality of data units and stores them in one payload. In the example of Fig. 48(a), a data unit header (DUH: Data Unit Header) and a data unit length (DUL: Data Unit Length) are attached to the head of each of the plurality of data units, and the data units to which the data unit header and the data unit length are attached are collectively stored in the payload.

[0468] Fig. 48(b) shows an example in which one data unit is stored in one payload. In the example of Fig. 48(b), a data unit header is attached to the head of the data unit and stored in the payload. Fig. 48(c) shows an example in which one data unit is divided, and a data unit header is attached to the divided data unit and stored in the payload.

[0469] The data units include types such as timed-MFU which is a media including information related to synchronization such as video, audio, or subtitles, non-timed-MFU which is a media not including information related to synchronization such as a file, MPU metadata, and MF metadata. The data unit header is determined according to the type of the data unit. Note that there is no data unit header in the MPU metadata and the MF metadata.

[0470] In addition, although the transmitting device generally cannot aggregate different types of data units, it may be specified that different types of data units can be aggregated. For example, when the size of MF metadata is small, such as when it is divided into movie fragments for each sample, aggregating the MF metadata and media data can reduce the number of packets and further reduce the transmission capacity.

[0471] When the data unit is an MFU, some information of the MPU, such as information for constructing the MPU (MP4), is stored as a header.

[0472] For example, the header of a timed-MFU includes movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter, etc., and the header of a non-timed-MFU includes item_iD. The meaning of each field is specified in standards such as ISO / IEC 23008-1 or ARIB STD-B60. Hereinafter, the meaning of each field specified in such standards will be described.

[0473] movie_fragment_sequence_number indicates the sequence number of the movie fragment to which the MFU belongs, and is also shown in ISO / IEC 14496-12.

[0474] sample_number indicates the sample number to which the MFU belongs, and is also shown in ISO / IEC 14496-12.

[0475] offset indicates, in bytes, the offset amount of the MFU in the sample to which the MFU belongs.

[0476] priority indicates the relative importance of the MFU in the MPU to which the MFU belongs, and an MFU with a larger priority number indicates that it is more important than an MFU with a smaller priority number.

[0477] The dependency_counter indicates the number of MFUs on which the decoding process depends (i.e., the number of MFUs for which the decoding process cannot be performed without decoding this MFU). For example, when the MFU is HEVC and a B picture or a P picture references an I picture, the B picture or the P picture cannot be decoded without decoding the I picture.

[0478] Therefore, when the MFU is in sample units, the dependency_counter in the MFU of the I picture indicates the number of pictures that reference the I picture. When the MFU is in NAL unit units, the dependency_counter in the MFU belonging to the I picture indicates the number of NAL units belonging to the pictures that reference the I picture. Further, in the case of a temporally hierarchical coded video signal, since the MFU of the enhancement layer depends on the MFU of the base layer, the dependency_counter in the MFU of the base layer indicates the number of MFUs of the enhancement layer. This field cannot be generated until the number of dependent MFUs is determined.

[0479] The item_iD indicates an identifier that uniquely identifies an item.

[0480] [MP4 non-support mode] As described with reference to FIGS. 19 and 21, as a method for the transmission device to transmit the MPU in MMT, there are a method of transmitting the MPU metadata or the MF metadata before or after the media data, and a method of transmitting only the media data. Also, in the receiving device, there are a method of performing decoding using a receiving device or a receiving method compliant with MP4, and a method of performing decoding without using a header.

[0481] As a method for transmitting data specialized for broadcast stream reproduction, for example, there is a transmission method that does not support MP4 reconstruction in the receiving device.

[0482] A transmission method that does not support MP4 reconstruction in a receiving device is, for example, a method of not transmitting metadata (MPU metadata and MF metadata) as shown in FIG. 21(b). In this case, the field value of the fragment type (information indicating the type of data unit) included in the MMTP packet is fixed at 2 (= MFU).

[0483] When metadata is not transmitted, as described above, in an MP4-compliant receiving device, etc., the received data cannot be decoded as MP4, but it is possible to decode without using metadata (headers).

[0484] Therefore, metadata is not necessarily essential information for broadcast stream decoding and playback. Similarly, the information in the data unit header in timed-MFU described in FIG. 48 is information for reconstructing MP4 in a receiving device. Since there is no need to reconstruct MP4 in broadcast stream playback, the information in the data unit header in timed-MFU (hereinafter also referred to as the timed-MFU header) is not necessarily essential information for broadcast stream playback.

[0485] The receiving device can easily reconstruct MP4 by using metadata and information for reconstructing MP4 in the data unit header (hereinafter also referred to as MP4 configuration information). However, even if only one of the metadata and the MP4 configuration information in the data unit header is transmitted, the receiving device cannot reconstruct MP4. There are few merits in transmitting only one of the metadata and the information for reconstructing MP4, and generating and transmitting unnecessary information causes an increase in processing and a decrease in transmission efficiency.

[0486] Therefore, the transmitting device controls the data structure and transmission of the MP4 configuration information using the following method. The transmitting device determines whether to indicate the MP4 configuration information in the data unit header based on whether the metadata is transmitted. Specifically, when the metadata is transmitted, the transmitting device indicates the MP4 configuration information in the data unit header, and when the metadata is not transmitted, the transmitting device does not indicate the MP4 configuration information in the data unit header.

[0487] As a method of not indicating the MP4 configuration information in the data unit header, for example, the following method can be used.

[0488] 1. The transmitting device sets the MP4 configuration information as reserved and does not operate it. Thereby, the processing amount on the transmission side (the processing amount of the transmitting device) for generating the MP4 configuration information can be reduced.

[0489] 2. The transmitting device deletes the MP4 configuration information and performs header compression. Thereby, the processing amount on the transmission side for generating the MP4 configuration information can be reduced, and the transmission capacity can also be reduced.

[0490] Note that when the transmitting device deletes the MP4 configuration information and performs header compression, it may indicate a flag indicating that the MP4 configuration information has been deleted (compressed). The flag is indicated in the header (MMTP packet header, MMTP payload header, data unit header) or control information, etc.

[0491] Also, the information on whether the metadata is transmitted may be determined in advance, or may be signaled to the receiving device separately in a header or control information.

[0492] For example, information on whether the metadata corresponding to the MFU is transmitted may be stored in the MFU header.

[0493] On the other hand, the receiving device can determine whether the MP4 configuration information is indicated based on whether the metadata is transmitted.

[0494] Here, when the transmission order of data (for example, in the order such as MPU metadata, MF metadata, and media data) is determined, the receiving device may make a determination based on whether the metadata is received before the media data.

[0495] When MP4 configuration information is indicated, the receiving device can use the MP4 configuration information for the reconstruction of MP4. Alternatively, the receiving device can use the MP4 configuration information for the detection of the start of other access units or NAL units.

[0496] Note that the MP4 configuration information may be all or part of the timed-MFU header.

[0497] Also, the transmitting device may similarly determine whether to indicate the item id in the non-timed-MFU header based on whether the metadata is transmitted in the non-timed-MFU header.

[0498] The transmitting device may indicate the MP4 configuration information only in either the timed-MFU or the non-timed-MFU. When indicating the MP4 configuration information only in either one, the transmitting device determines whether to indicate the MP4 configuration information based on whether the metadata is transmitted and whether it is a timed-MFU or a non-timed-MFU. The receiving device can determine whether the MP4 configuration information is indicated based on whether the metadata is transmitted and the timed / non-timed flag.

[0499] Note that in the above description, the transmitting device determines whether to indicate the MP4 configuration information based on whether the metadata (both MPU metadata and MF metadata) is transmitted. However, the transmitting device may not indicate the MP4 configuration information when a part of the metadata (either MPU metadata or MF metadata) is not transmitted.

[0500] Further, the transmission device may determine whether to indicate MP4 configuration information based on information other than metadata.

[0501] For example, modes such as an MP4 support mode / MP4 non-support mode are defined. In the case of the MP4 support mode, the transmission device may indicate MP4 configuration information in the data unit header, and in the case of the MP4 non-support mode, may not indicate MP4 configuration information in the data unit header. Further, in the case of the MP4 support mode, the transmission device may transmit metadata and indicate MP4 configuration information in the data unit header, and in the case of the MP4 non-support mode, may not transmit metadata and also may not indicate MP4 configuration information in the data unit header.

[0502] [Operation flow of the transmission device] Next, the operation flow of the transmission device will be described. FIG. 49 shows the operation flow of the transmission device.

[0503] First, the transmission device determines whether to transmit metadata (S1001). If the transmission device determines to transmit metadata (Yes in S1002), it proceeds to step S1003, generates MP4 configuration information, stores it in the header, and transmits it (S1003). In this case, the transmission device also generates and transmits metadata.

[0504] On the other hand, if the transmission device determines not to transmit metadata (No in S1002), it does not generate MP4 configuration information and does not store it in the header but transmits it (S1004). In this case, the transmission device does not generate or transmit metadata.

[0505] Note that whether to transmit metadata in step S1001 may be determined in advance, or may be determined based on whether metadata is generated inside the transmission device or whether metadata is being transmitted inside the transmission device.

[0506] [Operation Flow of Receiver Device] Next, the operation flow of the receiver device will be described. FIG. 50 shows the operation flow of the receiver device.

[0507] First, the receiver device determines whether metadata is being transmitted (S1101). Whether metadata is being transmitted can be determined by monitoring the fragment type in the MMTP packet payload. Also, whether it is being transmitted may be predefined.

[0508] If the receiver device determines that metadata is being transmitted (Yes in S1102), it reconstructs the MP4 and executes the decoding process using the MP4 configuration information (S1103). On the other hand, if it determines that metadata is not being transmitted (No in S1102), it does not perform the MP4 reconstruction process and executes the decoding process without using the MP4 configuration information (S1104).

[0509] Note that the receiver device can detect the random access point, the start of the access unit, the start of the NAL unit, etc. without using the MP4 configuration information by using the methods described so far, and can perform the decoding process, detect packet loss, and perform the process of recovery from packet loss.

[0510] For example, the start of the access unit is the first data of the MMT payload where the aggregation_flag value is 1. At this time, the Fragmentation_indicator value is 0.

[0511] Also, the start of the slice segment is the first data of the MMT payload where the aggregation_flag value is 0 and the fragmentation_indicator value is 00 or 01.

[0512] Based on the above information, the receiver device can detect the start of the access unit and detect the slice segment.

[0513] Note that the receiving device may analyze the NAL unit header in a packet including the start of a data unit with a fragmentation_indicator value of 00 or 01, and detect that the type of the NAL unit is an AU delimiter and that the type of the NAL unit is a slice segment.

[0514] [Broadcast Simple Mode] So far, as a data transmission method specialized for broadcast stream playback, a method that does not support MP4 configuration information in the receiving device has been described. However, the data transmission method specialized for broadcast stream playback is not limited to this.

[0515] As a data transmission method specialized for broadcast stream playback, for example, the following method may be used.

[0516] · In a fixed reception environment for broadcasting, the transmitting device may not use AL-FEC. When AL-FEC is not used, the FEC_type in the MMTP packet header is always fixed to 0.

[0517] · In a mobile reception environment for broadcasting and in the communication UDP transmission mode, the transmitting device may always use AL-FEC. When AL-FEC is used, the FEC_type in the MMTP packet header is always 0 or 1.

[0518] · The transmitting device may not perform bulk transmission of assets. When bulk transmission of assets is not performed, the location_info location indicating the number of transmission locations of the assets inside the MPT may be fixed to 1.

[0519] · The transmitting device may not perform hybrid transmission of assets, programs, and messages.

[0520] Also, for example, a broadcast simple mode is defined, and when the transmission device is in the broadcast simple mode, it may be set to the MP4 non-support mode, or the data transmission method specialized for broadcast stream playback shown above may be used. Whether it is in the broadcast simple mode may be determined in advance, or the transmission device may store a flag indicating that it is in the broadcast simple mode as control information and transmit it to the receiving device.

[0521] Also, based on whether it is in the MP4 non-support mode (whether metadata is being transmitted) as described with reference to FIG. 49, when the transmission device is in the MP4 non-support mode, it may use the data transmission method specialized for broadcast stream playback shown above as the broadcast simple mode.

[0522] When the receiving device is in the broadcast simple mode, it can perform decoding processing without reconstructing MP4, assuming it is in the MP4 non-support mode.

[0523] Also, when the receiving device is in the broadcast simple mode, it can determine that it has a function specialized for broadcasting and perform reception processing specialized for broadcasting.

[0524] As a result, when in the broadcast simple mode, by using only functions specialized for broadcasting, not only can unnecessary processing for the transmission device and the receiving device be reduced, but also transmission overhead can be reduced by compressing and not transmitting unnecessary information.

[0525] Note that when the MP4 non-support mode is used, hint information supporting an accumulation method other than the MP4 configuration may be shown.

[0526] Examples of accumulation methods other than the MP4 configuration include directly accumulating MMT packets or IP packets, and converting MMT packets into MPEG-2 TS packets.

[0527] Note that in the case of the MP4 non - support mode, a format that does not conform to the MP4 configuration may be used.

[0528] For example, the data stored in the MFU may be in a format with a byte start code instead of a format with the size of the NAL unit attached to the head of the NAL unit in MP4 format in the case of the MP4 non - support mode.

[0529] In MMT, the asset type indicating the type of the asset is described by a 4CC registered in MP4REG (http: / / www.mp4ra.org). When using HEVC as the video signal, 'HEV1' or 'HVC1' is used. 'HVC1' is a format that may include a parameter set in the sample, and 'HEV1' is a format that does not include a parameter set in the sample but includes a parameter set in the sample entry in the MPU metadata.

[0530] In the case of the broadcast simple mode or the MP4 non - support mode, when the MPU metadata and the MF metadata are not transmitted, it may be stipulated that a parameter set must be included in the sample. Also, regardless of whether 'HEV1' or 'HVC1' is indicated for the asset type, it may be stipulated that the 'HVC1' format must be taken.

[0531] [Supplement 1: Transmitting device] As described above, when the metadata is not transmitted, the MP4 configuration information is set as reserved, and the transmitting device that does not operate can also be configured as shown in FIG. 51. FIG. 51 is a diagram showing an example of the specific configuration of the transmitting device.

[0532] The transmitting device 300 includes an encoding unit 301, an attaching unit 302, and a transmitting unit 303. Each of the encoding unit 301, the attaching unit 302, and the transmitting unit 303 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0533] The symbolization unit 301 symbolizes a video signal or an audio signal to generate sample data. Specifically, the sample data is a data unit.

[0534] The adding unit 302 adds header information including MP4 configuration information to the sample data which is the data obtained by symbolizing the video signal or the audio signal. The MP4 configuration information is information for reconstructing the sample data as a file in the MP4 format on the receiving side, and the content thereof differs depending on whether the presentation time of the sample data is determined.

[0535] As described above, the adding unit 302 includes MP4 configuration information such as movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter in the header (header information) of a timed-MFU which is an example of sample data (sample data including synchronization-related information) for which the presentation time is determined.

[0536] On the other hand, the adding unit 302 includes MP4 configuration information such as item_id in the header (header information) of a timed-MFU which is an example of sample data (sample data not including synchronization-related information) for which the presentation time is not determined.

[0537] When the metadata corresponding to the sample data is not transmitted by the transmitting unit 303 (for example, in the case of (b) in FIG. 21), the adding unit 302 adds header information not including MP4 configuration information to the sample data according to whether the presentation time of the sample data is determined.

[0538] Specifically, when the presentation time of the sample data is determined, the adding unit 302 adds header information not including first MP4 configuration information to the sample data, and when the presentation time of the sample data is not determined, the adding unit 302 adds header information including second MP4 configuration information to the sample data.

[0539] For example, as shown in step S1004 of FIG. 49, when the metadata corresponding to the sample data is not transmitted by the transmission unit 303, the attachment unit 302 sets the MP4 configuration information as reserved (fixed value), thereby substantially not generating the MP4 configuration information and substantially not storing it in the header (header information). Note that the metadata includes MPU metadata and movie fragment metadata.

[0540] The transmission unit 303 transmits the sample data to which the header information is attached. More specifically, the transmission unit 303 packetizes and transmits the sample data to which the header information is attached in the MMT format.

[0541] As described above, in the transmission method and reception method specialized for the playback of a broadcast stream, it is not necessary for the receiving device to reconstruct the data unit into MP4. When it is not necessary for the receiving device to reconstruct it into MP4, the processing of the transmitting device is reduced by not generating unnecessary information such as MP4 configuration information.

[0542] On the other hand, the transmitting device must send the necessary information, but it is necessary to maintain compatibility with the standard so that it does not have to separately send unnecessary additional information.

[0543] According to the configuration such as the transmitting device 300, by setting the area where the MP4 configuration information is stored to a fixed value, etc., the MP4 configuration information is not transmitted, only the necessary information is transmitted based on the standard, and the effect of not having to transmit unnecessary additional information can be obtained. That is, the configuration of the transmitting device and the processing amount of the transmitting device can be reduced. Also, since unnecessary data is not transmitted, the transmission efficiency can be improved.

[0544] [Supplementary Note 2: Receiving Device] Also, the receiving device corresponding to the transmitting device 300 may be configured as shown in FIG. 52, for example. FIG. 52 is a diagram showing another example of the configuration of the receiving device.

[0545] The receiving device 400 includes a receiving unit 401 and a decoding unit 402. The receiving unit 401 and the decoding unit 402 are realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0546] The receiving unit 401 receives sample data which is data obtained by encoding a video signal or an audio signal, and to which header information including MP4 configuration information for reconstructing the sample data as a file in the MP4 format is attached.

[0547] When the metadata corresponding to the sample data is not received by the receiving unit and the presentation time of the sample data is determined, the decoding unit 402 decodes the sample data without using the MP4 configuration information.

[0548] For example, as shown in step S1104 of FIG. 50, when the metadata corresponding to the sample data is not received by the receiving unit 401, the decoding unit 402 executes a decoding process without using the MP4 configuration information.

[0549] Thereby, the configuration of the receiving device 400 and the processing amount in the receiving device 400 can be reduced.

[0550] (Embodiment 4) [Overview] In Embodiment 4, a method for storing non-timed media (which does not include information related to synchronization such as a file) in an MPU and a method for transmitting it in an MMTP packet will be described. Note that in Embodiment 4, the MPU in MMT will be described as an example, but the same MP4-based DASH is also applicable.

[0551] First, the details of the method for storing non-timed media (hereinafter also referred to as "asynchronous media data") in an MPU will be described with reference to FIG. 53. FIG. 53 is a diagram showing the method for storing non-timed media in an MPU and the method for transmitting it in an MMTP packet.

[0552] The MPU that stores non-timed media is composed of boxes such as ftyp, mmpu, moov, and meta, and stores information about the files stored in the MPU. Multiple idat boxes can be stored within the meta box, and one file is stored as an item in the idat box.

[0553] Some of the ftyp, mmpu, moov, and meta boxes constitute one data unit as MPU metadata, and the item or idat box constitutes a data unit as MFU.

[0554] After the data unit is aggregated or fragmented, a data unit header, an MMTP payload header, and an MMTP packet header are added and transmitted as an MMTP packet.

[0555] Note that in FIG. 53, an example is shown where File #1 and File #2 are stored in one MPU. The MPU metadata is not divided, and the MFU is divided and stored in MMTP packets, but this is not the only case, and it may be aggregated or fragmented according to the size of the data unit. Also, the MPU metadata may not be transmitted, in which case only the MFU is transmitted.

[0556] The header information such as the data unit header indicates an itemID (an identifier that uniquely identifies an item), and the MMTP payload header and the MMTP packet header include a packet sequence number (a sequence number for each packet) and an MPU sequence number (a sequence number of the MPU, a unique number within the asset).

[0557] Note that the data structures of the MMTP payload header and the MMTP packet header other than the data unit header are the same as those of the timed media (hereinafter also referred to as "synchronized media data") described so far, and include an aggregation_flag, a fragmentation_indicator, a fragment_counter, etc.

[0558] Next, a specific example of the header information when a file (= Item = MFU) is split and packetized will be described with reference to FIGS. 54 and 55.

[0559] FIGS. 54 and 55 are diagrams showing examples of packetizing and transmitting each of a plurality of split data obtained by splitting a file. Specifically, FIGS. 54 and 55 show information (packet sequence number, fragment counter, fragmentation indicator, MPU sequence number, item ID) included in any of the data unit header, MMTP payload header, and MMTP packet header, which is the header information for each split MMTP packet. Note that FIG. 54 shows an example in which File #1 is split into M (M <= 256) pieces, and FIG. 55 shows an example in which File #2 is split into N (256 < N) pieces.

[0560] The split data number indicates the index of the split data from the beginning of the file, and this information is not transmitted. That is, the split data number is not included in the header information. Also, the split data number is a number assigned to each packet corresponding to each of the plurality of split data obtained by splitting the file, and is a number assigned by incrementing by 1 in ascending order from the first packet.

[0561] The packet sequence number is the sequence number of packets having the same packet ID. In FIGS. 54 and 55, consecutive numbers are assigned from the split data at the beginning of the file as A to the split data at the end of the file. The packet sequence number is a number assigned by incrementing by 1 in ascending order from the split data at the beginning of the file, and is a number corresponding to the split data number.

[0562] The fragment counter indicates the number of a plurality of divided data that are after the divided data among the plurality of divided data obtained by dividing one file. Further, when the number of divided data, which is the number of the plurality of divided data obtained by dividing one file, exceeds 256, the fragment counter indicates the remainder obtained by dividing the number of divided data by 256. In the example of FIG. 54, since the number of divided data is 256 or less, the field value of the fragment counter is (M - divided data number). On the other hand, in the example of FIG. 55, since the number of divided data exceeds 256, it is the value obtained by dividing (N - divided data number) by 256 ((N - divided data number) % 256).

[0563] Note that the fragmentation indicator indicates the state of division of the data stored in the MMTP packet, and is a value indicating whether it is the first divided data, the last divided data, other divided data, or one or more data units that are not divided in the divided data unit. Specifically, the fragmentation indicator is "01" for the first divided data, "11" for the last divided data, "10" for the remaining divided data, and "00" for the data unit that is not divided.

[0564] In the present embodiment, when the number of divided data exceeds 256, it is described as indicating the remainder obtained by dividing the number of divided data by 256. However, the number of divided data is not limited to 256, and may be another number (predetermined number).

[0565] When a file is split as shown in FIGS. 54 and 55, and conventional header information is attached to each of the plurality of split data obtained by splitting the file and then transmitted, in the receiving device, there is no information on which split data (split data number) in the original file the data stored in the received MMTP packet is, and the number of split data of the file, or information from which the split data number and the number of split data can be derived. Therefore, with the conventional transmission method, even when an MMTP packet is received, the split data number and the number of split data of the data stored in the received MMTP packet cannot be uniquely detected.

[0566] For example, as shown in FIG. 54, when the number of split data is 256 or less and it is known in advance that the number of split data is 256 or less, it is possible to specify the split data number and the number of split data by referring to the fragment counter. However, when the number of split data is 256 or more, the split data number and the number of split data cannot be specified.

[0567] In addition, when restricting the number of split data of a file to 256 or less, if the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed, and in this case, the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes]. Therefore, when it is desired to transmit a file exceeding 1 [Mbytes], the number of split data of the file cannot be restricted to 256 or less.

[0568] Also, for example, by referring to the fragmentation indicator, it is possible to detect the split data at the beginning or the end of the file. Therefore, the number of MMTP packets can be counted until an MMTP packet containing the last split data of the file is received, or after receiving an MMTP packet containing the last split data of the file, the split data number and the number of split data can be calculated by combining them with the packet sequence number. Thus, it may be possible to signal the split data number and the number of split data by combining the fragmentation indicator and the packet sequence number. However, when starting reception from an MMTP packet containing split data in the middle of the file (i.e., split data that is neither the split data at the beginning nor the split data at the end of the file), the split data number and the number of split data of the split data cannot be specified. The split data number and the number of split data of the split data can be specified for the first time after receiving an MMTP packet containing the last split data of the file.

[0569] For the problems described in FIGS. 54 and 55, that is, when receiving packets containing split data of a file from the middle, the following method is used to uniquely determine the split data number and the number of split data of the split data of the file.

[0570] First, the split data number will be described.

[0571] Regarding the split data number, the packet sequence number in the split data at the beginning of the file (item) is signaled.

[0572] As a signaling method, it is stored in the control information for managing the file. Specifically, in FIGS. 54 and 55, the packet sequence number A of the split data at the beginning of the file is stored in the control information. At the receiving device, the value of A is obtained from the control information, and the split data number is calculated from the packet sequence number indicated in the packet header.

[0573] The divided data number of the divided data is obtained by subtracting the packet sequence number A of the first divided data from the packet sequence number of the divided data.

[0574] As control information for managing files, for example, there is an asset management table defined in ARIB STD-B60. In the asset management table, for each file, file size, version information, etc. are shown, and are stored in the data transmission message and transmitted. FIG. 56 is a diagram showing the syntax of the loop for each file in the asset management table.

[0575] When the area of the existing asset management table cannot be expanded, signaling may be performed using a 32-bit area of a part of the item_info_byte field indicating item information. A flag indicating whether the packet sequence number in the first divided data of the file (item) is shown may be shown in, for example, the reserved_future_use field of the control information in a part area of item_info_byte.

[0576] When repeatedly transmitting files such as a data carousel, a plurality of packet sequence numbers may be shown, or the packet sequence number of the head of the file to be transmitted immediately after may be shown.

[0577] Not limited to the packet sequence number of the divided data at the head of the file, any information that associates the divided data number of the file with the packet sequence number may be used.

[0578] Next, the number of divided data will be described.

[0579] The order of the loop for each file included in the asset management table may be defined as the file transfer order. As a result, since the packet sequence numbers at the beginning of two consecutive files in the transfer order are known, the number of split data of the previously transferred file can be identified by subtracting the packet sequence number at the beginning of the previously transferred file from the packet sequence number at the beginning of the subsequently transferred file. That is, for example, when File#1 shown in FIG. 54 and File#2 shown in FIG. 55 are consecutive files in this order, consecutive numbers are assigned to the last packet sequence number of File#1 and the packet sequence number at the beginning of File#2.

[0580] Also, by defining the file splitting method, it may be defined so that the number of split data of the file can be identified. For example, when the number of split data is N, the size of each of the 1st to (N - 1)th split data is L, and the size of the Nth split data is defined as the remainder (item_size - L * (N - 1)), the number of split data can be calculated inversely from the item_size shown in the asset management table. In this case, the integer value obtained by rounding up (item_size / L) is the number of split data. Note that the file splitting method is not limited to this.

[0581] Also, the number of split data may be directly stored in the asset management table.

[0582] In the receiving device, by using the above method, control information is received, and the number of split data is calculated based on the control information. Also, the packet sequence number corresponding to the split data number of the file can be calculated based on the control information. Note that when the timing of receiving the split data packets is earlier than the timing of receiving the control information, the split data number and the number of split data may be calculated at the timing of receiving the control information.

[0583] In addition, when signaling the split data number or the number of split data using the above method, the split data number and the number of split data are not specified based on the fragment counter, and the fragment counter becomes unnecessary data. Therefore, in the transmission of asynchronous media, when information capable of specifying the split data number and the number of split data is signaled using the above method or the like, the fragment counter may not be operated or may be header-compressed. Thereby, the processing amount of the transmission device and the reception device can be reduced, and the transmission efficiency can also be improved. That is, when transmitting asynchronous media, the fragment counter may be set to reserved (invalidated). Specifically, the value of the fragment counter may be set to a fixed value such as "0". Also, when receiving asynchronous media, the fragment counter may be ignored.

[0584] When storing synchronous media such as video and audio, the transmission order of MMTP packets in the transmission device matches the arrival order of MMTP packets in the reception device, and the packets are not retransmitted. In such a case, when it is not necessary to detect and reconstruct packet loss, the fragment counter may not be operated. In other words, in this case, the fragment counter may be set to reserved (invalidated).

[0585] Also, even without using the fragment counter, it is possible to detect a random access point, detect the head of an access unit, detect the head of an NAL unit, etc., and perform decoding processing, detect packet loss, and process recovery from packet loss.

[0586] In the transmission of real-time content such as live broadcasts, lower-latency transmission is required, and it is necessary to sequentially packetize and transmit the encoded data. However, in the transmission of real-time content, with the conventional fragment counter, the number of fragmented data cannot be determined at the time of transmitting the first fragmented data. Therefore, the transmission of the first fragmented data occurs after all the encoding of the data unit is completed and the number of fragmented data is determined, resulting in latency. Even in such a case, by using the above method and not operating the fragment counter, this latency can be reduced.

[0587] FIG. 57 is an operation flow for specifying the fragmented data number in the receiving device.

[0588] The receiving device acquires control information in which file information is described (S1201). The receiving device determines whether the packet sequence number at the head of the file is indicated in the control information (S1202). If the packet sequence number at the head of the file is indicated in the control information (Yes in S1202), the receiving device calculates the packet sequence number corresponding to the fragmented data number of the fragmented data of the file (S1203). Then, after acquiring the MMTP packet in which the fragmented data is stored, the receiving device specifies the fragmented data number of the file from the packet sequence number stored in the packet header of the acquired MMTP packet (S1204). On the other hand, if the packet sequence number at the head of the file is not indicated in the control information (No in S1202), after acquiring the MMTP packet including the last fragmented data of the file, the receiving device specifies the fragmented data number by using the fragment indicator and the packet sequence number stored in the packet header of the acquired MMTP packet (S1205).

[0589] FIG. 58 is an operation flow for specifying the number of fragmented data in the receiving device.

[0590] The receiving device acquires control information in which file information is described (S1301). The receiving device determines whether the control information includes information capable of calculating the number of divided data of the file (S1302). If it is determined that the information capable of calculating the number of divided data is included (Yes in S1302), the receiving device calculates the number of divided data based on the information included in the control information (S1303). On the other hand, if the receiving device determines that the number of divided data cannot be calculated (No in S1302), after acquiring the MMTP packet including the last divided data of the file, the receiving device uses the fragment indicator stored in the packet header of the acquired MMTP packet and the packet sequence number to identify the number of divided data (S1304).

[0591] FIG. 59 is an operation flow for determining whether to operate a fragment counter in a transmitting device.

[0592] First, the transmitting device determines whether the media to be transmitted (hereinafter, also referred to as "media data") is synchronous media or asynchronous media (S1401).

[0593] If the result of the determination in step S1401 is synchronous media (synchronous media in S1402), the transmitting device determines whether the MMTP packet order of transmission and reception matches in an environment where synchronous media is transmitted and whether packet reconstruction is unnecessary in the event of packet loss (S1403). If the transmitting device determines that it is unnecessary (Yes in S1403), the transmitting device does not operate the fragment counter (S1404). On the other hand, if the transmitting device determines that it is not unnecessary (No in S1403), the transmitting device operates the fragment counter (S1405).

[0594] If the result of the determination in step S1401 is asynchronous media (asynchronous media in S1402), the transmitting device determines whether to operate the fragment counter based on whether the split data number and the number of split data are signaled using the method described above. Specifically, when the split data number and the number of split data are signaled (Yes in S1406), the transmitting device does not operate the fragment counter (S1404). On the other hand, when the split data number and the number of split data are not signaled (No in S1406), the transmitting device operates the fragment counter (S1405).

[0595] Note that when the transmitting device does not operate the fragment counter, the value of the fragment counter may be set to reserved, or header compression may be performed.

[0596] Note that the transmitting device may determine whether to signal the split data number and the number of split data described above based on whether to operate the fragment counter.

[0597] Note that when the synchronous media does not operate the fragment counter, the transmitting device may signal the split data number and the number of split data using the method described above in the asynchronous media. Conversely, based on whether the asynchronous media operates the fragment counter, the operation of the synchronous media may be determined. In this case, in the synchronous media and the asynchronous media, whether to operate the fragment can be the same operation.

[0598] Next, a method for specifying the number of split data and the split data number (when using the fragment counter) will be described. FIG. 60 is a diagram for explaining a method for specifying the number of split data and the split data number (when using the fragment counter).

[0599] As described with reference to FIG. 54, when the number of divided data is 256 or less and it is known in advance that the number of divided data is 256 or less, the divided data number and the number of divided data can be specified by referring to the fragment counter.

[0600] When the number of divided data of a file is limited to 256 or less, if the data size that can be transmitted in one packet is x [bytes], the maximum size of the file that can be transmitted is limited to x * 256 [bytes]. For example, in broadcasting, x = 4k [bytes] is assumed. In this case, the maximum size of the file that can be transmitted is limited to 4k * 256 = 1M [bytes].

[0601] When the file size exceeds the maximum size of the file that can be transmitted, the file is divided in advance so that the size of the divided file is x * 256 [bytes] or less. Each of the plurality of divided files obtained by dividing the file is treated as one file (item), further divided within 256, and the divided data obtained by further division is stored in MMTP packets and transmitted.

[0602] Note that information indicating that the item is a divided file, the number of divided files, and the sequence number of the divided files may be stored in the control information and transmitted to the receiving device. Also, these information may be stored in the asset management table, or may be indicated using a part of the existing field item_info_byte.

[0603] When the receiving device receives one of the multiple split files obtained by splitting an item into a single file, it can identify the other split files and reconstruct the original file. Also, in the receiving device, by using the number of split files, the index of the split file, and the fragment counter in the control information, the number of split data and the split data number can be uniquely identified. Also, the number of split data and the split data number can be uniquely identified without using a packet sequence number or the like.

[0604] Here, it is desirable that the item_id of each of the multiple split files obtained by splitting one file is the same as each other. When assigning a different item_id, in order to uniquely refer to the file from other control information or the like, it may be assumed that the item_id of the first split file is indicated.

[0605] Also, the multiple split files may be assumed to always belong to the same MPU. When storing multiple files in the MPU, it is not necessary to store different types of files, and it may be assumed that only the files obtained by splitting one file are stored. The receiving device can detect file updates by checking the version information for each MPU without checking the version information for each item.

[0606] FIG. 61 is an operation flow of the transmitting device when utilizing the fragment counter.

[0607] First, the transmitting device checks the size of the file to be transmitted (S1501). Next, the transmitting device determines whether the file size exceeds x * 256 [bytes] (where x is the data size that can be transmitted in one packet. For example, the MTU size.) (S1502). If the file size exceeds x * 256 [bytes] (Yes in S1502), the file is split so that the size of the split file is less than x * 256 [bytes] (S1503). Then, the split files are transmitted as items, and information about the split files (for example, that they are split files, sequence numbers in the split files, etc.) is stored in the control information and transmitted (S1504). On the other hand, if the file size is less than x * 256 [bytes] (No in S1502), the file is transmitted as an item as usual (S1505).

[0608] Figure 62 shows the operation flow of the receiving device when utilizing the fragment counter.

[0609] First, the receiving device acquires and analyzes control information regarding the transmission of the file, such as the asset management table (S1601). Next, the receiving device determines whether the desired item is a split file (S1602). If the receiving device determines that the desired file is a split file (Yes in S1602), it acquires information for reconstructing the file, such as the split file and the index of the split file, from the control information (S1603). Then, the receiving device acquires the items that make up the split file and reconstructs the original file (S1604). On the other hand, if the receiving device determines that the desired file is not a split file (No in S1602), it acquires the file as usual (S1605).

[0610] In short, the transmitting device signals the packet sequence number of the split data at the beginning of the file. Also, the transmitting device signals information that can identify the number of split data. Alternatively, the transmitting device defines a splitting rule that can identify the number of split data. Also, the transmitting device operates without using the fragment counter and performs reserved or header compression.

[0611] When the packet sequence number of the data at the head of the file is signaled in the receiving device, the divided data number and the number of divided data are specified from the packet sequence number of the divided data at the head of the file and the packet sequence number of the MMTP packet.

[0612] From another perspective, the transmitting device divides a file, divides the data for each divided file, and transmits it. Information (such as sequence number, number of divisions, etc.) for associating the divided files is signaled.

[0613] The receiving device specifies the divided data number and the number of divided data based on the fragment counter and the sequence number of the divided file.

[0614] As a result, the divided data number and the divided data can be uniquely specified. Also, when receiving the intermediate divided data, the divided data number of the divided data can be specified, so the waiting time can be reduced and the memory can also be reduced.

[0615] Also, by not operating the fragment counter, the configuration of the transmitting and receiving devices can reduce the processing amount and improve the transmission efficiency.

[0616] FIG. 63 is a diagram showing a service configuration when transmitting the same program by a plurality of IP data flows. Here, a part (video / audio) of the data of the program with service ID = 2 is transmitted by an IP data flow using the MMT method, and data having the same service ID and different from the part of the data is transmitted by an IP data flow using the advanced BS data transmission method (in this example, the file transfer protocol is also different, but it may be the same protocol). An example is shown.

[0617] The transmitting device multiplexes the IP data so that the receiving device can guarantee that the data composed of a plurality of IP data flows is assembled by the decoding time.

[0618] The receiving device can realize guaranteed receiver operation by processing data composed of a plurality of IP data flows based on the decoding time.

[0619] [Supplementary Note: Transmitting Device and Receiving Device] As described above, the transmitting device that transmits data without operating the fragment counter can also be configured as shown in FIG. 64. Further, the receiving device that receives data without operating the fragment counter can also be configured as shown in FIG. 65. FIG. 64 is a diagram showing an example of the specific configuration of the transmitting device. FIG. 65 is a diagram showing an example of the specific configuration of the receiving device.

[0620] The transmitting device 500 includes a splitting unit 501, a configuring unit 502, and a transmitting unit 503. Each of the splitting unit 501, the configuring unit 502, and the transmitting unit 503 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0621] The receiving device 600 includes a receiving unit 601, a determining unit 602, and a configuring unit 603. Each of the receiving unit 601, the determining unit 602, and the configuring unit 603 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0622] Detailed descriptions of the components of the transmitting device 500 and the receiving device 600 will be given in the descriptions of the transmitting method and the receiving method, respectively.

[0623] First, the transmitting method will be described with reference to FIG. 66. FIG. 66 is the operation flow (transmitting method) by the transmitting device.

[0624] First, the splitting unit 501 of the transmitting device 500 splits the data into a plurality of split data (S1701).

[0625] Next, the configuring unit 502 of the transmitting device 500 configures a plurality of packets by attaching header information to each of the plurality of split data and packetizing them (S1702).

[0626] Then, the transmission unit 503 of the transmission device 500 transmits the configured plurality of packets (S1703). The transmission unit 503 transmits the divided data information and the value of the invalidated fragment counter. The divided data information is information for specifying the divided data number and the number of divided data. The divided data number is a number indicating which divided data among the plurality of divided data the divided data is. The number of divided data is the number of the plurality of divided data.

[0627] Thereby, the processing amount of the transmission device 500 can be reduced.

[0628] Next, the receiving method will be described with reference to FIG. 67. FIG. 67 is an operation flow (receiving method) by the receiving device.

[0629] First, the receiving unit 601 of the receiving device 600 receives a plurality of packets (S1801).

[0630] Next, the determination unit 602 of the receiving device 600 determines whether the divided data information is acquired from the received plurality of packets (S1802).

[0631] Then, when it is determined by the determination unit 602 that the divided data information is acquired (Yes in S1802), the configuration unit 603 of the receiving device 600 constructs data from the received plurality of packets without using the value of the fragment counter included in the header information (S1803).

[0632] On the other hand, when it is determined by the determination unit 602 that the divided data information is not acquired (No in S1802), the configuration unit 603 may construct data from the received plurality of packets using the value of the fragment counter included in the header information (S1804).

[0633] Thereby, the processing amount of the receiving device 600 can be reduced.

[0634] (Embodiment 5) [Overview] In Embodiment 5, a method for transmitting a transmission packet (TLV packet) when storing NAL units in multiplexing layers in the NAL size format will be described.

[0635] As described in Embodiment 1, when storing NAL units of H.264 or H.265 in multiplexing layers, there are the following two types of formats. One is a format called the byte stream format in which a start code consisting of a specific bit sequence is added immediately before the NAL unit header. The other is a format called the NAL size format in which a field indicating the size of the NAL unit is added. The byte stream format is used in MPEG-2 systems, RTP, etc., and the NAL size format is used in MP4, or in DASH, MMT, etc. that use MP4.

[0636] In the byte stream format, the start code is composed of 3 bytes, and an arbitrary byte (a byte with a value of 0) can also be added.

[0637] On the other hand, in the NAL size format in general MP4, the size information is indicated by either 1 byte, 2 bytes, or 4 bytes. This size information is indicated by the lengthSizeMinusOne field in the HEVC sample entry. When the value of the field is "0", it indicates 1 byte, when it is "1", it indicates 2 bytes, and when it is "3", it indicates 4 bytes.

[0638] Here, in ARIB STD - B60, "Media Transport Method by MMT in Digital Broadcasting" standardized in July 2014, when storing NAL units in multiplexing layers, if the output of the HEVC encoder is a byte stream, the byte start code is removed, and the size of the NAL unit in byte units indicated by 32 - bit (unsigned integer) is added as length information immediately before the NAL unit. Note that MPU metadata including HEVC sample entry is not transmitted, and the size information is fixed at 32 bits (4 bytes).

[0639] Also, in ARIB STD - B60, "Media Transport Method by MMT in Digital Broadcasting", in the reception buffer model considered during transmission to guarantee the buffer operation in the receiving device by the transmitting device, the pre - decoding buffer for the video signal is defined as a CPB.

[0640] However, there are the following problems. In the CPB in MPEG - 2 systems and the HRD in HEVC, it is defined on the premise that the video signal is in byte - stream format. For this reason, for example, when performing rate control of transmission packets on the premise that it is a byte - stream format with a 3 - byte start code attached, a receiving device that receives a transmission packet in NAL size format with a 4 - byte size area added may not be able to satisfy the reception buffer model in ARIB STD - B60. Also, since the specific buffer size and extraction rate are not shown in the reception buffer model in ARIB STD - B60, it is difficult to guarantee the buffer operation in the receiving device.

[0641] Therefore, in order to solve the above problems, a reception buffer model for guaranteeing the buffer operation in the receiver is defined as follows.

[0642] Figure 68 shows the reception buffer model, particularly when only a broadcast transmission path is used, based on the reception buffer model defined in ARIB STD - B60.

[0643] The reception buffer model includes a TLV packet buffer (first buffer), an IP packet buffer (second buffer), an MMTP buffer (third buffer), and a pre-decoding buffer (fourth buffer). Note that in the broadcast transmission path, since a digital buffer or a buffer for FEC is not necessary, it is omitted.

[0644] The TLV packet buffer receives a TLV packet (transmission packet) from the broadcast transmission path, and converts an IP packet composed of a variable-length packet header (IP packet header, full header at the time of IP packet compression, compressed header at the time of IP packet compression) and a variable-length payload stored in the received TLV packet into an IP packet having a fixed-length IP packet header with an extended header, and outputs the IP packet obtained by the conversion at a constant bit rate.

[0645] The IP packet buffer converts an IP packet into an MMTP packet (second packet) having a packet header and a variable-length payload, and outputs the MMTP packet obtained by the conversion at a constant bit rate. Note that the IP packet buffer may be merged into the MMTP buffer.

[0646] The MMTP buffer converts the output MMTP packet into a NAL unit, and outputs the NAL unit obtained by the conversion at a constant bit rate.

[0647] The pre-decoding buffer sequentially accumulates the output NAL units, generates an access unit from the plurality of accumulated NAL units, and outputs the generated access unit to the decoder at the timing of the decoding time corresponding to the access unit.

[0648] In the reception buffer model shown in FIG. 68, the MMTP buffer and the pre-decoding buffer, which are buffers other than the front-stage TLV packet buffer and IP packet buffer, are characterized by following the reception buffer model in MPEG-2 TS.

[0649] For example, the MMTP buffer (MMTP B1) in video is composed of a buffer corresponding to the transport buffer (TB) and the multiplexing buffer (MB) in MPEG-2 TS. Also, the MMTP buffer (MMTP Bn) in audio is composed of a buffer corresponding to the transport buffer (TB) in MPEG-2 TS.

[0650] The buffer size of the transport buffer is the same as that in MPEG-2 TS and is a fixed value. For example, it is set to n times the MTU size (n can be a decimal or an integer and is 1 or more).

[0651] Also, the MMTP packet size is defined so that the overhead rate of the MMTP packet header is smaller than the overhead rate of the PES packet header. As a result, the extraction rate from the transport buffer can directly apply the extraction rates RX1, RXn, and RXs of the transport buffer in MPEG-2 TS.

[0652] Also, the size and extraction rate of the multiplexing buffer are set to the MB size and RBX1 in MPEG-2 TS, respectively.

[0653] In addition to the above reception buffer model, the following constraints are provided to solve the problems.

[0654] The HRD specification of HEVC assumes a byte stream format, and MMT is a NAL size format that adds a 4-byte size area to the head of the NAL unit. Therefore, at the time of encoding, rate control is performed in the NAL size format to satisfy the HRD.

[0655] That is, in the transmission device, rate control of the transmission packet is performed based on the above reception buffer model and constraints.

[0656] In the reception device, by performing reception processing using the above signal, a decoding operation that does not cause underflow or overflow can be performed.

[0657] Note that even if the size area at the head of the NAL unit is not 4 bytes, rate control is performed so as to satisfy the HRD in consideration of the size area at the head of the NAL unit.

[0658] Note that the extraction rate of the TLV packet buffer (the bit rate when the TLV packet buffer outputs an IP packet) is set in consideration of the transmission rate after IP header extension.

[0659] That is, after inputting a TLV packet with a variable data size and performing removal of the TLV header and extension (restoration) of the IP header, the transmission rate of the output IP packet is considered. In other words, the increase or decrease amount of the header is considered with respect to the input transmission rate.

[0660] Specifically, since the data size is variable, packets with compressed IP headers and packets without compressed IP headers are mixed, and the size of the IP header varies depending on the packet type such as IPv4 and IPv6, the transmission rate of the output IP packet is not unique. For this reason, the average packet length of the variable data size is determined, and the transmission rate of the IP packet output from the TLV packet is determined.

[0661] Here, in order to define the maximum transmission speed after IP header extension, the transmission rate is determined assuming that the IP header is always compressed.

[0662] In addition, when IPv4 and IPv6 packet types are mixed, or when regulations are made without distinguishing packet types, the transmission rate is determined assuming an IPv6 packet with a large header size and a large increase rate after header extension.

[0663] For example, when the average packet length of TLV packets input to the TLV packet buffer is S, and all IP packets stored in the TLV packets are IPv6 packets and are header-compressed, the maximum output transmission rate after removing the TLV header and extending the IP header is Input rate × {S / (S + IPv6 header compression amount)} becomes.

[0664] More specifically, when setting the average packet length S of the TLV packet as S = 0.75 × 1500 (assuming 1500 is the maximum MTU size) as a reference, IPv6 header compression amount = TLV header length - IPv6 header length - UDP header length = 3 - 40 - 8 In this case, the maximum output transmission rate after removing the TLV header and extending the IP header is Input rate × 1.0417 ≒ Input rate × 1.05 becomes.

[0665] Figure 69 is a diagram showing an example of aggregating a plurality of data units and storing them in one payload.

[0666] In the MMT method, when aggregating data units, as shown in Figure 69, the data unit length and the data unit header are added before the data unit.

[0667] However, for example, when storing a video signal in the NAL size format as one data unit, as shown in FIG. 70, there are two fields indicating the size for one data unit, and the information is duplicated. FIG. 70 is an example of aggregating a plurality of data units and storing them in one payload, and is a diagram showing an example when a video signal in the NAL size format is used as one data unit. Specifically, both the size field at the beginning in the NAL size format (hereinafter referred to as the "size field" in the following description) and the data unit length field located before the data unit header in the MMTP payload header are fields indicating the size, and the information is duplicated. For example, when the length of the NAL unit is L bytes, L bytes are indicated in the size field, and in the data unit length field, L bytes + "the length of the size field" (byte) are indicated. Although the values indicated by the size field and the data unit length field do not exactly match, since one value can be easily calculated from the other value, it can be said that they are duplicated.

[0668] Thus, when storing data containing size information of data as data units and aggregating a plurality of such data units and storing them in one payload, there is a problem that the overhead is large and the transmission efficiency is poor because the size information is duplicated.

[0669] Therefore, in the transmitting device, when storing data containing size information of data as data units and aggregating a plurality of such data units and storing them in one payload, it can be considered to store them as shown in FIGS. 71 and 72.

[0670] As shown in FIG. 71, it can be considered to store the NAL unit including the size field as a data unit and not show the data unit length conventionally included in the MMTP payload header. FIG. 71 is a diagram showing the configuration of the payload of an MMTP packet in which the data unit length is not shown.

[0671] Also, as shown in FIG. 72, a flag indicating whether the data unit length is shown and information indicating the length of the size area may be newly stored in the header. The location for storing the flag and the information indicating the length of the size area may be indicated in units of data units such as the data unit header, or may be indicated in units obtained by aggregating a plurality of data units (packet units). FIG. 72 shows an example indicated in the extend area given in packet units. Note that the storage location of the newly shown information is not limited to this, and it may be the MMTP payload header, the MMTP packet header, or control information.

[0672] On the receiving side, when a flag indicating whether the data unit length is compressed indicates that the data unit length is compressed, the length information of the size area inside the data unit is acquired, and based on the length information of the size area, by acquiring the size area, the data unit length can be calculated using the acquired length information of the size area and the size area.

[0673] By the above method, the amount of data can be reduced on the sending side, and the transmission efficiency can be improved.

[0674] Note that instead of reducing the data unit length, the overhead may be reduced by reducing the size area. When reducing the size area, information indicating whether the size area is reduced and information indicating the length of the data unit length field may be stored.

[0675] Note that the MMTP payload header also includes length information.

[0676] When storing an NAL unit including a size area as a data unit, regardless of whether aggregation is performed or not, the payload size area in the MMTP payload header may be reduced.

[0677] Also, even when storing data that does not include the size area as a data unit, if it is aggregated and the data unit length is indicated, the payload size area in the MMTP payload header may be reduced.

[0678] When reducing the payload size area, similar to the above, a flag indicating whether it has been reduced, length information of the reduced size field, or length information of the size field that has not been reduced may be indicated.

[0679] FIG. 73 shows the operation flow of the receiving device.

[0680] In the transmitting device, as described above, it is assumed that the NAL unit including the size area is stored as a data unit, and the data unit length included in the MMTP payload header is not indicated in the MMTP packet.

[0681] Hereinafter, the case where it is indicated by a flag whether the data unit length is indicated or the length information of the size area is indicated in the MMTP packet will be described as an example.

[0682] The receiving device determines whether the data unit includes the size area and whether the data unit length has been reduced based on the information transmitted from the transmitting side (S1901).

[0683] If it is determined that the data unit length has been reduced (Yes in S1902), the length information of the size area inside the data unit is acquired, and then the size area inside the data unit is analyzed to obtain the data unit length by calculating it (S1903).

[0684] On the other hand, if it is determined that the data unit length has not been reduced (No in S1902), the data unit length is calculated from either the data unit length or the size area inside the data unit as usual (S1904).

[0685] Note that if the receiving device already knows in advance a flag indicating whether the data unit length has been reduced and the length information of the size area, there is no need to transmit them. In this case, the receiving device performs the process shown in FIG. 73 based on the predetermined information.

[0686] [Supplementary Explanation: Transmitting Device and Receiving Device] As described above, a transmitting device that performs rate control so as to satisfy the provisions of the reception buffer model during encoding can also be configured as shown in FIG. 74. Further, a receiving device that receives and decodes the transmission packets transmitted from the transmitting device can also be configured as shown in FIG. 75. FIG. 74 is a diagram showing an example of the specific configuration of the transmitting device. FIG. 75 is a diagram showing an example of the specific configuration of the receiving device.

[0687] The transmitting device 700 includes a generation unit 701 and a transmission unit 702. Each of the generation unit 701 and the transmission unit 702 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0688] The receiving device 800 includes a receiving unit 801, a first buffer 802, a second buffer 803, a third buffer 804, a fourth buffer 805, and a decoding unit 806. Each of the receiving unit 801, the first buffer 802, the second buffer 803, the third buffer 804, the fourth buffer 805, and the decoding unit (decoder) 806 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0689] Detailed explanations of the respective components of the transmitting device 700 and the receiving device 800 will be given in the explanations of the transmission method and the reception method, respectively.

[0690] First, the transmission method will be described with reference to FIG. 76. FIG. 76 is an operation flow (transmission method) by the transmitting device.

[0691] First, the generation unit 701 of the transmission device 700 generates an encoded stream by performing rate control so as to satisfy the regulations based on a reception buffer model determined in advance to guarantee the buffer operation of the reception device (S2001).

[0692] Next, the transmission unit 702 of the transmission device 700 packetizes the generated encoded stream and transmits the transmission packets obtained by packetization (S2002).

[0693] Note that since the reception buffer model used in the transmission device 700 has a configuration including the first to fourth buffers 802 to 805 of the reception device 800, the description thereof is omitted.

[0694] Thereby, when the transmission device 700 performs data transmission using a method such as MMT, the buffer operation of the reception device 800 can be guaranteed.

[0695] Next, the reception method will be described with reference to FIG. 77. FIG. 77 is an operation flow (reception method) by the reception device.

[0696] First, the reception unit 801 of the reception device 800 receives a transmission packet composed of a fixed-length packet header and a variable-length payload (S2101).

[0697] Next, the first buffer 802 of the reception device 800 converts a packet composed of the variable-length packet header and the variable-length payload stored in the received transmission packet into a first packet having a fixed-length packet header with an extended header, and outputs the first packet obtained by the conversion at a constant bit rate (S2102).

[0698] Next, the second buffer 803 of the reception device 800 converts the first packet obtained by the conversion into a second packet composed of a packet header and a variable-length payload, and outputs the second packet obtained by the conversion at a constant bit rate (S2103).

[0699] Next, the third buffer 804 of the receiving device 800 converts the output second packet into NAL units, and outputs the NAL units obtained by the conversion at a constant bit rate (S2104).

[0700] Next, the fourth buffer 805 of the receiving device 800 sequentially accumulates the output NAL units, generates access units from the accumulated multiple NAL units, and outputs the generated access units to the decoder at the timing of the decoding time corresponding to the access units (S2105).

[0701] Then, the decoding unit 806 of the receiving device 800 decodes the access unit output by the fourth buffer (S2106).

[0702] Thereby, the receiving device 800 can perform a decoding operation without underflow or overflow.

[0703] (Embodiment 6) [Overview] In Embodiment 6, a transmission method and a reception method in the case where leap second adjustment is performed on the time information serving as a reference for the reference clock in the MMT / TLV transmission method will be described.

[0704] FIG. 78 is a diagram showing a protocol stack of the MMT / TLV method defined in ARIB STD-B60.

[0705] In the MMT method, for packets, data such as video and audio is stored for each predetermined data unit such as a plurality of MPUs (Media Presentation Units) and MFUs (Media Fragment Units), and an MMTP packet header is added to generate an MMTP packet as a predetermined packet (MMTP packetization). Also, for control information such as control messages in MMTP, an MMTP packet header is added to generate an MMTP packet as a predetermined packet. The MMTP packet header is provided with a field for storing a 32-bit short format NTP (Network Time Protocol: specified in IETF RFC 5905), which can be used for QoS control of the communication line and the like.

[0706] Also, the reference clock of the transmission side (transmission device) is synchronized with the 64-bit long format NTP specified in RFC 5905, and based on the synchronized reference clock, time stamps such as PTS (Presentation Time Stamp) and DTS (Decode Time Stamp) are added to the synchronized media. Further, the reference clock information of the transmission side is transmitted to the reception side, and the reception device generates a system clock in the reception device based on the reference clock information received from the transmission side.

[0707] Specifically, PTS and DTS are stored in the MPU time stamp descriptor and the MPU extended time stamp descriptor, which are control information of MMTP, and are stored in the MP table for each asset. After being packetized into MMTP packets as control messages, they are transmitted.

[0708] The data packetized into MMTP is encapsulated into an IP packet with a UDP header and an IP header added. At this time, in the IP header and the UDP header, a set of packets with the same source IP address, destination IP address, source port number, destination port number, and protocol type is defined as an IP data flow. Note that since the headers of the IP packets in the same IP data flow are redundant, the headers are compressed in some IP packets.

[0709] Also, as reference clock information, a 64-bit NTP timestamp is stored in an NTP packet and then stored in an IP packet. At this time, in the IP packet storing the NTP packet, the source IP address, destination IP address, source port number, destination port number, and protocol type are fixed values, and the header of the IP packet is not compressed.

[0710] FIG. 79 is a diagram showing the configuration of a TLV packet.

[0711] As shown in FIG. 79, a TLV packet can include, as data, an IP packet, a compressed IP packet, and transmission control information such as an AMT (Address Map Table) and an NIT (Network Information Table). These data are identified using an 8-bit data type. Also, in the TLV packet, the data length (in bytes) is indicated using a 16-bit field, and then the value of the data is stored. Further, the TLV packet has 1-byte header information before the data type, and the header information is stored in a total 4-byte header area. Also, the TLV packet is mapped to a transmission slot in the advanced BS transmission method, and mapping information is stored in TMCC (Transmission and Multiplexing Configuration Control) control information.

[0712] FIG. 80 is a diagram showing an example of a block diagram of a receiving device.

[0713] In the receiving device, first, for the broadcast signal received by the tuner, the demodulation means decodes the transmission path encoded data, performs error correction, etc., and extracts the TLV packet. Then, the TLV / IP DEMUX means performs the DEMUX process of TLV and the DEMUX process of IP. The DEMUX process of TLV performs processing according to the data type of the TLV packet. For example, when the TLV packet has a compressed IP packet, the compressed header of the compressed IP packet is restored. In the IP DEMUX, processing such as header analysis of the IP packet and UDP packet is performed, and the MMTP packet and NTP packet are extracted.

[0714] The NTP clock generation means regenerates the NTP clock from the extracted NTP packet. The MMTP DEMUX performs filtering processing of components such as video and audio and control information based on the packet ID stored in the extracted MMTP packet header. The control information acquisition means acquires the time stamp descriptor stored in the MP table, and the PTS / DTS calculation means calculates the PTS and DTS for each access unit. Note that the time stamp descriptor includes both the MPU time stamp descriptor and the MPU extended time stamp descriptor.

[0715] The access unit reproduction means converts the video, audio, etc. filtered from the MMTP packet into data in units for presentation. Specifically, the data in units for presentation are, for example, the NAL unit of the video signal, the access unit, the audio frame, the presentation unit of the subtitle, etc. The decoding and presentation means decodes and presents the access unit at the time when the PTS / DTS of the access unit matches based on the reference time information of the NTP clock.

[0716] Note that the configuration of the receiving device is not limited to this.

[0717] Next, the time stamp descriptor will be described.

[0718] FIG. 81 is a diagram for explaining the time stamp descriptor.

[0719] PTS and DTS are stored in the MPU timestamp descriptor as the first control information and the MPU extended timestamp descriptor as the second control information, which are control information of MMT. They are stored in the MP table for each asset, and after being packetized into MMTP packets as control messages, they are transmitted.

[0720] Figure 81(a) is a diagram showing the configuration of the MPU timestamp descriptor defined in ARIB STD - B60. In the MPU timestamp descriptor, for each of a plurality of MPUs, the PTS (absolute value indicated by 64 - bit NTP) of the head (first) AU (hereinafter referred to as the "head AU") in the order of presentation among the plurality of AUs included in the MPU is stored. That is, the presentation time information of the MPU given to the MPU is stored in the control information of the MMTP packet and transmitted.

[0721] Figure 81(b) shows the configuration of the MPU extended timestamp descriptor. In the MPU extended timestamp descriptor, information for calculating the PTS and DTS of the AUs included in each of the plurality of MPUs is stored. The MPU extended timestamp descriptor includes relative information from the PTS of the head AU of the MPU stored in the MPU timestamp descriptor, and the PTS and DTS of the AUs included in the MPU can be calculated based on both the MPU timestamp descriptor and the MPU extended timestamp descriptor. That is, the PTS and DTS of the AUs other than the head AU included in the MPU can be calculated based on the PTS of the head AU stored in the MPU timestamp descriptor and the relative information stored in the MPU extended timestamp descriptor.

[0722] NTP is reference time information based on Coordinated Universal Time (UTC). UTC performs leap second adjustment (hereinafter referred to as "leap second adjustment") to adjust the difference from astronomical time based on the earth's rotation speed. Specifically, the leap second adjustment is carried out at 9:00 am in Japanese time, and it is an adjustment to insert or delete 1 second.

[0723] FIG. 82 is a diagram for explaining leap second adjustment.

[0724] (a) of FIG. 82 is a diagram showing an example of leap second insertion in Japanese time. As shown in (a) of FIG. 82, in leap second insertion, after 8:59:59 in Japanese time, it becomes 8:59:59 at the timing when it would originally be 9:00:00, and the 8:59:59 stage is repeated twice.

[0725] (b) of FIG. 82 is a diagram showing an example of leap second deletion in Japanese time. As shown in (b) of FIG. 82, in leap second deletion, after 8:59:58 in Japanese time, it becomes 9:00:00 at the timing when it would originally be 8:59:59, and one second of the 8:59:59 stage is deleted.

[0726] In the NTP packet, in addition to the 64-bit timestamp, a 2-bit leap_indicator is stored. The leap_indicator is a flag for notifying in advance that leap second adjustment is to be performed. When leap_indicator = 1, it indicates leap second insertion, and when leap_indicator = 2, it indicates leap second deletion. The advance notification can be done in ways such as starting from the beginning of the month when leap second adjustment is to be carried out, notifying 24 hours in advance, or starting the notification at any other time. Also, the leap_indicator becomes 0 at the time (9:00:00) when leap second adjustment ends. For example, when notifying 24 hours in advance, from 9:00 on the day before leap second adjustment is carried out in Japanese time until the moment immediately before leap second adjustment is carried out on the day when leap second adjustment is carried out (that is, in the case of leap second insertion, it is the time of the first 8:59:59 stage, and in the case of leap second deletion, it is the time of the 8:59:58 stage), it is shown that the leap_indicator is "1" or "2".

[0727] Next, the problems during leap second adjustment will be described.

[0728] FIG. 83 is a diagram showing the relationship among the NTP time, the MPU timestamp, and the MPU presentation timing. Note that the NTP time is the time indicated by NTP. Also, the MPU timestamp is the timestamp indicating the PTS of the first AU in the MPU. Further, the MPU presentation timing is the timing at which the receiving device should present the MPU according to the MPU timestamp. Specifically, FIGS. 83(a) to (c) are diagrams showing the relationships among the NTP time, the MPU timestamp, and the MPU presentation time in the cases where no leap second adjustment occurs, where a leap second is inserted, and where a leap second is deleted, respectively.

[0729] Here, a case where the NTP time (reference clock) on the transmission side is synchronized with the NTP server and the NTP time (system clock) on the receiving side is synchronized with the NTP time on the transmission side will be described as an example. In this case, the receiving device reproduces based on the timestamp stored in the NTP packet transmitted from the transmission side. Also, in this case, since both the NTP time on the transmission side and the NTP time on the receiving side are synchronized with the NTP server, an adjustment of ±1 second is made at the time of leap second adjustment. Also, it is assumed that the NTP time in FIG. 83 is common to the NTP time on the transmission side and the NTP time on the receiving side. Note that the description will be made assuming no transmission delay.

[0730] The MPU timestamp in FIG. 83 indicates the timestamp of the first AU in the order of presentation among the plurality of AUs included in each of the plurality of MPUs, and is generated (set) based on the NTP time indicated at the origin of the arrow. Specifically, the MPU presentation time information is generated by adding a predetermined time (for example, 1.5 seconds in FIG. 83) to the NTP time as the reference time information at the timing of generating the MPU presentation time information. The generated MPU timestamp is stored in the MPU timestamp descriptor.

[0731] The receiving device presents the MPU at the MPU presentation time based on the timestamp stored in the MPU timestamp descriptor.

[0732] In addition, in FIG. 83, the reproduction time of one MPU is described as 1 second, but the reproduction time of one MPU may be other reproduction times. For example, it may be 0.5 second or 0.1 second.

[0733] In the example of FIG. 83(a), the receiving device can present MPU#1-#5 in order based on the time stamps stored in the MPU time stamp descriptor.

[0734] However, in FIG. 83(b), the presentation times of MPU#2 and MPU#3 overlap due to leap second insertion. For this reason, if the receiving device presents the MPU based on the time stamps stored in the MPU time stamp descriptor, there will be two MPUs presented in the same time zone around 9:00:00, and it is impossible to determine which of the two MPUs should be presented. Also, since the MPU presentation time (8:59:59) indicated by the time stamp of MPU#1 exists twice in the receiving device, it is impossible to determine which of the two MPU presentation times should be used for presentation.

[0735] In addition, in FIG. 83(c), since the MPU presentation time (8:59:59) indicated by the MPU time stamp of MPU#3 does not exist in the NTP time due to the deletion of the leap second, the receiving device cannot present MPU#3.

[0736] When the receiving device performs the decoding process and presentation process of the MPU not based on the time stamp, the above problems can be solved. However, it is difficult for a receiving device that performs processing based on the time stamp to perform different processing (processing not based on the time stamp) only when a leap second occurs.

[0737] Next, a method for solving the problems during leap second adjustment by correcting the time stamp on the transmission side will be described.

[0738] FIG. 84 is a diagram for explaining a correction method for correcting a time stamp on the transmission side. Specifically, FIG. 84(a) shows an example of leap second insertion, and FIG. 84(b) is a diagram showing an example of leap second deletion.

[0739] First, the case of leap second insertion will be described.

[0740] As shown in FIG. 84(a), at the time of leap second insertion, the time up to immediately before the leap second insertion (that is, up to the first 8:59:59 range in NTP time) is defined as area A, and the time after the leap second insertion (that is, after the second 8:59:59 in NTP time) is defined as area B. Note that area A and area B are temporal areas, which are time zones or periods. The MPU time stamp in FIG. 84(a) is the same as the time stamp described in FIG. 83, and is a time stamp generated (set) based on the NTP time at the timing of assigning the MPU time stamp.

[0741] The method for correcting the time stamp on the transmission side in the case of leap second insertion will be specifically described.

[0742] On the transmission side (transmission device), the following processing is performed.

[0743] 1. When the timing of assigning the MPU time stamp (the timing indicated by the origin of the arrow in FIG. 84(a)) is included in area A and the MPU time stamp (the value of the MPU time stamp before correction) indicates after 9:00:00, the MPU time stamp is decremented by 1 second (-1 second correction), and the corrected MPU time stamp is stored in the MPU time stamp descriptor. That is, for an MPU time stamp generated based on the NTP time included in area A and when the MPU time stamp indicates after 9:00:00, the MPU time stamp is corrected by -1 second. Here, "9:00:00" is the time corresponding to the reference time for leap second adjustment in Japanese time (that is, the time calculated by adding 9 hours to the UTC time). Also, correction information indicating that correction has been performed separately is transmitted to the receiving device.

[0744] 2. When the timing for attaching the MPU timestamp (the timing indicated by the origin of the arrow in Fig. 84(a)) is in the B region, the MPU timestamp is not corrected. That is, if it is an MPU timestamp generated based on the NTP time included in the B region, the MPU timestamp is not corrected.

[0745] In the receiving device, the MPU is presented based on the MPU timestamp and correction information indicating whether the MPU timestamp has been corrected (that is, whether information indicating correction is included).

[0746] When the receiving device determines that the MPU timestamp has not been corrected (that is, when it determines that information indicating correction is not included), the MPU is presented at the time when the timestamp stored in the MPU timestamp descriptor matches the NTP time (including both before and after correction) of the receiving device. That is, in the case of the MPU timestamp stored in the MPU timestamp descriptor of an MPU transmitted before the MPU whose MPU timestamp is corrected, the MPU is presented at the timing when the MPU timestamp matches the NTP time before leap second insertion (that is, before the first 8:59:59). Also, in the case of the MPU timestamp stored in the MPU timestamp descriptor of the received MPU being a timestamp transmitted after the MPU timestamp to be corrected, the MPU is presented at the timing when it matches the NTP time after leap second insertion (that is, after the second 8:59:59).

[0747] Also, when the MPU timestamp stored in the MPU timestamp descriptor of the received MPU is corrected, the receiving device presents the MPU based on the timestamp stored in the MPU timestamp descriptor at the NTP time after leap second insertion (that is, after the second 8:59:59) in the receiving device.

[0748] Note that the information indicating that the MPU timestamp value has been corrected is stored and transmitted in control messages, descriptors, tables, MPU metadata, MF metadata, MMTP packet headers, etc.

[0749] Next, the case of leap second deletion will be described.

[0750] As shown in Fig. 84(b), at the time of leap second deletion, the time up to immediately before the leap second deletion (that is, immediately before 9:00:00 in NTP time) is defined as area C, and the time after the leap second deletion (that is, after 9:00:00 in NTP time) is defined as area D. Note that area C and area D are temporal areas, which are time zones or periods. The MPU timestamp in Fig. 84(b) is the same as the timestamp described in Fig. 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is assigned.

[0751] The method for correcting the timestamp on the transmission side in the case of leap second deletion will be specifically described.

[0752] On the transmission side (transmission device), the following processing is performed.

[0753] 1. When the timing for adding the MPU timestamp (the timing indicated by the origin of the arrow in Fig. 84(b)) is included in region C and the MPU timestamp (the value of the MPU timestamp before correction) indicates after 8:59:59, add 1 second to the MPU timestamp for +1 second correction and store the corrected MPU timestamp in the MPU timestamp descriptor. That is, for an MPU timestamp generated based on the NTP time included in region C and when the MPU timestamp indicates after 8:59:59, correct the MPU timestamp by +1 second. Here, "8:59:59" is the time obtained by subtracting 1 second from the time corresponding to the reference time for leap second adjustment in Japanese time (i.e., the time calculated by adding 9 hours to the UTC time). Also, transmit correction information, which is information indicating that separate correction has been performed, to the receiving device. Note that in this case, the correction information does not necessarily have to be transmitted.

[0754] 2. When the timing for adding the MPU timestamp (the timing indicated by the origin of the arrow in Fig. 84(b)) is in region D, do not correct the MPU timestamp. That is, for an MPU timestamp generated based on the NTP time included in region D, do not correct the MPU timestamp.

[0755] In the receiving device, present the MPU based on the MPU timestamp. If there is correction information indicating whether the MPU timestamp has been corrected, the MPU may be presented based on the MPU timestamp and the correction information.

[0756] Through the above processing, even when leap second adjustment is performed on the NTP time, in the receiving device, it is possible to normally present the MPU using the MPU timestamp stored in the MPU timestamp descriptor.

[0757] Note that it may be signaled and notified to the receiving side whether the timing of MPU timestamp assignment is in area A, area B, area C, or area D. That is, it may be signaled and notified to the receiving side whether the MPU timestamp was generated based on the NTP time included in area A, the NTP time included in area B, the NTP time included in area C, or the NTP time included in area D. In other words, identification information indicating whether the MPU timestamp (presentation time) was generated based on the reference time information (NTP time) before leap second adjustment may be transmitted. Since this identification information is assigned based on the leap_indicator included in the NTP packet, it is from 9:00 on the day before the leap second adjustment in Japanese time to the time immediately before the leap second adjustment on the day of the leap second adjustment (that is, in the case of leap second insertion, it is the time in the 8:59:59 range for the first time, and in the case of leap second deletion, it is the time in the 8:59:58 range). That is, the identification information is information indicating whether the MPU timestamp was generated based on the NTP time from a time a predetermined period (for example, 24 hours) before the time immediately before the leap second adjustment to the time immediately before the leap second adjustment.

[0758] Next, a method for solving the problems during leap second adjustment by correcting the timestamp in the receiving device will be described.

[0759] FIG. 85 is a diagram for explaining a correction method for correcting a timestamp in a receiving device. Specifically, FIG. 85(a) shows an example of leap second insertion, and FIG. 85(b) shows an example of leap second deletion.

[0760] First, the case of leap second insertion will be described.

[0761] As shown in Fig. 85(a), at the time of leap second insertion, similar to Fig. 84(a), the time up to immediately before leap second insertion (that is, up to the first 8:59:59 range in NTP time) is defined as Area A, and the time after leap second insertion (that is, after the second 8:59:59 in NTP time) is defined as Area B. Note that Area A and Area B are temporal areas, which are time zones or periods. The MPU timestamp in Fig. 85(a) is the same as the timestamp described in Fig. 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is given.

[0762] The method for correcting the timestamp at the receiving device in the case of leap second insertion will be specifically described.

[0763] On the transmitting side (transmitting device), the following processing is performed.

[0764] · Without correcting the generated MPU timestamp, store the MPU timestamp in the MPU timestamp descriptor and transmit it to the receiving device.

[0765] · Transmit to the receiving device, as identification information, information indicating whether the timing at which the MPU timestamp is given is in Area A or Area B. That is, transmit to the receiving device identification information indicating whether the MPU timestamp is generated based on the NTP time included in Area A or the NTP time included in Area B.

[0766] On the receiving device, the following processing is performed.

[0767] In the receiving device, the MPU timestamp is corrected based on the MPU timestamp and the identification information indicating whether the timing at which the MPU timestamp is given is in Area A or Area B.

[0768] Specifically, the following processing is performed.

[0769] 1. When the timing for assigning the MPU timestamp (the timing indicated by the origin of the arrow in Fig. 85(a)) is included in Area A and the MPU timestamp (the value of the MPU timestamp before correction) indicates a time after 9:00:00, the MPU timestamp is decremented by 1 second, i.e., corrected by -1 second. That is, for an MPU timestamp generated based on the NTP time included in Area A and when the MPU timestamp indicates a time after 9:00:00, the MPU timestamp is corrected by -1 second. Here, "9:00:00" refers to the time corresponding to the reference time for leap second adjustment in Japanese time (i.e., the time calculated by adding 9 hours to the UTC time).

[0770] 2. When the timing for assigning the MPU timestamp (the timing indicated by the origin of the arrow in Fig. 85(a)) is in Area B, the MPU timestamp is not corrected. That is, for an MPU timestamp generated based on the NTP time included in Area B, the MPU timestamp is not corrected.

[0771] When the MPU timestamp is not corrected, the MPU is presented at the time when the MPU timestamp stored in the MPU timestamp descriptor matches the NTP time (including before and after correction) of the receiving device.

[0772] That is, for an MPU timestamp received before the MPU timestamp to be corrected, the MPU is presented at the time that matches the NTP time before leap second insertion (i.e., before the first 8:59:59). Also, for an MPU timestamp received after the MPU timestamp to be corrected, the MPU is presented at the time that matches the NTP time after leap second insertion (after the second 8:59:59).

[0773] When the MPU timestamp is corrected, the MPU is presented based on the NTP time after leap second insertion (i.e., after the second 8:59:59) in the receiving device for the corrected MPU timestamp.

[0774] Note that the transmitting side stores the identification information indicating whether the timing when the MPU timestamp is added is in area A or area B in control messages, descriptors, tables, MPU metadata, MF metadata, MMTP packet headers, etc., and transmits it.

[0775] Next, the case of leap second deletion will be described.

[0776] As shown in Fig. 85(b), at the time of leap second deletion, similar to Fig. 84(a), the time until immediately before the leap second deletion (i.e., until immediately before 9:00:00 in NTP time) is set as area C, and the time after the leap second deletion (i.e., after 9:00:00: in NTP time) is set as area D. Note that areas C and D are temporal areas, i.e., time zones or periods. The MPU timestamp in Fig. 85(b) is the same as the timestamp described in Fig. 83, and is a timestamp generated (set) based on the NTP time at the timing when the MPU timestamp is added.

[0777] The method for correcting the timestamp at the receiving device in the case of leap second deletion will be specifically described.

[0778] The transmitting side (transmitting device) performs the following processing.

[0779] · Without correcting the generated MPU timestamp, store the MPU timestamp in the MPU timestamp descriptor and transmit it to the receiving device.

[0780] · Transmit to the receiving device, as identification information, the information indicating whether the timing when the MPU timestamp is added is in area C or area D. That is, transmit to the receiving device the identification information indicating whether the MPU timestamp is generated based on the NTP time included in area C or the NTP time included in area B.

[0781] The receiving device performs the following processing.

[0782] In the receiving device, the MPU timestamp is corrected based on the MPU timestamp and the identification information indicating whether the timing at which the MPU timestamp is added is in area C or area D.

[0783] Specifically, the following processing is performed.

[0784] 1. When the timing at which the MPU timestamp is added (the timing indicated by the origin of the arrow in (b) of FIG. 85) is included in area C and the MPU timestamp (the value of the MPU timestamp before correction) indicates after 8:59:59, the MPU timestamp value is incremented by 1 second for +1 second correction. That is, for an MPU timestamp generated based on the NTP time included in area C and when the MPU timestamp indicates after 8:59:59, the MPU timestamp is corrected by +1 second. Here, "8:59:59" is the time obtained by subtracting 1 second from the time corresponding to the time serving as the reference for leap second adjustment in Japanese time (that is, the time calculated by adding 9 hours to the UTC time).

[0785] 2. When the timing at which the MPU timestamp is added (the timing indicated by the origin of the arrow in (b) of FIG. 85) is in area D, the MPU timestamp is not corrected. That is, for an MPU timestamp generated based on the NTP time included in area D, the MPU timestamp is not corrected.

[0786] In the receiving device, the MPU is presented based on the MPU timestamp and the corrected MPU timestamp.

[0787] By the above processing, even when leap second adjustment is performed on the NTP time, in the receiving device, it is possible to normally present the MPU using the MPU timestamp stored in the MPU timestamp descriptor.

[0788] Even in this case, similar to the case of correcting the MPU timestamp on the transmission side described with reference to FIG. 84, it is also possible to signal whether the timing of assigning the MPU timestamp is in region A, region B, region C, or region D, and notify the receiving side. Since the details of the notification are the same, the description thereof is omitted.

[0789] Note that additional information (identification information) such as information indicating whether the MPU timestamp has been corrected and information indicating whether the timing of assigning the MPU timestamp is in region A, region B, region C, or region D, as described with reference to FIGS. 84 and 85, may be made effective when the leap_indicator of the NTP packet indicates deletion (leap_indicator = 2) or insertion (leap_indicator = 1) of a leap second. It may be made effective from a predetermined arbitrary time (for example, 3 seconds before the leap second adjustment), or may be made effective dynamically.

[0790] Also, the timing at which the validity period of the additional information ends and the timing at which the signaling of the additional information ends may be adjusted in accordance with the leap_indicator of the NTP packet, may be made effective from a predetermined arbitrary time (for example, 3 seconds before the leap second adjustment), or may be made effective dynamically.

[0791] Note that since the 32-bit timestamp stored in the MMTP packet and the timestamp information stored in the TMCC are also generated and assigned based on the NTP time, similar problems occur. Therefore, in the case of the 32-bit timestamp and the timestamp information stored in the TMCC, the timestamp can be corrected using the same method as in FIGS. 84 and 85, and the receiving device can perform processing based on the timestamp. For example, when storing the additional information for the 32-bit timestamp stored in the MMTP packet, it may be indicated using the extended area of the MMTP packet header. In this case, it is indicated that it is additional information in the extended type of the multi-header type. Also, for the time when the leap_indicator is set, additional information may be indicated using some bits of the 32-bit or 64-bit timestamp.

[0792] Note that in the example of FIG. 84, it was assumed that the corrected MPU timestamp was stored in the MPU timestamp descriptor, but both the pre-correction and post-correction MPU timestamps (i.e., the uncorrected MPU timestamp and the corrected MPU timestamp) may be transmitted to the receiving device. For example, the pre-correction MPU timestamp descriptor and the post-correction MPU timestamp descriptor may be stored in the same MPU timestamp descriptor, or they may be stored in two separate MPU timestamp descriptors. In this case, it may be possible to identify whether it is the pre-correction MPU timestamp or the post-correction MPU timestamp based on the arrangement order of the two MPU timestamp descriptors or the description order within the MPU timestamp descriptor. Also, it may be assumed that the pre-correction MPU timestamp is always stored in the MPU timestamp descriptor, and when correction is performed, the post-correction timestamp may be stored in the MPU extended timestamp descriptor.

[0793] In this embodiment, the time (9:00 am) in Japanese time is used as an example of the NTP time for explanation, but it is not limited to the time in Japanese time. The leap second adjustment is based on the UTC time and is corrected simultaneously worldwide. Japanese time is the time advanced by 9 hours based on the UTC time and is represented by the value of the time (+9) relative to the UTC time. Thus, it is also possible to adopt a time adjusted to different times according to the time difference depending on the location.

[0794] As described above, by correcting the time stamp at the transmitting side or the receiving device based on the information indicating the timing to which the time stamp is attached, normal reception processing using the time stamp becomes possible.

[0795] Note that a receiving device that continues the decoding process sufficiently before the leap second adjustment time may be able to perform decoding processing and presentation processing without using a time stamp. However, a receiving device that tunes in immediately before the leap second adjustment time may not be able to determine the time stamp and may not be able to present it until after the leap second adjustment is completed. Even in such a case, by using the correction method in this embodiment, reception processing using the time stamp becomes possible, and tuning in is also possible immediately before the leap second adjustment time.

[0796] The operation flow of the transmitting side (transmitting device) when correcting the MPU time stamp at the transmitting side (transmitting device) described with reference to FIG. 84 is shown in FIG. 86, and the operation flow of the receiving device is shown in FIG. 87.

[0797] First, the operation flow of the transmitting side (transmitting device) will be described with reference to FIG. 86.

[0798] When insertion or deletion of a leap second is performed, it is determined whether the timing of attaching the MPU time stamp is in area A, area B, area C, or area D (S2201). Note that the case where no insertion or deletion of a leap second is performed is not shown.

[0799] Here, areas A to D are defined as follows.

[0800] Area A: Time until just before the leap second is inserted (until the first around 8:59:59) Area B: Time after the leap second is inserted (after the second around 8:59:59) Area C: Time until just before the leap second is deleted (until 9:00:00) Area D: Time after the leap second is deleted (after 9:00:00)

[0801] In step S2201, if it is determined that it is Area A and the MPU timestamp indicates after 9:00:00, correct the MPU timestamp by -1 second and store the corrected MPU timestamp in the MPU timestamp descriptor (S2202).

[0802] Then, signal correction information indicating that the MPU timestamp has been corrected and transmit it to the receiving device (S2203).

[0803] Also, in step S2201, if it is determined that it is Area C and the MPU timestamp indicates after 8:59:59, correct the MPU timestamp by +1 second and store the corrected MPU timestamp in the MPU timestamp descriptor (S2205).

[0804] Also, in step S2201, if it is determined that it is Area B or Area D, store the MPU timestamp in the MPU timestamp descriptor without correction (S2204).

[0805] Next, the operation flow of the receiving device will be described with reference to FIG. 87.

[0806] Based on the information signaled from the transmitting side, determine whether the MPU timestamp has been corrected (S2301).

[0807] If it is determined that the MPU timestamp has been corrected (Yes in S2301), the MPU is presented based on the MPU timestamp at the NTP time after the leap second adjustment is performed in the receiving device (S2302).

[0808] If it is determined that the MPU timestamp has not been corrected (No in S2301), the MPU is presented based on the MPU timestamp at the NTP time.

[0809] Note that when correcting the MPU timestamp, the MPU corresponding to the MPU timestamp is presented in the section where leap seconds are inserted.

[0810] Conversely, when not correcting the MPU timestamp, the MPU corresponding to the MPU timestamp is not presented in the section where leap seconds are inserted, but in the section where no leap seconds are inserted.

[0811] The operation flow on the transmission side when correcting the MPU timestamp in the receiving device, as described in FIG. 85, is shown in FIG. 88, and the operation flow of the receiving device is shown in FIG. 89.

[0812] First, the operation flow of the transmission side (transmission device) will be described with reference to FIG. 88.

[0813] Based on the leap_indicator of the NTP packet, it is determined whether leap second adjustment (insertion or deletion) is performed (S2401).

[0814] If it is determined that leap second adjustment is performed (Yes in S2401), the timing of attaching the MPU timestamp is determined, identification information is signaled, and transmitted to the receiving device (S2402).

[0815] On the other hand, if it is determined that no leap second adjustment is performed (No in S2041), the process ends as normal operation.

[0816] Next, the operation flow of the receiving device will be described with reference to FIG. 89.

[0817] Based on the identification information signaled by the transmission side (transmission device), it is determined whether the timing of assigning the MPU timestamp is in area A, area B, area C, or area D (S2501). Here, since areas A to D are the same as those defined above, the description thereof is omitted. Note that, as in FIG. 87, cases where leap seconds are not inserted or deleted are not illustrated.

[0818] In step S2501, if it is determined that it is in area A and the MPU timestamp indicates after 9:00:00, the MPU timestamp is corrected by -1 second (S2502).

[0819] In step S2501, if it is determined that it is in area C and the MPU timestamp indicates after 8:59:59, the MPU timestamp is corrected by +1 second (S2504).

[0820] In step S2501, if it is determined that it is in area B or area D, the MPU timestamp is not corrected (S2503).

[0821] The receiving device further presents the MPU based on the corrected MPU timestamp in a process (not shown).

[0822] Note that when correcting the MPU timestamp, the MPU corresponding to the MPU timestamp is presented based on the MPU timestamp at the NTP time after the leap second adjustment.

[0823] Conversely, when not correcting the MPU timestamp, the MPU corresponding to the MPU timestamp is presented in a section where no leap second is inserted, and not in a section where a leap second is inserted.

[0824] In short, the transmitting side (transmitting device) determines, for each of the plurality of MPUs, the timing at which to assign an MPU timestamp corresponding to the MPU. Then, as a result of the determination, if the timing is a time up to immediately before the insertion of a leap second and the MPU timestamp indicates after 9:00:00, the MPU timestamp is corrected by -1 second. Further, correction information indicating that the MPU timestamp has been corrected is signaled and transmitted to the receiving device. Also, as a result of the determination, if the timing is a time up to immediately before the deletion of a leap second and the MPU timestamp indicates after 8:59:59, the MPU timestamp is corrected by +1 second.

[0825] Further, the receiving device presents the MPU based on the MPU timestamp at the NTP time after the leap second adjustment when it is indicated that the MPU timestamp is corrected based on the correction information signaled by the transmitting side (transmitting device). When it is not indicated that the MPU timestamp is corrected, the MPU is presented based on the MPU timestamp at the NTP time before the leap second adjustment.

[0826] Also, the transmitting side (transmitting device) determines, for each of the plurality of MPUs, the timing at which to assign an MPU timestamp corresponding to the MPU, and signals the determination result. The receiving device performs the following processing based on the information indicating the timing at which the MPU timestamp is assigned, signaled by the transmitting side. Specifically, when the information indicating the timing indicates a time up to immediately before the insertion of a leap second and the MPU timestamp indicates after 9:00:00, the MPU timestamp is corrected by -1 second. Also, when the information indicating the timing is a time up to immediately before the deletion of a leap second and the MPU timestamp indicates after 8:59:59, the MPU timestamp is corrected by +1 second.

[0827] When the MPU timestamp is corrected, the receiving device presents the MPU based on the MPU timestamp at the NTP time after the leap second adjustment. When the MPU timestamp is not corrected, the receiving device presents the MPU based on the MPU timestamp at the NTP time before the leap second adjustment.

[0828] In this way, by determining the timing of assigning the MPU timestamp in the leap second adjustment, the MPU timestamp can be corrected, and the receiving device can determine which MPU to present, enabling appropriate reception processing using the MPU timestamp descriptor and the MPU extended timestamp descriptor. That is, even when a leap second adjustment is made to the NTP time, the receiving device can present a normal MPU using the MPU timestamp stored in the MPU timestamp descriptor.

[0829] [Supplementary Note: Transmitting Device and Receiving Device] As described above, the transmitting device that stores the data constituting the encoded stream in the MPU and transmits it can also be configured as shown in FIG. 90. Further, the receiving device that receives the MPU in which the data constituting the encoded stream is stored can also be configured as shown in FIG. 91. FIG. 90 is a diagram showing an example of the specific configuration of the transmitting device. FIG. 91 is a diagram showing an example of the specific configuration of the receiving device.

[0830] The transmitting device 900 includes a generation unit 901 and a transmission unit 902. Each of the generation unit 901 and the transmission unit 902 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0831] The receiving device 1000 includes a reception unit 1001 and a reproduction unit 1002. Each of the reception unit 1001 and the reproduction unit 1002 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.

[0832] A detailed description of each component of the transmission device 900 and the reception device 1000 will be given in the descriptions of the transmission method and the reception method, respectively.

[0833] First, the transmission method will be described with reference to FIG. 92. FIG. 92 shows the operation flow (transmission method) by the transmission device.

[0834] First, the generation unit 901 of the transmission device 900 generates presentation time information (MPU timestamp) indicating the presentation time of the MPU as a predetermined data unit based on the NTP time as the reference time information received from the outside (S2601).

[0835] Next, the transmission unit 902 of the transmission device 900 transmits the MPU, the presentation time information generated by the generation unit 901, and identification information indicating whether or not the presentation time i...

Claims

1. A transmission method for storing data constituting an application in a predetermined data unit and transmitting the data, comprising the steps of: generating time designation information indicating an operation time of the application based on reference time information; (i) transmitting the predetermined data unit and (ii) control information indicating the generated time designation information; the control information stores the generated time designation information and identification information indicating whether the time designation information is time information before leap second adjustment; the time designation information of the predetermined data unit is time information obtained by adding a predetermined time to the time information, The time designation information is UTC (Universal Time Coordinated) time and NPT (Normal Play Time), The control information is a UTC-NPT reference descriptor that indicates the relationship between UTC time and NPT time. Transmission method.

2. A method for receiving a predetermined data unit in which data constituting an application is stored, comprising the steps of: (i) receiving the predetermined data unit; and (ii) control information storing time designation information indicating an operation time of the application and identification information indicating whether the time designation information is time information before leap second adjustment; Executing the application stored in the received predetermined data unit based on the received control information; the time designation information of the predetermined data unit is time information obtained by adding a predetermined time to the time information, The time designation information is UTC (Universal Time Coordinated) time and NPT (Normal Play Time), The control information is a UTC-NPT reference descriptor that indicates the relationship between UTC time and NPT time. Receiving method.

3. A transmitting device that transmits data constituting an application by storing the data in a predetermined data unit, a generating unit that generates time designation information indicating an operation time of the application based on reference time information; (i) a transmission unit that transmits the predetermined data unit and (ii) control information indicating the time designation information generated by the generation unit; the control information stores the time designation information generated by the generation unit and identification information indicating whether the time designation information is time information before leap second adjustment; the time designation information of the predetermined data unit is time information obtained by adding a predetermined time to the time information, The time designation information is UTC (Universal Time Coordinated) time and NPT (Normal Play Time), The control information is a UTC-NPT reference descriptor that indicates the relationship between UTC time and NPT time. Transmitting device.

4. A receiving device for receiving a predetermined data unit in which data constituting an application is stored, (i) a receiving unit that receives the predetermined data unit; and (ii) control information that stores time designation information that indicates an operation time of the application and identification information that indicates whether the time designation information is time information before leap second adjustment. Executing the application stored in the predetermined data unit received by the receiving unit based on the control information received by the receiving unit; the time designation information of the predetermined data unit is time information obtained by adding a predetermined time to the time information, The time designation information is UTC (Universal Time Coordinated) time and NPT (Normal Play Time), The control information is a UTC-NPT reference descriptor that indicates the relationship between UTC time and NPT time. Receiving device.

Citation Information

Patent Citations

  • Terminal device

    JP2010272071A

  • Self-positioning access point

    JP2011524725A

  • Leap second support in content timestamps

    US20140282791A1