Transmission method, reception method, transmission device, and reception device
The proposed method addresses the complexity and inefficiency of the MMT system for 8K video transmission by using MP4 configuration information and adaptive header information, resulting in reduced processing loads and optimized bandwidth usage.
Patent Information
- Application Number
- JP2024030300
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-08-04
- Filing Date
- 2024-02-29
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2035-07-10
AI Technical Summary
The existing MMT system for transmitting ultra-high definition video content, such as 8K, is complex and inefficient, leading to increased processing loads and unnecessary data transmission, which wastes bandwidth and complicates device configuration.
A transmission method that simplifies device configuration and reduces processing loads by using MP4 configuration information to reconfigure sample data on the receiving side, and by transmitting sample data with header information that includes or excludes MP4 configuration information based on the presentation time of the sample data.
This approach reduces the processing load and simplifies the configuration of transmitting and receiving devices, thereby optimizing bandwidth usage and ensuring efficient transmission of ultra-high definition video content.
Smart Images

Figure 0007672084000001 
Figure 0007672084000002 
Figure 0007672084000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a transmitting method, a receiving method, a transmitting device, and a receiving device. [Background technology]
[0002] As broadcasting and communication services become more advanced, the introduction of ultra-high definition video content such as 8K (7680×4320 pixels: hereinafter also referred to as 8K4K) and 4K (3840×2160 pixels: hereinafter also referred to as 4K2K) is being considered. A receiving device needs to decode the received encoded data of ultra-high definition video in real time and display it, but video with a resolution such as 8K in particular imposes a large processing load during decoding, making it difficult to decode such video in real time with a single decoder. Therefore, a method is being considered for parallelizing the decoding process using multiple decoders to reduce the processing load per decoder and achieve real-time processing.
[0003] Moreover, the encoded data is multiplexed based on a multiplexing method such as MPEG-2 TS (Transport Stream) or MMT (MPEG Media Transport) before being transmitted. For example, Non-Patent Document 1 discloses a technique for transmitting encoded media data packet by packet in accordance with MMT. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Information technology - High efficiency coding and media delivery in heterogeneous environment - Part1:MPEG media transport(MMT), ISO / IEC DIS 23008-1 Summary of the Invention [Problem to be solved by the invention]
[0005] MMT is a method that supports hybrid distribution using broadcasting and communication, and can transmit MP4 format media, and has various functions. However, if MMT is used to transmit data when playing back a broadcast stream on a receiving device, the configuration of the transmitting device and receiving device may become complex and the amount of processing may increase because MMT has more functions than necessary. In addition, transmission of unnecessary data may waste the transmission band.
[0006] The present invention provides a transmitting device and a receiving device that can simplify the device configuration and reduce the amount of processing performed by the device when transmitting data using a method such as MMT. [Means for solving the problem]
[0007] In order to achieve the above-mentioned object, a transmission method according to one embodiment of the present invention includes an assignment step of assigning header information to sample data, which is data in which a video signal or an audio signal is encoded, the header information including MP4 configuration information for reconstructing the sample data as an MP4 format file at the receiving side, the content of which differs depending on whether the presentation time of the sample data is specified or not, and a transmission step of transmitting the sample data to which the header information has been assigned, wherein in the assignment step, if metadata corresponding to the sample data is not transmitted in the transmission step, header information that does not include the MP4 configuration information is assigned to the sample data depending on whether the presentation time of the sample data is specified or not.
[0008] In addition, a receiving method according to one embodiment of the present invention includes a receiving step of receiving sample data, which is data in which a video signal or an audio signal is encoded, and which is provided with header information that does not include MP4 configuration information for reconstructing the sample data as a file in MP4 format, and a decoding step of decoding the sample data without using the MP4 configuration information if metadata corresponding to the sample data is not received in the receiving step and the presentation time of the sample data is determined.
[0009] Furthermore, these general or specific aspects may be realized by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. Effect of the Invention
[0010] The present invention can simplify the configuration of a device and reduce the amount of processing performed by the device when transmitting data using a method such as MMT. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing an example of dividing a picture into slice segments. [Diagram 2] FIG. 2 is a diagram showing an example of a PES packet sequence in which picture data is stored. [Diagram 3] FIG. 3 is a diagram showing an example of division of a picture according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of division of a picture according to a comparative example of the first embodiment. [Diagram 5] FIG. 5 is a diagram showing an example of data of an access unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a transmitting device according to the first embodiment. [Figure 7]FIG. 7 is a block diagram of a receiving device according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of an MMT packet according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating another example of the MMT packet according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of data input to each decoding unit according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 12] FIG. 12 is a diagram showing another example of data input to each decoding unit according to the first embodiment. In FIG. [Figure 13] FIG. 13 is a diagram showing an example of division of a picture according to the first embodiment. [Figure 14] FIG. 14 is a flowchart of the transmission method according to the first embodiment. [Figure 15] FIG. 15 is a block diagram of a receiving device according to the first embodiment. [Figure 16] FIG. 16 is a flowchart of the receiving method according to the first embodiment. [Figure 17] FIG. 17 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of an MMT packet and header information according to the first embodiment. [Figure 19] FIG. 19 is a diagram illustrating the configuration of an MPU. [Figure 20] FIG. 20 is a diagram showing the structure of the MF metadata. [Figure 21] FIG. 21 is a diagram for explaining the data transmission order. [Figure 22] FIG. 22 is a diagram showing an example of a method for performing decoding without using header information. [Diagram 23] FIG. 23 is a block diagram of a transmitting device according to the second embodiment. [Figure 24]FIG. 24 is a flowchart of a transmission method according to the second embodiment. [Diagram 25] FIG. 25 is a block diagram of a receiving device according to the second embodiment. [Figure 26] FIG. 26 is a flowchart of an operation for identifying an MPU start position and a NAL unit position. [Figure 27] FIG. 27 is a flowchart of an operation of obtaining initialization information based on a transmission order type and decoding media data based on the initialization information. [Figure 28] FIG. 28 is a flowchart of the operation of a receiving device when a low-delay presentation mode is provided. [Figure 29] FIG. 29 is a diagram illustrating an example of a transmission order of MMT packets when auxiliary data is transmitted. [Diagram 30] FIG. 30 is a diagram illustrating an example in which a transmitting device generates auxiliary data based on the configuration of moof. [Diagram 31] FIG. 31 is a diagram for explaining reception of auxiliary data. [Diagram 32] FIG. 32 is a flowchart of a receiving operation using auxiliary data. [Diagram 33] FIG. 33 is a diagram showing the configuration of an MPU that is made up of multiple movie fragments. [Diagram 34] FIG. 34 is a diagram for explaining the transmission order of MMT packets when the MPU having the configuration in FIG. 33 is transmitted. [Diagram 35] FIG. 35 is a first diagram for explaining an example of the operation of a receiving device when one MPU is composed of multiple movie fragments. [Diagram 36] FIG. 36 is a second diagram for explaining an example of the operation of the receiving device when one MPU is composed of multiple movie fragments. [Figure 37] FIG. 37 is a flowchart of the operation of the receiving method described in FIGS. [Figure 38]FIG. 38 is a diagram showing a case where non-VCL NAL units are individually treated as data units and aggregated. [Figure 39] FIG. 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. [Diagram 40] FIG. 40 is a flowchart showing the operation of the receiving device when a packet loss occurs. [Diagram 41] FIG. 41 is a flowchart of the receiving operation when the MPU is divided into multiple movie fragments. [Diagram 42] FIG. 42 is a diagram showing an example of a prediction structure of pictures for each TemporalId when implementing temporal scalability. [Diagram 43] FIG. 43 is a diagram showing the relationship between the decoding time (DTS) and the display time (PTS) in each picture in FIG. [Diagram 44] FIG. 44 is a diagram showing an example of a prediction structure of a picture that requires picture delay processing and reorder processing. [Diagram 45] Figure 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and an MMTP packet. [Diagram 46] FIG. 46 is a diagram for explaining the calculation method and problems of the PTS and DTS. [Figure 47] FIG. 47 is a flowchart of a receiving operation when the DTS is calculated using information for DTS calculation. [Figure 48] FIG. 48 is a diagram for explaining a method of storing a data unit in a payload in MMT. [Figure 49] FIG. 49 is an operational flow of the transmitting device according to the third embodiment. [Figure 50] FIG. 50 shows an operation flow of the receiving device according to the third embodiment. [Figure 51] FIG. 51 is a diagram illustrating an example of a specific configuration of a transmission device according to the third embodiment. In FIG. [Figure 52]FIG. 52 is a diagram illustrating an example of a specific configuration of a receiving device according to the third embodiment. In FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] A transmission method according to one embodiment of the present invention includes an assignment step of assigning header information to sample data, which is data in which a video signal or an audio signal is encoded, the header information including MP4 configuration information for reconstructing the sample data as an MP4 format file at the receiving side, the content of which differs depending on whether the presentation time of the sample data is specified or not, and a transmission step of transmitting the sample data to which the header information has been assigned.In the assignment step, if metadata corresponding to the sample data is not transmitted in the transmission step, header information that does not include the MP4 configuration information is assigned to the sample data depending on whether the presentation time of the sample data is specified or not.
[0013] This type of transmission method can simplify the device configuration and reduce the amount of processing required by the device when transmitting data using a method such as MMT.
[0014] In addition, in the assigning step, if metadata corresponding to the sample data is not transmitted in the transmitting step, and if the presentation time of the sample data is determined, header information that does not include the MP4 configuration information is assigned to the sample data, and if the presentation time of the sample data is not determined, header information that includes the MP4 configuration information is assigned to the sample data.
[0015] In addition, in the transmitting step, the sample data to which the header information has been added may be packetized in accordance with an MMT (MPEG Media Transport) standard and transmitted.
[0016] In addition, the header information assigned to the sample data for which a presentation time is defined may include at least one of movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter in an MMTP (MMT Protocol) payload as the MP4 configuration information, and the header information assigned to the sample data for which a presentation time is not defined may include item_id in an MP4 MMTP payload as the MP4 configuration information.
[0017] The sample data for which the presentation time is determined is called timed-MFU (Movie The sample data may be a non-timed MFU, and the presentation time of the sample data may not be determined.
[0018] The metadata may also include MPU (Media Processing Unit) metadata and movie fragment metadata.
[0019] In addition, a receiving method according to one embodiment of the present invention includes a receiving step of receiving sample data, which is data in which a video signal or an audio signal is encoded, and which is provided with header information that does not include MP4 configuration information for reconstructing the sample data as a file in MP4 format, and a decoding step of decoding the sample data without using the MP4 configuration information if metadata corresponding to the sample data is not received in the receiving step and the presentation time of the sample data is determined.
[0020] This type of receiving method can simplify the device configuration and reduce the amount of processing required by the device when transmitting data using a method such as MMT.
[0021] A transmitting device according to one embodiment of the present invention includes an assignment unit that assigns header information to sample data, which is data in which a video signal or an audio signal is encoded, the header information including MP4 configuration information for reconstructing the sample data as an MP4 format file at the receiving side, the content of which varies depending on whether the presentation time of the sample data is specified or not, and a transmitting unit that transmits the sample data to which the header information has been assigned.When metadata corresponding to the sample data is not transmitted by the transmitting unit, the assignment unit assigns header information that does not include the MP4 configuration information to the sample data, depending on whether the presentation time of the sample data is specified or not.
[0022] A receiving device according to one embodiment of the present invention includes a receiving unit that receives sample data, which is data in which a video signal or an audio signal has been encoded, and which is provided with header information that does not include MP4 configuration information for reconstructing the sample data as a file in MP4 format, and a decoding unit that decodes the sample data without using the MP4 configuration information when metadata corresponding to the sample data is not received by the receiving unit and the presentation time of the sample data has been determined.
[0023] These comprehensive or specific aspects may be realized in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or in any combination of the system, method, integrated circuit, computer program, or recording medium.
[0024] Hereinafter, the embodiment will be specifically described with reference to the drawings.
[0025] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.
[0026] (Findings on which the present invention is based) In recent years, the resolution of displays such as TVs, smartphones, and tablet devices has been increasing. In particular, 8K4K (resolution 8K x 4K) services are scheduled for broadcasting in Japan in 2020. Since it is difficult to decode ultra-high resolution moving images such as 8K4K in real time using a single decoder, methods of decoding in parallel using multiple decoders are being investigated.
[0027] Since the encoded data is multiplexed and transmitted based on a multiplexing method such as MPEG-2 TS or MMT, the receiving device needs to separate the encoded data of the moving image from the multiplexed data prior to decoding. Hereinafter, the process of separating the encoded data from the multiplexed data is called demultiplexing.
[0028] When parallelizing the decoding process, it is necessary to assign the encoded data to be decoded to each decoder. When assigning the encoded data, the encoded data itself needs to be analyzed, and since the bit rate is particularly high for 8K content, the processing load for analysis is large. Therefore, the demultiplexing part becomes a bottleneck, and there is an issue that real-time playback cannot be performed.
[0029] In video coding methods such as H.264 and H.265 standardized by MPEG and ITU, a transmitting device divides a picture into a plurality of areas called slices or slice segments, and can encode the divided areas so that each of the divided areas can be decoded independently. Therefore, for example, in the case of H.265, a receiving device that receives a broadcast can separate data for each slice segment from the received data and output the data for each slice segment to a separate decoder, thereby achieving parallel decoding processing.
[0030] 1 is a diagram showing an example of dividing one picture into four slice segments in HEVC. For example, a receiving device includes four decoders, and each decoder decodes one of the four slice segments.
[0031] In conventional broadcasting, a transmitting device stores one picture (access unit in the MPEG system standard) in one PES packet and multiplexes the PES packet into a TS packet sequence. Therefore, a receiving device needs to separate the payload of the PES packet, analyze the data of the access unit stored in the payload, separate each slice segment, and output the data of each separated slice segment to a decoder.
[0032] However, the present inventors have found that there is a problem in that it is difficult to perform this processing in real time because the amount of processing required to analyze access unit data and separate slice segments is large.
[0033] FIG. 2 is a diagram showing an example in which data of a picture divided into slice segments is stored in the payload of a PES packet.
[0034] 2, for example, data of a plurality of slice segments (slice segments 1 to 4) is stored in the payload of one PES packet, and the PES packet is multiplexed into a sequence of TS packets.
[0035] (Embodiment 1) In the following, an example will be described in which H.265 is used as the video encoding format, but this embodiment can also be applied to cases in which other encoding formats such as H.264 are used.
[0036] 3 is a diagram showing an example of dividing an access unit (picture) into division units in this embodiment. The access unit is divided into two equal parts horizontally and vertically into a total of four tiles by a function called a tile introduced by H.265. Also, slice segments and tiles are in one-to-one correspondence.
[0037] The reason for dividing the data into two equal parts horizontally and vertically will be explained below. First, when decoding, a line memory for storing one horizontal line of data is generally required, but when it comes to ultra-high resolution such as 8K4K, the horizontal size becomes large, so the size of the line memory increases. In implementing a receiving device, it is desirable to be able to reduce the size of the line memory. In order to reduce the size of the line memory, vertical division is required. A data structure called a tile is required for vertical division. For these reasons, tiles are used.
[0038] On the other hand, since images generally have high correlation in the horizontal direction, the coding efficiency improves if a wider range can be referenced in the horizontal direction. Therefore, from the viewpoint of coding efficiency, it is desirable to divide the access unit in the horizontal direction.
[0039] By dividing the access unit into two equal parts horizontally and vertically, these two characteristics can be achieved simultaneously, and both implementation and coding efficiency can be taken into consideration. If a single decoder can decode a 4K2K video image in real time, the receiving device can decode the 8K4K image in real time by dividing the 8K4K image into four equal parts and dividing each slice segment into 4K2K.
[0040] Next, the reason for establishing one-to-one correspondence between tiles obtained by dividing an access unit in the horizontal and vertical directions and slice segments will be described. In H.265, an access unit is composed of a plurality of units called NAL (Network Adaptation Layer) units.
[0041] The payload of a NAL unit stores any of the following: an access unit delimiter indicating the start position of an access unit, an SPS (Sequence Parameter Set) which is initialization information used commonly in sequence units during decoding, a PPS (Picture Parameter Set) which is initialization information used commonly in a picture during decoding, SEI (Supplemental Enhancement Information) which is not required for the decoding process itself but is required for processing and displaying the decoding result, and coded data of a slice segment. The header of a NAL unit includes type information for identifying the data stored in the payload.
[0042] Here, the transmitting device transmits the encoded data in MPEG-2 TS, MMT (MPEG Media Transport), MPEG DASH (Dynamic Adaptive Streaming) or other formats. When multiplexing using a multiplexing format such as HTTP (Streaming over HTTP) or RTP (Real-time Transport Protocol), the basic unit can be set to the NAL unit. In order to store one slice segment in one NAL unit, it is desirable to divide the access unit into slice segment units when dividing the access unit into regions. For this reason, the transmitting device sets a one-to-one correspondence between tiles and slice segments.
[0043] As shown in Fig. 4, the transmitting device can also set tiles 1 to 4 together in one slice segment. In this case, however, all tiles are stored in one NAL unit, making it difficult for the receiving device to separate the tiles in the multiplexing layer.
[0044] Note that there are two types of slice segments: independent slice segments that can be decoded independently, and reference slice segments that refer to independent slice segments. Here, a case will be described in which an independent slice segment is used.
[0045] Fig. 5 is a diagram showing an example of data of an access unit divided so that the boundaries between tiles and slice segments coincide as shown in Fig. 3. The data of the access unit includes a NAL unit in which an access unit delimiter placed at the beginning is stored, followed by NAL units of SPS, PPS, and SEI, and data of a slice segment in which data from tile 1 to tile 4 is stored. Note that the data of the access unit does not have to include some or all of the NAL units of SPS, PPS, and SEI.
[0046] Next, a configuration of the transmitting device 100 according to the present embodiment will be described. Fig. 6 is a block diagram showing an example of the configuration of the transmitting device 100 according to the present embodiment. This transmitting device 100 includes an encoding unit 101, a multiplexing unit 102, a modulating unit 103, and a transmitting unit 104.
[0047] The encoding unit 101 generates encoded data by encoding an input image according to, for example, H.265. In addition, the encoding unit 101 divides an access unit into four slice segments (tiles) and encodes each slice segment, for example, as shown in FIG. 3.
[0048] The multiplexing unit 102 multiplexes the encoded data generated by the encoding unit 101. The modulation unit 103 modulates the data obtained by the multiplexing. The transmission unit 104 transmits the modulated data as a broadcast signal.
[0049] Next, a configuration of the receiving device 200 according to the present embodiment will be described. Fig. 7 is a block diagram showing a configuration example of the receiving device 200 according to the present embodiment. The receiving device 200 includes a tuner 201, a demodulation unit 202, a demultiplexing unit 203, a plurality of decoding units 204A to 204D, and a display unit 205.
[0050] The tuner 201 receives a broadcast signal. The demodulation unit 202 demodulates the received broadcast signal. The demodulated data is input to the demultiplexing unit 203.
[0051] The demultiplexing unit 203 separates the demodulated data into division units, and outputs the data for each division unit to the decoding units 204A to 204D. Here, the division unit is a division area obtained by dividing an access unit, for example, a slice segment in H.265. Here, an 8K4K image is divided into four 4K2K images. Therefore, there are four decoding units 204A to 204D.
[0052] The multiple decoding units 204A to 204D operate in synchronization with one another based on a predetermined reference clock. Each decoding unit decodes the coded data in division units according to a DTS (Decoding Time Stamp) of the access unit, and outputs the decoding result to the display unit 205.
[0053] The display unit 205 generates an 8K4K output image by integrating the multiple decoding results output from the multiple decoding units 204A to 204D. The display unit 205 displays the generated output image according to a separately acquired PTS (Presentation Time Stamp) of the access unit. When integrating the decoding results, the display unit 205 may perform filtering such as a deblocking filter in a boundary area between adjacent division units, such as a tile boundary, so that the boundary becomes visually less noticeable.
[0054] In the above, the transmitting device 100 and the receiving device 200 that transmit or receive broadcasts have been described as an example, but the content may be transmitted and received via a communication network. When the receiving device 200 receives the content via a communication network, the receiving device 200 separates the multiplexed data from the IP packets received via a network such as Ethernet.
[0055] In broadcasting, the transmission path delay from when a broadcast signal is transmitted until it reaches the receiving device 200 is constant. On the other hand, in a communication network such as the Internet, due to the influence of congestion, the transmission path delay from when data transmitted from a server reaches the receiving device 200 is not constant. Therefore, the receiving device 200 often does not perform strictly synchronous playback based on a reference clock such as PCR in the MPEG-2 TS of broadcasting. Therefore, the receiving device 200 may display the 8K4K output image on the display unit according to the PTS without strictly synchronizing each decoding unit.
[0056] Also, due to congestion in the communication network, the decoding process of all division units may not be completed at the time indicated by the PTS of the access unit. In this case, the receiving device 200 skips displaying the access unit, or delays displaying until decoding of at least four division units is completed and generation of the 8K4K image is completed.
[0057] The content may be transmitted and received by combining broadcasting and communication. The present method is also applicable to playing back multiplexed data stored in a recording medium such as a hard disk or a memory.
[0058] Next, a method of multiplexing an access unit divided into slice segments when MMT is used as the multiplexing scheme will be described.
[0059] 8 is a diagram showing an example of packetizing data of an access unit of HEVC into MMT. Although SPS, PPS, SEI, and the like do not necessarily need to be included in an access unit, a case in which they exist is illustrated here.
[0060] NAL units that are located before the first slice segment in an access unit, such as an access unit delimiter, SPS, PPS, and SEI, are stored together in MMT packet #1. Subsequent slice segments are stored in separate MMT packets for each slice segment.
[0061] As shown in FIG. 9, a NAL unit that is arranged before the first slice segment in an access unit may be stored in the same MMT packet as the first slice segment.
[0062] Furthermore, when NAL units such as End-of-Sequence or End-of-Bitstream, which indicate the end of a sequence or stream, are added after the last slice segment, they are stored in the same MMT packet as the last slice segment. However, since NAL units such as End-of-Sequence or End-of-Bitstream are inserted at the end point of the decoding process or the connection point of two streams, it may be desirable for the receiving device 200 to easily obtain these NAL units in the multiplexing layer. In this case, these NAL units may be stored in an MMT packet different from the slice segment. This allows the receiving device 200 to easily separate these NAL units in the multiplexing layer.
[0063] In addition, TS, DASH, RTP, etc. may be used as the multiplexing method. In these methods, the transmitting device 100 stores different slice segments in different packets. This ensures that the receiving device 200 can separate the slice segments in the multiplexing layer.
[0064] For example, when TS is used, the encoded data is packetized in slice segment units in PES format. When RTP is used, the encoded data is packetized in slice segment units in RTP format. Even in these cases, the NAL unit and the slice segment that are located before the slice segment may be packetized separately, as in MMT packet #1 shown in FIG. 8.
[0065] When TS is used, the transmitting device 100 indicates the unit of data stored in the PES packet by using a data alignment descriptor or the like. In addition, since DASH is a method of downloading MP4 format data units called segments by HTTP or the like, the transmitting device 100 does not packetize the encoded data when transmitting. For this reason, the transmitting device 100 may create subsamples in slice segment units and store information indicating the storage position of the subsamples in the MP4 header so that the receiving device 200 can detect slice segments in the multiplexing layer in MP4.
[0066] The MMT packetization of slice segments will be described in detail below.
[0067] As shown in Fig. 8, by packetizing the encoded data, data commonly referred to when decoding all slice segments in an access unit, such as SPS and PPS, is stored in MMT packet #1. In this case, receiving device 200 concatenates payload data of MMT packet #1 with data of each slice segment, and outputs the obtained data to the decoding unit. In this way, receiving device 200 can easily generate input data to the decoding unit by concatenating payloads of multiple MMT packets.
[0068] FIG. 10 is a diagram showing an example in which input data to the decoding units 204A to 204D is generated from the MMT packets shown in FIG. 8. The demultiplexing unit 203 generates data necessary for the decoding unit 204A to decode the slice segment 1 by linking the payload data of the MMT packet #1 and the MMT packet #2. The demultiplexing unit 203 similarly generates input data for the decoding units 204B to 204D. That is, the demultiplexing unit 203 links the payload data of the MMT packet #1 and the MMT packet #3 to generate input data for the decoding unit 204B. The demultiplexing unit 203 links the payload data of the MMT packet #1 and the MMT packet #4 to generate input data for the decoding unit 204C. The demultiplexing unit 203 links the payload data of the MMT packet #1 and the MMT packet #5 to generate input data for the decoding unit 204D.
[0069] In addition, the demultiplexing unit 203 may remove NAL units that are not necessary for the decoding process, such as the access unit delimiter and SEI, from the payload data of MMT packet #1, and separate only the SPS and PPS NAL units that are necessary for the decoding process and add them to the slice segment data.
[0070] When the encoded data is packetized as shown in FIG. 9, the demultiplexer 203 demultiplexes the MMT packet #1 including the head data of the access unit in the multiplexing layer as the first packet. The demultiplexing unit 203 also analyzes the MMT packet including the first data of the access unit in the multiplexing layer, separates the NAL units of the SPS and PPS, and adds the separated NAL units of the SPS and PPS to each of the data of the second and subsequent slice segments to generate input data for each of the second and subsequent decoding units.
[0071] Furthermore, it is desirable that the receiving device 200 can identify the type of data stored in the MMT payload and the index number of the slice segment in the access unit when the slice segment is stored in the payload, using information included in the header of the MMT packet. Here, the type of data is either data before the slice segment (NAL units arranged before the first slice segment in the access unit are collectively referred to as such) or data of the slice segment. When storing a unit obtained by fragmenting an MPU such as a slice segment in an MMT packet, a mode for storing an MFU (Media Fragment Unit) is used. When using this mode, the transmitting device 100 can set, for example, a Data Unit, which is a basic unit of data in an MFU, to a sample (a data unit in MMT, equivalent to an access unit) or a subsample (a unit obtained by dividing a sample).
[0072] At this time, the header of the MMT packet includes a field called a fragmentation indicator and a field called a fragment counter.
[0073] The fragmentation indicator indicates whether the data stored in the payload of the MMT packet is a fragment of a data unit, and if it is a fragment, whether the fragment is the first or last fragment in the data unit, or a fragment that is neither the first nor the last. In other words, the fragmentation indicator included in the header of a certain packet is identification information that indicates whether (1) only the packet is included in the data unit, which is the basic data unit, (2) the data unit is divided and stored into multiple packets, and the packet is the first packet of the data unit, (3) the data unit is divided and stored into multiple packets, and the packet is a packet other than the first or last packet of the data unit, or (4) the data unit is divided and stored into multiple packets, and the packet is the last packet of the data unit.
[0074] The fragment counter is an index number indicating which fragment in the data unit the data stored in the MMT packet corresponds to.
[0075] Therefore, the transmitting device 100 sets the samples in the MMT to the Data unit, and sets the data before the slice segment and each slice segment to a fragment unit of the Data unit, so that the receiving device 200 can identify the type of data stored in the payload by using the information included in the header of the MMT packet. That is, the demultiplexing unit 203 can generate input data to each of the decoding units 204A to 204D by referring to the header of the MMT packet.
[0076] FIG. 11 is a diagram showing an example in which a sample is set in a data unit, and data before a slice segment and a slice segment are packetized as fragments of the data unit.
[0077] The data before the slice segment and the slice segment are divided into five fragments, fragment #1 to fragment #5. Each fragment is stored in an individual MMT packet. At this time, the values of the fragmentation indicator and fragment counter included in the header of the MMT packet are as shown in the figure.
[0078] For example, the fragment indicator is a 2-bit binary value. The fragment indicator of MMT packet #1 at the beginning of the data unit, the fragment indicator of MMT packet #5 at the end, and the fragment indicators of MMT packet #2 to MMT packet #4, which are packets between them, are each set to a different value. Specifically, the fragment indicator of MMT packet #1 at the beginning of the data unit is set to 01, the fragment indicator of MMT packet #5 at the end is set to 11, and the fragment indicators of MMT packet #2 to MMT packet #4, which are packets between them, are set to 10. Note that when only one MMT packet is included in a data unit, the fragment indicator is set to 00.
[0079] In addition, the fragment counter is 4 in MMT packet #1, which is the total number of fragments (5) minus 1, and is decremented by 1 in the subsequent packets, until it is 0 in the final MMT packet #5.
[0080] Therefore, receiving apparatus 200 can identify an MMT packet that stores pre-slice segment data by using either the fragment indicator or the fragment counter. Also, receiving apparatus 200 can identify an MMT packet that stores the N-th slice segment by referring to the fragment counter.
[0081] The header of the MMT packet additionally includes a sequence number within the MPU of the Movie Fragment to which the Data unit belongs, a sequence number of the MPU itself, and a sequence number within the Movie Fragment of the sample to which the Data unit belongs. By referring to these, the demultiplexer 203 can uniquely determine the sample to which the Data unit belongs.
[0082] Furthermore, since demultiplexing unit 203 can determine the index number of the fragment in the data unit from the fragment counter or the like, it can uniquely identify the slice segment stored in the fragment even if a packet loss occurs. For example, even if demultiplexing unit 203 cannot acquire fragment #4 shown in Fig. 11 due to packet loss, it can know that the fragment received next after fragment #3 is fragment #5, and therefore can correctly output slice segment 4 stored in fragment #5 to decoding unit 204D instead of decoding unit 204C.
[0083] When a transmission path that guarantees no packet loss is used, the demultiplexer 203 may periodically process the arrived packets without determining the type of data stored in the MMT packet or the index number of the slice segment by referring to the header of the MMT packet. For example, when an access unit is transmitted by a total of five MMT packets including the pre-slice data and four slice segments, the receiving device 200 can sequentially acquire the pre-slice data and the data of the four slice segments by sequentially processing the received MMT packets after determining the pre-slice data of the access unit from which decoding is to be started.
[0084] A variation of the packetization will now be described.
[0085] Slice segments do not necessarily have to be divided both horizontally and vertically within the plane of an access unit; as shown in FIG. 1, they may be divided only horizontally or only vertically within the plane of an access unit.
[0086] Also, if an access unit is divided only horizontally, tiles do not need to be used.
[0087] Furthermore, the number of divisions within an access unit is arbitrary and is not limited to 4. However, the area size of slice segments and tiles needs to be equal to or larger than the lower limit of coding standards such as H.265.
[0088] The transmitting device 100 may store identification information indicating a division method in the plane of the access unit in an MMT message, a TS descriptor, or the like. For example, information indicating the number of divisions in the horizontal direction and the vertical direction in the plane may be stored. Alternatively, unique identification information may be assigned to the division method, such as dividing the access unit into two equal parts in the horizontal direction and the vertical direction as shown in FIG. 3, or dividing the access unit into four equal parts in the horizontal direction as shown in FIG. 1. For example, when the access unit is divided as shown in FIG. 3, the identification information indicates mode 1, and when the access unit is divided as shown in FIG. 1, the identification information indicates mode 1.
[0089] Also, information indicating constraints on coding conditions related to the division method within a plane may be included in the multiplex layer. For example, information indicating that one slice segment is composed of one tile may be used. Or, information indicating that a reference block when performing motion compensation during decoding of a slice segment or tile is limited to a slice segment or tile at the same position in a screen, or limited to a block within a predetermined range in an adjacent slice segment may be used.
[0090] Furthermore, the transmitting device 100 may switch whether to divide an access unit into a plurality of slice segments according to the resolution of the video. For example, the transmitting device 100 may not perform intra-plane division when the video to be processed has a resolution of 4K2K, and may divide the access unit into four when the video to be processed has a resolution of 8K4K. By predefining the division method for 8K4K video, the receiving device 200 can acquire the resolution of the video to be received, determine whether to perform intra-plane division and the division method, and switch the decoding operation.
[0091] Furthermore, receiving device 200 can detect whether or not a plane is divided by referring to the header of the MMT packet. For example, if an access unit is not divided, if the Data unit of MMT is set to a sample, the Data unit is not fragmented. Therefore, receiving device 200 can determine that an access unit is not divided if the value of the Fragment counter included in the header of the MMT packet is always zero. Alternatively, receiving device 200 may detect whether the value of the Fragmentation indicator is always 01. Receiving device 200 can also determine that an access unit is not divided if the value of the Fragmentation indicator is always 01.
[0092] In addition, the receiving device 200 can also handle a case where the number of divisions in a plane in an access unit does not match the number of decoding units. For example, when the receiving device 200 includes two decoding units 204A and 204B that can decode 8K2K encoded data in real time, the demultiplexing unit 203 outputs two of the four slice segments that constitute the 8K4K encoded data to the decoding unit 204A.
[0093] Fig. 12 is a diagram showing an example of operation in the case where data packetized as MMT as shown in Fig. 8 is input to two decoders 204A and 204B. Here, it is desirable that the receiving device 200 can directly integrate and output the decoding results in the decoders 204A and 204B. Therefore, the demultiplexer 203 selects slice segments to be output to each of the decoders 204A and 204B so that the decoding results of the decoders 204A and 204B are spatially continuous.
[0094] Also, the demultiplexer 203 may select a decoding unit to be used depending on the resolution or frame rate of the encoded data of the video. For example, when the receiving device 200 has four 4K2K decoding units, if the resolution of the input image is 8K4K, the receiving device 200 performs the decoding process using all four decoding units. Also, if the resolution of the input image is 4K2K, the receiving device 200 performs the decoding process using only one decoding unit. Alternatively, even if the plane is divided into four, if 8K4K can be decoded in real time by a single decoding unit, the demultiplexer 203 integrates all the division units and outputs them to one decoding unit.
[0095] Furthermore, the receiving device 200 may determine the decoding unit to be used in consideration of the frame rate. For example, when the receiving device 200 has two decoding units with an upper limit of 60 fps of the frame rate that can be decoded in real time when the resolution is 8K4K, there is a case where 120 fps encoded data in 8K4K is input. In this case, if the plane is composed of four division units, slice segment 1 and slice segment 2 are input to the decoding unit 204A, and slice segment 3 and slice segment 4 are input to the decoding unit 204B, as in the example of FIG. 12. Since each of the decoding units 204A and 204B can decode up to 120 fps in real time if it is 8K2K (resolution is half of 8K4K), the decoding process is performed by these two decoding units 204A and 204B.
[0096] In addition, even if the resolution and frame rate are the same, the processing amount differs if the profile or level in the encoding method, or the encoding method itself, such as H.264 or H.265, is different. Therefore, the receiving device 200 may select a decoding unit to be used based on these pieces of information. Note that, when the receiving device 200 cannot decode all of the encoded data received by broadcasting or communication, or when all of the slice segments or tiles constituting the area selected by the user cannot be decoded, the receiving device 200 may automatically determine slice segments or tiles that can be decoded within the processing range of the decoding unit. Alternatively, the receiving device 200 may provide a user interface for the user to select the area to be decoded. At this time, the receiving device 200 may display a warning message indicating that all areas cannot be decoded, or may display information indicating the number of areas, slice segments, or tiles that can be decoded.
[0097] In addition, the above method can also be applied to cases where MMT packets storing slice segments of the same encoded data are transmitted and received using multiple transmission paths such as broadcasting and communication.
[0098] Furthermore, the transmitting device 100 may perform coding so that the areas of the slice segments overlap in order to make the boundaries between the division units less noticeable. In the example shown in FIG. 13, an 8K4K picture is divided into four slice segments 1 to 4. Each of the slice segments 1 to 3 is, for example, 8K×1.1K, and the slice segment 4 is 8K×1K. Adjacent slice segments overlap each other. In this way, motion compensation during coding can be efficiently performed at the boundaries when the picture is divided into four, as indicated by the dotted lines, improving the image quality of the boundary portions. In this way, image quality degradation at the boundary portions is reduced.
[0099] In this case, the display unit 205 cuts out an 8K×1K area from the 8K×1.1K area and integrates the obtained area. Note that the transmission device 100 may include information indicating whether the slice segments are coded with overlap and the extent of the overlap in the multiplex layer or the coded data and transmit the information separately.
[0100] Note that a similar approach can be applied when tiles are used.
[0101] The following describes the flow of operations of the transmission device 100. FIG.
[0102] First, the encoding unit 101 divides a picture (access unit) into a plurality of slice segments (tiles), which are a plurality of regions (S101). Next, the encoding unit 101 generates encoded data corresponding to each of the plurality of slice segments by encoding each of the plurality of slice segments so that the slice segments can be decoded independently (S102). Note that the encoding unit 101 may encode the plurality of slice segments using a single encoding unit, or may process the slice segments in parallel using a plurality of encoding units.
[0103] Next, the multiplexing unit 102 multiplexes the multiple pieces of encoded data generated by the encoding unit 101 by storing the multiple pieces of encoded data in multiple MMT packets (S103). Specifically, as shown in Figs. 8 and 9, the multiplexing unit 102 stores the multiple pieces of encoded data in multiple MMT packets so that encoded data corresponding to different slice segments is not stored in one MMT packet. Also, as shown in Fig. 8, the multiplexing unit 102 stores control information commonly used for all decoding units in a picture in an MMT packet #1 that is different from the multiple MMT packets #2 to #5 in which the multiple pieces of encoded data are stored. Here, the control information includes at least one of an access unit delimiter, an SPS, a PPS, and an SEI.
[0104] Note that multiplexing unit 102 may store the control information in the same MMT packet as any one of the multiple MMT packets in which the multiple encoded data are stored. For example, as shown in Fig. 9, multiplexing unit 102 may store the control information in the first MMT packet (MMT packet #1 in Fig. 9) of the multiple MMT packets in which the multiple encoded data are stored.
[0105] Finally, the transmitting device 100 transmits a plurality of MMT packets. Specifically, the modulating unit 103 modulates the data obtained by multiplexing, and the transmitting unit 104 transmits the modulated data (S104).
[0106] Fig. 15 is a block diagram showing an example of the configuration of receiving device 200, and is a diagram showing in detail the configuration of demultiplexing unit 203 and the subsequent stages shown in Fig. 7. As shown in Fig. 15, receiving device 200 further includes a decoding command unit 206. In addition, demultiplexing unit 203 includes a type discrimination unit 211, a control information acquisition unit 212, a slice information acquisition unit 213, and a decoded data generation unit 214.
[0107] The following describes the flow of operations of receiving device 200. Fig. 16 is a flowchart showing an example of operations of receiving device 200. Here, the operation for one access unit is shown. When decoding processes for a plurality of access units are executed, the process of this flowchart is repeated.
[0108] First, the receiving device 200 receives, for example, a plurality of packets (MMT packets) generated by the transmitting device 100 (S201).
[0109] Next, the type discrimination unit 211 analyzes the header of the received packet to obtain the type of the encoded data stored in the received packet (S202).
[0110] Next, the type discrimination unit 211 determines whether the data stored in the received packet is pre-slice segment data or slice segment data, based on the acquired type of encoded data (S203).
[0111] If the data stored in the received packet is pre-slice segment data (Yes in S203), the control information acquisition unit 212 acquires the pre-slice segment data of the access unit to be processed from the payload of the received packet and stores the pre-slice segment data in memory (S204).
[0112] On the other hand, if the data stored in the received packet is data of a slice segment (No in S203), the receiving device 200 uses the header information of the received packet to determine which of the multiple areas the data stored in the received packet is encoded data of. Specifically, the slice information acquisition unit 213 analyzes the header of the received packet to acquire an index number Idx of the slice segment stored in the received packet (S205). Specifically, the index number Idx is an index number in a Movie Fragment of an access unit (a sample in MMT).
[0113] The process of step S205 may be performed collectively in step S202.
[0114] Next, the decoding data generation unit 214 determines a decoding unit that decodes the slice segment (S206). Specifically, the index number Idx and a plurality of decoding units are associated in advance, and the decoding data generation unit 214 determines the decoding unit that corresponds to the index number Idx acquired in step S205 as the decoding unit that decodes the slice segment.
[0115] 12, the decoded data generation unit 214 may determine a decoding unit to decode the slice segment based on at least one of the resolution of the access unit (picture), the division method of the access unit into a plurality of slice segments (tiles), and the processing capabilities of a plurality of decoding units included in the receiving device 200. For example, the decoded data generation unit 214 determines the division method of the access unit based on identification information in a descriptor such as an MMT message or a TS section.
[0116] Next, the decoded data generating unit 214 generates a plurality of input data (combined data) to be input to a plurality of decoding units by combining control information, which is included in any of the plurality of packets and is used in common for all decoding units in the picture, with each of the plurality of encoded data of the plurality of slice segments. Specifically, the decoded data generating unit 214 acquires data of the slice segment from the payload of the received packet. The decoded data generating unit 214 generates input data to the decoding unit determined in step S206 by combining the pre-slice segment data stored in the memory in step S204 with the acquired slice segment data (S207).
[0117] After step S204 or S207, if the data of the received packet is not the final data of the access unit (No in S208), the process from step S201 onward is performed again. That is, the above process is repeated until input data for the multiple decoding units 204A to 204D corresponding to all slice segments included in the access unit is generated.
[0118] The timing at which the packets are received is not limited to the timing shown in FIG. 16, and a plurality of packets may be received in advance or sequentially and stored in a memory or the like.
[0119] On the other hand, if the data of the received packet is the final data of the access unit (Yes in S208), the decoding command unit 206 outputs the multiple pieces of input data generated in step S207 to the corresponding decoding units 204A to 204D (S209).
[0120] Next, the multiple decoding units 204A to 204D decode the multiple input data in parallel in accordance with the DTS of the access unit, thereby generating multiple decoded images (S210).
[0121] Finally, the display unit 205 generates a display image by combining the multiple decoded images generated by the multiple decoding units 204A to 204D, and displays the display image in accordance with the PTS of the access unit (S211).
[0122] The receiving device 200 obtains the DTS and PTS of the access unit by analyzing the header information of the MPU or the payload data of the MMT packet that stores the header information of the Movie Fragment. When the multiplexing method is TS, the receiving device 200 obtains the DTS and PTS of the access unit from the header of the PES packet. When the multiplexing method is RTP, the receiving device 200 obtains the DTS and PTS of the access unit from the header of the RTP packet.
[0123] Furthermore, when integrating the decoding results of the multiple decoding units, the display unit 205 may perform a filter process such as a deblocking filter at the boundary between adjacent division units. Note that, since the filter process is not necessary when displaying the decoding results of a single decoding unit, the display unit 205 may switch the process depending on whether or not to perform a filter process at the boundary between the decoding results of the multiple decoding units. Whether or not the filter process is necessary may be specified in advance depending on the presence or absence of division. Alternatively, information indicating whether or not the filter process is necessary may be stored separately in the multiplexing layer. Also, information necessary for the filter process such as a filter coefficient may be stored in the SPS, PPS, SEI, or slice segment. The decoding units 204A to 204D or the demultiplexing unit 203 acquire these pieces of information by analyzing the SEI, and output the acquired information to the display unit 205. The display unit 205 performs a filter process using these pieces of information. Note that, when these pieces of information are stored in the slice segment, it is preferable that the decoding units 204A to 204D acquire these pieces of information.
[0124] In the above description, the type of data stored in the fragment is two types, i.e., pre-slice segment data and slice segments, but the type of data may be three or more. In this case, the case is classified according to the type in step S203.
[0125] In addition, when the data size of a slice segment is large, the transmitting device 100 may fragment the slice segment and store it in the MMT packet. That is, the transmitting device 100 may fragment the slice segment pre-data and the slice segment. In this case, if the access unit and the data unit are set to be equal as in the example of packetization shown in FIG. 11, the following problem occurs.
[0126] For example, when slice segment 1 is divided into three fragments, slice segment 1 is divided and transmitted into three packets with fragment counter values of 1 to 3. From slice segment 2 onwards, the fragment counter value becomes 4 or more, and it becomes impossible to associate the fragment counter value with the data stored in the payload. Therefore, receiving device 200 cannot identify the packet that stores the first data of the slice segment from the information in the header of the MMT packet.
[0127] In such a case, receiving device 200 may analyze data in the payload of the MMT packet to identify the start position of the slice segment. Here, there are two types of formats for storing NAL units in a multiplexing layer in H.264 or H.265: a format called a byte stream format in which a start code consisting of a specific bit string is added immediately before the NAL unit header, and a format called a NAL size format in which a field indicating the size of the NAL unit is added.
[0128] The byte stream format is used in MPEG-2 systems, RTP, etc. The NAL size format is used in MP4, as well as DASH and MMT that use MP4, etc.
[0129] When the byte stream format is used, the receiving device 200 analyzes whether the leading data of the packet matches the start code. If the leading data of the packet matches the start code, the receiving device 200 can detect whether the data included in the packet is data of a slice segment by obtaining the type of the NAL unit from the subsequent NAL unit header.
[0130] On the other hand, in the case of the NAL size format, the receiving device 200 cannot detect the start position of the NAL unit based on the bit string. Therefore, in order to obtain the start position of the NAL unit, the receiving device 200 needs to shift the pointer by reading data by the size of the NAL unit, starting from the first NAL unit of the access unit.
[0131] However, in the case where the size of the subsample unit is indicated in the header of the MPU or Movie Fragment in MMT and the subsample corresponds to pre-slice data or a slice segment, the receiving device 200 can identify the start position of each NAL unit based on the size information of the subsample. Therefore, the transmitting device 100 may include information indicating whether the information of the subsample unit exists in the MPU or Movie Fragment in information that the receiving device 200 acquires when starting to receive data, such as the MPT in MMT.
[0132] The MPU data is an extension of the MP4 format. In MP4, there are a mode in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and a mode in which they cannot be stored. Information for identifying this mode is indicated as the entry name of SampleEntry. When a mode in which the parameter sets can be stored is used and the parameter sets are included in the samples, the receiving device 200 acquires the parameter sets by the above-mentioned method.
[0133] On the other hand, when a mode that cannot be stored is used, the parameter set is stored as Decoder Specific Information in SampleEntry, or is stored using a stream for the parameter set. Here, since the stream for the parameter set is not generally used, it is desirable that the transmitting device 100 stores the parameter set in the Decoder Specific Information. In this case, the receiving device 200 analyzes the SampleEntry transmitted as metadata of the MPU or metadata of the Movie Fragment in the MMT packet to obtain the parameter set referred to by the access unit.
[0134] When a parameter set is stored as sample data, the receiving device 200 can acquire a parameter set required for decoding by referring only to the sample data without referring to the SampleEntry. In this case, the transmitting device 100 does not need to store a parameter set in the SampleEntry. In this way, the transmitting device 100 can use the same SampleEntry in different MPUs, so that the processing load of the transmitting device 100 when generating an MPU can be reduced. Furthermore, there is an advantage that the receiving device 200 does not need to refer to the parameter set in the SampleEntry.
[0135] Alternatively, the transmitting device 100 may store one default parameter set in SampleEntry, and store the parameter set referenced by the access unit in the sample data. In conventional MP4, since it was common to store a parameter set in SampleEntry, there is a possibility that a receiving device may stop playback if a parameter set does not exist in SampleEntry. This problem can be solved by using the above method.
[0136] Alternatively, transmitting device 100 may store a parameter set in sample data only when a parameter set different from the default parameter set is used.
[0137] In addition, since it is possible to store a parameter set in a SampleEntry in both modes, transmitting device 100 may always store a parameter set in a VisualSampleEntry, and receiving device 200 may always obtain the parameter set from the VisualSampleEntry.
[0138] In the MMT standard, MP4 header information such as Moov and Moof is transmitted as MPU metadata or movie fragment metadata, but the transmitting device 100 does not necessarily have to transmit the MPU metadata and movie fragment metadata. Furthermore, the receiving device 200 can also determine whether or not the SPS and PPS are stored in the sample data based on the service, asset type, or the presence or absence of transmission of MPU meta of the ARIB (Association of Radio Industries and Businesses) standard.
[0139] FIG. 17 is a diagram showing an example in which pre-slice-segment data and each slice segment are set to different Data units.
[0140] 17, the data sizes of the data before the slice segment and slice segment 1 to slice segment 4 are Length #1 to Length #5, respectively. The field values of the Fragmentation indicator, Fragment counter, and Offset included in the header of the MMT packet are as shown in the figure.
[0141] Here, Offset is offset information indicating the bit length (offset) from the beginning of the encoded data of the sample (access unit or picture) to which the payload data belongs to to the first byte of the payload data (encoded data) included in the MMT packet. Note that, although the value of the Fragment counter is described as starting from a value obtained by subtracting 1 from the total number of fragments, it may start from another value.
[0142] Fig. 18 is a diagram showing an example of a case where a Data unit is fragmented. In the example shown in Fig. 18, slice segment 1 is divided into three fragments, which are stored in MMT packets #2 to #4, respectively. In this case, if the data size of each fragment is Length#2_1 to Length#2_3, respectively, the values of each field are as shown in the figure.
[0143] In this way, when a data unit such as a slice segment is set to Data unit, the start of the access unit and the start of the slice segment can be determined as follows based on the field values of the MMT packet header.
[0144] The start of the payload in a packet whose Offset value is 0 is the start of the access unit.
[0145] The start of the payload of a packet in which the Offset value is a value other than 0 and the Fragmentation indicator value is 00 or 01 is the start of the slice segment.
[0146] In addition, if no fragmentation of the data unit occurs and no packet loss occurs, the receiving device 200 can identify the index number of the slice segment to be stored in the MMT packet based on the number of slice segments obtained after detecting the beginning of the access unit.
[0147] In addition, even if the data unit of the data before the slice segment is fragmented, the receiving device 200 can similarly detect the beginning of the access unit and the slice segment.
[0148] Furthermore, even if a packet loss occurs or if the SPS, PPS, and SEI included in the slice segment pre-data are set in different Data units, the receiving device 200 can identify the MMT packet that stores the start data of the slice segment based on the analysis result of the MMT header, and then analyze the header of the slice segment to identify the start position of the slice segment or tile in the picture (access unit). The amount of processing involved in analyzing the slice header is small, and the processing load is not a problem.
[0149] In this way, each of the multiple coded data of the multiple slice segments is in one-to-one correspondence with a basic data unit, which is a unit of data stored in one or more packets. Also, each of the multiple coded data is stored in one or more MMT packets.
[0150] The header information of each MMT packet includes a fragmentation indicator (identification information) and an offset (offset information).
[0151] Receiving apparatus 200 determines that the start of payload data included in a packet having header information including a Fragmentation indicator whose value is 00 or 01 is the start of encoded data of each slice segment. Specifically, receiving apparatus 200 determines that the start of payload data included in a packet having header information including an Offset whose value is not 0 and a Fragmentation indicator whose value is 00 or 01 is the start of encoded data of each slice segment.
[0152] 17, the start of a Data Unit is either the start of an access unit or the start of a slice segment, and the value of the Fragmentation indicator is 00 or 01. Furthermore, receiving device 200 can detect the start of an access unit or the start of a slice segment without referring to an Offset by referring to the type of NAL unit and determining whether the start of a Data Unit is an access unit delimiter or a slice segment.
[0153] In this way, the transmitting device 100 performs packetization so that the beginning of the NAL unit always starts from the beginning of the payload of the MMT packet, and thus the receiving device 200 can detect the beginning of the access unit or slice segment by analyzing the fragmentation indicator and the NAL unit header, including the case where the data before the slice segment is divided into a plurality of Data units. The type of the NAL unit exists in the first byte of the NAL unit header. Therefore, when analyzing the header part of the MMT packet, the receiving device 200 can obtain the type of the NAL unit by analyzing an additional 1 byte of data. In the case of audio, the receiving device 200 only needs to detect the beginning of the access unit, and can make a determination based on whether the value of the fragmentation indicator is 00 or 01.
[0154] Furthermore, as described above, when storing coded data that has been coded so that it can be divided and decoded in a PES packet of MPEG-2 TS, the transmitting device 100 can use a data alignment descriptor. An example of a method for storing coded data in a PES packet will be described in detail below.
[0155] For example, in HEVC, the transmission device 100 can indicate whether the data stored in the PES packet is an access unit, a slice segment, or a tile by using a data alignment descriptor. The alignment types in HEVC are specified as follows.
[0156] Alignment type=8 indicates an HEVC slice segment. Alignment type=9 indicates an HEVC slice segment or access unit. Alignment type=12 indicates an HEVC slice segment or tile.
[0157] Therefore, the transmitting device 100 can indicate that the data of the PES packet is either a slice segment or pre-slice segment data by using, for example, type 9. Since a type indicating a slice instead of a slice segment is also separately defined, the transmitting device 100 may use a type indicating a slice instead of a slice segment.
[0158] Furthermore, the DTS and PTS included in the header of a PES packet are set only in a PES packet that includes the first data of an access unit. Therefore, if the type is 9 and the PES packet includes a DTS or PTS field, receiving device 200 can determine that the PES packet stores the entire access unit or the first division unit of the access unit.
[0159] Furthermore, the transmitting device 100 may enable the receiving device 200 to distinguish the data contained in the packet by using a field such as transport_priority indicating the priority of a TS packet storing a PES packet including the first data of an access unit. The receiving device 200 may also determine the data contained in the packet by analyzing whether the payload of the PES packet is an access unit delimiter. Furthermore, the data_alignment_indicator in the PES packet header indicates whether data is stored in the PES packet according to these types. If this flag (data_alignment_indicator) is set to 1, it is guaranteed that the data stored in the PES packet complies with the type indicated in the data alignment descriptor.
[0160] Furthermore, the transmitting device 100 may use the data alignment descriptor only when PES packetization is performed in a unit that can be divided and decoded, such as a slice segment. As a result, when the data alignment descriptor is present, the receiving device 200 can determine that the coded data is PES packetized in a unit that can be divided and decoded, and when the data alignment descriptor is not present, the receiving device 200 can determine that the coded data is PES packetized in an access unit. Note that when the data_alignment_indicator is set to 1 and the data alignment descriptor is not present, it is stipulated in the MPEG-2 TS standard that the unit of PES packetization is the access unit.
[0161] If the data alignment descriptor is included in the PMT, the receiving device 200 can determine that the PES packetization is performed in a unit that can be divided and decoded, and can generate input data to each decoding unit based on the packetized unit. If the data alignment descriptor is not included in the PMT and the receiving device 200 determines that parallel decoding of the encoded data is necessary based on the program information or information of other descriptors, the receiving device 200 generates input data to each decoding unit by analyzing the slice header of the slice segment. If the encoded data can be decoded by a single decoding unit, the receiving device 200 decodes the data of the entire access unit by the decoding unit. If information indicating whether the encoded data is composed of a unit that can be divided and decoded, such as a slice segment or a tile, is separately indicated by a descriptor of the PMT, the receiving device 200 may determine whether the encoded data can be decoded in parallel based on the analysis result of the descriptor.
[0162] Furthermore, since the DTS and PTS included in the header of a PES packet are set only in the PES packet including the first data of an access unit, when an access unit is divided and packetized as PES packets, the second and subsequent PES packets do not include information indicating the DTS and PTS of the access unit. Therefore, when performing decoding processes in parallel, each of the decoding units 204A-204D and display unit 205 uses the DTS and PTS stored in the header of the PES packet including the first data of an access unit.
[0163] (Embodiment 2) In the second embodiment, a method for storing data in the NAL size format in an MPU based on the MP4 format in MMT will be described. Note that, as an example, a method for storing data in an MPU used in MMT will be described below, but such a method for storing data can also be applied to DASH, which is based on the same MP4 format.
[0164] [How to store in MPU] In the MP4 format, multiple access units are stored together in one MP4 file. In the MPU used in MMT, data for each media is stored in one MP4 file, and the data can contain any number of access units. Since the MPU is a unit that can be decoded independently, for example, the MPU stores access units in units of GOPs.
[0165] 19 is a diagram showing the configuration of an MPU. At the beginning of an MPU are ftyp, mmpu, and moov, which are collectively defined as MPU metadata. In moov, initialization information common to files and an MMT hint track are stored.
[0166] In addition, moof stores initialization information and size for each sample and subsample, information (sample_duration, sample_size, sample_composition_time_offset) that can specify the presentation time (PTS) and decoding time (DTS), and data_offset that indicates the position of the data.
[0167] In addition, each of the access units is stored as a sample in mdat (mdat box). The data in moof and mdat excluding the samples is defined as movie fragment metadata (hereinafter, referred to as MF metadata), and the sample data in mdat is defined as media data.
[0168] Fig. 20 is a diagram showing the structure of MF metadata. As shown in Fig. 20, the MF metadata is composed of, in more detail, the type, length, and data of a moof box (moof), and the type and length of an mdat box (mdat).
[0169] When storing access units in MP4 data, there are a mode in which parameter sets such as SPS and PPS of H.264 or H.265 can be stored as sample data, and a mode in which they cannot be stored.
[0170] Here, in the above-mentioned mode in which the parameter set cannot be stored, the parameter set is stored in the Decoder Specific Information of the SampleEntry in the moov, and in the above-mentioned mode in which the parameter set can be stored, the parameter set is included in the sample.
[0171] The MPU metadata, MF metadata, and media data are each stored in the MMT payload, and a fragment type (FT) is stored in the header of the MMT payload as an identifier that can identify these data. FT=0 indicates MPU metadata, FT=1 indicates MF metadata, and FT=2 indicates media data.
[0172] In addition, in Fig. 19, an example in which MPU metadata units and MF metadata units are stored in the MMT payload as data units is illustrated, but units such as ftyp, mmpu, moov, and moof may be stored in the MMT payload as data units in data unit units. Similarly, in Fig. 19, an example in which sample units are stored in the MMT payload as data units is illustrated. However, data units may be configured in sample units or NAL unit units, and such data units may be stored in the MMT payload in data unit units. Such data units may be further fragmented and stored in the MMT payload.
[0173] [Conventional transmission methods and issues] Conventionally, when multiple access units are encapsulated in the MP4 format, moov and moof are created when all samples to be stored in the MP4 are available.
[0174] When MP4 format is transmitted in real time using broadcasting, for example, if samples stored in one MP4 file are in GOP units, delays occur due to encapsulation because moov and moof are created after GOP unit time samples are accumulated. This encapsulation on the transmitting side always lengthens the end-to-end delay by the GOP unit time. This makes it difficult to provide services in real time, and leads to deterioration of the service for viewers, especially when live content is transmitted.
[0175] Fig. 21 is a diagram for explaining the data transmission order. When MMT is applied to broadcasting, as shown in Fig. 21(a), if data is loaded onto MMT packets and transmitted in the order of the MPU configuration (transmitting MMT packets #1, #2, #3, #4, #5, and #6 in that order), a delay occurs in the transmission of the MMT packets due to encapsulation.
[0176] To prevent this delay due to encapsulation, a method has been proposed in which MPU header information such as MPU metadata and MF metadata is not sent (packets #1 and #2 are not sent, and packets #3-#6 are sent in this order), as shown in (b) of Fig. 21. Also, a method can be considered in which media data is sent first without waiting for the creation of MPU header information, and the MPU header information is sent after the media data has been sent (sending in the order #3-#6, #1, #2), as shown in (c) of Fig. 20.
[0177] If the MPU header information is not transmitted, the receiving device decodes without using the MPU header information. Also, if the MPU header information is sent later than the media data, the receiving device waits until it obtains the MPU header information before decoding.
[0178] However, in conventional MP4-compliant receiving devices, decoding without using MPU header information is not guaranteed. Also, if a receiving device performs decoding without using an MPU header by special processing, the decoding process becomes complicated when using a conventional transmission method, and there is a high possibility that real-time decoding becomes difficult. Also, if a receiving device waits for acquisition of MPU header information before decoding, media data needs to be buffered until the receiving device acquires the header information, but a buffer model has not been specified and decoding has not been guaranteed.
[0179] Therefore, the transmitting device according to the second embodiment transmits the MPU metadata before the media data by storing only common information in the MPU metadata as shown in (d) of Fig. 20. The transmitting device according to the second embodiment transmits the MF metadata, the generation of which is delayed, after the media data. This provides a transmitting method or receiving method that can guarantee the decoding of the media data.
[0180] The receiving method when each of the transmission methods (a) to (d) in FIG. 21 is used will be described below.
[0181] In each transmission method shown in FIG. 21, first, MPU data is configured in the following order: MPU metadata, MFU metadata, and media data.
[0182] After constructing the MPU data, if the transmitting device transmits data in the order of MPU metadata, MF metadata, and media data, as shown in (a) of Figure 21, the receiving device can perform decoding using either of the following methods (A-1) and (A-2).
[0183] (A-1) After acquiring the MPU header information (MPU metadata and MF metadata), the receiving device decodes the media data using the MPU header information.
[0184] (A-2) The receiving device decodes the media data without using the MPU header information.
[0185] Although these methods all cause delays due to encapsulation on the sending side, they have the advantage that the receiving device does not need to buffer the media data to obtain the MPU header. If buffering is not performed, there is no need to install memory for buffering, and furthermore, no buffering delay occurs. Also, method (A-1) is applicable to conventional receiving devices because it uses MPU header information for decoding.
[0186] When the transmitting device transmits only media data as shown in (b) of FIG. 21, the receiving device can perform decoding using the following method (B-1).
[0187] (B-1) The receiving device decodes the media data without using the MPU header information.
[0188] Although not shown, if MPU metadata is transmitted prior to the transmission of the media data in FIG. 21(b), decoding can be performed using the following method (B-2).
[0189] (B-2) The receiving device decodes the media data using the MPU metadata.
[0190] The advantages of both methods (B-1) and (B-2) above are that no delay due to encapsulation occurs on the sending side, and there is no need to buffer media data to obtain the MPU header. However, both methods (B-1) and (B-2) may require special processing for decoding, since they do not use MPU header information for decoding.
[0191] When the transmitting device transmits data in the order of media data, MPU metadata, and MF metadata, as shown in (c) of Figure 21, the receiving device can perform decoding using either of the following methods (C-1) and (C-2).
[0192] (C-1) The receiving device acquires the MPU header information (MPU metadata and MF metadata) and then decodes the media data.
[0193] (C-2) The receiving device decodes the media data without using the MPU header information.
[0194] When the above method (C-1) is used, it is necessary to buffer the media data in order to obtain the MPU header information. In contrast, when the above method (C-2) is used, it is not necessary to buffer the media data in order to obtain the MPU header information.
[0195] In addition, neither method (C-1) nor (C-2) causes delays due to encapsulation on the sending side. Also, method (C-2) may require special processing because it does not use MPU header information.
[0196] When the transmitting device transmits data in the order of MPU metadata, media data, and MF metadata, as shown in (d) of Figure 21, the receiving device can perform decoding using either method (D-1) or (D-2) below.
[0197] (D-1) After acquiring the MPU metadata, the receiving device further acquires the MF metadata, and then decodes the media data.
[0198] (D-2) After acquiring the MPU metadata, the receiving device decodes the media data without using the MF metadata.
[0199] When the above method (D-1) is used, it is necessary to buffer the media data in order to obtain the MF metadata, but when the above method (D-2) is used, it is not necessary to perform buffering in order to obtain the MF metadata.
[0200] The above method (D-2) does not use MF metadata for decoding, and therefore may require special processing.
[0201] As described above, when decoding is possible using MPU metadata and MF metadata, there is an advantage that decoding can also be performed by a conventional MP4 receiving device.
[0202] In Fig. 21, the MPU data is structured in the order of MPU metadata, MFU metadata, and media data, and in the moof, the position information (offset) for each sample and subsample is determined based on this structure. In addition, the MF metadata also includes data other than the media data in the mdat box (box size and type).
[0203] Therefore, when a receiving device identifies media data based on MF metadata, the receiving device reconstructs the data in the order in which the MPU data was constructed, regardless of the order in which the data was transmitted, and then decodes it using the moov of the MPU metadata or the moof of the MF metadata.
[0204] In addition, in FIG. 21, the MPU data is configured in the order of MPU metadata, MFU metadata, and media data, but the MPU data may be configured in an order different from that in FIG. 21, and the position information (offset) may be determined.
[0205] For example, MPU data may be configured in the order of MPU metadata, media data, and MF metadata, and negative position information (offset) may be indicated in the MF metadata. In this case, regardless of the order in which the data is transmitted, the receiving device reconstructs the data in the order in which the MPU data was configured on the transmitting side, and then performs decoding using moov or moof.
[0206] In addition, the transmitting device may signal information indicating the order in which the MPU data is constructed, and the receiving device may reconstruct the data based on the signaled information.
[0207] As described above, the receiving device receives packetized MPU metadata, packetized media data (sample data), and packetized MF metadata in this order, as shown in (d) of Fig. 21. Here, the MPU metadata is an example of first metadata, and the MF metadata is an example of second metadata.
[0208] Next, the receiving device reconstructs MPU data (MP4 format file) including the received MPU metadata, the received MF metadata, and the received sample data. Then, the receiving device decodes the sample data included in the reconstructed MPU data using the MPU metadata and MF metadata. The MF metadata is metadata including data (e.g., length stored in an mbox) that can be generated only after the sample data is generated on the transmitting side.
[0209] In addition, the operation of the receiving device is performed by each component constituting the receiving device in more detail. For example, the receiving device includes a receiving unit that receives the data, a reconstructing unit that reconstructs the MPU data, and a decoding unit that decodes the MPU data. Each of the receiving unit, the generating unit, and the decoding unit is realized by a microcomputer, a processor, a dedicated circuit, etc.
[0210] [Method of decrypting without using header information] Next, a method of decoding without using header information will be described. Here, a method of decoding without using header information in a receiving device will be described regardless of whether or not the transmitting side sends header information. That is, this method is applicable to any of the transmission methods described with reference to FIG. 21. However, some of the decoding methods are applicable only to specific transmission methods.
[0211] Fig. 22 is a diagram showing an example of a method for decoding without using header information. In Fig. 22, only MMT payloads and MMT packets including only media data are shown, and MMT payloads and MMT packets including MPU metadata and MF metadata are not shown. In the following description of Fig. 22, it is assumed that media data belonging to the same MPU are transmitted continuously. In addition, a case where samples are stored in the payload as media data will be described as an example, but in the following description of Fig. 22, it is natural that NAL units or fragmented NAL units may be stored.
[0212] To decode media data, the receiving device must first obtain initialization information required for decoding. If the media is video, the receiving device must obtain initialization information for each sample, identify the start position of the MPU, which is a random access unit, and obtain the start positions of the samples and NAL units. The receiving device must also identify the decode time (DTS) and presentation time (PTS) of each sample.
[0213] Therefore, the receiving device can perform decoding without using header information, for example, by using the following method. Note that, when NAL unit units or fragmented NAL units are stored in the payload, "sample" in the following description can be read as "NAL unit in sample."
[0214] <Random access (=identify the first sample of the MPU)> When header information is not transmitted, the receiving device can identify the first sample of the MPU using the following methods 1 and 2. Note that when header information is transmitted, method 3 can be used.
[0215] [Method 1] The receiving device acquires a sample contained in an MMT packet with 'RAP_flag=1' in the MMT packet header.
[0216] [Method 2] The receiving device obtains samples with 'sample number=0' in the MMT payload header.
[0217] [Method 3] When at least one of MPU metadata and MF metadata is transmitted before or after the media data, the receiving device acquires a sample contained in the MMT payload in which the fragment type (FT) in the MMT payload header has been switched to media data.
[0218] In Methods 1 and 2, when a single payload contains multiple samples belonging to different MPUs, it is impossible to determine which NAL unit is a random access point (RAP_flag = 1 or sample number = 0). For this reason, it is necessary to impose a constraint such as not mixing samples from different MPUs in a single payload, or, when samples from different MPUs are mixed in a single payload, to impose a constraint such as setting RAP_flag to 1 if the last (or first) sample is a random access point.
[0219] Furthermore, in order for the receiving device to obtain the start position of the NAL unit, it is necessary to shift the data read pointer by the size of the NAL unit, starting from the first NAL unit of the sample.
[0220] If the data is fragmented, the receiving device can identify the data unit by referring to the fragment_indicator and fragment_number.
[0221] <Determining the DTS of a sample> There are two methods for determining the DTS of a sample: Method 1 and Method 2 below.
[0222] [Method 1] The receiver determines the DTS of the first sample based on the prediction structure. However, this method requires analysis of the encoded data, and may be difficult to decode in real time, so the following method 2 is preferable.
[0223] [Method 2] The receiving device transmits the DTS of the first sample separately and acquires the transmitted DTS of the first sample. The method of transmitting the DTS of the first sample includes, for example, transmitting the DTS of the MPU first sample using MMT-SI, or transmitting the DTS for each sample using the MMT packet header extension area. Note that the DTS may be an absolute value or a relative value to the PTS. Also, the transmitting side may signal whether the DTS of the first sample is included.
[0224] In both methods 1 and 2, the DTS of subsequent samples is calculated assuming a fixed frame rate.
[0225] As a method for storing the DTS for each sample in the packet header, other than using the extension field, there is a method for storing the DTS of the sample contained in the MMT packet in the 32-bit NTP timestamp field in the MMT packet header. If the DTS cannot be expressed by the number of bits (32 bits) of one packet header, the DTS may be expressed using multiple packet headers. Also, the DTS may be expressed by combining the NTP timestamp field and the extension field of the packet header. If DTS information is not included, a known value (for example, ALL0) is used.
[0226] <Determining the PTS of a sample> The receiving device obtains the PTS of the first sample from the MPU timestamp descriptor for each asset included in the MPU. The receiving device calculates the PTS of subsequent samples from parameters indicating the display order of samples such as POC, assuming a fixed frame rate. In this way, transmission at a fixed frame rate is essential to calculate the DTS and PTS without using header information.
[0227] In addition, when MF metadata is transmitted, the receiving device can calculate the absolute values of the DTS and PTS from the relative time information of the DTS and PTS from the first sample indicated in the MF metadata and the absolute value of the timestamp of the MPU first sample indicated in the MPU timestamp descriptor.
[0228] When analyzing the encoded data to calculate the DTS and PTS, the receiving device may perform the calculations using SEI information included in the access unit.
[0229] <Initialization information (parameter set)> [For video] In the case of video, the parameter set is stored in the sample data. Also, if the MPU metadata and the MF metadata are not transmitted, it is guaranteed that the parameter set required for decoding can be obtained by referring only to the sample data.
[0230] Also, when MPU metadata is transmitted before media data, as in (a) and (d) of Figure 21, it may be specified that parameter sets are not stored in SampleEntry. In this case, the receiving device refers to only the parameter sets in the sample, without referring to the parameter sets in SampleEntry.
[0231] In addition, when MPU metadata is transmitted before media data, a parameter set common to MPUs or a default parameter set is stored in SampleEntry, and a receiving device may refer to the parameter set in SampleEntry and the parameter set in the sample. Storing a parameter set in SampleEntry enables decoding even in a conventional receiving device that cannot play back data unless a parameter set is present in SampleEntry.
[0232] [For audio] For audio, a LATM header is required for decoding, and in MP4, it is mandatory that the LATM header is included in the sample entry. However, if the header information is not transmitted, it is difficult for the receiving device to obtain the LATM header, so the LATM header is included separately in control information such as SI. The LATM header may be included in a message, table, or descriptor. The LATM header may also be included in the sample.
[0233] The receiving device acquires the LATM header from the SI or the like before starting decoding, and starts decoding the audio. Alternatively, as shown in (a) and (d) of Figure 21, if the MPU metadata is transmitted before the media data, the receiving device can receive the LATM header before the media data. Therefore, if the MPU metadata is transmitted before the media data, decoding can be performed even using a conventional receiving device.
[0234] <Other> The transmission order and the type of transmission order may be notified as control information such as an MMT packet header or a payload header, or an MPT or other table, message, or descriptor. Note that the type of transmission order here refers to, for example, the four types of transmission order shown in (a) to (d) of Fig. 21, and an identifier for identifying each type may be stored in a location where it can be obtained before decoding begins.
[0235] Also, the transmission order type may be different for audio and video, or a common type may be used for audio and video. Specifically, for example, audio may be transmitted in the order of MPU metadata, MF metadata, and media data as shown in (a) of Fig. 21, and video may be transmitted in the order of MPU metadata, media data, and MF metadata as shown in (d) of Fig. 21.
[0236] The above-described method allows the receiving device to decode without using header information. Also, if the MPU metadata is transmitted before the media data (FIG. 21(a) and FIG. 21(d)), decoding is possible even with a conventional receiving device.
[0237] In particular, by transmitting the MF metadata after the media data (FIG. 21(d)), delay due to encapsulation is not generated, and decoding can be performed even by a conventional receiving device.
[0238] [Configuration and operation of transmitting device] Next, the configuration and operation of a transmission device will be described. Fig. 23 is a block diagram of a transmission device according to the second embodiment, and Fig. 24 is a flowchart of a transmission method according to the second embodiment.
[0239] As shown in FIG. 23, the transmission device 15 includes an encoding unit 16, a multiplexing unit 17, and a transmission unit .
[0240] The encoding unit 16 generates encoded data by encoding the video or audio to be encoded according to, for example, H.265 (S10).
[0241] The multiplexing unit 17 multiplexes (packetizes) the encoded data generated by the encoding unit 16 (S11). Specifically, the multiplexing unit 17 packetizes each of the sample data, MPU metadata, and MF metadata that constitute an MP4 format file. The sample data is data in which a video signal or an audio signal is encoded, the MPU metadata is an example of the first metadata, and the MF metadata is an example of the second metadata. Both the first metadata and the second metadata are metadata used to decode the sample data, but the difference between them is that the second metadata includes data that can be generated only after the sample data is generated.
[0242] Here, the data that can be generated only after the generation of the sample data is, for example, data other than the sample data stored in mdat in the MP4 format (data in the header of mdat, that is, the type and length shown in FIG. 20). Here, the second metadata only needs to include the length, which is at least a part of this data.
[0243] The transmitting unit 18 transmits the packetized MP4 format file (S12). The transmitting unit 18 transmits the MP4 format file, for example, by the method shown in (d) of Fig. 21. That is, the transmitting unit 18 transmits the packetized MPU metadata, the packetized sample data, and the packetized MF metadata in this order.
[0244] Each of the encoding unit 16, the multiplexing unit 17, and the transmitting unit 18 is realized by a microcomputer, a processor, or a dedicated circuit.
[0245] [Receiver configuration] Next, the configuration and operation of a receiving device will be described. Fig. 25 is a block diagram of a receiving device according to the second embodiment.
[0246] As shown in FIG. 25, the receiving device 20 includes a packet filtering unit 21, a transmission order type discrimination unit 22, a random access unit 23, a control information acquisition unit 24, a data acquisition unit 25, a PTS, DTS calculation unit 26, an initialization information acquisition unit 27, a decoding command unit 28, a decoding unit 29, and a presentation unit 30.
[0247] [Receiver operation 1] First, an operation of the receiving device 20 for identifying the MPU start position and the NAL unit position when the media is video will be described. Fig. 26 is a flowchart of such an operation of the receiving device 20. Note that it is assumed here that the transmission order type of the MPU data is stored in the SI information by the transmitting device 15 (multiplexing unit 17).
[0248] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type discrimination unit 22 analyzes the SI information obtained by the packet filtering, and acquires the transmission order type of the MPU data (S21).
[0249] Next, the transmission order type discrimination unit 22 judges (discriminates) whether or not the data after packet filtering contains MPU header information (at least one of MPU metadata and MF metadata) (S22). If the MPU header information is included (Yes in S22), the random access unit 23 detects that the fragment type of the MMT payload header is switched to media data, thereby identifying the MPU first sample (S23).
[0250] On the other hand, if the MPU header information is not included (No in S22), the random access unit 23 identifies the MPU first sample based on the RAP_flag in the MMT packet header or the sample number in the MMT payload header (S24).
[0251] Furthermore, the transmission order type discrimination unit 22 judges whether or not MF metadata is included in the packet-filtered data (S25). If it is judged that MF metadata is included (Yes in S25), the data acquisition unit 25 acquires the NAL units by reading the NAL units based on the sample, subsample offset, and size information included in the MF metadata (S26). On the other hand, if it is judged that MF metadata is not included (No in S25), the data acquisition unit 25 acquires the NAL units by reading data of the size of the NAL units in order from the first NAL unit of the sample (S27).
[0252] Note that even if it is determined in step S22 that the MPU header information is included, the receiving device 20 may identify the MPU first sample using the process of step S24 instead of step S23. Furthermore, when it is determined that the MPU header information is included, the process of step S23 and the process of step S24 may be used in combination.
[0253] Furthermore, even if it is determined in step S25 that MF metadata is included, the receiving device 20 may acquire the NAL unit using the process of step S27 without using the process of step S26. Furthermore, when it is determined that MF metadata is included, the process of step S23 and the process of step S24 may be used in combination.
[0254] Also, when it is determined in step S25 that MF metadata is included, it is assumed that the MF data is transmitted after the media data. In this case, the receiving device 20 may buffer the media data and wait until the MF metadata is acquired before performing the process of step S26, or the receiving device 20 may determine whether or not to perform the process of step S27 without waiting for the MF metadata to be acquired.
[0255] For example, the receiving device 20 may determine whether to wait for acquisition of MF metadata based on whether a buffer with a buffer size capable of buffering media data is held. The receiving device 20 may also determine whether to wait for acquisition of MF metadata based on whether the end-to-end delay is reduced. The receiving device 20 may also perform the decoding process mainly using the process of step S26, and use the process of step S27 in the case of a processing mode when a packet loss or the like occurs.
[0256] In addition, if the transmission order type is predetermined, steps S22 and S26 may be omitted, and in this case, the receiving device 20 may determine the method of identifying the MPU first sample and the method of identifying the NAL unit, taking into account the buffer size and the end-to-end delay.
[0257] If the transmission order type is known in advance, the transmission order type discriminator 22 in the receiving device 20 is not necessary.
[0258] 26, a decode command unit 28 outputs the data acquired by the data acquisition unit to a decode unit 29 based on the PTS and DTS calculated by the PTS, DTS calculation unit 26 and the initialization information acquired by the initialization information acquisition unit 27. The decode unit 29 decodes the data, and a presentation unit 30 presents the decoded data.
[0259] [Receiver operation 2] Next, an operation of the receiving device 20 to obtain the initialization information based on the transmission order type and to decode the media data based on the initialization information will be described. FIG. 27 is a flowchart of such an operation.
[0260] First, the packet filtering unit 21 performs packet filtering on the received file. The transmission order type discrimination unit 22 analyzes the SI information obtained by the packet filtering, and acquires the transmission order type (S301).
[0261] Next, the transmission order type discrimination unit 22 judges whether or not MPU metadata has been transmitted (S302). If it is judged that MPU metadata has been transmitted (Yes in S302), the transmission order type discrimination unit 22 judges whether or not MPU metadata has been transmitted prior to media data as a result of the analysis in step S301 (S303). If MPU metadata has been transmitted prior to media data (Yes in S303), the initialization information acquisition unit 27 decodes the media data based on the common initialization information included in the MPU metadata and the initialization information of the sample data (S304).
[0262] On the other hand, if it is determined that the MPU metadata was transmitted after the media data (No in S303), the data acquisition unit 25 buffers the media data until the MPU metadata is acquired (S305), and performs the processing of step S304 after the MPU metadata is acquired.
[0263] Furthermore, if it is determined in step S302 that the MPU metadata has not been transmitted (No in S302), the initialization information acquisition unit 27 decodes the media data based only on the initialization information of the sample data (S306).
[0264] If the decoding of the media data is guaranteed only based on the initialization information of the sample data on the transmitting side, the processes based on the determinations in steps S302 and S303 are not performed, and the process in step S306 is used.
[0265] Furthermore, before step S305, receiving device 20 may determine whether or not to buffer the media data. In this case, if receiving device 20 determines to buffer the media data, it proceeds to the process of step S305, and if receiving device 20 determines not to buffer the media data, it proceeds to the process of step S306. The determination of whether or not to buffer the media data may be made based on the buffer size and occupancy of receiving device 20, or may be made taking into account the end-to-end delay, for example, by selecting the buffer with the smaller end-to-end delay.
[0266] [Receiver operation 3] Here, we will explain the details of the transmission method and reception method when MF metadata is transmitted after the media data ((c) of Figure 21 and (d) of Figure 21). The following explains the case of (d) of Figure 21 as an example. Note that in transmission, only the method of (d) of Figure 21 is used, and no signaling of the transmission order type is performed.
[0267] As mentioned above, when data is transmitted in the order of MPU metadata, media data, and MF metadata, as shown in (d) of FIG. 21, (D-1) The receiving device 20 acquires the MPU metadata, and then acquires the MF metadata, and then decodes the media data. (D-2) After acquiring the MPU metadata, the receiving device 20 decodes the media data without using the MF metadata. There are two possible decoding methods:
[0268] Here, D-1 requires buffering of media data to obtain MF metadata, but since decoding can be performed using MPU header information, it can be decoded by a conventional MP4-compliant receiving device. Meanwhile, D-2 does not require buffering of media data to obtain MF metadata, but since decoding cannot be performed using MF metadata, special processing is required for decoding.
[0269] Moreover, the method of FIG. 21(d) has the advantage that since the MF metadata is transmitted after the media data, no delay occurs due to encapsulation, and the end-to-end delay can be reduced.
[0270] The receiving device 20 can select one of the above two decoding methods depending on the capabilities of the receiving device 20 and the quality of service that the receiving device 20 provides.
[0271] The transmitting device 15 must guarantee that the decoding operation in the receiving device 20 can be performed with reduced occurrence of buffer overflow and underflow. For example, the following parameters can be used as elements for defining the decoder model when decoding using the D-1 method.
[0272] Buffer size for reconfiguring the MPU (MPU buffer) For example, buffer size = maximum rate x maximum MPU time x α, where the maximum rate is the upper limit rate of the profile and level of the encoded data + the overhead of the MPU header, and the maximum MPU time is the maximum time length of a GOP when 1 MPU = 1 GOP (video).
[0273] Here, audio may be in GOP units common to the video, or may be in other units. α is a margin to prevent overflow, and may be multiplied or added to the maximum rate x maximum MPU time. When multiplied, α≧1, and when added, α≧0.
[0274] The upper limit of the decoding delay time from when data is input to the MPU buffer until it is decoded. (TSTD_delay in the MPEG-TS STD) For example, at the time of transmission, the DTS is set so that the acquisition completion time of the MPU data at the receiver is less than or equal to the DTS, taking into consideration the maximum MPU time and the upper limit of the decoding delay time.
[0275] In addition, the transmitting device 15 may provide the DTS and PTS according to a decoder model for decoding using the D-1 method, thereby ensuring that the receiving device performs decoding using the D-1 method, and may also transmit auxiliary information required when decoding is performed using the D-2 method.
[0276] For example, the transmitting device 15 can guarantee the operation of a receiving device that decodes using the D-2 method by signaling the pre-buffering time in the decoder buffer when decoding using the D-2 method.
[0277] The pre-buffering time may be included in SI control information such as a message, a table, or a descriptor, or may be included in the header of an MMT packet or an MMT payload. Also, the SEI in the encoded data may be overwritten. The DTS and PTS for decoding using the D-1 method may be stored in the MPU timestamp descriptor and SampleEntry, and the DTS and PTS for decoding using the D-2 method or the pre-buffering time may be described in the SEI.
[0278] If the receiving device 20 only supports MP4-compliant decoding operations using an MPU header, it may select decoding method D-1, and if it supports both D-1 and D-2, it may select either one of them.
[0279] The transmitting device 15 may provide a DTS and a PTS so as to guarantee the decoding operation of one of the streams (D-1 in this example), and may also transmit auxiliary information for supporting the decoding operation of the other stream.
[0280] Furthermore, when the D-2 method is used, the end-to-end delay is more likely to be large due to the delay caused by pre-buffering of MF metadata than when the D-1 method is used. Therefore, when the receiving device 20 wants to reduce the end-to-end delay, it may select the D-2 method for decoding. For example, the receiving device 20 may always use the D-2 method when it wants to reduce the end-to-end delay. Furthermore, the receiving device 20 may use the D-2 method only when operating in a low-delay presentation mode in which it is desired to present live content, channel selection, zapping, and the like with low delay.
[0281] FIG. 28 is a flow chart of such a receiving method.
[0282] First, the receiving device 20 receives an MMT packet and acquires MPU data (S401). Then, the receiving device 20 (transmission order type determination unit 22) determines whether to present the program in a low-delay presentation mode (S402).
[0283] If the program is not presented in the low-latency presentation mode (No in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using the header information (S405). Also, the receiving device 20 (PTS, DTS calculation unit 26, decoding command unit 28, decoding unit 29, presentation unit 30) performs decoding and presentation processing based on the PTS and DTS assigned by the transmitting side (S406).
[0284] On the other hand, when the program is presented in low-latency presentation mode (Yes in S402), the receiving device 20 (random access unit 23 and initialization information acquisition unit 27) acquires random access and initialization information using a decoding method that does not use header information (S403). The receiving device 20 also performs decoding and presentation processing based on auxiliary information for decoding without using the PTS, DTS, and header information added by the transmitting side (S404). Note that in steps S403 and S404, processing may be performed using MPU metadata.
[0285] [Transmission and reception method using auxiliary data] The above describes the transmission and reception operations in the cases where MF metadata is transmitted after the media data (cases (c) and (d) of Figure 21). Next, a method will be described in which the transmitting device 15 transmits auxiliary data having some of the functions of the MF metadata, thereby enabling decoding to begin earlier and reducing end-to-end delay. Here, an example will be described in which auxiliary data is further transmitted based on the transmission method shown in (d) of Figure 21, but the method using auxiliary data is also applicable to the transmission methods shown in (a) to (c) of Figure 21.
[0286] Fig. 29(a) is a diagram showing an MMT packet transmitted using the method shown in Fig. 21(d). That is, data is transmitted in the order of MPU metadata, media data, and MF metadata.
[0287] Here, sample #1, sample #2, sample #3, and sample #4 are samples included in the media data. Note that, although an example in which the media data is stored in the MMT packet in units of samples is described here, the media data may be stored in the MMT packet in units of NAL units, or may be stored in units obtained by dividing the NAL units. Note that there are also cases in which a plurality of NAL units are aggregated and stored in the MMT packet.
[0288] As described above in D-1, in the case of the method shown in (d) of Fig. 21, that is, when data is transmitted in the order of MPU metadata, media data, and MF metadata, there is a method in which the MPU metadata is acquired, then the MF metadata is acquired, and then the media data is decoded. Such a method of D-1 requires buffering of the media data to acquire the MF metadata, but has the advantage that the method of D-1 can be applied to conventional MP4-compliant receiving devices because the decoding is performed using the MPU header information. On the other hand, it has the disadvantage that the receiving device 20 must wait to start decoding until the MF metadata is acquired.
[0289] In contrast, as shown in (b) of FIG. 29, in the method using auxiliary data, the auxiliary data is transmitted before the MF metadata.
[0290] MF metadata includes information indicating the DTS, PTS, offsets, and sizes of all samples included in a movie fragment, while ancillary data includes information indicating the DTS, PTS, offsets, and sizes of some of the samples included in a movie fragment.
[0291] For example, the MF metadata includes information on all samples (sample #1-sample #4), whereas the auxiliary data includes information on some samples (sample #1-sample #2).
[0292] In the case shown in (b) of Fig. 29, the auxiliary data is used to enable decoding of sample #1 and sample #2, so that the end-to-end delay is smaller than that of the transmission method of D-1. Note that the auxiliary data may include any combination of sample information, and the auxiliary data may be transmitted repeatedly.
[0293] 29(c), when transmitting auxiliary information at timing A, transmitting device 15 includes information of sample #1 in the auxiliary information, and when transmitting auxiliary information at timing B, transmitting device 15 includes information of sample #1 and sample #2 in the auxiliary information. When transmitting auxiliary information at timing C, transmitting device 15 includes information of sample #1, sample #2, and sample #3 in the auxiliary information.
[0294] The MF metadata includes information on sample #1, sample #2, sample #3, and sample #4 (information on all samples in the movie fragment).
[0295] The auxiliary data does not necessarily have to be transmitted immediately after it is generated.
[0296] In addition, in the header of an MMT packet or an MMT payload, a type is specified that indicates that auxiliary data is stored.
[0297] For example, when auxiliary data is stored in the MMT payload using the MPU mode, a data type indicating that it is auxiliary data is specified as a fragment_type field value (e.g., FT=3). The auxiliary data may be data based on the moof configuration, or may have another configuration.
[0298] When auxiliary data is stored as a control signal (descriptor, table, message) in the MMT payload, a descriptor tag, table ID, message ID, etc. indicating that it is auxiliary data are specified.
[0299] In addition, the PTS or DTS may be stored in the header of the MMT packet or the MMT payload.
[0300] [Example of generating auxiliary data] An example in which a transmission device generates auxiliary data based on the configuration of moof will be described below. Fig. 30 is a diagram for explaining an example in which a transmission device generates auxiliary data based on the configuration of moof.
[0301] In normal MP4, a moof is created for a movie fragment as shown in Fig. 20. The moof contains information indicating the DTS, PTS, offset, and size of the samples contained in the movie fragment.
[0302] Here, the transmitting device 15 composes an MP4 (MP4 file) using only a portion of the sample data that constitutes the MPU, and generates auxiliary data.
[0303] For example, as shown in (a) of FIG. 30, the transmitting device 15 generates an MP4 using only sample #1 of samples #1-#4 that make up the MPU, and treats the header of moof+mdat as auxiliary data.
[0304] Next, as shown in (b) of Figure 30, the transmitting device 15 generates an MP4 using samples #1 and #2 of samples #1-#4 that make up the MPU, and treats the header of moof+mdat as the next auxiliary data.
[0305] Next, as shown in (c) of Figure 30, the transmitting device 15 generates an MP4 using samples #1, #2, and #3 of the samples #1-#4 that make up the MPU, and treats the moof+mdat header as the next auxiliary data.
[0306] Next, as shown in (d) of FIG. 30, the transmitting device 15 generates all MP4s from samples #1-#4 that make up the MPU, and the header of moof+mdat among them becomes movie fragment metadata.
[0307] Although the transmitting device 15 generates auxiliary data for each sample here, it may generate auxiliary data for every N samples. The value of N is an arbitrary number, and for example, when transmitting auxiliary data M times when transmitting one MPU, N may be set to total samples / M.
[0308] In addition, the information indicating the offset of a sample in moof may be an offset value after the sample entry area of the subsequent samples is secured as a NULL area.
[0309] The auxiliary data may be generated so as to fragment the MF metadata.
[0310] [Example of reception using auxiliary data] The following describes reception of auxiliary data generated as described in Fig. 30. Fig. 31 is a diagram for explaining reception of auxiliary data. In Fig. 31(a), the number of samples constituting the MPU is 30, and auxiliary data is generated and transmitted every 10 samples.
[0311] In (a) of FIG. 30, auxiliary data #1 includes sample information for samples #1-#10, auxiliary data #2 includes sample information for samples #1-#20, and MF metadata includes sample information for samples #1-#30.
[0312] Although samples #1-#10, samples #11-#20, and samples #21-#30 are stored in one MMT payload, they may be stored in sample units or NAL units, or in fragment or aggregate units.
[0313] The receiving device 20 receives the packets of MPU meta, samples, MF meta and auxiliary data, respectively.
[0314] The receiving device 20 concatenates the sample data in the order in which they are received (to the end), and after receiving the latest auxiliary data, updates the previous auxiliary data. In addition, the receiving device 20 can configure a complete MPU by finally replacing the auxiliary data with MF metadata.
[0315] Upon receiving auxiliary data #1, receiving device 20 concatenates the data to construct an MP4 as shown in the upper part of (b) of Fig. 31. This enables receiving device 20 to parse samples #1-#10 using the MPU metadata and information in auxiliary data #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.
[0316] Furthermore, upon receiving auxiliary data #2, receiving device 20 concatenates the data as shown in the middle part of (b) of Fig. 31 to construct an MP4 file. This enables receiving device 20 to parse samples #1-#20 using the MPU metadata and information on auxiliary data #2, and to perform decoding based on the PTS, DTS, offset, and size information included in the auxiliary data.
[0317] Furthermore, upon receiving the MF metadata, the receiving device 20 concatenates the data as shown in the lower part of (b) of Fig. 31 to construct an MP4. This enables the receiving device 20 to parse samples #1-#30 using the MPU metadata and the MF metadata, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF metadata.
[0318] In the absence of auxiliary data, receiving device 20 can obtain sample information only after receiving MF metadata, and therefore needs to start decoding after receiving MF metadata. However, by transmitting device 15 generating and transmitting auxiliary data, receiving device 20 can obtain sample information using the auxiliary data without waiting for reception of MF metadata, and therefore the decoding start time can be advanced. Furthermore, by transmitting device 15 generating auxiliary data based on moof described with reference to Fig. 30, receiving device 20 can parse using a conventional MP4 parser as is.
[0319] Furthermore, the newly generated auxiliary data and MF metadata contain sample information that overlaps with auxiliary data transmitted in the past. Therefore, even if the previous auxiliary data cannot be acquired due to packet loss, etc., it is possible to reconstruct the MP4 and acquire sample information (PTS, DTS, size, and offset) by using the newly acquired auxiliary data and MF metadata.
[0320] In addition, the auxiliary data does not necessarily have to include information on past sample data. For example, the auxiliary data #1 may correspond to the sample data #1-#10, and the auxiliary data #2 may correspond to the sample data #11-#20. For example, as shown in (c) of Fig. 31, the transmission device 15 may transmit complete MF metadata as a data unit, and fragment units of the data unit as auxiliary data in sequence.
[0321] Furthermore, in order to deal with packet loss, the transmitting device 15 may repeatedly transmit the auxiliary data or the MF metadata.
[0322] In addition, the MMT packet and MMT payload in which auxiliary data is stored include an MPU sequence number and an asset ID, as well as the MPU metadata, MF metadata, and sample data.
[0323] The above-mentioned receiving operation using auxiliary data will be described with reference to the flowchart in Fig. 32. Fig. 32 is a flowchart of the receiving operation using auxiliary data.
[0324] First, the receiving device 20 receives an MMT packet and analyzes the packet header and payload header (S501). Next, the receiving device 20 analyzes whether the fragment type is auxiliary data or MF metadata (S502), and if the fragment type is auxiliary data, overwrites and updates the previous auxiliary data (S503). At this time, if there is no previous auxiliary data for the same MPU, the receiving device 20 treats the received auxiliary data as new auxiliary data as it is. Then, the receiving device 20 acquires samples based on the MPU metadata, auxiliary data, and sample data, and performs decoding (S507).
[0325] On the other hand, if the fragment type is MF metadata, the receiving device 20 overwrites the previous auxiliary data with the MF metadata in step S505 (S505).Then, the receiving device 20 obtains the sample in the form of a complete MPU based on the MPU metadata, MF metadata, and sample data, and performs decoding (S506).
[0326] Although not shown in Figure 32, in step S502, if the fragment type is MPU metadata, the receiving device 20 stores the data in a buffer, and if the fragment type is sample data, the receiving device 20 stores the data concatenated at the end for each sample in the buffer.
[0327] If auxiliary data cannot be obtained due to packet loss, the receiving device 20 can decode the sample by overwriting it with the latest auxiliary data or by using the previous auxiliary data.
[0328] The transmission cycle and the number of transmissions of the auxiliary data may be predetermined values. Information on the transmission cycle and the number of times (count, count down) may be transmitted together with the data. For example, the transmission cycle, the number of times of transmission, and time stamps such as initial_cpb_removal_delay may be stored in the data unit header.
[0329] By sending auxiliary data including information on the first sample of the MPU at least once before initial_cpb_removal_delay, it is possible to comply with the CPB buffer model. In this case, the MPU timestamp descriptor is set to a value based on the picture timing SEI.
[0330] In addition, the transmission method in the receiving operation in which such auxiliary data is used is not limited to the MMT method, but can be applied to the streaming transmission of packets configured in the ISOBMFF file format, such as MPEG-DASH.
[0331] [Transmission method when one MPU consists of multiple movie fragments] In the above explanation from Fig. 19 onwards, one MPU is composed of one movie fragment, but here, we will explain the case where one MPU is composed of multiple movie fragments. Fig. 33 is a diagram showing the configuration of an MPU composed of multiple movie fragments.
[0332] In Figure 33, samples (#1-#6) stored in one MPU are divided into two movie fragments. The first movie fragment is generated based on samples #1-#3, and the corresponding moof boxes are generated. The second movie fragment is generated based on samples #4-#6, and the corresponding moof boxes are generated.
[0333] The headers of the moof box and mdat box in the first movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #1. Meanwhile, the headers of the moof box and mdat box in the second movie fragment are stored in the MMT payload and MMT packet as movie fragment metadata #2. In FIG. 33, the MMT payload in which the movie fragment metadata is stored is hatched.
[0334] The number of samples constituting the MPU and the number of samples constituting the movie fragment are arbitrary. For example, the number of samples constituting the MPU may be the number of samples in a GOP unit, and two movie fragments may be constituted by using half the number of samples in the GOP unit as the movie fragment.
[0335] Note that, although an example in which one MPU contains two movie fragments (a moof box and an mdat box) is shown here, one MPU may contain three or more movie fragments, rather than two. Also, the samples stored in a movie fragment do not have to be divided into equal samples, and may be divided into any number of samples.
[0336] In addition, in Fig. 33, the MPU metadata unit and the MF metadata unit are each stored as a data unit in the MMT payload. However, the transmitting device 15 may store units such as ftyp, mmpu, moov, and moof as data units in the MMT payload in data unit units, or may store the data units in the MMT payload in fragmented units. In addition, the transmitting device 15 may store the data units in aggregated units in the MMT payload.
[0337] Also, in Fig. 33, samples are stored in the MMT payload in sample units. However, the transmitting device 15 may configure a data unit in NAL unit units or units that aggregate a plurality of NAL units, instead of in sample units, and store the data unit units in the MMT payload. Also, the transmitting device 15 may store the data unit in the MMT payload in fragmented units, or may store the data unit in aggregated units in the MMT payload.
[0338] In Fig. 33, the MPU is configured in the order of moof#1, mdat#1, moof#2, and mdat#2, and an offset is added to moof#1 assuming that the corresponding mdat#1 is attached at the end. However, an offset may be added to mdat#1 assuming that it is attached before moof#1. In this case, however, movie fragment metadata cannot be generated in the form of moof+mdat, and the headers of moof and mdat are transmitted separately.
[0339] Next, a description will be given of the transmission order of MMT packets when transmitting an MPU having the configuration described in Fig. 33. Fig. 34 is a diagram for explaining the transmission order of MMT packets.
[0340] Figure 34(a) shows the transmission order when transmitting MMT packets in the configuration order of the MPUs shown in Figure 33. Figure 34(a) specifically shows an example in which MPU meta, MF meta #1, media data #1 (samples #1-#3), MF meta #2, and media data #2 (samples #4-#6) are transmitted in this order.
[0341] FIG. 34(b) shows an example in which MPU meta, media data #1 (samples #1-#3), MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in this order.
[0342] FIG. 34(c) shows an example in which media data #1 (samples #1-#3), MPU meta, MF meta #1, media data #2 (samples #4-#6), and MF meta #2 are transmitted in that order.
[0343] MF meta #1 is generated using samples #1-#3, and MF meta #2 is generated using samples #4-#6. Therefore, when the transmission method of Fig. 34(a) is used, a delay occurs in the transmission of sample data due to encapsulation.
[0344] In contrast, when the transmission methods of Figures 34(b) and 34(c) are used, samples can be transmitted without waiting for the MF meta to be generated, so no delay due to encapsulation occurs and end-to-end delay can be reduced.
[0345] Also, in the (a) transmission order of Figure 34, one MPU is divided into multiple movie fragments, and the number of samples stored in the MF meta is smaller than in the case of Figure 19, so the amount of delay due to encapsulation can be made smaller than in the case of Figure 19.
[0346] In addition to the method shown here, for example, the transmitting device 15 may concatenate MF meta #1 and MF meta #2 and transmit them together at the end of the MPU. In this case, MF meta of different movie fragments may be aggregated and stored in one MMT payload. Also, MF meta of different MPUs may be aggregated and stored in the MMT payload.
[0347] [How to receive when one MPU consists of multiple movie fragments] Here, an example of operation of the receiving device 20 that receives and decodes MMT packets transmitted in the transmission order described in (b) of Fig. 34 will be described. Figs. 35 and 36 are diagrams for explaining such an example of operation.
[0348] Receiving device 20 receives MMT packets including MPU meta, samples, and MF meta, each transmitted in the transmission order as shown in Fig. 35. The sample data is concatenated in the order in which it is received.
[0349] At T1, which is the time when MF meta #1 is received, receiving device 20 concatenates the data as shown in (1) of Fig. 36 to construct an MP4. This enables receiving device 20 to acquire samples #1-#3 based on the MPU metadata and information on MF meta #1, and to perform decoding based on the PTS, DTS, offset, and size information included in the MF meta.
[0350] Furthermore, at T2, which is the time when MF meta #2 is received, receiving device 20 concatenates the data as shown in (2) of Fig. 36 to construct an MP4. This allows receiving device 20 to acquire samples #4-#6 based on the MPU metadata and information in MF meta #2, and to perform decoding based on the PTS, DTS, offset, and size information in the MF meta. Receiving device 20 may also acquire samples #1-#6 based on information in MF meta #1 and MF meta #2 by concatenating the data as shown in (3) of Fig. 36 to construct an MP4.
[0351] By dividing a single MPU into multiple movie fragments, the time it takes for the MPU to obtain the first MF meta is shortened, so the decoding start time can be advanced. Also, the buffer size for storing samples before decoding can be reduced.
[0352] The transmitting device 15 may set the division unit of the movie fragment so that the time from transmitting (or receiving) the first sample in the movie fragment to transmitting (or receiving) the MF meta corresponding to the movie fragment is shorter than the initial_cpb_removal_delay specified by the encoder. By setting in this way, the receiving buffer can follow the cpb buffer, and low-delay decoding can be realized. In this case, absolute time based on the initial_cpb_removal_delay can be used for the PTS and DTS.
[0353] The transmitting device 15 may also divide the movie fragments at equal intervals, or divide the succeeding movie fragments at shorter intervals than the preceding movie fragments, so that the receiving device 20 can always receive the MF meta including the information of the sample before decoding the sample, enabling continuous decoding.
[0354] The absolute time of the PTS and DTS can be calculated using the following two methods.
[0355] (1) The absolute times of the PTS and DTS are determined based on the reception time (T1 or T2) of the MF meta #1 or MF meta #2, and the relative times of the PTS and DTS included in the MF meta.
[0356] (2) The absolute time of the PTS and DTS is determined based on the absolute time signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative time of the PTS and DTS included in the MF meta.
[0357] Also, (2-A) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on the initial_cpb_removal_delay specified by the encoder.
[0358] Also, (2-B) the absolute time signaled by the transmitting device 15 may be an absolute time calculated based on a predicted value of the receiving time of the MF meta.
[0359] Note that MF meta #1 and MF meta #2 may be transmitted repeatedly. By repeatedly transmitting MF meta #1 and MF meta #2, even if the receiving device 20 is unable to acquire the MF meta due to packet loss or the like, it is possible to acquire the MF meta again.
[0360] An identifier indicating the order of the movie fragments can be stored in the payload header of an MFU including samples constituting a movie fragment. On the other hand, an identifier indicating the order of the MF meta constituting a movie fragment is not included in the MMT payload. Therefore, the receiving device 20 identifies the order of the MF meta by packet_sequence_number. Alternatively, the transmitting device 15 may store and signal an identifier indicating which movie fragment the MF meta belongs to in control information (message, table, descriptor), MMT header, MMT payload header, or data unit header.
[0361] The transmitting device 15 may transmit the MPU meta, MF meta, and samples in a predetermined transmission order, and the receiving device 20 may perform the receiving process based on the predetermined transmission order. Also, the transmitting device 15 may signal the transmission order, and the receiving device 20 may select (determine) the receiving process based on the signaling information.
[0362] The above-mentioned receiving method will be explained with reference to Fig. 37. Fig. 37 is a flowchart of the operation of the receiving method explained with Figs.
[0363] First, the receiving device 20 determines (identifies) whether the data included in the payload is MPU metadata, MF metadata, or sample data (MFU) based on the fragment type indicated in the MMT payload (S601, S602). If the data is sample data, the receiving device 20 buffers the sample and waits for reception of MF metadata corresponding to the sample and for the start of decoding (S603).
[0364] On the other hand, in step S602, if the data is MF metadata, the receiving device 20 obtains sample information (PTS, DTS, position information, and size) from the MF metadata, obtains a sample based on the obtained sample information, and decodes and presents the sample based on the PTS and DTS (S604).
[0365] Although not shown, if the data is MPU metadata, the MPU metadata contains initialization information necessary for decoding. Therefore, the receiving device 20 stores this information and uses it to decode the sample data in step S604.
[0366] When the receiving device 20 stores the received MPU data (MPU metadata, MF metadata, and sample data) in a storage device, the receiving device 20 stores the data after rearranging it into the MPU configuration described in Figure 19 or Figure 33.
[0367] In addition, on the transmitting side, a packet sequence number is assigned to an MMT packet having the same packet ID. At this time, the packet sequence number may be assigned after the MMT packets including the MPU metadata, MF metadata, and sample data are rearranged in the transmission order, or the packet sequence number may be assigned in the order before rearrangement.
[0368] If packet sequence numbers are assigned in the order before rearrangement, the data can be rearranged in the receiving device 20 based on the packet sequence numbers into the configuration order of the MPU, facilitating storage.
[0369] [Method for detecting the beginning of an access unit and the beginning of a slice segment] A method for detecting the start of an access unit or the start of a slice segment based on information in the MMT packet header and the MMT payload header will be described.
[0370] Here, two examples are shown: a case where non-VCL NAL units (such as access unit delimiters, VPS, SPS, PPS, and SEI) are collectively stored as data units in an MMT payload, and a case where each non-VCL NAL unit is treated as a data unit, and the data units are aggregated and stored in a single MMT payload.
[0371] FIG. 38 is a diagram showing a case where non-VCL NAL units are individually treated as data units and aggregated.
[0372] 38, the head of the access unit is an MMT packet whose fragment_type value is MFU, and is the head data of an MMT payload including a data unit whose aggregation_flag value is 1 and whose offset value is 0. At this time, the fragmentation_indicator value is 0.
[0373] Also, in the case of Figure 38, the start of the slice segment is an MMT packet whose fragment_type value is MFU, and is the start data of an MMT payload whose aggregation_flag value is 0 and whose fragmentation_indicator value is 00 or 01.
[0374] 39 is a diagram showing a case where non-VCL NAL units are grouped together into a data unit. Note that the field values of the packet header are as shown in FIG. 17 (or FIG. 18).
[0375] In the case of FIG. 39, the head of the access unit is the head data of the payload in a packet with an Offset value of 0.
[0376] Also, in the case of FIG. 39, the start of a slice segment is the first data of the payload of a packet whose offset value is a value other than 0 and whose fragmentation indicator value is 00 or 01.
[0377] [Reception process when packet loss occurs] Typically, when transmitting MP4 format data in an environment where packet loss occurs, the receiving device 20 restores packets using ALFEC (Application Layer FEC), packet retransmission control, or the like.
[0378] However, if packet loss occurs in streaming such as broadcasting when AL-FEC cannot be used, the packets cannot be restored.
[0379] After data is lost due to packet loss, the receiving device 20 needs to resume decoding of video and audio again. To do so, the receiving device 20 needs to detect the beginning of an access unit or NAL unit and start decoding from the beginning of the access unit or NAL unit.
[0380] However, since there is no start code at the beginning of an NAL unit in MP4 format, even if the receiving device 20 analyzes the stream, it cannot detect the beginning of an access unit or an NAL unit.
[0381] FIG. 40 is a flowchart of the operation of the receiving device 20 when a packet loss occurs.
[0382] The receiving device 20 detects packet loss using a packet sequence number, a packet counter, a fragment counter, and the like in the header of an MMT packet or an MMT payload (S701), and determines which packet has been lost based on the context (S702).
[0383] If it is determined that no packet loss has occurred (No in S702), the receiving device 20 constructs an MP4 file and decodes the access units or NAL units (S703).
[0384] When it is determined that a packet loss has occurred (Yes in S702), the receiving device 20 generates a NAL unit corresponding to the NAL unit that has experienced the packet loss using dummy data, and constructs an MP4 file (S704). When putting dummy data into a NAL unit, the receiving device 20 indicates that the NAL unit type is dummy data.
[0385] In addition, the receiving device 20 can resume decoding by detecting the start of the next access unit or NAL unit and inputting the start data into the decoder based on the methods described in Figures 17, 18, 38, and 39 (S705).
[0386] In addition, if packet loss occurs, the receiving device 20 may resume decoding from the beginning of the access unit and NAL unit based on information detected based on the packet header, or may resume decoding from the beginning of the access unit and NAL unit based on header information of the reconstructed MP4 file, which includes a NAL unit of dummy data.
[0387] When storing an MP4 file (MPU), the receiving device 20 may separately acquire and store (replace) packet data (such as NAL units) lost due to packet loss from broadcasting or communication.
[0388] At this time, when receiving device 20 obtains the lost packet from the communication, it notifies the server of the information of the lost packet (packet ID, MPU sequence number, packet sequence number, IP data flow number, IP address, etc.) and obtains the packet. Receiving device 20 may obtain not only the lost packet, but also a group of packets before and after the lost packet at the same time.
[0389] [How to configure a movie fragment] Here we will explain in detail how to configure movie fragments.
[0390] As described in Fig. 33, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are arbitrary. For example, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU may be a fixed number or may be dynamically determined.
[0391] Here, by configuring the movie fragment on the transmitting side (transmitting device 15) so as to satisfy the following conditions, low-delay decoding in the receiving device 20 can be guaranteed.
[0392] The conditions are as follows:
[0393] The transmitting device 15 generates and transmits MF meta in the form of movie fragments, which are units obtained by dividing sample data, so that the receiving device 20 can always receive MF meta containing information about any sample (Sample(i)) before the decoding time (DTS(i)) of that sample.
[0394] Specifically, the transmitting device 15 constructs a movie fragment using samples (including the i-th sample) that have been encoded prior to DTS(i).
[0395] To ensure low-latency decoding, the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU can be dynamically determined by, for example, the following method.
[0396] (1) At the start of decoding, the decoding time DTS(0) of the first sample Sample(0) of the GOP is based on initial_cpb_removal_delay. The transmitting device constructs a first movie fragment using samples that have already been encoded at a time before DTS(0). The transmitting device 15 also generates MF metadata corresponding to the first movie fragment and transmits it at a time before DTS(0).
[0397] (2) The transmitting device 15 constructs movie fragments so that the above conditions are satisfied for subsequent samples as well.
[0398] For example, if the first sample of a movie fragment is the kth sample, the MF meta of the movie fragment including the kth sample is transmitted by the decoding time DTS(k) of the kth sample. If the encoding completion time of the lth sample is before DTS(k) and the encoding completion time of the (l+1)th sample is after DTS(k), the transmitting device 15 constructs a movie fragment using the kth sample to the lth sample.
[0399] In addition, the transmitting device 15 may construct a movie fragment using the kth sample through to the lth sample.
[0400] (3) After completing the encoding of the last MPU sample, the transmission device 15 constructs a movie fragment using the remaining samples, generates MF metadata corresponding to the movie fragment, and transmits it.
[0401] In addition, the transmitting device 15 may construct a movie fragment using only a portion of the samples that have been completely coded, rather than using all of the samples that have been completely coded.
[0402] In the above, an example has been shown in which the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU are dynamically determined based on the above conditions so as to ensure low-latency decoding. However, the method of determining the number of samples and the number of movie fragments is not limited to such a method. For example, the number of movie fragments constituting one MPU may be fixed to a predetermined value, and the number of samples may be determined so as to satisfy the above conditions. Also, the number of movie fragments constituting one MPU and the time at which the movie fragments are divided (or the amount of code of the movie fragments) may be fixed to a predetermined value, and the number of samples may be determined so as to satisfy the above conditions.
[0403] In addition, if the MPU is divided into multiple movie fragments, information indicating whether the MPU is divided into multiple movie fragments, attributes of the divided movie fragments, or MF meta attributes for the divided movie fragments may be transmitted.
[0404] Here, the attribute of a movie fragment is information indicating whether the movie fragment is the first movie fragment of an MPU, the last movie fragment of an MPU, or some other movie fragment.
[0405] In addition, the attributes of the MF meta are information that indicates whether the MF meta corresponds to the first movie fragment of the MPU, the last movie fragment of the MPU, or the MF meta corresponding to some other movie fragment.
[0406] The transmitting device 15 may store and transmit the number of samples constituting a movie fragment and the number of movie fragments constituting one MPU as control information.
[0407] [Operation of receiving device] The operation of the receiving device 20 based on the movie fragment configured as above will be described.
[0408] The receiving device 20 determines the absolute times of the PTS and DTS based on the absolute times signaled from the transmitting side, such as the MPU timestamp descriptor, and the relative times of the PTS and DTS included in the MF meta.
[0409] Based on information on whether the MPU is divided into multiple movie fragments, the receiving device 20 performs the following processing based on the attributes of the divided movie fragments if the MPU is divided.
[0410] (1) When the movie fragment is the first movie fragment of an MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the absolute time of the PTS of the first sample included in the MPU timestamp descriptor and the relative times of the PTS and DTS included in the MF meta.
[0411] (2) If the movie fragment is not the first movie fragment of the MPU, the receiving device 20 generates the absolute times of the PTS and DTS using the relative times of the PTS and DTS included in the MF meta, without using the information in the MPU timestamp descriptor.
[0412] (3) When the movie fragment is the last movie fragment of the MPU, the receiving device 20 calculates the absolute times of the PTS and DTS of all samples, and then resets the calculation process of the PTS and DTS (addition process of the relative times). Note that the reset process may be performed on the first movie fragment of the MPU.
[0413] The receiving device 20 may determine whether a movie fragment is divided as follows: The receiving device 20 may also acquire attribute information of the movie fragment as follows.
[0414] For example, the receiving device 20 may determine whether the movie fragment has been divided based on the movie_fragment_sequence_number field value, which is an identifier indicating the order of the movie fragments indicated in the MMTP (MMT Protocol) payload header.
[0415] Specifically, the receiving device 20 may determine that an MPU is divided into multiple movie fragments if the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 1, and the field value is 2 or greater.
[0416] In addition, the receiving device 20 may determine that an MPU is divided into multiple movie fragments when the number of movie fragments contained in one MPU is 1, the movie_fragment_sequence_number field value is 0, and the field value is a value other than 0.
[0417] The attribute information of the movie fragment may also be determined based on the movie_fragment_sequence_number.
[0418] In addition, without using movie_freagment_sequence_number, it is also possible to determine whether a movie fragment is divided and the attribute information of the movie fragment by counting the transmission of movie fragments and MF meta contained in one MPU.
[0419] With the above-described configurations of the transmitting device 15 and the receiving device 20, the receiving device 20 can receive movie fragment metadata at intervals shorter than that of the MPU, and can start decoding with low delay. Also, it becomes possible to perform decoding with low delay by using a decoding process based on the MP4 parsing method.
[0420] The reception operation when the MPU is divided into multiple movie fragments as described above will be described using a flowchart. Figure 41 is a flowchart of the reception operation when the MPU is divided into multiple movie fragments. This flowchart illustrates the operation of step S604 in Figure 37 in more detail.
[0421] First, the receiving device 20 acquires MF metadata if the data type is MF meta based on the data type indicated in the MMTP payload header (S801).
[0422] Next, the receiving device 20 determines whether the MPU is divided into multiple movie fragments (S802), and if the MPU is divided into multiple movie fragments (Yes in S802), determines whether the received MF metadata is the first metadata in the MPU (S803). If the received MF metadata is the first MF metadata in the MPU (Yes in S803), the receiving device 20 calculates the absolute times of the PTS and DTS from the absolute time of the PTS indicated in the MPU timestamp descriptor and the relative times of the PTS and DTS indicated in the MF metadata (S804), and determines whether the metadata is the last metadata in the MPU (S805).
[0423] On the other hand, if the received MF metadata is not the first MF metadata in the MPU (No in S803), the receiving device 20 does not use the information in the MPU timestamp descriptor, but calculates the absolute times of the PTS and DTS using the relative times of the PTS and DTS indicated in the MF metadata (S808), and proceeds to processing of step S805.
[0424] In step S805, if it is determined that this is the last MF metadata of the MPU (Yes in S805), the receiving device 20 calculates the absolute times of the PTS and DTS of all samples, and then resets the PTS and DTS calculation process. If it is determined in step S805 that this is not the last MF metadata of the MPU (No in S805), the receiving device 20 ends the process.
[0425] Also, if it is determined in step S802 that the MPU is not divided into multiple movie fragments (No in S802), the receiving device 20 obtains sample data based on the MF metadata transmitted after the MPU and determines the PTS and DTS (S807).
[0426] Then, although not shown, the receiving device 20 finally performs decoding and presentation processing based on the determined PTS and DTS.
[0427] [Issues that arise when splitting movie fragments and their solutions] So far, we have explained how to reduce the end-to-end delay by splitting movie fragments. From here, we will explain the new issues that arise when splitting movie fragments and how to solve them.
[0428] First, as background, a picture structure in encoded data will be described. Fig. 42 is a diagram showing an example of a prediction structure of a picture for each TemporalId when implementing temporal scalability.
[0429] In coding formats such as MPEG-4 AVC and High Efficiency Video Coding (HEVC), temporal scalability can be achieved by using B pictures (bidirectional reference predictive pictures) that can be referenced from other pictures.
[0430] TemporalId shown in (a) of FIG. 42 is an identifier of a hierarchy of the coding structure, and the larger the value of TemporalId, the deeper the hierarchy. The rectangular blocks indicate pictures, and Ix in the blocks indicates an I picture (intra-screen predicted picture), Px indicates a P picture (forward reference predicted picture), and Bx and bx indicate B pictures (bidirectional reference predicted pictures). The x in Ix / Px / Bx indicates a display order, which indicates the order in which pictures are displayed. The arrows between pictures indicate a reference relationship, and for example, picture B4 indicates that a predicted image is generated using I0 and B8 as reference images. Here, it is prohibited for one picture to use another picture with a TemporalId larger than its own TemporalId as a reference image. The hierarchies are defined to allow for temporal scalability. For example, decoding all pictures in Figure 42 results in 120 fps (frames per second) video, but decoding only the hierarchies with TemporalId from 0 to 3 results in 60 fps video.
[0431] Figure 43 shows the relationship between the decoding time (DTS) and the display time (PTS) for each picture in Figure 42. For example, picture I0 shown in Figure 43 is displayed after the decoding of B4 is completed so that there is no gap in the decoding and display.
[0432] As shown in Figure 43, when the prediction structure includes a B picture, the decoding order and the display order are different, so that after decoding the picture in the receiving device 20, picture delay processing and picture reordering processing are required.
[0433] Although an example of a prediction structure of a picture in time scalability has been described above, even if time scalability is not used, delay processing and reorder processing of a picture may be required depending on the prediction structure. Figure 44 is a diagram showing an example of a prediction structure of a picture that requires delay processing and reorder processing of a picture. Note that the numbers in Figure 44 indicate the decoding order.
[0434] As shown in Figure 44, depending on the prediction structure, the first sample in the decoding order may be different from the first sample in the presentation order, and in Figure 44, the first sample in the presentation order is the fourth sample in the decoding order. Note that Figure 44 shows an example of a prediction structure, and the prediction structure is not limited to this structure. In other prediction structures, the first sample in the decoding order may be different from the first sample in the presentation order.
[0435] Fig. 45 is a diagram showing an example in which an MPU in MP4 format is divided into multiple movie fragments and stored in an MMTP payload and an MMTP packet, similar to Fig. 33. The number of samples constituting an MPU and the number of samples constituting a movie fragment are arbitrary. For example, the number of samples constituting an MPU may be the number of samples in a GOP unit, and two movie fragments may be constituted by using half the number of samples in a GOP unit as a movie fragment. One sample may be one movie fragment, or the samples constituting an MPU may not be divided.
[0436] In FIG. 45, an example is shown in which one MPU includes two movie fragments (a moof box and an mdat box), but the number of movie fragments included in one MPU does not have to be two. The number of movie fragments included in one MPU may be three or more, or the number of samples included in the MPU. Furthermore, the samples stored in a movie fragment may not be an equal number of samples, but may be divided into any number of samples.
[0437] Movie fragment metadata (MF metadata) contains information on the PTS, DTS, offset, and size of the samples contained in the movie fragment, and when the receiving device 20 decodes a sample, it extracts the PTS and DTS from the MF meta which contains information about the sample, and determines the decoding timing and presentation timing.
[0438] From here on, for the sake of detailed explanation, the absolute value of the decoding time of the i sample will be written as DTS(i) and the absolute value of the presentation time will be written as PTS(i).
[0439] The information of the i-th sample among the timestamp information stored in the moof in the MF meta is specifically the relative value of the decoding time of the i-th sample and the (i+1)-th sample, and the relative value of the decoding time of the i-th sample and the presentation time, which will be hereinafter referred to as DT(i) and CT(i).
[0440] Movie fragment metadata #1 contains DT(i) and CT(i) for samples #1-#3, and movie fragment metadata #2 contains DT(i) and CT(i) for samples #4-#6.
[0441] Furthermore, the PTS absolute value of the access unit at the beginning of the MPU is stored in the MPU timestamp descriptor or the like, and the receiving device 20 calculates the PTS and DTS based on the PTS_MPU of the access unit at the beginning of the MPU, and the CT and DT.
[0442] FIG. 46 is a diagram for explaining a method of calculating PTS and DTS and problems involved when an MPU is configured from samples #1 to #10.
[0443] (a) of Figure 46 shows an example where the MPU is not divided into movie fragments, (b) of Figure 46 shows an example where the MPU is divided into two movie fragments of 5 sample units, and (c) of Figure 46 shows an example where the MPU is divided into 10 movie fragments of sample units.
[0444] As described in Fig. 45, when the PTS and DTS are calculated using the MPU timestamp descriptor and the timestamp information (CT and DT) in the MP4, the first sample in the presentation order in Fig. 44 is the fourth sample in the decoding order. Therefore, the PTS stored in the MPU timestamp descriptor is the PTS (absolute value) of the fourth sample in the decoding order. Note that hereinafter, this sample will be referred to as sample A. Also, the first sample in the decoding order will be referred to as sample B.
[0445] Since the only absolute time information related to the timestamp is the information in the MPU timestamp descriptor, the receiving device 20 cannot calculate the PTS (absolute time) and DTS (absolute time) of other samples until the arrival of the A sample. The receiving device 20 cannot calculate the PTS and DTS of the B sample either.
[0446] 46(a), sample A is included in the same movie fragment as sample B and is stored in one MF meta, so that receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.
[0447] 46(b), sample A is included in the same movie fragment as sample B and stored in one MF meta, so that receiving device 20 can determine the DTS of sample B immediately after receiving the MF meta.
[0448] In the example of (c) in Figure 46, sample A is included in a movie fragment different from sample B. For this reason, receiving device 20 cannot determine the DTS of sample B until it receives MF meta including the CT and DT of the movie fragment including sample A.
[0449] Therefore, in the example of FIG. 46(c), the receiving device 20 cannot start decoding immediately after the arrival of the B sample.
[0450] In this way, if a movie fragment containing a B sample does not contain an A sample, the receiving device 20 cannot start decoding the B sample until it has received the MF meta for the movie fragment containing the A sample.
[0451] This issue occurs when the first sample in presentation order does not match the first sample in decoding order, and the movie fragment is split to the point where samples A and B are no longer stored in the same movie fragment. This issue occurs regardless of whether the MF meta is forward or backward.
[0452] In this way, when the first sample in presentation order does not match the first sample in decoding order, if sample A and sample B are not stored in the same movie fragment, the DTS cannot be determined immediately after receiving sample B. Therefore, transmitting device 15 separately transmits the DTS (absolute value) of sample B or information that allows the receiving side to calculate the DTS (absolute value) of sample B. Such information may be transmitted using control information, a packet header, etc.
[0453] Using such information, the receiving device 20 calculates the DTS (absolute value) of the B sample. Figure 47 is a flowchart of the receiving operation when the DTS is calculated using such information.
[0454] Receiving device 20 receives the movie fragment at the beginning of the MPU (S901), and determines whether sample A and sample B are stored in the same movie fragment (S902). If they are stored in the same movie fragment (Yes in S902), receiving device 20 calculates the DTS using only the information in the MF meta, without using the DTS (absolute time) of sample B, and starts decoding (S904). Note that in step S904, receiving device 20 may determine the DTS using the DTS of sample B.
[0455] On the other hand, if sample A and sample B are not stored in the same movie fragment in step S902 (No in S902), receiving device 20 obtains the DTS (absolute time) of sample B, determines the DTS, and starts decoding (S903).
[0456] In the above description, an example has been described in which the absolute value of the decoded time and the absolute value of the presentation time of each sample are calculated using the MF meta (time stamp information stored in the moof in MP4 format) in the MMT standard, but it goes without saying that the MF meta may be replaced with any control information that can be used to calculate the absolute value of the decoded time and the absolute value of the presentation time of each sample. Examples of such control information include control information in which the above-mentioned relative value CT(i) of the decoded time of the i-th sample and the (i+1)-th sample is replaced with the relative value of the presentation time of the i-th sample and the (i+1)-th sample, and control information that includes both the relative value CT(i) of the decoded time of the i-th sample and the (i+1)-th sample and the relative value of the presentation time of the i-th sample and the (i+1)-th sample.
[0457] (Embodiment 3) [overview] In the third embodiment, a content transmission method and data structure when content such as video, audio, subtitles, and data broadcasting is transmitted by broadcasting will be described. That is, a content transmission method and data structure specialized for playing back a broadcast stream will be described.
[0458] In the third embodiment, an example will be described in which the MMT method (hereinafter, also simply referred to as MMT) is used as the multiplexing method, but other multiplexing methods such as MPEG-DASH or RTP may also be used.
[0459] First, a method of storing a data unit (DU) in a payload in MMT will be described in detail. Fig. 48 is a diagram for explaining a method of storing a data unit in a payload in MMT.
[0460] In MMT, the transmitting device stores part of the data that constitutes the MPU in the MMTP payload as a data unit, and transmits it with a header. The header includes an MMTP payload header and an MMTP packet header. The data unit may be in NAL unit units or sample units. When an MMTP packet is scrambled, the payload is the target of scrambling.
[0461] Fig. 48(a) shows an example in which a transmitting device aggregates multiple data units and stores them in one payload. In the example of Fig. 48(a), a data unit header (DUH: Data Unit Header) and a data unit length (DUL: Data Unit Length) are added to the beginning of each of the multiple data units, and multiple data units with the data unit header and data unit length added are stored together in the payload.
[0462] Figure 48(b) shows an example of storing one data unit in one payload. In the example of Figure 48(b), a data unit header is added to the beginning of the data unit and stored in the payload. Figure 48(c) shows an example of dividing one data unit, adding a data unit header to the divided data units and storing them in the payload.
[0463] There are various types of data units, such as timed-MFU, which is a media including information on synchronization such as video, audio, or subtitles, non-timed-MFU, which is a media including no information on synchronization such as a file, MPU metadata, and MF metadata, and a data unit header is determined according to the type of data unit. Note that there is no data unit header in MPU metadata and MF metadata. In addition, although a transmitting device cannot aggregate different types of data units in principle, it may be specified to be able to aggregate different types of data units. For example, when the size of MF metadata is small, such as when it is divided into movie fragments for each sample, the number of packets can be reduced by aggregating MF metadata and media data, and further, the transmission capacity can be reduced.
[0464] If the data unit is an MFU, some information about the MPU, such as information for configuring the MPU (MP4), is stored as a header.
[0465] For example, the header of a timed-MFU includes movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter, while the header of a non-timed-MFU includes item_iD. The meaning of each field is specified in standards such as ISO / IEC23008-1 or ARIB STD-B60. The meaning of each field specified in such standards will be described below.
[0466] The movie_fragment_sequence_number indicates the sequence number of the movie fragment to which the MFU belongs, and is also specified in ISO / IEC14496-12.
[0467] The sample_number indicates the sample number to which the MFU belongs, and is also indicated in ISO / IEC14496-12.
[0468] The offset indicates the offset amount of the MFU in the sample to which the MFU belongs, in bytes.
[0469] The priority indicates the relative importance of the MFU in the MPU to which the MFU belongs, and an MFU with a higher priority number is more important than an MFU with a lower priority number.
[0470] The dependency_counter indicates the number of MFUs whose decoding process depends on the MFU (i.e., the number of MFUs whose decoding process cannot be performed without decoding the MFU). For example, when the MFU is HEVC and a B picture or a P picture refers to an I picture, the B picture or the P picture cannot be decoded without decoding the I picture.
[0471] Therefore, when the MFU is in sample units, the dependency_counter in the MFU of an I-picture indicates the number of pictures that refer to the I-picture. When the MFU is in NAL unit units, the dependency_counter in the MFU belonging to the I-picture indicates the number of NAL units that belong to the picture that refers to the I-picture. Furthermore, in the case of a video signal that is hierarchically coded in the time direction, the MFU of the enhancement layer depends on the MFU of the base layer, so the dependency_counter in the MFU of the base layer indicates the number of MFUs of the enhancement layer. This field can only be generated after the number of dependent MFUs has been determined.
[0472] The item_iD indicates an identifier that uniquely identifies the item.
[0473] The method of storing control information in the payload in MMT is similar to the method of storing data units in the payload, and can be explained by replacing the data units in Figure 48 with the control information and the data unit length with the control information length. Note that there is no information equivalent to the data unit header.
[0474] [MP4 non-support mode] As described in Figures 19 and 21, the method of transmitting MPU in MMT by the transmitting device includes a method of transmitting MPU metadata or MF metadata before or after the media data, and a method of transmitting only the media data. In addition, the receiving device may decode using a receiving device or receiving method that complies with MP4, or may decode without using a header.
[0475] An example of a data transmission method specialized for broadcast stream playback is a transmission method that does not support MP4 reconstruction in a receiving device.
[0476] A transmission method that does not support MP4 reconstruction in a receiving device is, for example, a method that does not transmit metadata (MPU metadata and MF metadata) as shown in (b) of Fig. 21. In this case, the field value of the fragment type (information indicating the type of data unit) included in the MMTP packet is fixed to 2 (=MFU).
[0477] If metadata is not transmitted, as explained above, an MP4-compliant receiving device cannot decode the received data as MP4, but it is possible to decode it without using the metadata (header).
[0478] Therefore, metadata is not necessarily required information for broadcast stream decoding and playback. Similarly, the information of the data unit header in the timed-MFU described in Fig. 48 is information for reconstructing MP4 in the receiving device. Since there is no need to reconstruct MP4 for broadcast stream playback, the information of the data unit header in the timed-MFU (hereinafter also referred to as the timed-MFU header) is not necessarily required information for broadcast stream playback.
[0479] A receiving device can easily reconstruct an MP4 by using metadata and information for reconstructing an MP4 in a data unit header (hereinafter also referred to as MP4 configuration information). However, a receiving device cannot easily reconstruct an MP4 even if only one of metadata and MP4 configuration information in a data unit header is transmitted. There is little advantage to transmitting only one of metadata and information for reconstructing an MP4, and generating and transmitting unnecessary information leads to increased processing and reduced transmission efficiency.
[0480] Therefore, the transmitting device controls the data structure and transmission of the MP4 configuration information using the following method. The transmitting device determines whether to indicate the MP4 configuration information in the data unit header based on whether metadata is transmitted. Specifically, the transmitting device indicates the MP4 configuration information in the data unit header when metadata is transmitted, and does not indicate the MP4 configuration information in the data unit header when metadata is not transmitted.
[0481] As a method for not indicating MP4 configuration information in the data unit header, for example, the following method can be used.
[0482] 1. The transmitting device sets the MP4 configuration information as reserved and does not use it. This makes it possible to reduce the amount of processing on the sending side (the amount of processing on the transmitting device) that generates the MP4 configuration information.
[0483] 2. The transmitting device deletes the MP4 configuration information and compresses the header. This makes it possible to reduce the amount of processing on the sending side that generates the MP4 configuration information and to reduce the transmission capacity.
[0484] In addition, when deleting the MP4 configuration information and compressing the header, the transmitting device may indicate a flag indicating that the MP4 configuration information has been deleted (compressed). The flag is indicated in the header (MMTP packet header, MMTP payload header, data unit header) or control information.
[0485] In addition, information as to whether metadata is transmitted may be determined in advance, or may be signaled separately in a header or control information and transmitted to the receiving device.
[0486] For example, information as to whether metadata corresponding to the MFU has been transmitted may be stored in the MFU header.
[0487] On the other hand, the receiving device can determine whether MP4 configuration information is indicated based on whether metadata is transmitted.
[0488] Here, if the data transmission order (e.g., MPU metadata, MF metadata, media data, etc.) is fixed, the receiving device may make a determination based on whether metadata is received before media data.
[0489] If the MP4 configuration information is indicated, the receiving device can use the MP4 configuration information to reconstruct the MP4, or the receiving device can use the MP4 configuration information to detect the beginning of other access units or NAL units, etc.
[0490] The MP4 configuration information may be the entirety or a part of the timed-MFU header.
[0491] Similarly, the transmitting device determines whether metadata is transmitted in the non-timed-MFU header. It may decide whether to show the id.
[0492] The transmitting device may indicate the MP4 configuration information in only one of the timed-MFU and non-timed-MFU. When indicating the MP4 configuration information in only one of the two, the transmitting device determines whether to indicate the MP4 configuration information based on whether metadata is transmitted and whether it is timed-MFU or non-timed-MFU. The receiving device can determine whether the MP4 configuration information is indicated based on whether metadata is transmitted and the timed / non-timed flag.
[0493] In the above description, the transmitting device determines whether to indicate the MP4 configuration information based on whether the metadata (both the MPU metadata and the MF metadata) are transmitted. However, the transmitting device may not indicate the MP4 configuration information when part of the metadata (either the MPU metadata or the MF metadata) is not transmitted.
[0494] The sending device may also determine whether to indicate MP4 configuration information based on other information other than metadata.
[0495] For example, modes such as an MP4 support mode / MP4 non-support mode may be defined, and the transmitting device may indicate MP4 configuration information in a data unit header in the MP4 support mode, and may not indicate MP4 configuration information in the data unit header in the MP4 non-support mode. Also, the transmitting device may transmit metadata and indicate MP4 configuration information in a data unit header in the MP4 support mode, and may not transmit metadata and may not indicate MP4 configuration information in the data unit header in the MP4 non-support mode.
[0496] [Transmitter operation flow] Next, the operation flow of the transmitting device will be described with reference to FIG.
[0497] The transmitting device first determines whether to transmit metadata (S1001). If the transmitting device determines to transmit metadata (Yes in S1002), the transmitting device proceeds to step S1003, generates MP4 configuration information, stores it in a header, and transmits it (S1003). In this case, the transmitting device also generates and transmits metadata.
[0498] On the other hand, when the transmitting device determines not to transmit metadata (No in S1002), the transmitting device transmits the MP4 configuration information without generating it and storing it in the header (S1004). In this case, the transmitting device does not generate or transmit metadata.
[0499] In addition, whether or not to transmit metadata in step S1001 may be determined in advance, or may be determined based on whether metadata has been generated within the transmitting device or whether metadata has been transmitted within the transmitting device.
[0500] [Operation flow of receiving device] Next, the operation flow of the receiving device will be described with reference to FIG.
[0501] First, the receiving device determines whether metadata is being transmitted (S1101). Whether metadata is being transmitted can be determined by monitoring the fragmentation type in the MMTP packet payload. Also, whether metadata is being transmitted may be determined in advance.
[0502] When the receiving device determines that metadata has been transmitted (Yes in S1102), it reconstructs the MP4 and executes a decoding process using the MP4 configuration information (S1103).On the other hand, when the receiving device determines that metadata has not been transmitted (No in S1102), it does not reconstruct the MP4 and executes a decoding process without using the MP4 configuration information (S1104).
[0503] In addition, using the method described above, the receiving device is able to detect random access points, the start of an access unit, the start of a NAL unit, etc., without using MP4 configuration information, and can perform decoding processes, packet loss detection, and recovery from packet loss.
[0504] 38, when non-VCL NAL units are aggregated as individual data units, the head of the access unit is the head data of the MMT payload whose aggregation_flag value is 1. In this case, the fragmentation_indicator value is 0.
[0505] In addition, the beginning of a slice segment is the beginning data of an MMT payload in which the aggregation_flag value is 0 and the fragmentation_indicator value is 00 or 01.
[0506] Based on the above information, the receiving device can detect the beginning of an access unit and detect a slice segment.
[0507] In addition, the receiving device may analyze the NAL unit header in a packet including the beginning of a data unit whose fragmentation_indicator value is 00 or 01, and detect that the type of the NAL unit is an AU delimiter and that the type of the NAL unit is a slice segment.
[0508] [Broadcast Simple Mode] So far, we have described a method of transmitting data specialized for broadcast stream playback that does not support MP4 configuration information in the receiving device, but the method of transmitting data specialized for broadcast stream playback is not limited to this.
[0509] As a method for transmitting data specialized for broadcast stream playback, for example, the following method may be used.
[0510] A transmitting device does not need to use AL-FEC in a fixed broadcast reception environment. If AL-FEC is not used, the FEC_type in the MMTP packet header is always fixed to 0.
[0511] The transmitting device may always use AL-FEC in a mobile broadcast reception environment and in the communication UDP transmission mode. When AL-FEC is used, the FEC_type in the MMTP packet header is always 0 or 1.
[0512] The transmitting device may not transmit assets in bulk. If assets are not transmitted in bulk, location_infolocation, which indicates the number of transmission locations of the assets within the MPT, may be fixed to 1.
[0513] · The transmitting device does not require hybrid transmission of assets, programs and messages.
[0514] Also, for example, a broadcast simple mode may be specified, and when the transmitting device is in the broadcast simple mode, the transmitting device may be in the MP4 non-support mode, or may use the above-mentioned method of transmitting data specialized for broadcast stream playback. Whether or not the transmitting device is in the broadcast simple mode may be determined in advance, or the transmitting device may store a flag indicating that the transmitting device is in the broadcast simple mode as control information and transmit it to the receiving device.
[0515] In addition, the transmitting device may determine whether the mode is MP4 non-support mode as described in Figure 49 (whether metadata is transmitted) and, if the mode is MP4 non-support mode, use the data transmitting method specialized for broadcast stream playback shown above as broadcast simple mode.
[0516] When the receiving device is in broadcast simple mode, it is considered to be in MP4 non-support mode and can perform decoding processing without reconstructing the MP4.
[0517] Furthermore, when the receiving device is in the simple broadcast mode, it can determine that the functions are specialized for broadcasting and perform reception processing specialized for broadcasting.
[0518] As a result, when in simple broadcast mode, by using only functions specialized for broadcasting, not only can unnecessary processing be reduced for the transmitting device and receiving device, but transmission overhead can also be reduced by not compressing and transmitting unnecessary information.
[0519] In addition, when the MP4 non-support mode is used, hint information supporting a storage method other than the MP4 format may be indicated.
[0520] Examples of storage methods other than the MP4 format include directly storing MMT packets or IP packets, and converting MMT packets into MPEG-2 TS packets.
[0521] In addition, in the case of an MP4 non-support mode, a format that does not conform to the MP4 structure may be used.
[0522] For example, in the case of an MP4 non-support mode, data stored in the MFU may be in a format with a byte start code attached to the beginning of the NAL unit, rather than the MP4 format in which the size of the NAL unit is attached to the beginning of the NAL unit.
[0523] In MMT, the asset type indicating the type of asset is described in 4CC registered in MP4REG (http: / / www.mp4ra.org), and when HEVC is used as the video signal, 'HEV1' or 'HVC1' is used. 'HVC1' is a format that may include a parameter set in a sample, while 'HEV1' is a format that does not include a parameter set in a sample, but includes a parameter set in the sample entry in the MPU metadata.
[0524] In the case of broadcast simple mode or MP4 non-support mode, if MPU metadata and MF metadata are not transmitted, it may be specified that a parameter set is always included in the sample. Also, it may be specified that the format is always 'HVC1' regardless of whether 'HEV1' or 'HVC1' is indicated in the asset type.
[0525] [Supplement 1: Transmitter] As described above, when metadata is not transmitted, the MP4 configuration information is set to "reserved," and a transmitting device that is not in operation can be configured as shown in Fig. 51. Fig. 51 is a diagram showing an example of a specific configuration of a transmitting device.
[0526] The transmission device 300 includes an encoding unit 301, an attachment unit 302, and a transmission unit 303. Each of the encoding unit 301, the attachment unit 302, and the transmission unit 303 is realized by, for example, a microcomputer, a processor, or a dedicated circuit.
[0527] The encoding unit 301 generates sample data by encoding a video signal or an audio signal. The sample data is specifically a data unit.
[0528] The adding unit 302 adds header information including MP4 configuration information to sample data, which is data obtained by encoding a video signal or an audio signal. The MP4 configuration information is information for reconstructing the sample data as a file in MP4 format on the receiving side, and the content of the information varies depending on whether the presentation time of the sample data is determined or not.
[0529] As described above, the adding unit 302 includes MP4 configuration information such as movie_fragment_sequence_number, sample_number, offset, priority, and dependency_counter in the header (header information) of a timed-MFU, which is an example of sample data with a defined presentation time (sample data including information regarding synchronization).
[0530] On the other hand, the attachment unit 302 includes MP4 configuration information such as item_id in the header (header information) of a timed-MFU, which is an example of sample data for which the presentation time is not determined (sample data that does not include information regarding synchronization).
[0531] Then, when the transmitting unit 303 does not transmit metadata corresponding to the sample data (for example, in the case of (b) in Figure 21), the adding unit 302 adds header information that does not include MP4 configuration information to the sample data, depending on whether the presentation time of the sample data has been specified.
[0532] Specifically, when the presentation time of the sample data is determined, the attachment unit 302 attaches header information that does not include the first MP4 configuration information to the sample data, and when the presentation time of the sample data is not determined, the attachment unit 302 attaches header information that includes the second MP4 configuration information to the sample data.
[0533] For example, as shown in step S1004 in Fig. 49, when the transmitting unit 303 does not transmit metadata corresponding to the sample data, the adding unit 302 sets the MP4 configuration information to reserved (fixed value), so that the MP4 configuration information is not actually generated and is not actually stored in the header (header information). Note that the metadata includes MPU metadata and movie fragment metadata.
[0534] The transmitting unit 303 transmits the sample data to which the header information has been added. More specifically, the transmitting unit 303 packetizes the sample data to which the header information has been added in accordance with the MMT method and transmits the packetized data.
[0535] As described above, in the transmission method and the reception method specialized for playing back a broadcast stream, the receiving device does not need to reconstruct the data units into MP4. If the receiving device does not need to reconstruct the data into MP4, the processing load of the transmitting device is reduced by not generating unnecessary information such as MP4 configuration information.
[0536] On the other hand, while the transmitting device must transmit the necessary information, it must maintain compatibility with the standard so as to avoid having to transmit unnecessary additional information separately.
[0537] According to the configuration of the transmitting device 300, by setting the area in which the MP4 configuration information is stored to a fixed value, the MP4 configuration information is not transmitted, and only necessary information is transmitted based on the standard, and unnecessary additional information does not need to be transmitted. In other words, the configuration of the transmitting device and the amount of processing of the transmitting device can be reduced. Furthermore, since unnecessary data is not transmitted, the transmission efficiency can be improved.
[0538] [Supplement 2: Receiving device] Moreover, a receiving device corresponding to the transmitting device 300 may be configured, for example, as shown in Fig. 52. Fig. 52 is a diagram showing another example of the configuration of the receiving device.
[0539] The receiving device 400 includes a receiving unit 401 and a decoding unit 402. The receiving unit 401 and the decoding unit 402 are realized by, for example, a microcomputer, a processor, or a dedicated circuit.
[0540] The receiving unit 401 receives sample data, which is data in which a video signal or an audio signal has been encoded, and which is provided with header information including MP4 configuration information for reconstructing the sample data as a file in MP4 format.
[0541] If the receiving unit does not receive metadata corresponding to the sample data and the presentation time of the sample data is determined, the decoding unit 402 decodes the sample data without using the MP4 configuration information.
[0542] For example, as shown in step S1104 in FIG. 50, if the receiving unit 401 does not receive metadata corresponding to the sample data, the decoding unit 402 executes the decoding process without using the MP4 configuration information.
[0543] This makes it possible to reduce the configuration of the receiving device 400 and the amount of processing in the receiving device 400.
[0544] (Other embodiments) Although the transmitting method, receiving method, transmitting device, and receiving device according to the embodiment have been described above, the present invention is not limited to this embodiment.
[0545] Furthermore, each processing unit included in the transmitting device and the receiving device according to the above-described embodiment is typically realized as an LSI, which is an integrated circuit. These may be individually implemented as single chips, or some or all of them may be included in a single chip.
[0546] The integrated circuit is not limited to an LSI, but may be realized by a dedicated circuit or a general-purpose processor. A field programmable gate array (FPGA) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connections and settings of the circuit cells inside the LSI may also be used.
[0547] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0548] In other words, the transmitting device and the receiving device include a processing circuitry and a storage electrically connected to the processing circuitry (accessible from the processing circuitry). The processing circuitry includes at least one of dedicated hardware and a program execution unit. In addition, when the processing circuitry includes a program execution unit, the storage unit stores a software program executed by the program execution unit. The processing circuitry uses the storage unit to execute the transmitting method or the receiving method according to the above-mentioned embodiment.
[0549] Furthermore, the present invention may be the above-mentioned software program, or a non-transitory computer-readable recording medium on which the above-mentioned program is recorded. Needless to say, the above-mentioned program can be distributed via a transmission medium such as the Internet.
[0550] Furthermore, all the numbers used above are merely examples for the purpose of specifically explaining the present invention, and the present invention is not limited to the numbers exemplified.
[0551] In addition, the division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as one functional block, one functional block may be divided into multiple blocks, or some functions may be transferred to another functional block. Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in a time-sharing manner by a single piece of hardware or software.
[0552] The order in which the steps included in the above-mentioned transmission method or reception method are executed is merely an example for specifically explaining the present invention, and an order other than the above may be used. Also, some of the above steps may be executed simultaneously (in parallel) with other steps.
[0553] The above describes the transmission method, reception method, transmission device, and reception device according to one or more aspects of the present invention based on the embodiments, but the present invention is not limited to these embodiments. As long as it does not deviate from the spirit of the present invention, various modifications conceived by a person skilled in the art to the present embodiments and forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects of the present invention. [Industrial Applicability]
[0554] The present invention can be applied to devices or equipment that transport media such as video data and audio data. [Explanation of symbols]
[0555] 15, 100, 300 Transmitting device 16, 101, 301 Encoding section 17, 102 Multiplexing section 18, 104 Transmitter 20, 200, 400 Receiver 21 Packet Filtering Section 22 Transmission order type discrimination section 23 Random Access Section 24, 212 Control information acquisition unit 25 Data Acquisition Section 26 Calculation section 27 Initialization information acquisition unit 28, 206 Decoding instruction section 29, 204A, 204B, 204C, 204D, 402 Decoding section 30 Presentation section 201 Tuner 202 Demodulation section 203 Demultiplexer 205 Display section 211 Type Identification Section 213 Slice information acquisition unit 214 Decoding data generation unit 302 Granting Department 303 Transmitter 401 Receiving unit
Claims
1. An adding step of adding header information to the sample data, which is the encoded data; a transmitting step of transmitting the sample data to which the header information is added, If the presentation time of the sample data is not specified, In the providing step, the header information including MP4 configuration information for reconstructing the sample data as an MP4 format file on a receiving side is provided to the sample data; When the presentation time of the sample data is determined, transmitting control information related to the decoding of the sample data separately from the sample data; In the adding step, header information in which the MP4 configuration information is invalidated is added to the sample data, In the transmitting step, the sample data to which the header information is added is packetized in accordance with an MMT (MPEG Media Transport) method and transmitted; The sample data with a presentation time is a timed-MFU (Movie Fragment Unit). Transmission method.
2. A receiving step of receiving sample data, which is encoded data, and which has header information added thereto; and decoding the sample data. If the presentation time of the sample data is not specified, In the decoding step, the sample data is decoded using MP4 configuration information included in the header information for reconstructing the sample data as a file in MP4 format; When the presentation time of the sample data is determined, In the decoding step, a decoding process of the sample data is performed using control information related to decoding of the sample data, the control information being received separately from the sample data; The sample data to which the header information is added is packetized in an MMT format, The sample data with a presentation time is timed-MFU. Receiving method.
3. an adding unit that adds header information to sample data that is encoded data; a transmission unit that transmits the sample data to which the header information is added, If the presentation time of the sample data is not specified, The adding unit adds the header information to the sample data, the header information including MP4 configuration information for reconstructing the sample data as an MP4 format file on a receiving side; When the presentation time of the sample data is determined, transmitting control information related to the decoding of the sample data separately from the sample data; The adding unit adds header information in which the MP4 configuration information is invalidated to the sample data, The transmission unit packetizes the sample data to which the header information is added in an MMT format and transmits the packetized data, The sample data with a presentation time is timed-MFU. Transmitting device.
4. A receiving unit that receives sample data that is encoded data and has header information added thereto; A decoding unit that decodes the sample data, If the presentation time of the sample data is not specified, The decoding unit performs a decoding process on the sample data by using MP4 configuration information for reconstructing the sample data as a file in MP4 format, the MP4 configuration information being included in the header information; When the presentation time of the sample data is determined, the decoding unit performs a decoding process on the sample data by using control information related to decoding of the sample data, the control information being received separately from the sample data; The sample data to which the header information is added is packetized in an MMT format, The sample data with a presentation time is timed-MFU. Receiving device.
Citation Information
Patent Citations
Reproduction device, distribution device, data structure, reproduction method, distribution method, control program, and recording medium
JP2013229689A
Compression of TS packet headers
JP2013520035A
Method of configuring and transmitting an MMT transport packet
US20130094563A1
Method and apparatus for encapsulation of motion picture experts group media transport assets in international organization for standardization base media files
WO2014084643A1