Information processor, method for processing information, and program
By transmitting a meta file with identification information for metadata presence, the technology ensures easy recognition and extraction of metadata from audio streams, addressing inefficiencies in metadata processing.
Patent Information
- Application Number
- JP2025064128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-09-12
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2035-09-07
AI Technical Summary
Existing technologies fail to easily recognize and process metadata inserted into audio streams, leading to inefficiencies in metadata extraction and utilization.
A transmission unit transmits a meta file with meta information for acquiring an audio stream containing metadata, and an information insertion unit inserts identification information into the meta file to indicate metadata presence, enabling easy recognition and extraction on the receiving side.
Enables easy recognition and reliable extraction of metadata from audio streams, facilitating efficient processing and utilization without waste.
Smart Images

Figure 2025108523000001_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, it has been proposed to insert metadata into an audio stream and transmit it (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Metadata is defined in, for example, the user data area of the audio stream. However, metadata is not inserted into all audio streams.
[0005] An object of the present technology is to make it easily recognizable on the receiving side that metadata is inserted into an audio stream and to facilitate processing.
Means for Solving the Problems
[0006] The concept of the present technology is a transmission unit that transmits a meta file having meta information for acquiring, by a receiving device, an audio stream into which metadata is inserted; and an information insertion unit that inserts, into the meta file, identification information indicating that the metadata is inserted into the audio stream. The above configuration is provided in a transmission device.
[0007] In this technology, a meta file having meta information for a receiving device to acquire an audio stream with metadata inserted therein is transmitted by a transmitting unit. For example, the metadata may be access information for connecting to a predetermined network service. In this case, for example, the metadata may be a character code indicating URI information.
[0008] Also, for example, the transmitting unit may be configured to transmit the meta file through an RF transmission path or a communication network transmission path. Also, for example, the transmitting unit may be further configured to transmit a container in a predetermined format including an audio stream with metadata inserted therein. In this case, for example, the container may be MP4 (ISO / IEC 14496-14:2003).
[0009] An information insertion unit inserts identification information indicating that metadata is inserted into an audio stream into the meta file. For example, the meta file may be an MPD (Media Presentation Description) file. In this case, for example, the information insertion unit may insert the identification information into the meta file using "Supplementary Descriptor".
[0010] Thus, in this technology, identification information indicating that metadata is inserted into an audio stream is inserted into a meta file having meta information for a receiving device to acquire the audio stream with metadata inserted therein. Therefore, on the receiving side, it can be easily recognized that metadata is inserted into the audio stream. And, for example, it is also possible to perform an extraction process of the metadata inserted into the audio stream based on this recognition, and it becomes possible to surely acquire the metadata without waste.
[0011] Also, other concepts of this technology are A receiving unit that receives a meta file having meta information for obtaining an audio stream inserted with metadata, The meta file has inserted therein identification information indicating that the metadata is inserted into the audio stream, The apparatus further includes a transmitting unit that transmits the audio stream, together with identification information indicating that metadata is inserted into the audio stream, to an external device via a predetermined transmission path. It is in a receiving device.
[0012] In the present technology, the receiving unit receives a meta file having meta information for obtaining an audio stream inserted with metadata. For example, the metadata may be access information for connecting to a predetermined network service. The meta file has inserted therein identification information indicating that the metadata is inserted into the audio stream.
[0013] For example, the metadata may be access information for connecting to a predetermined network service. Also, for example, the meta file may be an MPD file, and the identification information may be inserted into this meta file by "Supplementary Descriptor".
[0014] The transmitting unit transmits the audio stream, together with identification information indicating that the metadata is inserted into the audio stream, to an external device via a predetermined transmission path. For example, the transmitting unit may insert the audio stream and the identification information during a blanking period of the image data and transmit the image data to the external device, thereby transmitting the audio stream and the identification information to the external device. Also, for example, the predetermined transmission path may be an HDMI cable.
[0015] Thus, in the present technology, an audio stream with metadata inserted is transmitted to an external device together with identification information indicating that metadata has been inserted into the audio stream. Therefore, on the external device side, it can be easily recognized that metadata has been inserted into the audio stream. And, for example, based on this recognition, it is also possible to perform an extraction process of the metadata inserted into the audio stream, and it becomes possible to reliably obtain the metadata without waste.
[0016] Also, another concept of the present technology is a receiving unit that receives a metadata file having meta information for acquiring an audio stream with metadata inserted, wherein the metadata file has identification information inserted therein indicating that the metadata has been inserted into the audio stream, a metadata extraction unit that decodes the audio stream and extracts the metadata based on the identification information, and a processing unit that further performs processing using the metadata in a receiving device.
[0017] In the present technology, the receiving unit receives a metadata file having meta information for acquiring an audio stream with metadata inserted. The metadata file has identification information inserted therein indicating that the metadata has been inserted into the audio stream. For example, the metadata file may be an MPD file, and the identification information may be inserted into this metadata file by "Supplementary Descriptor".
[0018] The metadata extraction unit decodes the audio stream and extracts the metadata based on the identification information. And the processing unit performs processing using this metadata. For example, the metadata may be access information for connecting to a predetermined network service, and the processing unit may access a predetermined server on the network based on the network access information.
[0019] In this way, in the present technology, based on identification information indicating that metadata is inserted into the audio stream, which is inserted into the metafile, the metadata is extracted from the audio stream and used for processing. Therefore, the metadata inserted into the audio stream can be surely obtained without waste, and the processing using the metadata can be appropriately executed.
[0020] Further, another concept of the present technology is a stream generation unit that generates an audio stream into which metadata including network access information is inserted, and a transmission unit that transmits a container in a predetermined format having the audio stream. This is in a transmission device.
[0021] In the present technology, the stream generation unit generates an audio stream into which metadata including network access information is inserted. For example, the audio stream is generated by encoding audio data with AAC, AC3, AC4, MPEGH (3D audio), etc., and the metadata is embedded in the user data area thereof.
[0022] The transmission unit transmits a container in a predetermined format having the audio stream. Here, the container in a predetermined format is, for example, MP4, MPEG2-TS, etc. For example, the metadata may be a character code indicating URI information.
[0023] In this way, in the present technology, the metadata including network access information is embedded in the audio stream and transmitted. Therefore, for example, from a broadcasting station, a distribution server, etc., the network access information can be easily transmitted with the audio stream as a container and made available for use on the receiving side.
Effects of the Invention
[0024] According to this technology, it becomes possible for the receiving side to easily recognize that metadata is inserted into an audio stream. Note that the effects described in this specification are merely examples and are not limited, and there may be additional effects.
Brief Description of Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Embodiments for Carrying Out the Invention
[0026] Hereinafter, embodiments for carrying out the invention (hereinafter referred to as "embodiments") will be described. The description will be made in the following order. 1. Embodiments 2. Variations
[0027] <1. Embodiments> [Overview of an MPEG-DASH-Based Stream Delivery System] First, an overview of an MPEG-DASH-based stream delivery system to which this technology can be applied will be described.
[0028] FIG. 1(a) shows an example of the configuration of an MPEG-DASH-based stream delivery system 30A. In this configuration example, a media stream and an MPD file are transmitted through a communication network transmission path. This stream delivery system 30A has a configuration in which N receiving systems 33-1, 33-2, ···, 33-N are connected to a DASH stream file server 31 and a DASH MPD server 32 via a CDN (Content Delivery Network) 34.
[0029] The DASH stream file server 31 generates stream segments (hereinafter, appropriately referred to as "DASH segments") in accordance with the DASH specification based on media data (such as video data, audio data, subtitle data, etc.) of predetermined content, and sends out the segments in response to an HTTP request from a receiving system. This DASH stream file server 31 may be a dedicated streaming server, or may also be used as a web server.
[0030] Also, the DASH stream file server 31, in response to a request for segments of a predetermined stream sent from the receiving system 33 (33-1, 33-2, ···, 33-N) via the CDN 34, transmits the segments of that stream to the requesting receiver via the CDN 34. In this case, the receiving system 33 refers to the rate value described in the MPD (Media Presentation Description) file and selects and requests the optimal rate stream according to the state of the network environment where the client is located.
[0031] The DASH MPD server 32 is a server that generates an MPD file for obtaining the DASH segments generated in the DASH stream file server 31. An MPD file is generated based on the content metadata from a content management server (not shown) and the address (URL) of the segments generated in the DASH stream file server 31. Note that the DASH stream file server 31 and the DASH MPD server 32 may be physically the same.
[0032] In the MPD format, for each stream such as video and audio, an element called Representation is used to describe each attribute. For example, in the MPD file, for each of multiple video data streams with different rates, the Representations are separated and each rate is described. In the receiving system 33, referring to the value of that rate, as described above, the optimal stream can be selected according to the state of the network environment where the receiving system 33 is located.
[0033] Figure 1(b) shows a configuration example of an MPEG-DASH based stream delivery system 30B. In this configuration example, the media stream and the MPD file are transmitted through the RF transmission path. This stream delivery system 30B is composed of a broadcast transmission system 36 to which a DASH stream file server 31 and a DASH MPD server 32 are connected, and M receiving systems 35-1, 35-2, ···, 35-M.
[0034] In the case of this stream delivery system 30B, the broadcast transmission system 36 transmits the DASH specification stream segment (DASH segment) generated by the DASH stream file server 31 and the MPD file generated by the DASH MPD server 32 on the broadcast wave.
[0035] Figure 2 shows an example of the relationship between each structure hierarchically arranged in the MPD file. As shown in Figure 2(a), in the Media Presentation as a whole of the MPD file, there are multiple Periods separated by time intervals. For example, the start of the first Period is from 0 seconds, the start of the next Period is from 100 seconds, and so on.
[0036] As shown in FIG. 2(b), there are multiple Representations in a period. Among these multiple Representations, there is a group of Representations related to media streams of the same content with different stream attributes, such as rate, grouped by an AdaptationSet.
[0037] As shown in FIG. 2(c), a Representation includes SegmentInfo. As shown in FIG. 2(d), this SegmentInfo includes an Initialization Segment and multiple Media Segments in which information for each Segment that further divides the period is described. In the Media Segment, there is information such as the address (url) for actually obtaining segment data such as video and audio.
[0038] Note that between multiple Representations grouped by an AdaptationSet, stream switching can be freely performed. Thereby, according to the state of the network environment where the receiving system is located, a stream with an optimal rate can be selected, enabling seamless delivery.
[0039] [Configuration of Transmission / Reception System] FIG. 3 shows a configuration example of a transmission / reception system as an embodiment. The transmission / reception system 10 in FIG. 3(a) includes a service transmission system 100, a set-top box (STB) 200, and a television receiver (TV) 300. The set-top box 200 and the television receiver 300 are connected via an HDMI (High Definition Multimedia Interface) cable 400. Note that "HDMI" is a registered trademark.
[0040] In this transmission and reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31 and the DASH MPD server 32 of the stream distribution system 30A shown in FIG. 1(a) described above. Also, in this transmission and reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31, the DASH MPD server 32, and the broadcast transmission system 36 of the stream distribution system 30B shown in FIG. 1(b) described above.
[0041] In this transmission and reception system 10, the set-top box (STB) 200 and the television receiver (TV) 300 correspond to the reception system 33 (33-1, 33-2, ···, 33-N) of the stream distribution system 30A shown in FIG. 1(a) described above. Also, in this transmission and reception system 10, the set-top box (STB) 200 and the television receiver (TV) 300 correspond to the reception system 35 (35-1, 35-2, ···, 35-M) of the stream distribution system 30B shown in FIG. 1(b) described above.
[0042] Also, the transmission and reception system 10' in FIG. 3(b) has the service transmission system 100 and the television receiver (TV) 300. In this transmission and reception system 10', the service transmission system 100 corresponds to the DASH stream file server 31 and the DASH MPD server 32 of the stream distribution system 30A shown in FIG. 1(a) described above. Also, in this transmission and reception system 10', the service transmission system 100 corresponds to the DASH stream file server 31, the DASH MPD server 32, and the broadcast transmission system 36 of the stream distribution system 30B shown in FIG. 1(b) described above.
[0043] In this transmission and reception system 10´, the television receiver (TV) 300 corresponds to the reception system 33 (33-1, 33-2, ···, 33-N) of the stream distribution system 30A shown in FIG. 1(a) described above. Also, in this transmission and reception system 10´, the television receiver (TV) 300 corresponds to the reception system 35 (35-1, 35-2, ···, 35-M) of the stream distribution system 30B shown in FIG. 1(b) described above.
[0044] The service transmission system 100 transmits DASH / MP4, that is, an MPD file as a metadata file and an MP4 containing media streams (media segments) such as video and audio, through an RF transmission path or a communication network transmission path. The service transmission system 100 inserts metadata into the audio stream. Examples of this metadata include access information for connecting to a predetermined network service, predetermined content information, and the like. In this embodiment, access information for connecting to a predetermined network service is inserted.
[0045] The service transmission system 100 inserts identification information indicating that metadata has been inserted into the audio stream into the MPD file. The service transmission system 100 inserts, for example, identification information indicating that metadata has been inserted into the audio stream using "Supplementary Descriptor".
[0046] FIG. 4 shows an example of an MPD file description. " <adaptationset mimetype=""audio / mp4”" group=""1”">According to the description of "」, there is an AdaptationSet for the audio stream, and the audio stream is supplied in the MP4 file structure, indicating that Group 1 is assigned.
[0047] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:AudioMetaContained”" value=""true” / ">The description of "」 indicates that metadata is inserted into the audio stream. With "SupplementaryDescriptor", "schemeIdUri" can be newly defined for broadcasting and other applications separately from the predefined in the conventional standard. As shown in FIG. 5, "schemeIdUri = \"urn:brdcst:AudiometaContained\"" indicates that audio meta information is included, that is, metadata is inserted into the audio stream. For example, when "value" is "true", it indicates that audio meta information is included. When "value" is "false", it indicates that no audio meta information is included.
[0048] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description of "」 indicates that the codec of the audio stream is MPEGH (3D audio). As shown in Figure 5, "schemeIdUri = \"urn:brdcst:codecType\"" indicates the type of codec. For example, "value" is set to "mpegh", "AAC", "AC3", "AC4", etc.
[0049] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:coordinatedControl”" value=""true” / ">The description of "..." indicates that the information necessary for network connection is emphasized and supplied among multiple media streams. As shown in FIG. 5, "schemeIdUri = \"urn:brdcst:coordinatedControl\"" indicates that the information necessary for network connection is supplied in coordination among multiple media streams. For example, when "value" is "true", it indicates that the network connection information is supplied in coordination with the streams of other adaptation sets. When "value" is "false", it indicates that the network connection information is supplied only by the streams of this adaptation set.
[0050] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:type”" value=""netlink” / ">The description of "」 indicates that the type of meta service is a network connection. As shown in Figure 5, the type of meta service is indicated by "schemeIdUri = "urn:brdcst:type". For example, when "value" is "netlink", it indicates that the type of meta service is a network connection.
[0051] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:metaInsertionFrequency”" value=""1” / ">The description of "..." indicates that meta-information is supplied in access unit units. As shown in FIG. 5, "schemeIdUri = \"urn:brdcst:metaInsertionFrequency\"" indicates the frequency at which meta-information is supplied in access unit units. For example, when "value" is "1", it indicates that one user data entry occurs in one access unit. When "value" is "2", it indicates that multiple user data entries occur in one access unit. When "value" is "3", it indicates that one or more user data entries occur during a period delimited by random access points.
[0052] FIG. 6(a) shows an example of the arrangement of video and audio access units contained in MP4. "VAU" indicates a video access unit. "AAU" indicates an audio access unit. FIG. 6(b) shows the case where "frequency_type = 1", and one user data entry (metadata) is inserted into each audio access unit.
[0053] FIG. 6(c) shows the case where "frequency_type = 2", and multiple user data (metadata) are inserted into one audio access unit. FIG. 6(d) shows the case where "frequency_type = 3", and at least one user data (metadata) is inserted into the first audio access unit for each group including random access points.
[0054] Also, returning to FIG. 4, " <representation id=""11”" bandwidth=""128000”">According to the description of "」, the existence of an audio stream with a bit rate of 128 kbps is indicated as "Representation id="11". And, " <baseurl>audio / jp / 128.mp4< / baseurl> According to the description of "」, the location destination of the audio stream is indicated as "audio / jp / 128.mp4".
[0055] Also, " <adaptationset mimetype=""video / mp4”" group=""2”">According to the description of "」, there is an AdaptationSet for the video stream, and the video stream is supplied in the MP4 file structure, indicating that Group 2 is assigned.
[0056] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:VideoMetaContained”" value=""true” / ">The description indicates that metadata is inserted into the video stream. As shown in FIG. 5, "schemeIdUri = \"urn:brdcst:VideoMetaContained\"" indicates that video meta information is included, that is, metadata is inserted into the video stream. For example, when "value" is "true", it indicates that video meta information is included. When "value" is "false", it indicates that video meta information is not included.
[0057] Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""hevc” / ">」The description shows that the codec of the video stream is HEVC. Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:coordinatedControl”" value=""true” / ">The description of "..." indicates that information necessary for network connection is supplied in an emphasized manner among a plurality of media streams.
[0058] Also, "..." <supplementarydescriptor schemeiduri=""urn:brdcst:type”" value=""netlink” / ">The description of "」 indicates that the type of service by Meta is a network connection. Also, " <supplementarydescriptor schemeiduri=""urn:brdcst:metaInsertionFrequency”" value=""1” / ">The description of "..." indicates that meta information is supplied in units of access units.
[0059] Also, "..." <representation id=""21”" bandwidth=""20000000”">According to the description of "", the existence of a video stream with a bitrate of 20 Mbps is indicated as "Representation id="21". And, " <baseurl>video / jp / 20000000.mp4< / baseurl> According to the description of "", the location destination of the video stream is indicated as "video / jp / 20000000.mp4".
[0060] Here, " <baseurl>The physical media file at the location indicated by "」" will be described. In the case of non-fragmented MP4 (Non-Fragmented MP4), for example, as shown in Fig. 7(a), it may be defined as "url 1". In this case, a "ftyp" box that describes the file type first is placed. This "ftyp" box indicates that it is an MP4 file that has not been fragmented. Subsequently, a "moov" box and an "mdat" box are placed. The "moov" box contains all metadata, such as the header information of each track, the meta description of the content, time information, etc. The "mdat" box contains the media data body.
[0061] Also, in the case of fragmented MP4 (Fragmented MP4), for example, as shown in Fig. 7(b), it may be defined as "url 2". In this case, a "styp" box that describes the segment type first is placed. Subsequently, a "sidx" box that describes the segment index is placed. Then, a predetermined number of movie fragments are placed. Here, a movie fragment is composed of a "moof" box into which control information enters and an "mdat" box into which the media data body enters. Since the "mdat" box of one movie fragment contains the fragments obtained by fragmenting the transmission media, the control information that enters the "moof" box is the control information regarding that fragment.
[0062] Also, the combination of the above "url 1" and "url 2" is also conceivable. In this case, for example, "url 1" can be used as an initialization segment, and it is also possible to make "url 1" and "url 2" an MP4 for one service. Alternatively, "url 1" and "url 2" can be combined into one and defined as "url 3" as shown in Fig. 7(c).
[0063] The set-top box 200 receives DASH / MP4, that is, an MPD file as a metadata file, and MP4 including media streams (media segments) such as video and audio, which are sent from the service transmission system 100 through an RF transmission line or a communication network transmission line. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream included in the MP4. Further, identification information indicating that metadata is inserted into the audio stream is inserted into the MPD file by "Supplementary Descriptor".
[0064] The set-top box 200 transmits the audio stream, together with identification information indicating that metadata is inserted into this audio stream, to the television receiver 300 via the HDMI cable 400.
[0065] Here, the set-top box 200 inserts the audio stream and the identification information during the blanking period of the image data obtained by decoding the video stream, and transmits this image data to the television receiver 300, thereby transmitting the audio stream and the identification information to the television receiver 300. The set-top box 200 inserts this identification information into, for example, an Audio InfoFrame packet.
[0066] In the transmission / reception system 10 shown in Fig. 3(a), the television receiver 300 receives the audio stream, together with identification information indicating that metadata is inserted into this audio stream, from the set-top box 200 via the HDMI cable 400. That is, the television receiver 300 receives the image data in which the audio stream and the identification information are inserted during the blanking period from the set-top box 200.
[0067] Then, based on the identification information, the television receiver 300 decodes the audio stream to extract metadata and performs processing using this metadata. In this case, the television receiver 300 accesses a predetermined server on the network based on predetermined network service information as metadata.
[0068] Also, in the transmission / reception system 10' shown in Fig. 3(b), the television receiver 300 receives DASH / MP4, that is, an MPD file as a metadata file, and an MP4 containing media streams (media segments) such as video and audio, which are sent from the service transmission system 100 through an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream included in the MP4. Further, identification information indicating that metadata is inserted into the audio stream is inserted into the MPD file by "Supplementary Descriptor".
[0069] Then, based on the identification information, the television receiver 300 decodes the audio stream to extract metadata and performs processing using this metadata. In this case, the television receiver 300 accesses a predetermined server on the network based on predetermined network service information as metadata.
[0070] [DASH / MP4 Generation Unit of Service Transmission System] Fig. 8 shows a configuration example of the DASH / MP4 generation unit 110 included in the service transmission system 100. This DASH / MP4 generation unit 110 has a control unit 111, a video encoder 112, an audio encoder 113, and a DASH / MP4 formatter 114.
[0071] The control unit 111 includes a CPU 111a and controls each part of the DASH / MP4 generation unit 110. The video encoder 112 performs encoding on the image data SV, such as MPEG2, H.264 / AVC, H.265 / HEVC, etc., to generate a video stream (video elementary stream). The image data SV is, for example, image data reproduced from a recording medium such as an HDD, or live image data obtained by a video camera, etc.
[0072] The audio encoder 113 performs encoding on the audio data SA using a compression format such as AAC, AC3, AC4, MPEGH (3D audio), etc., to generate an audio stream (audio elementary stream). The audio data SA is audio data corresponding to the above-mentioned image data SV, such as audio data reproduced from a recording medium such as an HDD, or live audio data obtained by a microphone, etc.
[0073] The audio encoder 113 has an audio encoding block unit 113a and an audio framing unit 113b. Encoding blocks are generated in the audio encoding block unit 113a, and framing is performed in the audio framing unit 113b. In this case, depending on the compression format, the encoding blocks are different and the framing is also different.
[0074] The audio encoder 113 inserts the metadata MD into the audio stream under the control of the control unit 111. In this embodiment, the metadata MD is access information for connecting to a predetermined network service. Here, all services such as a music network service and an audio-video network service are targeted as the predetermined network service. Here, the metadata MD is embedded in the user data area of the audio stream.
[0075] The DASH / MP4 formatter 114 generates an MP4 that includes media streams (media segments) such as video and audio that are the content, based on the video stream output from the video encoder 112 and the audio stream output from the audio encoder 113. Also, it generates an MPD file using content metadata, segment URL information, etc. Here, identification information indicating that metadata is inserted into the audio stream, etc. is inserted into the MPD file (see FIG. 4).
[0076] The operation of the DASH / MP4 generation unit 110 shown in FIG. 8 will be briefly described. The image data SV is supplied to the video encoder 112. In this video encoder 112, the image data SV is encoded using H.264 / AVC, H.265 / HEVC, etc., and a video stream including encoded video data is generated.
[0077] Also, the audio data SA is supplied to the audio encoder 113. In this audio encoder 113, the audio data SA is encoded using AAC, AC3, AC4, MPEGH (3D audio), etc., and an audio stream is generated.
[0078] At this time, metadata MD is supplied from the control unit 111 to the audio encoder 113, and size information for embedding this metadata MD in the user data area is supplied. Then, the audio encoder 113 embeds the metadata MD in the user data area of the audio stream.
[0079] The video stream generated by the video encoder 112 is supplied to the DASH / MP4 formatter 114. Also, the audio stream generated by the audio encoder 113, in which metadata MD is embedded in the user data area, is supplied to the DASH / MP4 formatter 114. Then, this DASH / MP4 formatter 114 generates an MP4 containing media streams (media segments) such as video and audio that are the content. Also, this DASH / MP4 formatter 114 generates an MPD file using content metadata, segment URL information, etc. At this time, identification information indicating that metadata is inserted into the audio stream is inserted into the MPD file.
[0080] [Details of Insertion of Metadata MD in Each Compression Format] "In the Case of AAC" First, the case where the compression format is AAC (Advanced Audio Coding) will be described. Figure 9 shows the structure of an AAC audio frame. This audio frame consists of a plurality of elements. At the head of each element, there is a 3-bit identifier (ID) of "id_syn_ele", and the element content can be identified.
[0081] When "id_syn_ele" is "0x4", it is shown that it is a DSE (Data Stream Element), an element where user data can be placed. When the compression format is AAC, metadata MD is inserted into this DSE. Figure 10 shows the configuration (Syntax) of the DSE (Data Stream Element()).
[0082] The 4-bit field of 「element_instance_tag」 indicates the data type in the DSE. However, when using the DSE as unified user data, this value may be set to "0". 「Data_byte_align_flag」 is set to "1" so that the entire DSE is byte-aligned. The value of 「count」, or 「esc_count」 which means the additional byte count, is determined as appropriate according to the size of the user data. 「metadata ()」 is inserted into the field of 「data_stream_byte」.
[0083] Figure 11(a) shows the configuration (Syntax) of 「metadata ()」, and Figure 11(b) shows the content (semantics) of the main information in this configuration. The 32-bit field of 「userdata_identifier」 indicates that it is audio user data by setting the values of a predefined array. When 「userdata_identifier」 indicates user data with "AAAA", there is an 8-bit field of 「metadata_type」. This field indicates the type of metadata. For example, "0x08" is access information for connecting to a predetermined network service and indicates that it is included in the ATSC's 「SDO_payload()」. When it is "0x08", 「SDO_payload()」 exists. Here, it is set as "ATSC", but it can also be used by other standardization organizations.
[0084] Figure 12 shows the configuration (Syntax) of 「SDO_payload()」. When the command ID (cmdID) is less than "0x05", there is a field of 「URI_character」. Character codes indicating URI information for connecting to a predetermined network service are inserted into this field. Figure 13 shows the meaning of the value of the command ID (cmdID). Note that this 「SDO_payload()」 is standardized by the ATSC (Advanced Television Systems Committee standards).
[0085] "In the case of AC3" Next, the case where the compression format is AC3 will be described. FIG. 14 shows the structure of an AC3 frame (AC3 Synchronization Frame). The audio data SA is encoded so that the total size of "mantissa data" of "Audblock 5", "AUX", and "CRC" does not exceed 3 / 8 of the whole. When the compression format is AC3, metadata MD is inserted into the "AUX" area. FIG. 15 shows the syntax of AC3 auxiliary data.
[0086] When "auxdatae" is "1", "aux data" is enabled, and data with a size indicated by 14 bits (in bit units) of "auxdatal" is defined in "auxbits". The size of "auxbits" at that time is described in "nauxbits". In this technology, the field of "auxbits" is defined as "metadata()". That is, the "metadata()" shown in FIG. 11(a) above is inserted into this "auxbits" field, and in the "data_byte" field thereof, the "SDO_payload()" of ATSC (see FIG. 12) having access information for connecting to a predetermined network service according to the syntax structure shown in FIG. 11(a) is placed.
[0087] "In the case of AC4" Next, the case where the compression format is AC4 will be described. This AC4 is regarded as one of the next-generation audio encoding formats of AC3. Figure 16(a) shows the structure of the Simple Transport layer of AC4. There are a syncWord field, a frame Length field, a "RawAc4Frame" field as the encoded data field, and a CRC field. As shown in Figure 16(b), in the "RawAc4Frame" field, there is a TOC (Table Of Content) field at the beginning, and then there are fields of a predetermined number of substreams.
[0088] As shown in Figure 17(b), in the substream (ac4_substream_data()), there is a metadata area, and in it, a field of "umd_payloads_substream()" is provided. In this "umd_payloads_substream()" field, the ATSC's "SDO_payload()" (see Figure 12), which has access information for connecting to a predetermined network service, is placed.
[0089] Note that as shown in Figure 17(a), in the TOC (ac4_toc()), there is a field of "ac4_presentation_info()", and further in it, there is a field of "umd_info()", and it is shown that there is an insertion of metadata in the above-mentioned "umd_payloads_substream())" field in it.
[0090] Figure 18 shows the syntax of "umd_info()". The "umd_version" field indicates the version number. The "substream_index" field indicates the index value. A combination of values of the version number and the index value is defined to indicate that there is metadata insertion in the field of "umd_payloads_substream()".
[0091] Figure 19 shows the syntax of "umd_payloads_substream()". The 5-bit field of "umd_payload_id" is set to a value other than "0". The 32-bit field of "umd_userdata_identifier" is set to a value in a predefined array to indicate that it is audio user data. The 16-bit field of "umd_payload_size" indicates the number of subsequent bytes. When "umd_userdata_identifier" indicates user data with "AAAA", the 8-bit field of "umd_metadata_type" exists. This field indicates the type of metadata. For example, "0x08" is access information for connecting to a predetermined network service and is indicated to be included in "SDO_payload()" of ATSC. When it is "0x08", "SDO_payload()" (see Figure 12) exists.
[0092] "In the case of MPEGH" Next, the case where the compression format is MPEGH (3D audio) will be described. Figure 20 shows the structure of an audio frame (1024 samples) in the transmission data of MPEGH (3D audio). This audio frame consists of a plurality of MPEG audio stream packets. Each MPEG audio stream packet is composed of a header and a payload.
[0093] The header holds information such as Packet Type, Packet Label, Packet Length, etc. In the payload, the information defined by the packet type of the header is arranged. This payload information includes "SYNC" corresponding to the synchronization start code, "Frame" which is the actual data of the 3D audio transmission data, and "Config" which indicates the configuration of this "Frame".
[0094] "Frame" includes the channel-encoded data and object-encoded data that make up the 3D audio transmission data. Here, the channel-encoded data is composed of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), LFE (Low Frequency Element). Also, the object-encoded data is composed of the encoded sample data of SCE (Single Channel Element) and metadata for mapping it to a speaker existing at an arbitrary position for rendering. This metadata is included as an extension element (Ext_element).
[0095] Here, the configuration information (config) of each "Frame" included in "Config" and the correspondence with each "Frame" are maintained as follows. That is, as shown in FIG. 21, the configuration information (config) of each "Frame" is registered with ID (elemIdx) in "Config", but each "Frame" is transmitted in the order of ID registration. Note that the value of the packet label (PL) is set to the same value for "Config" and each corresponding "Frame".
[0096] Returning to FIG. 20, in this embodiment, an element (Ext_userdata) including user data (userdata) as an extension element (Ext_element) is newly defined. Along with this, configuration information (userdataConfig) of the element (Ext_userdata) is newly defined in "Config".
[0097] FIG. 22 shows the correspondence between the type (ExElementType) of the extension element (Ext_element) and its value (Value). Currently, 0 to 7 are defined. Since it is extensible up to beyond 128 for non-MPEG, for example, 128 is newly defined as the value of the type "ID_EXT_ELE_userdata".
[0098] FIG. 23 shows the syntax of "userdataConfig()". The 32-bit field of "userdata_identifier" indicates that it is audio user data by setting the value of a predefined array. The 16-bit field of "userdata_frameLength" indicates the number of bytes of "audio_userdata()". FIG. 24 shows the syntax of "audio_userdata()". When "userdata_identifier" in "userdataConfig()" indicates user data with "AAAA", an 8-bit field of "metadataType" exists. This field indicates the type of metadata. For example, "0x08" is access information for connecting to a predetermined network service and indicates that it is included in "SDO_payload()" of ATSC. When it is "0x08", "SDO_payload()" (see FIG. 12) exists.
[0099] [Configuration Example of Set-Top Box] FIG. 25 shows a configuration example of the set-top box 200. This set-top box 200 includes a receiving unit 204, a DASH / MP4 analysis unit 205, a video decoder 206, an audio framing unit 207, an HDMI transmission unit 208, and an HDMI terminal 209. The set-top box 200 also includes a CPU 211, a flash ROM 212, a DRAM 213, an internal bus 214, a remote control receiving unit 215, and a remote control transmitter 216.
[0100] The CPU 211 controls the operations of each part of the set-top box 200. The flash ROM 212 stores control software and stores data. The DRAM 213 constitutes the work area of the CPU 211. The CPU 211 expands the software and data read from the flash ROM 212 onto the DRAM 213 to start the software and controls each part of the set-top box 200.
[0101] The remote control receiving unit 215 receives the remote control signal (remote control code) transmitted from the remote control transmitter 216 and supplies it to the CPU 211. The CPU 211 controls each part of the set-top box 200 based on this remote control code. The CPU 211, the flash ROM 212, and the DRAM 213 are connected to the internal bus 214.
[0102] The receiving unit 204 receives DASH / MP4, that is, an MPD file as a metadata file, and an MP4 including media streams (media segments) such as video and audio, which are sent from the service transmission system 100 through an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream included in the MP4. Also, identification information indicating that metadata is inserted into the audio stream is inserted into the MPD file by "Supplementary Descriptor".
[0103] The DASH / MP4 analysis unit 205 analyzes the MPD file and MP4 received by the reception unit 204. The DASH / MP4 analysis unit 205 extracts the MPD information contained in the MPD file and sends it to the CPU 211. Here, this MPD information also includes identification information indicating that metadata is inserted into the audio stream. The CPU 211 controls the acquisition process of video and audio streams based on this MPD information. Also, the DASH / MP4 analysis unit 205 extracts metadata from the MP4, such as header information of each track, meta-description of the content, time information, etc., and sends it to the CPU 211.
[0104] The DASH / MP4 analysis unit 205 extracts the video stream from the MP4 and sends it to the video decoder 206. The video decoder 206 performs decoding processing on the video stream to obtain uncompressed image data. Also, the DASH / MP4 analysis unit 205 extracts the audio stream from the MP4 and sends it to the audio framing unit 207. The audio framing unit 207 performs framing on the audio stream.
[0105] The HDMI transmission unit 208 sends the uncompressed image data obtained by the video decoder 206 and the audio stream after being framed by the audio framing unit 207 from the HDMI terminal 209 through communication compliant with HDMI. Since the HDMI transmission unit 208 transmits through the TMDS channel of HDMI, it packs the image data and the audio stream and outputs them to the HDMI terminal 209.
[0106] The HDMI transmission unit 208 inserts identification information indicating that metadata is inserted into the audio stream under the control of the CPU 211. The HDMI transmission unit 208 inserts the audio stream and the identification information during the blanking period of the image data. The details of this HDMI transmission unit 209 will be described later.
[0107] In this embodiment, the HDMI transmitter 208 inserts identification information into an Audio InfoFrame packet that is arranged in the blanking period of the image data. This Audio InfoFrame packet is arranged in the data island section.
[0108] FIG. 26 shows an example of the structure of an Audio InfoFrame packet. In HDMI, with this Audio InfoFrame packet, it is possible to transmit additional information related to audio from the source device to the sink device.
[0109] "Packet Type" indicating the type of data packet is defined in the 0th byte, and the Audio InfoFrame packet is "0x84". Version information of the packet data definition is described in the 1st byte. Information representing the packet length is described in the 2nd byte. In this embodiment, 1-bit flag information of "userdata_presence_flag" is defined in the 5th bit of the 5th byte. When the flag information is "1", it indicates that metadata is inserted into the audio stream.
[0110] Also, when the flag information is "1", various information is defined in the 9th byte. The 7th bit to the 5th bit are fields of "metadata_type", the 4th bit is a field of "coordinated_control_flag", and the 2nd bit to the 0th bit are fields of "frequency_type". Although detailed description is omitted, each of these fields indicates the same information as the information added to the MPD file shown in FIG. 4.
[0111] The operation of the set-top box 200 will be briefly described. In the receiving unit 204, DASH / MP4, that is, an MPD file as a metadata file, and an MP4 including media streams (media segments) such as video and audio are received from the service transmission system 100 through an RF transmission line or a communication network transmission line. The MPD file and MP4 received in this way are supplied to the DASH / MP4 analysis unit 205.
[0112] In the DASH / MP4 analysis unit 205, the MPD file and MP4 are analyzed. Then, in the DASH / MP4 analysis unit 205, the MPD information included in the MPD file is extracted and sent to the CPU 211. Here, this MPD information also includes identification information indicating that metadata is inserted into the audio stream. Also, in the DASH / MP4 analysis unit 205, metadata such as header information of each track, meta description of the content, time information, etc. is extracted from the MP4 and sent to the CPU 211.
[0113] Also, in the DASH / MP4 analysis unit 205, a video stream is extracted from the MP4 and sent to the video decoder 206. In the video decoder 206, the video stream is decoded to obtain uncompressed image data. This image data is supplied to the HDMI transmission unit 208. Also, in the DASH / MP4 analysis unit 205, an audio stream is extracted from the MP4. This audio stream is framed by the audio framing unit 207 and then supplied to the HDMI transmission unit 208. Then, in the HDMI transmission unit 208, the image data and the audio stream are packed and sent out from the HDMI terminal 209 to the HDMI cable 400.
[0114] In the HDMI transmission unit 208, under the control of the CPU 211, identification information indicating that metadata is inserted into the audio stream is inserted into the audio info frame packet arranged in the blanking period of the image data. Thereby, the set-top box 200 transmits the identification information indicating that metadata is inserted into the audio stream to the HDMI television receiver 300.
[0115] [Configuration Example of Television Receiver] FIG. 27 shows a configuration example of the television receiver 300. This television receiver 300 includes a receiving unit 306, a DASH / MP4 analysis unit 307, a video decoder 308, a video processing circuit 309, a panel driving circuit 310, and a display panel 311.
[0116] Further, the television receiver 300 includes an audio decoder 312, an audio processing circuit 313, an audio amplification circuit 314, a speaker 315, an HDMI terminal 316, an HDMI receiving unit 317, and a communication interface 318. The television receiver 300 also includes a CPU 321, a flash ROM 322, a DRAM 323, an internal bus 324, a remote control receiving unit 325, and a remote control transmitter 326.
[0117] The CPU 321 controls the operations of each part of the television receiver 300. The flash ROM 322 stores control software and stores data. The DRAM 323 constitutes the work area of the CPU 321. The CPU 321 expands the software and data read from the flash ROM 322 onto the DRAM 323 to start the software and controls each part of the television receiver 300.
[0118] The remote control receiving unit 325 receives the remote control signal (remote control code) transmitted from the remote control transmitter 326 and supplies it to the CPU 321. The CPU 321 controls each part of the television receiver 300 based on this remote control code. The CPU 321, the flash ROM 322, and the DRAM 323 are connected to the internal bus 324.
[0119] Under the control of the CPU 321, the communication interface 318 communicates with a server existing on a network such as the Internet. This communication interface 318 is connected to the internal bus 324.
[0120] The receiving unit 306 receives DASH / MP4, that is, an MPD file as a metadata file, and MP4 including media streams (media segments) such as video and audio, which are sent from the service transmission system 100 through an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream included in the MP4. Further, in the MPD file, identification information indicating that metadata is inserted into the audio stream is inserted by "Supplementary Descriptor".
[0121] The DASH / MP4 analysis unit 307 analyzes the MPD file and MP4 received by the receiving unit 306. The DASH / MP4 analysis unit 307 extracts the MPD information included in the MPD file and sends it to the CPU 321. The CPU 321 controls the acquisition process of video and audio streams based on this MPD information. Further, the DASH / MP4 analysis unit 307 extracts metadata from the MP4, for example, header information of each track, meta description of the content, time information, etc., and sends it to the CPU 321.
[0122] The DASH / MP4 analysis unit 307 extracts a video stream from the MP4 and sends it to the video decoder 308. The video decoder 308 performs a decoding process on the video stream to obtain uncompressed image data. Further, the DASH / MP4 analysis unit 307 extracts an audio stream from the MP4 and sends it to the audio decoder 312.
[0123] The HDMI receiving unit 317 receives the image data and audio stream supplied to the HDMI terminal 316 via the HDMI cable 400 by communication compliant with HDMI. Also, the HDMI receiving unit 317 extracts various control information inserted during the blanking period of the image data and transmits it to the CPU 321. Here, this control information includes identification information (see FIG. 26) inserted into the audio information frame packet indicating that metadata is inserted into the audio stream. Details of this HDMI receiving unit 317 will be described later.
[0124] The video processing circuit 309 performs scaling processing, compositing processing, etc. on the image data obtained by the video decoder 308, or the image data obtained by the HDMI receiving unit 316, and further the image data received from a server on the network by the communication interface 318, etc. to obtain display image data.
[0125] The panel driving circuit 310 drives the display panel 311 based on the display image data obtained by the video processing circuit 308. The display panel 311 is composed of, for example, an LCD (Liquid Crystal Display), an organic EL display (organic electroluminescence display), etc.
[0126] The audio decoder 312 performs decoding processing on the audio stream extracted by the DASH / MP4 analysis unit 307 or obtained by the HDMI receiving unit 317 to obtain uncompressed audio data. Also, the audio decoder 312 extracts the metadata inserted into the audio stream under the control of the CPU 321 and sends it to the CPU 321. In this embodiment, the metadata is access information for connecting to a predetermined network service (see FIG. 12). The CPU 321 appropriately causes each part of the television receiver 300 to perform processing using the metadata.
[0127] Note that the CPU 321 is supplied with MPD information from the DASH / MP4 analysis unit 307. Based on the identification information included in this MPD information, the CPU 321 can recognize in advance that metadata is inserted into the audio stream and can control the audio decoder 312 so that the metadata is extracted.
[0128] The audio processing circuit 313 performs necessary processing such as D / A conversion on the audio data obtained by the audio decoder 312. The audio amplification circuit 314 amplifies the audio signal output from the audio processing circuit 313 and supplies it to the speaker 315.
[0129] The operation of the television receiver 300 shown in FIG. 27 will be briefly described. In the receiving unit 306, DASH / MP4, that is, an MPD file as a metafile and MP4 including media streams (media segments) such as video and audio, sent from the service transmission system 100 through the RF transmission path or the communication network transmission path, is received. The MPD file and MP4 received in this way are supplied to the DASH / MP4 analysis unit 307.
[0130] In the DASH / MP4 analysis unit 307, the MPD file and MP4 are analyzed. Then, in the DASH / MP4 analysis unit 307, the MPD information included in the MPD file is extracted and sent to the CPU 321. Here, this MPD information also includes identification information indicating that metadata is inserted into the audio stream. Further, in the DASH / MP4 analysis unit 307, metadata such as the header information of each track, the meta description of the content, and time information is extracted from the MP4 and sent to the CPU 321.
[0131] Also, in the DASH / MP4 analysis unit 307, a video stream is extracted from the MP4 and sent to the video decoder 308. In the video decoder 308, the video stream is decoded to obtain uncompressed image data. This image data is supplied to the video processing circuit 309. Also, in the DASH / MP4 analysis unit 307, an audio stream is extracted from the MP4. This audio stream is supplied to the audio decoder 312.
[0132] In the HDMI receiver unit 317, image data and an audio stream supplied to the HDMI terminal 316 via the HDMI cable 400 are received through communication compliant with HDMI. The image data is supplied to the video processing circuit 309. Also, the audio stream is supplied to the audio decoder 312.
[0133] Also, in the HDMI receiver unit 317, various control information inserted in the blanking period of the image data is extracted and sent to the CPU 321. This control information includes identification information inserted in the audio information frame packet indicating that metadata is inserted in the audio stream. Therefore, the CPU 321 can control the operation of the audio decoder 312 based on this identification information and extract metadata from the audio stream.
[0134] In the video processing circuit 309, scaling processing, compositing processing, etc. are performed on the image data obtained by the video decoder 308, or the image data obtained by the HDMI receiver unit 317, or further the image data received from a server on the network by the communication interface 318, etc., to obtain display image data. Here, when receiving and processing a TV broadcast signal, in the video processing circuit 309, the image data obtained by the video decoder 308 is handled. On the other hand, when the set-top box 200 is connected by an HDMI interface, in the video processing circuit 309, the image data obtained by the HDMI receiver unit 317 is handled.
[0135] The image data for display obtained by the image processing circuit 309 is supplied to the panel driving circuit 310. In the panel driving circuit 310, the display panel 311 is driven based on the image data for display. As a result, an image corresponding to the image data for display is displayed on the display panel 311.
[0136] In the audio decoder 312, decoding processing is performed on the audio stream obtained by the DASH / MP4 analysis unit 307 or the HDMI reception unit 316 to obtain uncompressed audio data. Here, when receiving and processing a television broadcast signal, the audio decoder 312 processes the audio stream obtained by the DASH / MP4 analysis unit 307. On the other hand, when the set-top box 200 is connected by an HDMI interface, the audio decoder 312 processes the audio stream obtained by the HDMI reception unit 317.
[0137] The audio data obtained by the audio decoder 312 is supplied to the audio processing circuit 313. In the audio processing circuit 313, necessary processing such as D / A conversion is performed on the audio data. This audio data is amplified by the audio amplifier circuit 314 and then supplied to the speaker 315. Therefore, audio corresponding to the display image of the display panel 311 is output from the speaker 315.
[0138] Also, in the audio decoder 312, the metadata inserted into the audio stream is extracted. For example, as described above, this metadata extraction process is surely performed without waste by the CPU 321 grasping that metadata is inserted into the audio stream based on the identification information and controlling the operation of the audio decoder 312.
[0139] The metadata extracted by the audio decoder 312 in this way is sent to the CPU 321. Then, under the control of the CPU 321, processing using the metadata is appropriately performed in each part of the television receiver 300. For example, image data is acquired from a server on the network and multi-screen display is performed.
[0140] [Configuration Examples of HDMI Transmitter and HDMI Receiver] FIG. 28 shows a configuration example of the HDMI transmitter (HDMI source) 208 of the set-top box 200 shown in FIG. 25 and the HDMI receiver (HDMI sink) 317 of the television receiver 300 shown in FIG. 27.
[0141] In the valid image interval (hereinafter, also referred to as the active video interval as appropriate), the HDMI transmitter 208 transmits, in a plurality of channels, differential signals corresponding to the pixel data of an uncompressed image for one screen in one direction to the HDMI receiver 317. Here, the valid image interval is an interval from a certain vertical synchronization signal to the next vertical synchronization signal, excluding the horizontal blanking interval and the vertical blanking interval. Also, in the horizontal blanking interval or the vertical blanking interval, the HDMI transmitter 208 transmits, in a plurality of channels, differential signals corresponding to at least audio data, control data, and other auxiliary data associated with the image in one direction to the HDMI receiver 317.
[0142] The transmission channels of the HDMI system composed of the HDMI transmitter 208 and the HDMI receiver 317 are as follows. That is, there are three TMDS channels #0 to #2 as transmission channels for serially transmitting pixel data and audio data from the HDMI transmitter 208 to the HDMI receiver 317 in one direction synchronized with the pixel clock. Also, there is a TMDS clock channel as a transmission channel for transmitting the pixel clock.
[0143] The HDMI transmitter 208 has an HDMI transmitter 81. The transmitter 81 converts, for example, the pixel data of an uncompressed image into corresponding differential signals and serially transmits them in one direction to the HDMI receiver 317 connected via the HDMI cable 400 on three TMDS channels #0, #1, and #2, which are a plurality of channels.
[0144] Further, the transmitter 81 converts the audio data associated with the uncompressed image, as well as necessary control data and other auxiliary data, etc. into corresponding differential signals, and serially transmits them in one direction to the HDMI receiver 317 via the three TMDS channels #0, #1, and #2.
[0145] Furthermore, the transmitter 81 transmits a pixel clock synchronized with the pixel data transmitted via the three TMDS channels #0, #1, and #2 to the HDMI receiver 317 connected via the HDMI cable 400 via the TMDS clock channel. Here, in one TMDS channel #i (i = 0, 1, 2), 10-bit pixel data is transmitted during one clock of the pixel clock.
[0146] The HDMI receiver 317 receives, in a plurality of channels, the differential signals corresponding to the pixel data transmitted in one direction from the HDMI transmitter 208 in the active video period. Also, this HDMI receiver 317 receives, in a plurality of channels, the differential signals corresponding to the audio data and control data transmitted in one direction from the HDMI transmitter 208 in the horizontal blanking period or the vertical blanking period.
[0147] That is, the HDMI receiver 317 has an HDMI receiver 82. This HDMI receiver 82 receives the differential signals corresponding to the pixel data transmitted in one direction from the HDMI transmitter 208 via the TMDS channels #0, #1, and #2, and the differential signals corresponding to the audio data and control data. In this case, it receives in synchronization with the pixel clock transmitted from the HDMI transmitter 208 via the TMDS clock channel.
[0148] In addition to the above-described TMDS channels #0 to #2 and the TMDS clock channel, the transmission channel of the HDMI system has transmission channels called DDC (Display Data Channel) 83 and CEC line 84. The DDC 83 consists of two signal lines (not shown) included in the HDMI cable 400. The DDC 83 is used by the HDMI transmitter 208 to read the E-EDID (Enhanced Extended Display Identification Data) from the HDMI receiver 317.
[0149] In addition to the HDMI receiver 317, the HDMI receiver has an EDID ROM (Read Only Memory) 85 that stores the E-EDID, which is performance information regarding its own performance (Configuration / capability). The HDMI transmitter 208 reads the E-EDID from the HDMI receiver 317 connected via the HDMI cable 400 via the DDC 83, for example, in response to a request from the CPU 211 (see FIG. 20).
[0150] The HDMI transmitter 208 sends the read E-EDID to the CPU 211. The CPU 211 stores this E-EDID in the flash ROM 212 or the DRAM 213.
[0151] The CEC line 84 consists of one signal line (not shown) included in the HDMI cable 400 and is used to perform two-way communication of control data between the HDMI transmitter 208 and the HDMI receiver 317. This CEC line 84 constitutes a control data line.
[0152] In addition, the HDMI cable 400 includes a line (HPD line) 86 connected to a pin called HPD (Hot Plug Detect). The source device can detect the connection of the sink device by using this line 86. Note that this HPD line 86 is also used as a HEAC-line that constitutes a bidirectional communication path. Further, the HDMI cable 400 includes a power line 87 used to supply power from the source device to the sink device. Furthermore, the HDMI cable 400 includes a utility line 88. This utility line 88 is also used as a HEAC+ line that constitutes a bidirectional communication path.
[0153] FIG. 29 shows sections of various transmission data when image data of 1920 pixels × 1080 lines in the horizontal × vertical direction is transmitted in the TMDS channels #0, #1, and #2. In a video field in which transmission data is transmitted through the three TMDS channels #0, #1, and #2 of HDMI, there are three types of sections: a video data period 17, a data island period 18, and a control period 19, according to the type of transmission data.
[0154] Here, the video field section is a section from the rising edge of a certain vertical synchronization signal to the rising edge of the next vertical synchronization signal, and is divided into a horizontal blanking period 15, a vertical blanking period 16, and an active video section 14 which is a section of the video field section excluding the horizontal blanking period and the vertical blanking period.
[0155] The video data section 17 is allocated to the effective pixel section 14. In this video data section 17, data of 1920 pixels (picture elements) × 1080 lines of effective pixels (Active Pixel) that constitute uncompressed image data for one screen is transmitted. The data island section 18 and the control section 19 are allocated to the horizontal blanking period 15 and the vertical blanking period 16. In this data island section 18 and control section 19, auxiliary data is transmitted.
[0156] That is, the data island section 18 is allocated to a part of the horizontal blanking period 15 and the vertical blanking period 16. In this data island section 18, among the auxiliary data, data that has nothing to do with control, such as packets of audio data, is transmitted. The control section 19 is allocated to the other parts of the horizontal blanking period 15 and the vertical blanking period 16. In this control section 19, among the auxiliary data, data related to control, such as vertical synchronization signals, horizontal synchronization signals, control packets, etc., is transmitted.
[0157] Next, with reference to FIG. 30, a specific example of processing using metadata in the television receiver 300 will be described. The television receiver 300 acquires, for example, as metadata, an initial server URL, network service identification information, a target file name, session start / end commands, media recording / playback commands, and the like. Note that although it was stated above that the metadata is access information for connecting to a predetermined network service, here it is assumed that other necessary information is also included in the metadata.
[0158] The television receiver 300, which is a network client, accesses the primary server using the initial server URL. Then, the television receiver 300 acquires information such as a streaming server URL, a target file name, a MIME type indicating the type of the file, and media playback time information from the primary server.
[0159] Then, the TV receiver 300 accesses the streaming server using the streaming server URL. Then, the TV receiver 300 specifies the target file name. Here, when receiving services via multicast, the services of the program are specified using network identification information and service identification information.
[0160] Then, the TV receiver 300 starts or ends a session with the streaming server using session start / end commands. Also, during the session with the streaming server, the TV receiver 300 acquires media data from the streaming server using media recording / playback commands.
[0161] Note that in the example of FIG. 30, the primary server and the streaming server exist separately. However, these servers may be integrally configured.
[0162] FIG. 31 shows an example of the transition of the screen display when the TV receiver 300 accesses a network service based on metadata. FIG. 31(a) shows a state where no image is displayed on the display panel 311. FIG. 31(b) shows a state where the broadcast reception has started and the main content related to this broadcast reception is displayed full-screen on the display panel 311.
[0163] FIG. 31(c) shows a state where access to the service based on the metadata has occurred and the session between the TV receiver 300 and the server has started. In this case, the main content related to the broadcast reception changes from full-screen display to partial-screen display.
[0164] FIG. 31(d) shows a state in which media playback from the server is performed, and net service content 1 is displayed on the display panel 311 in parallel with the display of the main content. And FIG. 31(e) shows a state in which media playback from the server is performed, and net service content 2 is superimposed on the main content display together with the display of net service content 1 on the display panel 311 in parallel with the display of the main content.
[0165] FIG. 31(f) shows a state in which the playback of service content from the network has ended and the session between the television receiver 300 and the server has ended. In this case, the display panel 311 returns to a state in which the main content related to broadcast reception is displayed in full screen.
[0166] Note that the television receiver 300 shown in FIG. 27 includes a speaker 315, and as shown in FIG. 32, the audio data obtained by the audio decoder 312 is supplied to the speaker 315 via the audio processing circuit 313 and the audio amplifier circuit 314, and audio is output from this speaker 315.
[0167] However, as shown in FIG. 33, the television receiver 300 may be configured not to include a speaker and to supply the audio stream obtained by the DASH / MP4 analysis unit 307 or the HDMI reception unit 317 to the external speaker system 350 from the interface unit 331. The interface unit 331 is a digital interface such as HDMI (High-Definition Multimedia Interface), SPDIF (Sony Philips Digital Interface), MHL (Mobile High-definition Link), for example.
[0168] In this case, the audio stream is decoded by the audio decoder 351a of the external speaker system 350, and audio is output from this external speaker system 350. Even when the television receiver 300 includes the speaker 315 (see FIG. 32), a configuration in which the audio stream is supplied from the interface unit 331 to the external speaker system 350 (see FIG. 33) is further conceivable.
[0169] As described above, in the transmission / reception systems 10 and 10' shown in FIGS. 3(a) and 3(b), the service transmission system 100 inserts identification information indicating that metadata is inserted into the audio stream into the MPD file. Therefore, on the receiving side (the set-top box 200 and the television receiver 300), it is possible to easily recognize that metadata is inserted into the audio stream.
[0170] Also, in the transmission / reception system 10 shown in FIG. 3(a), the set-top box 200 transmits the audio stream with metadata inserted to the television receiver 300 via HDMI together with identification information indicating that metadata is inserted into this audio stream. Therefore, the television receiver 300 can easily recognize that metadata is inserted into the audio stream, and by performing extraction processing of the metadata inserted into the audio stream based on this recognition, the metadata can be surely obtained and used without waste.
[0171] Also, in the transmission / reception system 10' shown in FIG. 3(b), the television receiver 300 extracts metadata from the audio stream based on the identification information inserted into the MPD file and uses it for processing. Therefore, the metadata inserted into the audio stream can be surely obtained without waste, and processing using the metadata can be appropriately executed.
[0172] <2. Modification Example> In the above-described embodiment, an example of handling DASH / MP4 is shown as the transmission / reception systems 10 and 10'. However, an example of handling MPEG2-TS can be considered in the same way.
[0173] [Configuration of Transmission / Reception System] FIG. 34 shows a configuration example of a transmission / reception system that handles MPEG2-TS. The transmission / reception system 10A in FIG. 34(a) includes a service transmission system 100A, a set-top box (STB) 200A, and a television receiver (TV) 300A. The set-top box 200A and the television receiver 300A are connected via an HDMI (High Definition Multimedia Interface) cable 400. The transmission / reception system 10A' in FIG. 3(b) includes a service transmission system 100A and a television receiver (TV) 300A.
[0174] The service transmission system 100A transmits the transport stream TS of MPEG2-TS through an RF transmission path or a communication network transmission path. The service transmission system 100A inserts metadata into the audio stream. Examples of this metadata include access information for connecting to a predetermined network service, predetermined content information, etc. Here, as in the above-described embodiment, it is assumed that access information for connecting to a predetermined network service is inserted.
[0175] The service transmission system 100A inserts identification information indicating that metadata is inserted into the audio stream into the layer of the container. The service transmission system 100A inserts this identification information as a descriptor, for example, into the audio elementary stream loop under a program map table (PMT).
[0176] The set-top box 200A receives a transport stream TS sent from the service transmission system 100A through an RF transmission line or a communication network transmission line. This transport stream TS includes a video stream and an audio stream, and metadata is inserted into the audio stream.
[0177] The set-top box 200A transmits the audio stream, together with identification information indicating that metadata is inserted into this audio stream, to the television receiver 300A via the HDMI cable 400.
[0178] Here, the set-top box 200A inserts the audio stream and the identification information during the blanking period of the image data obtained by decoding the video stream, and transmits this image data to the television receiver 300A, thereby transmitting the audio stream and the identification information to the television receiver 300A. The set-top box 200A inserts this identification information into, for example, an Audio InfoFrame packet (see FIG. 26).
[0179] In the transmission / reception system 10A shown in FIG. 34(a), the television receiver 300A receives the audio stream, together with identification information indicating that metadata is inserted into this audio stream, from the set-top box 200A via the HDMI cable 400. That is, the television receiver 300A receives the image data in which the audio stream and the identification information are inserted during the blanking period from the set-top box 200A.
[0180] Then, the television receiver 300A decodes the audio stream based on the identification information to extract the metadata, and performs processing using this metadata. In this case, the television receiver 300A accesses a predetermined server on the network based on predetermined network service information as the metadata.
[0181] Also, in the transmission / reception system 10A' shown in FIG. 34(b), the television receiver 300A receives a transport stream TS sent from the service transmission system 100A through an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream included in this transport stream TS. Further, identification information indicating that metadata is inserted into the audio stream is inserted into the container layer.
[0182] Then, the television receiver 300A decodes the audio stream based on the identification information to extract the metadata, and performs processing using this metadata. In this case, the television receiver 300A accesses a predetermined server on the network based on the predetermined network service information as metadata.
[0183] [TS generation unit of service transmission system] FIG. 35 shows a configuration example of the TS generation unit 110A included in the service transmission system 100A. In this FIG. 35, parts corresponding to FIG. 8 are denoted by the same reference numerals. This TS generation unit 110A includes a control unit 111, a video encoder 112, an audio encoder 113, and a TS formatter 114A.
[0184] The control unit 111 includes a CPU 111a and controls each part of the TS generation unit 110A. The video encoder 112 performs encoding such as MPEG2, H.264 / AVC, H.265 / HEVC on the image data SV to generate a video stream (video elementary stream). The image data SV is, for example, image data reproduced from a recording medium such as an HDD, or live image data obtained by a video camera.
[0185] The audio encoder 113 encodes the audio data SA using a compression format such as AAC, AC3, AC4, MPEG-H (3D audio), etc., and generates an audio stream (audio elementary stream). The audio data SA is audio data corresponding to the above-described video data SV, and is audio data reproduced from a recording medium such as an HDD, or live audio data obtained by a microphone, etc.
[0186] The audio encoder 113 has an audio encoding block section 113a and an audio framing section 113b. An encoding block is generated in the audio encoding block section 113a, and framing is performed in the audio framing section 113b. In this case, depending on the compression format, the encoding blocks are different and the framing is also different.
[0187] Under the control of the control unit 111, the audio encoder 113 inserts metadata MD into the audio stream. As this metadata MD, for example, access information for connecting to a predetermined network service, predetermined content information, etc. can be considered. Here, as in the above-described embodiment, it is assumed that access information for connecting to a predetermined network service is inserted.
[0188] This metadata MD is inserted into the user data area of the audio stream. Although detailed description is omitted, the insertion of the metadata MD in each compression format is performed in the same manner as in the case of the DASH / MP4 generation unit 110 in the above-described embodiment, and "SDO_payload()" is inserted as the metadata MD (see FIGS. 8 - 24).
[0189] The TS formatter 114A packetizes the video stream output from the video encoder 112 and the audio stream output from the audio encoder 113 into PES packets, further packetizes them into transport packets and multiplexes them to obtain a transport stream TS as a multiplexed stream.
[0190] Also, the TS formatter 114A inserts identification information indicating that metadata MD is inserted into the audio stream, under the program map table (PMT). To insert this identification information, an audio_userdata_descriptor is used. Details of this descriptor will be described later.
[0191] The operation of the TS generation unit 110A shown in FIG. 35 will be briefly described. The image data SV is supplied to the video encoder 112. In this video encoder 112, the image data SV is encoded using H.264 / AVC, H.265 / HEVC, etc., and a video stream including encoded video data is generated.
[0192] Also, the audio data SA is supplied to the audio encoder 113. In this audio encoder 113, the audio data SA is encoded using AAC, AC3, AC4, MPEGH (3D audio), etc., and an audio stream is generated.
[0193] At this time, the metadata MD is supplied from the control unit 111 to the audio encoder 113, and size information for embedding this metadata MD in the user data area is supplied. Then, the audio encoder 113 embeds the metadata MD in the user data area of the audio stream.
[0194] The video stream generated by the video encoder 112 is supplied to the TS formatter 114A. Also, the audio stream in which the metadata MD is embedded in the user data area, generated by the audio encoder 113, is supplied to the TS formatter 114A.
[0195] In this TS formatter 114A, the streams supplied from each encoder are packetized, multiplexed, and a transport stream TS is obtained as transmission data. Also, in this TS formatter 114A, identification information indicating that metadata MD is inserted into the audio stream is inserted under the program map table (PMT).
[0196] [Details of Audio User Data Descriptor] FIG. 36 shows an example of the structure (Syntax) of an audio_userdata_descriptor. Also, FIG. 37 shows the content (Semantics) of the main information in the structure example.
[0197] The 8-bit field of "descriptor_tag" indicates the descriptor type. Here, it indicates that it is an audio_userdata_descriptor. The 8-bit field of "descriptor_length" indicates the length (size) of the descriptor, and indicates the number of subsequent bytes as the length of the descriptor.
[0198] The 8-bit field of "audio_codec_type" indicates the audio encoding method (compression format). For example, "1" indicates "MPEGH", "2" indicates "AAC", "3" indicates "AC3", and "4" indicates "AC4". By adding this information, on the receiving side, the encoding method of the audio data in the audio stream can be easily grasped.
[0199] The 3-bit field of "metadata_type" indicates the type of metadata. For example, "1" indicates that the ATSC's "SDO_payload()" having access information for connecting to a predetermined network service is placed in the "userdata()" field. By adding this information, on the receiving side, the type of metadata, that is, what kind of metadata it is, can be easily grasped, and for example, it is also possible to make a judgment on whether to acquire it or not.
[0200] The 1-bit flag information of "coordinated_control_flag" indicates whether the metadata is inserted only into the audio stream. For example, "1" indicates that it is also inserted into the streams of other components, and "0" indicates that it is inserted only into the audio stream. By adding this information, the receiving side can easily grasp whether the metadata is inserted only into the audio stream.
[0201] The 3-bit field of "frequency_type" indicates the type of the insertion frequency of the metadata for the audio stream. For example, "1" indicates that one piece of user data (metadata) is inserted into each audio access unit. "2" indicates that multiple pieces of user data (metadata) are inserted into one audio access unit. Further, "3" indicates that at least one piece of user data (metadata) is inserted into the first audio access unit for each group including the random access point. By adding this information, the receiving side can easily grasp the insertion frequency of the metadata for the audio stream.
[0202] [Configuration of Transport Stream TS] Figure 38 shows a configuration example of the transport stream TS. In this configuration example, there is a PES packet "video PES" of the video stream identified by PID1 and a PES packet "audio PES" of the audio stream identified by PID2. The PES packet consists of a PES header (PES_header) and a PES payload (PES_payload). Time stamps of DTS and PTS are inserted into the PES header. There is a user data area containing metadata in the PES payload of the PES packet of the audio stream.
[0203] In addition, the transport stream TS includes a Program Map Table (PMT) as Program Specific Information (PSI). PSI is information that describes which program each elementary stream included in the transport stream belongs to. In the PMT, there is a program loop that describes information related to the entire program.
[0204] In addition, the PMT has an elementary stream loop that holds information related to each elementary stream. In this configuration example, there is a video elementary stream loop (video ES loop) corresponding to the video stream and an audio elementary stream loop (audio ES loop) corresponding to the audio stream.
[0205] In the video elementary stream loop (video ES loop), information such as the stream type and PID (packet identifier) is arranged corresponding to the video stream, and a descriptor that describes information related to the video stream is also arranged. The value of "Stream_type" of this video stream is set to "0x24", and the PID information indicates PID1 assigned to the PES packet "video PES" of the video stream as described above. As one of the descriptors, an HEVC descriptor is arranged.
[0206] In addition, in an audio elementary stream loop, information such as a stream type and a PID (packet identifier) is arranged corresponding to an audio stream, and a descriptor for describing information related to the audio stream is also arranged. The value of "Stream_type" of this audio stream is set to "0x11", and the PID information indicates PID2 assigned to the PES packet "audio PES" of the audio stream as described above. As one of the descriptors, the above-described audio user data descriptor is arranged.
[0207] [Configuration Example of Set-Top Box] FIG. 39 shows a configuration example of a set-top box 200A. In this FIG. 39, the parts corresponding to FIG. 25 are denoted by the same reference numerals. The receiving unit 204A receives a transport stream TS sent from the service transmission system 100A through an RF transmission path or a communication network transmission path.
[0208] The TS analysis unit 205A extracts video stream packets from the transport stream TS and sends them to the video decoder 206. The video decoder 206 reconstructs a video stream from the video packets extracted by the demultiplexer 205 and performs decoding processing to obtain uncompressed image data. In addition, the TS analysis unit 205A extracts audio stream packets from the transport stream TS and reconstructs an audio stream. The audio framing unit 207 performs framing on the audio stream reconstructed in this way.
[0209] Note that it is also possible to send the audio stream transferred from the TS analysis unit 205A to the audio framing unit 207 and decode it with an audio decoder (not shown) to output audio in parallel.
[0210] Also, the TS analysis unit 205A extracts various descriptors and the like from the transport stream TS and transmits them to the CPU 211. Here, the descriptor also includes an audio user data descriptor (see FIG. 36) as identification information indicating that metadata is inserted into the audio stream.
[0211] Although detailed description is omitted, the other parts of the set-top box 200A shown in this FIG. 39 are configured in the same manner as the set-top box 200 shown in FIG. 25 and perform the same operations.
[0212] [Configuration Example of Television Receiver] FIG. 40 shows a configuration example of a television receiver 300A. In this FIG. 40, the parts corresponding to FIG. 27 are denoted by the same reference numerals. The receiving unit 306A receives the transport stream TS sent from the service transmission system 100A through the RF transmission path or the communication network transmission path.
[0213] The TS analysis unit 307A extracts packets of the video stream from the transport stream TS and sends them to the video decoder 308. The video decoder 308 reconstructs the video stream from the video packets extracted by the demultiplexer 205 and performs decoding processing to obtain non-compressed image data. Also, the TS analysis unit 307A extracts packets of the audio stream from the transport stream TS and reconstructs the audio stream.
[0214] Also, the TS analysis unit 307A extracts packets of the audio stream from the transport stream TS and reconstructs the audio stream. Also, the TS analysis unit 307A extracts various descriptors and the like from the transport stream TS and transmits them to the CPU 321. Here, this descriptor also includes an audio user data descriptor (see FIG. 36) as identification information indicating that metadata is inserted into the audio stream.
[0215] Although detailed description is omitted, the other parts of the television receiver 300A shown in FIG. 40 are configured and operate in the same manner as the television receiver 300 shown in FIG. 27.
[0216] As described above, in the image display systems 10A and 10A' shown in FIGS. 34(a) and 34(b), the service transmission system 100A inserts metadata into the audio stream and inserts identification information indicating that metadata is inserted into the audio stream into the container layer. Therefore, on the receiving side (set-top box 200A, television receiver 300A), it is possible to easily recognize that metadata is inserted into the audio stream.
[0217] Also, in the image display system 10A shown in FIG. 34(a), the set-top box 200A transmits the audio stream with metadata inserted thereto to the television receiver 300A via HDMI together with the identification information indicating that metadata is inserted into the audio stream. Therefore, the television receiver 300A can easily recognize that metadata is inserted into the audio stream, and by performing the extraction process of the metadata inserted into the audio stream based on this recognition, the metadata can be surely obtained and used without waste.
[0218] Also, in the image display system 10A' shown in FIG. 34(b), the television receiver 300A extracts metadata from the audio stream based on the identification information received together with the audio stream and uses it for processing. Therefore, the metadata inserted into the audio stream can be surely obtained without waste, and the processing using the metadata can be appropriately executed.
[0219] In the above-described embodiment, the set-top box 200 is configured to transmit image data and an audio stream to the television receiver 300. However, instead of the television receiver 300, a configuration for transmitting to a monitor device, a projector, or the like is also conceivable. Further, instead of the set-top box 200, a configuration using a recorder with a reception function, a personal computer, or the like is also conceivable.
[0220] In the above-described embodiment, the set-top box 200 and the television receiver 300 are connected by an HDMI cable 400. However, it goes without saying that the present invention can be similarly applied when these are connected by wire using a digital interface similar to HDMI, or even when they are connected wirelessly.
[0221] Note that the present technology can also have the following configuration. (1) A transmission unit that transmits a meta file having meta information for acquiring an audio stream with inserted metadata by a receiving device, and an information insertion unit that inserts identification information indicating that the metadata is inserted into the audio stream into the meta file. A transmission device. (2) The transmission device according to (1) above, wherein the metadata is access information for connecting to a predetermined network service. The transmission device according to (1) above. (3) The transmission device according to (2) above, wherein the metadata is a character code indicating URI information. The transmission device according to (2) above. (4) The transmission device according to any one of (1) to (3) above, wherein the meta file is an MPD file. The transmission device according to any one of (1) to (3) above. (5) The transmission device according to (4) above, wherein the information insertion unit inserts the identification information into the meta file using "Supplementary Descriptor". The transmission device according to (4) above. (6) The transmission unit according to (4) above, Transmit the above metadata file through an RF transmission path or a communication network transmission path The transmission device according to any one of (1) to (5) above (7) The above transmission unit Further transmit a container in a predetermined format including an audio stream into which the above metadata is inserted The transmission device according to any one of (1) to (6) above (8) The above container is an MP4 The transmission device according to (7) above (9) A transmission step of transmitting, by the transmission unit, a metadata file having meta information for acquiring, by a receiving device, an audio stream into which metadata is inserted; An information insertion step of inserting, into the above metadata file, identification information indicating that the above metadata is inserted into the above audio stream Transmission method (10) A receiving unit that receives a metadata file having meta information for acquiring an audio stream into which metadata is inserted; Identification information indicating that the above metadata is inserted into the above audio stream is inserted into the above metadata file; The apparatus further comprises a transmission unit that transmits the above audio stream, together with identification information indicating that metadata is inserted into the above audio stream, to an external device via a predetermined transmission path Receiving device (11) The above metadata is access information for connecting to a predetermined network service The receiving device according to (10) above (12) The above metadata file is an MPD file; The above identification information is inserted into the above metadata file by "Supplementary Descriptor" The receiving device according to (10) or (11) above (13) The above transmission unit Insert the audio stream and the identification information during the blanking period of the image data, and transmit the image data to the external device, thereby transmitting the audio stream and the identification information to the external device. The receiving device according to any one of (10) to (12) above. (14) The predetermined transmission path is an HDMI cable. The receiving device according to any one of (10) to (13) above. (15) A receiving step of receiving, by a receiving unit, a meta file having meta information for acquiring an audio stream into which metadata is inserted. The meta file has identification information inserted therein indicating that the metadata is inserted into the audio stream. The method further includes a transmitting step of transmitting the audio stream to an external device via a predetermined transmission path together with identification information indicating that metadata is inserted into the audio stream. Receiving method. (16) A receiving unit that receives a meta file having meta information for acquiring an audio stream into which metadata is inserted. The meta file has identification information inserted therein indicating that the metadata is inserted into the audio stream. A metadata extraction unit that decodes the audio stream based on the identification information and extracts the metadata. The apparatus further includes a processing unit that performs processing using the metadata. Receiving device. (17) The meta file is an MPD file. In the meta file, the identification information is inserted by "Supplementary Descriptor". The receiving device according to (16) above. (18) The metadata is access information for connecting to a predetermined network service. The processing unit Access a predetermined server on the network based on the network access information The receiving device according to (16) or (17) above. (19) A receiving step of receiving, by a receiving unit, a meta file having meta information for acquiring an audio stream into which metadata is inserted The meta file has identification information inserted therein indicating that the metadata is inserted into the audio stream A metadata extraction step of decoding the audio stream based on the identification information and extracting the metadata And a processing step of performing processing using the metadata Receiving method. (20) A stream generation unit that generates an audio stream into which metadata including network access information is inserted And a transmission unit that transmits a container in a predetermined format having the audio stream Transmission device.
[0222] The main feature of this technology is that when inserting metadata into an audio stream in DASH / MP4 distribution, by inserting identification information indicating that metadata is inserted into the audio stream into the MPD file, on the receiving side, it is made easy to recognize that metadata is inserted into the audio stream (see FIGS. 3 and 4).
Explanation of Signs
[0223] 10, 10´, 10A, 10A´ ··· Transmission / reception system 14 ··· Effective pixel section 15 ··· Horizontal blanking period 16 ··· Vertical blanking period 17 ··· Video data section 18 ··· Data island section 19 ··· Control section 30A, 30B ··· MPEG-DASH-based stream distribution system 31... DASH Stream File Server 32... DASH MPD Server 33, 33-1 to 33-N... Receiving System 34... CDN 35, 35-1 to 35-M... Receiving System 36... Broadcasting Transmission System 81... HDMI Transmitter 82... HDMI Receiver 83... DDC 84... CEC Line 85... EDID ROM 100, 100A... Service Transmission System 110... DASH / MP4 Generation Unit 110A... TS Generation Unit 111... Control Unit 111a... CPU 112... Video Encoder 113... Audio Encoder 113a... Audio Encoding Block Unit 113b... Audio Framing Unit 114... DASH / MP4 Formatter 114A... TS Formatter 200, 200A... Set Top Box (STB) 204, 204A... Receiver 205... DASH / MP4 Analysis Unit 205A... TS Analysis Unit 206... Video Decoder 207... Audio Framing Unit 208... HDMI Transmitting Unit 209... HDMI Terminal 211... CPU211 212... Flash ROM 213... DRAM 214... Internal Bus 215... Remote Control Receiver 216... Remote Control Transmitter 300, 300A ··· TV receiver 306, 306A ··· Receiver section 307 ··· DASH / MP4 analysis section 307A ··· TS analysis section 308 ··· Video decoder 309 ··· Video processing circuit 310 ··· Panel drive circuit 311 ··· Display panel 312 ··· Audio decoder 313 ··· Audio processing circuit 314 ··· Audio amplifier circuit 315 ··· Speaker 316 ··· HDMI terminal 317 ··· HDMI receiver section 318 ··· Communication interface 321 ··· CPU 322 ··· Flash ROM 323 ··· DRAM 324 ··· Internal bus 325 ··· Remote control receiver section 326 ··· Remote control transmitter 350 ··· External speaker system 400 ··· HDMI cable< / baseurl> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / baseurl>
Claims
1. a receiving unit that receives an audio stream; a decoding unit that performs a decoding process on the audio stream to obtain audio data; an audio processing unit that processes the audio data and outputs it to an output unit, and the audio stream is composed of MPEG Audio Stream Packets, the MPEG Audio Stream Packet is composed of a header and a payload, the payload includes "SYNC" corresponding to a synchronization start code, "Frame" which is actual data of transmission data of 3D audio, and "Config" indicating the configuration of "Frame", the "Frame" includes channel-encoded data or object-encoded data constituting transmission data of 3D audio, the object-encoded data is composed of encoded sample data of SCE (Single Channel Element) and metadata for mapping the encoded sample data to a speaker existing at an arbitrary position for rendering, the metadata is included as an extension element (Ext_element), An information processing apparatus.
2. The type (ExElementType) of the extension element corresponding to the case where the value of the extension element (Ext_element) is 0 is ID_EXT_ELE_FILL, The information processing apparatus according to claim 1.
3. The type (ExElementType) of the extension element corresponding to the case where the value of the extension element (Ext_element) is 1 is ID_EXT_ELE_MPEGS, The information processing apparatus according to claim 1.
4. The type (ExElementType) of the extension element corresponding to the case where the value of the extension element (Ext_element) is 2 is ID_EXT_ELE_SAOC, The information processing apparatus according to claim 1.
5. The type (ExElementType) of the extension element corresponding to the case where the value of the extension element (Ext_element) is 3 is ID_EXT_ELE_AUDIOPREROLL, The information processing apparatus according to claim 1.
6. When the value of the extension element (Ext_element) is 4, the type (ExElementType) of the extension element is ID_EXT_ELE_UNI_DRC, The information processing apparatus according to claim 1.
7. When the value of the extension element (Ext_element) is 5, the type (ExElementType) of the extension element is ID_EXT_ELE_OBJ_METADATA, The information processing apparatus according to claim 1.
8. When the value of the extension element (Ext_element) is 6, the type (ExElementType) of the extension element is ID_EXT_ELE_SAOC_3D, The information processing apparatus according to claim 1.
9. When the value of the extension element (Ext_element) is 7, the type (ExElementType) of the extension element is ID_EXT_ELE_HOA, The information processing apparatus according to claim 1.
10. When the value of the extension element (Ext_element) is 128 or more, it is extensible up to non-MPEG, The information processing apparatus according to claim 1.
11. The audio stream has a structure in MPEG-H (3D audio), The information processing apparatus according to claim 1.
12. The header includes information on packet type, packet label, and packet length, The information processing apparatus according to claim 1.
13. The audio stream is included in the transport stream of MPEG2-TS, The information processing apparatus according to claim 1.
14. The audio stream further includes identification information indicating that the audio stream includes the metadata, The information processing apparatus according to claim 1.
15. A procedure for receiving an audio stream, A procedure for performing decoding processing on the audio stream to obtain audio data, A procedure for processing the audio data and outputting it to an output unit, The audio stream is composed of MPEG Audio Stream Packets, The MPEG Audio Stream Packet is composed of a Header and a Payload, The Payload includes "SYNC" corresponding to a synchronization start code, "Frame" which is the actual data of the transmission data of 3D audio, and "Config" indicating the configuration of the "Frame", The "Frame" includes channel-encoded data or object-encoded data constituting the transmission data of 3D audio, The object-encoded data is composed of encoded sample data of SCE (Single Channel Element) and metadata for mapping the encoded sample data to speakers existing at arbitrary positions for rendering, The metadata is included as an extension element (Ext_element), Information processing method.
16. A procedure for receiving an audio stream, A procedure for performing a decoding process on the audio stream to obtain audio data, A procedure for processing the audio data and outputting it to an output unit, and has, The audio stream is composed of MPEG Audio Stream Packets, The MPEG Audio Stream Packet is composed of a Header and a Payload, The Payload includes "SYNC" corresponding to a synchronization start code, "Frame" which is the actual data of the transmission data of 3D audio, and "Config" indicating the configuration of the "Frame", The "Frame" includes channel-encoded data or object-encoded data constituting the transmission data of 3D audio, The object-encoded data is composed of encoded sample data of SCE (Single Channel Element) and metadata for mapping the encoded sample data to speakers existing at arbitrary positions for rendering, The metadata is included as an extension element (Ext_element), A program for causing a computer to execute an information processing method.
Citation Information
Patent Citations
Object-Oriented Audio Streaming System
JP2013502183A
System and method for adaptive audio signal generation, coding, and rendering
JP2014522155A
Transmission device, transmission method, receiving device, receiving method, program, and broadcasting system
WO2012133064A1
Mapping virtual speakers to physical speakers
WO2014124268A1
Transmitter, transmission method, receiver, reception method and transmission / reception system
JP2012010311A