Information processing device, information processing method, and program

By inserting metadata into audio streams with identification information, the technology enhances metadata recognition and processing, ensuring efficient acquisition and network connectivity for receiving devices.

JP7852776B2Active Publication Date: 2026-04-28SONY GROUP CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2025-05-09
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Metadata is not consistently inserted into audio streams, making it difficult for receiving devices to recognize and process the metadata effectively.

Method used

A transmission unit inserts metadata into audio streams and includes identification information in a metadata file, allowing receiving devices to easily identify and extract the metadata using a 'Supplementary Descriptor' in formats like MP4 or MPD files.

Benefits of technology

Ensures efficient and reliable metadata acquisition by enabling receiving devices to recognize and process metadata inserted into audio streams, facilitating seamless network connectivity and service access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852776000001
    Figure 0007852776000001
  • Figure 0007852776000002
    Figure 0007852776000002
  • Figure 0007852776000003
    Figure 0007852776000003
Patent Text Reader

Abstract

To allow a reception side to easily recognize insertion of meta data in an audio stream.SOLUTION: A meta file is sent which has meta information for a reception device to acquire an audio stream with meta data inserted therein. Identification information showing that the meta data is inserted in the audio stream is inserted into the meta file. The reception side can easily recognize insertion of the meta data in the audio stream from identification information inserted in the meta file.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, it has been proposed to insert metadata into an audio stream and transmit it (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Metadata is defined, for example, in a user data area of an audio stream. However, metadata is not inserted into all audio streams.

[0005] An object of the present technology is to make it easy for a receiving side to recognize that metadata is inserted into an audio stream and to facilitate processing.

Means for Solving the Problems

[0006] The concept of the present technology is a transmission unit that transmits a meta file having meta information for acquiring an audio stream into which metadata is inserted by a receiving device, and an information insertion unit that inserts identification information indicating that the metadata is inserted into the audio stream into the meta file. in a transmission device.

[0007] In this technology, the transmitting unit transmits a metadata file containing metadata information for the receiving device to acquire an audio stream into which metadata has been inserted. For example, the metadata may be access information for connecting to a predetermined network service. In this case, for example, the metadata may be a character code indicating URI information.

[0008] Furthermore, for example, the transmitting unit may transmit the metadata file via an RF transmission path or a communication network transmission path. Alternatively, for example, the transmitting unit may further transmit a container in a predetermined format containing an audio stream with the metadata inserted. In this case, for example, the container may be MP4 (ISO / IEC 14496-14:2003).

[0009] The information insertion unit inserts identification information into the metafile to indicate that metadata has been inserted into the audio stream. For example, the metafile may be an MPD (Media Presentation Description) file. In this case, for example, the information insertion unit may use a “Supplementary Descriptor” to insert the identification information into the metafile.

[0010] In this technology, an identifier indicating that metadata has been inserted into the audio stream is inserted into a metadata file containing metadata for the receiving device to acquire the audio stream into which metadata has been inserted. Therefore, the receiving device can easily recognize that metadata has been inserted into the audio stream. Furthermore, it is possible to perform metadata extraction processing based on this recognition, enabling efficient and reliable metadata acquisition.

[0011] Furthermore, other concepts of this technology include: It includes a receiving unit that receives a metafile containing metadata for obtaining an audio stream with metadata inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. The system further includes a transmitting unit that transmits the above audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to an external device via a predetermined transmission path. It is located in the receiving device.

[0012] In this technology, the receiving unit receives a metafile containing metadata for acquiring an audio stream into which metadata has been inserted. For example, the metadata may be access information for connecting to a predetermined network service. The metafile contains identification information indicating that metadata has been inserted into the audio stream.

[0013] For example, metadata may be access information for connecting to a specific network service. Alternatively, for example, the metafile may be an MPD file, and this metafile may have identification information inserted by a “Supplementary Descriptor”.

[0014] The transmitting unit transmits the audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to an external device via a predetermined transmission path. For example, the transmitting unit may transmit the audio stream and identification information to the external device by inserting the audio stream and identification information during the blanking period of the image data and transmitting this image data to the external device. Alternatively, the predetermined transmission path may be, for example, an HDMI cable.

[0015] In this technology, the audio stream with the inserted metadata is transmitted to an external device along with identification information indicating that metadata has been inserted into the audio stream. Therefore, the external device can easily recognize that metadata has been inserted into the audio stream. Based on this recognition, it is possible to perform a process to extract the metadata inserted into the audio stream, enabling efficient and reliable metadata acquisition.

[0016] Furthermore, other concepts of this technology include: It includes a receiving unit that receives a metafile containing metadata for obtaining an audio stream with metadata inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. A metadata extraction unit decodes the audio stream based on the above identification information and extracts the above metadata, The system further comprises a processing unit that performs processing using the above metadata. It is located in the receiving device.

[0017] In this technology, the receiving unit receives a metafile containing metadata for acquiring an audio stream into which metadata has been inserted. The metafile contains identification information indicating that metadata has been inserted into the audio stream. For example, the metafile may be an MPD file, and this metafile may contain identification information inserted by a “Supplementary Descriptor”.

[0018] The metadata extraction unit decodes the audio stream based on the identification information and extracts the metadata. Then, the processing unit performs processing using this metadata. For example, the metadata may be access information for connecting to a predetermined network service, and the processing unit may access a predetermined server on the network based on the network access information.

[0019] Thus, in the present technology, based on the identification information indicating that metadata is inserted into the audio stream, which is inserted into the metafile, the metadata is extracted from the audio stream and used for processing. Therefore, the metadata inserted into the audio stream can be surely obtained without waste, and the processing using the metadata can be appropriately executed.

[0020] Further, another concept of the present technology is a stream generation unit that generates an audio stream into which metadata including network access information is inserted, and a transmission unit that transmits a container in a predetermined format having the audio stream. This is in the transmission device.

[0021] In the present technology, the stream generation unit generates an audio stream into which metadata including network access information is inserted. For example, the audio stream is generated by encoding audio data with AAC, AC3, AC4, MPEGH (3D audio), etc., and the metadata is embedded in the user data area thereof.

[0022] The transmission unit transmits a container in a predetermined format having the audio stream. Here, the container in a predetermined format is, for example example, example, MP4, MPEG2-TS, etc. For example, the metadata may be a character code indicating URI information.

[0023] Thus, in the present technology, the metadata including network access information is embedded in the audio stream and transmitted. Therefore, for example, from a broadcasting station, a distribution server, etc., the network access information can be easily transmitted with the audio stream as a container and made available for use on the receiving side.

Advantages of the Invention

[0024] This technology makes it easy for the receiving end to recognize that metadata has been inserted into the audio stream. The effects described herein are merely illustrative and not limited to those described herein, and additional effects may also exist. [Brief explanation of the drawing]

[0025] [Figure 1] This block diagram shows an example configuration of an MPEG-DASH-based streaming system. [Figure 2] This diagram shows an example of the hierarchical relationships between the various structures in an MPD file. [Figure 3] This block diagram shows an example configuration of a transmission and reception system as an embodiment. [Figure 4] This figure shows an example of an MPD file description. [Figure 5] This figure shows an example of a definition of "schemeIdUri" using "SupplementaryDescriptor". [Figure 6] This diagram illustrates an example of the placement of video and audio access units in a transport stream, and the frequency of metadata insertion into the audio stream. [Figure 7] " <baseurl>This is a diagram to explain the actual media file at the location indicated by "[ ]". [Figure 8] This block diagram shows an example configuration of the DASH / MP4 generation unit included in the service transmission system. [Figure 9] This diagram shows the structure of an AAC audio frame. [Figure 10] This diagram shows the configuration of a "DSE (data stream element)" into which metadata MD is inserted when the compression format is AAC. [Figure 11] This diagram shows the structure of "metadata()" and the contents of its main constituent information. [Figure 12] This diagram shows the structure of "SDO_payload()". [Figure 13] This diagram shows the meaning of the command ID (cmdID) value. [Figure 14] This diagram shows the structure of the AC3 frame (AC3 Synchronization Frame). [Figure 15] This diagram shows the structure of the AC3 Auxiliary Data. [Figure 16] This diagram shows the layer structure of the AC4 Simple Transport. [Figure 17] This diagram shows the schematic structure of the TOC (ac4_toc()) and substream (ac4_substream_data()). [Figure 18] This diagram shows the structure of "umd_info()" within TOC(ac4_toc()). [Figure 19] This diagram shows the structure of "umd_payloads_substream()" within the substream (ac4_substream_data()). [Figure 20] This diagram shows the structure of an audio frame (1024 samples) in MPEGH (3D audio) transmission data. [Figure 21] This diagram illustrates how the correspondence between the configuration information (config) of each "Frame" contained within "Config" and each "Frame" is maintained. [Figure 22] This diagram shows the correspondence between the type (ExElementType) and value (Value) of an extension element (Ext_element). [Figure 23] This diagram shows the structure of "userdataConfig()". [Figure 24] This diagram shows the structure of "userdata()". [Figure 25] This block diagram shows an example configuration of a set-top box that makes up a transmission / reception system. [Figure 26] This diagram shows an example of the structure of an audio infoframe packet placed in a data island section. [Figure 27] This block diagram shows an example configuration of a television receiver that makes up a transmission and reception system. [Figure 28] This is a block diagram showing an example configuration of the HDMI transmitter of a set-top box and the HDMI receiver of a television receiver. [Figure 29] This diagram shows the various transmission data intervals when image data is transmitted over a TMDS channel. [Figure 30] This diagram illustrates a specific example of metadata processing in a television receiver. [Figure 31] This diagram shows an example of screen display transitions when accessing network services based on metadata using a television receiver. [Figure 32] This is a block diagram showing the configuration of the audio output system in a television receiver according to an embodiment. [Figure 33] This block diagram shows another example of an audio output system configuration in a television receiver. [Figure 34] This block diagram shows other configuration examples of the transmission and reception system. [Figure 35] This block diagram shows an example configuration of the TS generation unit included in the service transmission system. [Figure 36] This diagram shows an example of the structure of an audio user data descriptor. [Figure 37] This diagram shows the contents of key information in an example of the structure of an audio user data descriptor. [Figure 38] This figure shows an example of a transport stream configuration. [Figure 39] This block diagram shows an example configuration of a set-top box that makes up a transmission / reception system. [Figure 40] This block diagram shows an example configuration of a television receiver that makes up a transmission and reception system. [Modes for carrying out the invention]

[0026] The following describes embodiments for carrying out the invention. The description will be given in the following order. 1. Embodiment 2. Variations

[0027] <1. Embodiment> [Overview of MPEG-DASH-based streaming systems] First, we will describe an overview of an MPEG-DASH-based streaming system to which this technology can be applied.

[0028] Figure 1(a) shows an example configuration of an MPEG-DASH-based stream distribution system 30A. In this example configuration, media streams and MPD files are transmitted through a communication network transmission path. This stream distribution system 30A consists of a DASH stream file server 31 and a DASH MPD server 32, to which N receiving systems 33-1, 33-2, ..., 33-N are connected via a CDN (Content Delivery Network) 34.

[0029] The DASH stream file server 31 generates a DASH-compliant stream segment (hereinafter referred to as "DASH segment" as appropriate) based on the media data (video data, audio data, subtitle data, etc.) of the specified content, and sends the segment in response to an HTTP request from the receiving system. This DASH stream file server 31 may be a server dedicated to streaming, or it may also be used as a web server.

[0030] Furthermore, the DASH stream file server 31 responds to requests for predetermined stream segments sent from receiving systems 33 (33-1, 33-2, ..., 33-N) via CDN 34, and transmits those stream segments to the requesting receiver via CDN 34. In this case, the receiving system 33 refers to the rate value described in the MPD (Media Presentation Description) file and selects the stream with the optimal rate according to the network environment where the client is located, and makes the request.

[0031] The DASH MPD server 32 is a server that generates MPD files for obtaining DASH segments generated by the DASH stream file server 31. It generates MPD files based on content metadata from the content management server (not shown) and the address (url) of the segment generated by the DASH stream file server 31. Note that the DASH stream file server 31 and the DASH MPD server 32 may be physically the same.

[0032] In the MPD format, each stream, such as video and audio, uses an element called "Representation" to describe its attributes. For example, an MPD file contains separate representations for each of several video data streams with different rates, each describing its rate. The receiving system 33 can use these rate values ​​as a reference to select the optimal stream according to the network environment in which the receiving system 33 is located, as described above.

[0033] Figure 1(b) shows an example configuration of an MPEG-DASH-based stream distribution system 30B. In this example configuration, media streams and MPD files are transmitted via an RF transmission path. This stream distribution system 30B consists of a broadcast transmission system 36 to which a DASH stream file server 31 and a DASH MPD server 32 are connected, and M receiving systems 35-1, 35-2, ..., 35-M.

[0034] In this stream distribution system 30B, the broadcast transmission system 36 transmits the DASH-compliant stream segments (DASH segments) generated by the DASH stream file server 31 and the MPD files generated by the DASH MPD server 32 on the broadcast wave.

[0035] Figure 2 shows an example of the hierarchical relationships between the various structures in an MPD file. As shown in Figure 2(a), the Media Presentation as a whole MPD file contains multiple periods separated by time intervals. For example, the first period starts at 0 seconds, the next period starts at 100 seconds, and so on.

[0036] As shown in Figure 2(b), a period has multiple representations. These multiple representations include groups of representations related to stream attributes, such as media streams with the same content but different rates, which are grouped by an AdaptationSet.

[0037] As shown in Figure 2(c), the representation includes SegmentInfo. As shown in Figure 2(d), this SegmentInfo contains an Initialization Segment and multiple Media Segments, each containing information for a segment further divided by periods. The Media Segments contain information such as addresses (urls) for actually retrieving segment data such as video and audio.

[0038] Furthermore, it is possible to freely switch between streams among multiple representations grouped in an adaptation set. This allows the system to select the optimal rate stream according to the network environment in which the receiving system is located, enabling uninterrupted delivery.

[0039] [Configuration of the transmission / reception system] Figure 3 shows an example configuration of a transmission / reception system as an embodiment. The transmission / reception system 10 in Figure 3(a) includes a service transmission system 100, a set-top box (STB) 200, and a television receiver (TV) 300. The set-top box 200 and the television receiver 300 are connected via an HDMI (High Definition Multimedia Interface) cable 400. Note that "HDMI" is a registered trademark.

[0040] In this transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31 and DASH MPD server 32 of the stream distribution system 30A shown in Figure 1(a) above. In addition, in this transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31, DASH MPD server 32 and broadcast transmission system 36 of the stream distribution system 30B shown in Figure 1(b) above.

[0041] In this transmission / reception system 10, the set-top box (STB) 200 and the television receiver (TV) 300 correspond to the receiving systems 33 (33-1, 33-2, ..., 33-N) of the stream distribution system 30A shown in Figure 1(a) above. In addition, in this transmission / reception system 10, the set-top box (STB) 200 and the television receiver (TV) 300 correspond to the receiving systems 35 (35-1, 35-2, ..., 35-M) of the stream distribution system 30B shown in Figure 1(b) above.

[0042] Furthermore, the transmission / reception system 10' in Figure 3(b) includes a service transmission system 100 and a television receiver (TV) 300. In this transmission / reception system 10', the service transmission system 100 corresponds to the DASH stream file server 31 and DASH MPD server 32 of the stream distribution system 30A shown in Figure 1(a) above. Also, in this transmission / reception system 10', the service transmission system 100 corresponds to the DASH stream file server 31, DASH MPD server 32 and broadcast transmission system 36 of the stream distribution system 30B shown in Figure 1(b) above.

[0043] In this transmission / reception system 10', the television receiver (TV) 300 corresponds to the receiving system 33 (33-1, 33-2, ..., 33-N) of the stream distribution system 30A shown in Figure 1(a) above. Also, in this transmission / reception system 10', the television receiver (TV) 300 corresponds to the receiving system 35 (35-1, 35-2, ..., 35-M) of the stream distribution system 30B shown in Figure 1(b) above.

[0044] The service transmission system 100 transmits DASH / MP4, that is, an MPD file as a metadata file, and an MP4 containing media streams (media segments) such as video and audio, via an RF transmission path or a communication network transmission path. The service transmission system 100 inserts metadata into the audio stream. This metadata could include, for example, access information for connecting to a predetermined network service, or predetermined content information. In this embodiment, access information for connecting to a predetermined network service is inserted.

[0045] The service transmission system 100 inserts identification information into the MPD file indicating that metadata has been inserted into the audio stream. The service transmission system 100 inserts identification information indicating that metadata has been inserted into the audio stream, for example, using a "Supplementary Descriptor".

[0046] Figure 4 shows an example of an MPD file description. <adaptationset mimetype=""audio / mp4”" group=""1”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 1.

[0047] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:AudioMetaContained”" value=""true” / ">The description " indicates that metadata is inserted into the audio stream. The "SupplementaryDescriptor" allows for a new definition of "schemeIdUri" for broadcast and other applications, separate from the predefined definitions in previous standards. As shown in Figure 5, "schemeIdUri="urn:brdcst:AudiometaContained"" indicates that audio metadata is included, that is, that metadata is inserted into the audio stream. For example, when "value" is "true", it indicates that audio metadata is included. When "value" is "false", it indicates that audio metadata is not included.

[0048] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is MPEGH (3D audio). As shown in Figure 5, "schemeIdUri="urn:brdcst:codecType"" indicates the type of codec. For example, "value" can be "mpegh", "AAC", "AC3", "AC4", etc.

[0049] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:coordinatedControl”" value=""true” / ">The description "" indicates that the information necessary for network connectivity is supplied in coordination across multiple media streams. As shown in Figure 5, "schemeIdUri="urn:brdcst:coordinatedControl"" indicates that the information necessary for network connectivity is supplied in coordination across multiple media streams. For example, when "value" is "true", it indicates that network connectivity information is supplied in coordination with streams from other adaptation sets. When "value" is "false", it indicates that network connectivity information is supplied only by the streams of this adaptation set.

[0050] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:type”" value=""netlink” / ">The description indicates that the type of service provided by the metadata is network connectivity. As shown in Figure 5, the type of service provided by the metadata is indicated by "schemeIdUri="urn:brdcst:type"". For example, when "value" is "netlink", it indicates that the type of service provided by the metadata is network connectivity.

[0051] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:metaInsertionFrequency”" value=""1” / ">The description indicates that metadata is supplied on an access unit basis. As shown in Figure 5, "schemeIdUri="urn:brdcst:metaInsertionFrequency"" indicates the frequency at which metadata is supplied on an access unit basis. For example, when "value" is "1", it indicates that one user data entry occurs in one access unit. When "value" is "2", it indicates that multiple user data entries occur in one access unit. When "value" is "3", it indicates that one or more user data entries occur during a period separated by random access points.

[0052] Figure 6(a) shows an example of the arrangement of video and audio access units containerized in MP4. "VAU" indicates a video access unit. "AAU" indicates an audio access unit. Figure 6(b) shows the case where "frequency_type = 1" and one user data entry (metadata) is inserted into each audio access unit.

[0053] Figure 6(c) shows the case where "frequency_type = 2", indicating that multiple user data (metadata) are inserted into a single audio access unit. Figure 6(d) shows the case where "frequency_type = 3", indicating that for each group containing random access points, at least one user data (metadata) is inserted into the first audio access unit.

[0054] Returning to Figure 4, <representation id=""11”" bandwidth=""128000”">The description " " indicates the existence of an audio stream with a bitrate of 128kbps, referred to as "Representation id="11"". And, <baseurl> audio / jp / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp / 128.mp4".

[0055] Also," <adaptationset mimetype=""video / mp4”" group=""2”">The description indicates that an AdaptationSet exists for the video stream, that the video stream is supplied in an MP4 file structure, and that it is assigned to group 2.

[0056] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:VideoMetaContained”" value=""true” / ">The description indicates that metadata is inserted into the video stream. As shown in Figure 5, "schemeIdUri="urn:brdcst:VideoMetaContained"" indicates that video metadata is included, that is, that metadata is inserted into the video stream. For example, when "value" is "true", it indicates that video metadata is included. When "value" is "false", it indicates that video metadata is not included.

[0057] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""hevc” / ">The description " indicates that the video stream's codec is HEVC. Also, <supplementarydescriptor schemeiduri=""urn:brdcst:coordinatedControl”" value=""true” / ">The description indicates that the information necessary for network connectivity is supplied in a coordinated manner across multiple media streams.

[0058] Also," <supplementarydescriptor schemeiduri=""urn:brdcst:type”" value=""netlink” / ">The description " indicates that the type of service provided by the metadata is network connectivity. <supplementarydescriptor schemeiduri=""urn:brdcst:metaInsertionFrequency”" value=""1” / ">The description indicates that metadata is supplied on an access unit basis.

[0059] Also," <representation id=""21”" bandwidth=""20000000”">The description " indicates the existence of a video stream with a bitrate of 20Mbps, referred to as "Representation id="21"". <baseurl> video / jp / 20000000.mp4< / baseurl> The description indicates that the location of the video stream is "video / jp / 20000000.mp4".

[0060] Here, <baseurl>This section describes the media file entity at the location indicated by ". In the case of Non-Fragmented MP4, for example, it may be defined as "url 1" as shown in Figure 7(a). In this case, the "ftyp" box, which describes the file type, is placed first. This "ftyp" box indicates that it is an unfragmented MP4 file. Next, the "moov" box and the "mdat" box are placed. The "moov" box contains all metadata, such as header information and content metadata for each track, and time information. The "mdat" box contains the media data itself.

[0061] In the case of Fragmented MP4, for example, it may be defined as "url 2" as shown in Figure 7(b). In this case, the "styp" box, which describes the segment type, is placed first. Next, the "sidx" box, which describes the segment index, is placed. Following that, a predetermined number of movie fragments are placed. Here, a movie fragment consists of a "moof" box containing control information and an "mdat" box containing the media data itself. Since the "mdat" box of a single movie fragment contains fragments obtained by fragmenting the transmission media, the control information placed in the "moof" box is the control information for that fragment.

[0062] Furthermore, the combination of "url 1" and "url 2" mentioned above is also possible. In this case, for example, "url 1" can be used as the initialization segment, and "url 1" and "url 2" can be combined into a single MP4 service. Alternatively, "url 1" and "url 2" can be combined and defined as "url 3," as shown in Figure 7(c).

[0063] The set-top box 200 receives DASH / MP4 files, i.e., MPD files as metadata, and MP4 files containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path. The audio streams contained in the MP4 files have access information for connecting to a predetermined network service inserted as metadata. In addition, the MPD files have identification information inserted by a "Supplementary Descriptor" to indicate that metadata has been inserted into the audio stream.

[0064] The set-top box 200 transmits the audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to the television receiver 300 via the HDMI cable 400.

[0065] Here, the set-top box 200 transmits the audio stream and identification information to the television receiver 300 by inserting the audio stream and identification information into the blanking period of the image data obtained by decoding the video stream and transmitting this image data to the television receiver 300. The set-top box 200 inserts this identification information, for example, into an Audio InfoFrame packet.

[0066] In the transmission / reception system 10 shown in Figure 3(a), the television receiver 300 receives an audio stream from the set-top box 200 via an HDMI cable 400, along with identification information indicating that metadata has been inserted into the audio stream. In other words, the television receiver 300 receives the audio stream and image data from the set-top box 200, in which the identification information has been inserted during the blanking period.

[0067] The television receiver 300 then decodes the audio stream based on the identification information, extracts metadata, and performs processing using this metadata. In this case, the television receiver 300 accesses a predetermined server on the network based on predetermined network service information as metadata.

[0068] Furthermore, in the transmission / reception system 10' shown in Figure 3(b), the television receiver 300 receives DASH / MP4, that is, an MPD file as a metadata file, and an MP4 containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path. The audio stream contained in the MP4 has access information for connecting to a predetermined network service inserted as metadata. In addition, the MPD file has identification information inserted by a "Supplementary Descriptor" to indicate that metadata has been inserted into the audio stream.

[0069] The television receiver 300 then decodes the audio stream based on the identification information, extracts metadata, and performs processing using this metadata. In this case, the television receiver 300 accesses a predetermined server on the network based on predetermined network service information as metadata.

[0070] [DASH / MP4 generation unit of the service transmission system] Figure 8 shows an example configuration of the DASH / MP4 generation unit 110 included in the service transmission system 100. This DASH / MP4 generation unit 110 includes a control unit 111, a video encoder 112, an audio encoder 113, and a DASH / MP4 formatter 114.

[0071] The control unit 111 is equipped with a CPU 111a and controls each part of the DASH / MP4 generation unit 110. The video encoder 112 encodes the image data SV using MPEG2, H.264 / AVC, H.265 / HEVC, etc., to generate a video stream (video elementary stream). The image data SV is, for example, image data played back from a recording medium such as an HDD, or live image data obtained from a video camera.

[0072] The audio encoder 113 encodes the audio data SA using a compression format such as AAC, AC3, AC4, or MPEGH (3D audio) to generate an audio stream (audio elementary stream). The audio data SA is the audio data corresponding to the image data SV mentioned above, and can be audio data played back from a recording medium such as an HDD, or live audio data obtained from a microphone.

[0073] The audio encoder 113 has an audio encoding block section 113a and an audio framing section 113b. The audio encoding block section 113a generates encoding blocks, and the audio framing section 113b performs framing. In this case, the encoding blocks and framing will differ depending on the compression format.

[0074] The audio encoder 113 inserts metadata MD into the audio stream under the control of the control unit 111. In this embodiment, metadata MD is access information for connecting to a predetermined network service. Here, the predetermined network service can be any service, such as a music network service or an audio-video network service. Here, metadata MD is embedded in the user data area of ​​the audio stream.

[0075] The DASH / MP4 formatter 114 generates an MP4 file containing media streams (media segments) such as video and audio content, based on the video stream output from the video encoder 112 and the audio stream output from the audio encoder 113. It also generates an MPD file using content metadata and segment URL information. Here, the MPD file includes identification information indicating that metadata has been inserted into the audio stream (see Figure 4).

[0076] The operation of the DASH / MP4 generation unit 110 shown in Figure 8 will be briefly explained. Image data SV is supplied to the video encoder 112. The video encoder 112 encodes the image data SV using H.264 / AVC, H.265 / HEVC, etc., and generates a video stream containing encoded video data.

[0077] The audio data SA is also supplied to the audio encoder 113. The audio encoder 113 encodes the audio data SA using AAC, AC3, AC4, MPEGH (3D audio), etc., to generate an audio stream.

[0078] At this time, the control unit 111 supplies metadata MD to the audio encoder 113, along with size information for embedding this metadata MD into the user data area. The audio encoder 113 then embeds the metadata MD into the user data area of ​​the audio stream.

[0079] The video stream generated by the video encoder 112 is supplied to the DASH / MP4 formatter 114. Similarly, the audio stream generated by the audio encoder 113, with metadata MD embedded in the user data area, is also supplied to the DASH / MP4 formatter 114. The DASH / MP4 formatter 114 then generates an MP4 file containing the media streams (media segments) such as video and audio content. The DASH / MP4 formatter 114 also generates an MPD file using the content metadata and segment URL information. At this stage, the MPD file includes identification information indicating that metadata has been inserted into the audio stream.

[0080] [Details on inserting metadata MD in each compression format] "In the case of AAC" First, let's explain the case where the compression format is AAC (Advanced Audio Coding). Figure 9 shows the structure of an AAC audio frame. This audio frame consists of multiple elements. At the beginning of each element, there is a 3-bit identifier (ID) "id_syn_ele", which makes the content of the element identifiable.

[0081] When "id_syn_ele" is "0x4", it indicates that it is a Data Stream Element (DSE), which is an element where user data can be placed. If the compression format is AAC, metadata MD is inserted into this DSE. Figure 10 shows the structure (syntax) of a Data Stream Element (DSE).

[0082] The 4-bit field "element_instance_tag" indicates the data type within the DSE, but if the DSE is used as unified user data, this value can be set to "0". "Data_byte_align_flag" is set to "1" to ensure the entire DSE is byte-aligned. "count", or "esc_count" which represents the number of additional bytes, is determined appropriately depending on the size of the user data. "metadata ()" is inserted into the "data_stream_byte" field.

[0083] Figure 11(a) shows the structure (syntax) of "metadata()", and Figure 11(b) shows the content (semantics) of the main information in that structure. The 32-bit field "userdata_identifier" is set to a value from a predefined array to indicate that it is audio user data. When "userdata_identifier" is "AAAA" indicating user data, the 8-bit field "metadata_type" exists. This field indicates the type of metadata. For example, "0x08" indicates access information for connecting to a specified network service and is included in ATSC's "SDO_payload()". When it is "0x08", "SDO_payload()" exists. Note that although "ATSC" is used here, it can also be used by other standardizing organizations.

[0084] Figure 12 shows the structure (syntax) of "SDO_payload()". When the command ID (cmdID) is less than "0x05", the "URI_character" field exists. A character code indicating the URI information for connecting to a specified network service is inserted into this field. Figure 13 shows the meaning of the command ID (cmdID) value. Note that "SDO_payload()" is standardized by ATSC (Advanced Television Systems Committee standards).

[0085] "In the case of AC3" Next, we will explain the case where the compression format is AC3. Figure 14 shows the structure of an AC3 frame (AC3 Synchronization Frame). The audio data SA is encoded so that the total size of the "mantissa data" of "Audblock 5", "AUX", and "CRC" does not exceed 3 / 8 of the total size. When the compression format is AC3, metadata MD is inserted into the "AUX" area. Figure 15 shows the structure (syntax) of the AC3 auxiliary data.

[0086] When "auxdatae" is "1", "aux data" is enabled, and data of the size indicated by 14 bits (in bits) of "auxdatal" is defined in "auxbits". The size of "auxbits" at that time is recorded in "nauxbits". In this technology, the field of "auxbits" is defined as "metadata()". In other words, the "metadata ()" shown in Figure 11(a) above is inserted into this "auxbits" field, and in its "data_byte" field, the ATSC "SDO_payload()" (see Figure 12), which has access information for connecting to a predetermined network service, is placed according to the syntax structure shown in Figure 11(a).

[0087] "In the case of AC4" Next, we will explain the case where the compression format is AC4. AC4 is considered one of the next-generation audio encoding formats after AC3. Figure 16(a) shows the layer structure of AC4's Simple Transport. There is a syncWord field, a frame Length field, a "RawAc4Frame" field as the encoded data field, and a CRC field. As shown in Figure 16(b), the "RawAc4Frame" field has a TOC (Table of Contents) field at the beginning, followed by a predetermined number of substream fields.

[0088] As shown in Figure 17(b), a metadata area exists within the substream (ac4_substream_data()), and within it is a field called "umd_payloads_substream()". In this "umd_payloads_substream()" field, the ATSC's "SDO_payload()" (see Figure 12), which contains access information for connecting to a specified network service, is placed.

[0089] As shown in Figure 17(a), the TOC (ac4_toc()) contains a field called "ac4_presentation_info()", which in turn contains a field called "umd_info()", and within that, metadata is inserted into the aforementioned "umd_payloads_substream())" field.

[0090] Figure 18 shows the structure (syntax) of "umd_info()". The "umd_version" field indicates the version number. The "substream_index" field indicates the index value. Certain combinations of version numbers and index values ​​are defined to indicate that metadata has been inserted into the "umd_payloads_substream()" field.

[0091] Figure 19 shows the structure (syntax) of "umd_payloads_substream()". The 5-bit field "umd_payload_id" is set to a value other than "0". The 32-bit field "umd_userdata_identifier" is set to a value from a predefined array to indicate that it is audio user data. The 16-bit field "umd_payload_size" indicates the number of bytes remaining. If "umd_userdata_identifier" is "AAAA" indicating user data, then the 8-bit field "umd_metadata_type" exists. This field indicates the type of metadata. For example, "0x08" indicates access information for connecting to a given network service and is included in ATSC's "SDO_payload()". When it is "0x08", "SDO_payload()" (see Figure 12) exists.

[0092] "In the case of MPEGH" Next, we will explain the case where the compression format is MPEGH (3D audio). Figure 20 shows the structure of an audio frame (1024 samples) in MPEGH (3D audio) transmission data. This audio frame consists of multiple MPEG audio stream packets. Each MPEG audio stream packet consists of a header and a payload.

[0093] The header contains information such as the packet type, packet label, and packet length. The payload contains the information defined by the packet type in the header. This payload information includes "SYNC," which corresponds to the synchronization start code, "Frame," which is the actual data of the 3D audio transmission, and "Config," which indicates the structure of this "Frame."

[0094] A "Frame" contains channel encoding data and object encoding data that make up the transmission data for 3D audio. Here, the channel encoding data consists of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), and LFE (Low Frequency Element). The object encoding data consists of encoded sample data for SCE (Single Channel Element) and metadata for mapping it to speakers located at arbitrary positions and rendering it. This metadata is included as an extension element (Ext_element).

[0095] Here, the correspondence between the configuration information (config) of each "Frame" contained in "Config" and each "Frame" is maintained as follows: That is, as shown in Figure 21, the configuration information (config) of each "Frame" is registered as an ID (elemIdx) in "Config", and each "Frame" is transmitted in the order in which the IDs were registered. Note that the packet label (PL) value is the same for "Config" and each corresponding "Frame".

[0096] Returning to Figure 20, in this embodiment, a new element (Ext_userdata) containing user data (userdata) is defined as an extension element (Ext_element). Accordingly, configuration information (userdataConfig) for that element (Ext_userdata) is newly defined in "Config".

[0097] Figure 22 shows the correspondence between the type (ExElementType) and the value (Value) of an extension element (Ext_element). Currently, values ​​from 0 to 7 are defined. Since values ​​from 128 onwards can be extended to include formats other than MPEG, for example, 128 can be newly defined as a value of type "ID_EXT_ELE_userdata".

[0098] Figure 23 shows the structure (syntax) of "userdataConfig()". The 32-bit field "userdata_identifier" is set to a value from a predefined array to indicate that it is audio user data. The 16-bit field "userdata_frameLength" indicates the number of bytes in "audio_userdata()". Figure 24 shows the structure (syntax) of "audio_userdata()". When "userdata_identifier" in "userdataConfig()" is "AAAA" indicating user data, the 8-bit field "metadataType" exists. This field indicates the type of metadata. For example, "0x08" indicates access information for connecting to a given network service and is included in ATSC's "SDO_payload()". When it is "0x08", "SDO_payload()" (see Figure 12) exists.

[0099] [Example of a set-top box configuration] Figure 25 shows an example configuration of the set-top box 200. This set-top box 200 includes a receiver 204, a DASH / MP4 analysis unit 205, a video decoder 206, an audio framing unit 207, an HDMI transmitter 208, and an HDMI terminal 209. The set-top box 200 also includes a CPU 211, a flash ROM 212, a DRAM 213, an internal bus 214, a remote control receiver 215, and a remote control transmitter 216.

[0100] The CPU 211 controls the operation of each part of the set-top box 200. The flash ROM 212 stores the control software and data. The DRAM 213 constitutes the work area of ​​the CPU 211. The CPU 211 loads the software and data read from the flash ROM 212 onto the DRAM 213, starts the software, and controls each part of the set-top box 200.

[0101] The remote control receiver 215 receives the remote control signal (remote control code) transmitted from the remote control transmitter 216 and supplies it to the CPU 211. The CPU 211 controls the various parts of the set-top box 200 based on this remote control code. The CPU 211, flash ROM 212, and DRAM 213 are connected to the internal bus 214.

[0102] The receiving unit 204 receives DASH / MP4, i.e., an MPD file as a metadata file, and an MP4 containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream contained in the MP4. In addition, identification information indicating that metadata has been inserted into the audio stream is inserted by a "Supplementary Descriptor" into the MPD file.

[0103] The DASH / MP4 analysis unit 205 analyzes the MPD file and MP4 received by the receiving unit 204. The DASH / MP4 analysis unit 205 extracts MPD information contained in the MPD file and sends it to the CPU 211. This MPD information includes identification information indicating that metadata has been inserted into the audio stream. The CPU 211 controls the acquisition process of the video and audio streams based on this MPD information. The DASH / MP4 analysis unit 205 also extracts metadata from the MP4, such as header information for each track, metadata descriptions of the content, and time information, and sends it to the CPU 211.

[0104] The DASH / MP4 analysis unit 205 extracts the video stream from the MP4 and sends it to the video decoder 206. The video decoder 206 performs a decoding process on the video stream to obtain uncompressed image data. The DASH / MP4 analysis unit 205 also extracts the audio stream from the MP4 and sends it to the audio framing unit 207. The audio framing unit 207 performs framing on the audio stream.

[0105] The HDMI transmitter 208 transmits uncompressed image data obtained by the video decoder 206 and the audio stream framed by the audio framing unit 207 from the HDMI terminal 209 using HDMI-compliant communication. The HDMI transmitter 208 packs the image data and audio stream for transmission via the HDMI TMDS channel and outputs them to the HDMI terminal 209.

[0106] The HDMI transmitter 208, under the control of the CPU 211, inserts identification information into the audio stream to indicate that metadata has been inserted. The HDMI transmitter 208 inserts the audio stream and identification information during the blanking period of the image data. Details of this HDMI transmitter 209 will be described later.

[0107] In this embodiment, the HDMI transmission unit 208 inserts identification information into an Audio InfoFrame packet placed during the blanking period of the image data. This Audio InfoFrame packet is placed in the data island section.

[0108] Figure 26 shows an example of the structure of an audio info frame packet. In HDMI, this audio info frame packet allows for the transmission of supplementary audio information from the source device to the sink device.

[0109] The 0th byte defines "Packet Type," which indicates the type of data packet; for audio infoframe packets, this is "0x84." The 1st byte describes the version information of the packet data definition. The 2nd byte describes information representing the packet length. In this embodiment, the 5th bit of the 5th byte defines a 1-bit flag information called "userdata_presence_flag." When the flag information is "1," it indicates that metadata has been inserted into the audio stream.

[0110] Furthermore, when the flag information is "1", various pieces of information are defined in the 9th byte. Bits 7 through 5 are designated as the "metadata_type" field, bit 4 as the "coordinated_control_flag" field, and bits 2 through 0 as the "frequency_type" field. A detailed explanation is omitted, but each of these fields represents the same information as the information attached to the MPD file shown in Figure 4.

[0111] The operation of the set-top box 200 will be briefly explained. The receiving unit 204 receives DASH / MP4 files, which are MPD files as metafiles and MP4 files containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path. The MPD files and MP4 files received in this way are supplied to the DASH / MP4 analysis unit 205.

[0112] The DASH / MP4 analysis unit 205 analyzes MPD files and MP4 files. The DASH / MP4 analysis unit 205 then extracts MPD information contained in the MPD files and sends it to the CPU 211. This MPD information also includes identification information indicating that metadata has been inserted into the audio stream. The DASH / MP4 analysis unit 205 also extracts metadata from the MP4 files, such as header information for each track, metadata descriptions of the content, and time information, and sends it to the CPU 211.

[0113] Furthermore, the DASH / MP4 analysis unit 205 extracts the video stream from the MP4 and sends it to the video decoder 206. The video decoder 206 decodes the video stream to obtain uncompressed image data. This image data is supplied to the HDMI transmission unit 208. The DASH / MP4 analysis unit 205 also extracts the audio stream from the MP4. This audio stream is framed by the audio framing unit 207 and then supplied to the HDMI transmission unit 208. The HDMI transmission unit 208 then packs the image data and audio stream and sends them from the HDMI terminal 209 to the HDMI cable 400.

[0114] In the HDMI transmitter 208, under the control of the CPU 211, identification information indicating that metadata has been inserted into the audio stream is inserted into the audio infoframe packets placed during the blanking period of the image data. This allows the set-top box 200 to transmit identification information indicating that metadata has been inserted into the audio stream to the HDMI television receiver 300.

[0115] [Example of a TV receiver configuration] Figure 27 shows an example configuration of a television receiver 300. This television receiver 300 includes a receiving unit 306, a DASH / MP4 analysis unit 307, a video decoder 308, an image processing circuit 309, a panel driving circuit 310, and a display panel 311.

[0116] The television receiver 300 also includes an audio decoder 312, an audio processing circuit 313, an audio amplification circuit 314, a speaker 315, an HDMI terminal 316, an HDMI receiver 317, and a communication interface 318. Furthermore, the television receiver 300 includes a CPU 321, a flash ROM 322, a DRAM 323, an internal bus 324, a remote control receiver 325, and a remote control transmitter 326.

[0117] The CPU 321 controls the operation of each part of the television receiver 300. The flash ROM 322 stores the control software and data. The DRAM 323 constitutes the work area of ​​the CPU 321. The CPU 321 loads the software and data read from the flash ROM 322 onto the DRAM 323, starts the software, and controls each part of the television receiver 300.

[0118] The remote control receiver 325 receives the remote control signal (remote control code) transmitted from the remote control transmitter 326 and supplies it to the CPU 321. The CPU 321 controls various parts of the television receiver 300 based on this remote control code. The CPU 321, flash ROM 322, and DRAM 323 are connected to the internal bus 324.

[0119] The communication interface 318 communicates with servers located on a network such as the Internet, under the control of the CPU 321. This communication interface 318 is connected to the internal bus 324.

[0120] The receiving unit 306 receives DASH / MP4, i.e., an MPD file as a metadata file, and an MP4 containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream contained in the MP4. In addition, identification information indicating that metadata has been inserted into the audio stream is inserted by a "Supplementary Descriptor" into the MPD file.

[0121] The DASH / MP4 analysis unit 307 analyzes the MPD file and MP4 received by the receiving unit 306. The DASH / MP4 analysis unit 307 extracts the MPD information contained in the MPD file and sends it to the CPU 321. The CPU 321 controls the acquisition process of video and audio streams based on this MPD information. The DASH / MP4 analysis unit 307 also extracts metadata from the MP4, such as header information for each track, metadata descriptions of the content, and time information, and sends it to the CPU 321.

[0122] The DASH / MP4 analysis unit 307 extracts the video stream from the MP4 and sends it to the video decoder 308. The video decoder 308 performs a decoding process on the video stream to obtain uncompressed image data. The DASH / MP4 analysis unit 307 also extracts the audio stream from the MP4 and sends it to the audio decoder 312.

[0123] The HDMI receiver 317 receives image data and audio streams supplied to the HDMI terminal 316 via the HDMI cable 400 using HDMI-compliant communication. The HDMI receiver 317 also extracts various control information inserted during the blanking period of the image data and transmits it to the CPU 321. This control information includes identification information (see Figure 26) that indicates metadata has been inserted into the audio stream, which is inserted into the audio info frame packet. Details of the HDMI receiver 317 will be described later.

[0124] The video processing circuit 309 performs scaling and compositing on image data obtained by the video decoder 308 or the HDMI receiver 316, as well as image data received from a server on the network via the communication interface 318, to obtain image data for display.

[0125] The panel driving circuit 310 drives the display panel 311 based on the display image data obtained by the video processing circuit 308. The display panel 311 is composed of, for example, an LCD (Liquid Crystal Display) or an organic EL display (organic electroluminescence display).

[0126] The audio decoder 312 performs decoding on the audio stream extracted by the DASH / MP4 analysis unit 307 or obtained by the HDMI receiver unit 317 to obtain uncompressed audio data. The audio decoder 312 also extracts metadata inserted into the audio stream under the control of the CPU 321 and sends it to the CPU 321. In this embodiment, the metadata is access information for connecting to a predetermined network service (see Figure 12). The CPU 321 causes various parts of the television receiver 300 to perform processing using the metadata as appropriate.

[0127] Furthermore, MPD information is supplied to the CPU 321 from the DASH / MP4 analysis unit 307. The CPU 321 can recognize in advance that metadata has been inserted into the audio stream based on the identification information contained in this MPD information, and can control the audio decoder 312 so that metadata is extracted.

[0128] The audio processing circuit 313 performs necessary processing, such as D / A conversion, on the audio data obtained by the audio decoder 312. The audio amplification circuit 314 amplifies the audio signal output from the audio processing circuit 313 and supplies it to the speaker 315.

[0129] The operation of the television receiver 300 shown in Figure 27 will be briefly explained. The receiving unit 306 receives DASH / MP4, that is, an MPD file as a metafile, and an MP4 containing media streams (media segments) such as video and audio, which are sent from the service transmission system 100 via an RF transmission line or a communication network transmission line. The MPD file and MP4 received in this way are supplied to the DASH / MP4 analysis unit 307.

[0130] The DASH / MP4 analysis unit 307 analyzes MPD files and MP4 files. The DASH / MP4 analysis unit 307 then extracts MPD information contained in the MPD files and sends it to the CPU 321. This MPD information also includes identification information indicating that metadata has been inserted into the audio stream. The DASH / MP4 analysis unit 307 also extracts metadata from the MP4 files, such as header information for each track, metadata descriptions of the content, and time information, and sends it to the CPU 321.

[0131] Furthermore, the DASH / MP4 analysis unit 307 extracts the video stream from the MP4 file and sends it to the video decoder 308. The video decoder 308 decodes the video stream to obtain uncompressed image data. This image data is supplied to the video processing circuit 309. The DASH / MP4 analysis unit 307 also extracts the audio stream from the MP4 file. This audio stream is supplied to the audio decoder 312.

[0132] The HDMI receiver 317 receives image data and audio streams supplied to the HDMI terminal 316 via the HDMI cable 400 using HDMI-compliant communication. The image data is supplied to the video processing circuit 309, and the audio stream is supplied to the audio decoder 312.

[0133] Furthermore, the HDMI receiver 317 extracts various control information inserted during the blanking period of the image data and sends it to the CPU 321. This control information includes identification information inserted into the audio info frame packet, indicating that metadata has been inserted into the audio stream. Therefore, the CPU 321 can control the operation of the audio decoder 312 based on this identification information and extract metadata from the audio stream.

[0134] In the video processing circuit 309, scaling and compositing are performed on image data obtained by the video decoder 308 or the HDMI receiver 317, as well as image data received from a server on the network via the communication interface 318, to obtain image data for display. When processing a television broadcast signal, the video processing circuit 309 handles the image data obtained by the video decoder 308. On the other hand, when the set-top box 200 is connected via an HDMI interface, the video processing circuit 309 handles the image data obtained by the HDMI receiver 317.

[0135] Image data for display obtained by the video processing circuit 309 is supplied to the panel driving circuit 310. The panel driving circuit 310 drives the display panel 311 based on the image data for display. As a result, the display panel 311 displays an image corresponding to the image data for display.

[0136] The audio decoder 312 performs decoding on the audio stream obtained by the DASH / MP4 analysis unit 307 or the HDMI receiver 316 to obtain uncompressed audio data. When receiving and processing a television broadcast signal, the audio decoder 312 handles the audio stream obtained by the DASH / MP4 analysis unit 307. On the other hand, when the set-top box 200 is connected via an HDMI interface, the audio decoder 312 handles the audio stream obtained by the HDMI receiver 317.

[0137] The audio data obtained by the audio decoder 312 is supplied to the audio processing circuit 313. The audio processing circuit 313 performs necessary processing on the audio data, such as D / A conversion. This audio data is amplified by the audio amplification circuit 314 and then supplied to the speaker 315. As a result, the speaker 315 outputs audio corresponding to the image displayed on the display panel 311.

[0138] Furthermore, the audio decoder 312 extracts metadata inserted into the audio stream. For example, as described above, this metadata extraction process is carried out efficiently and reliably by the CPU 321 recognizing, based on identification information, that metadata has been inserted into the audio stream and controlling the operation of the audio decoder 312.

[0139] The metadata extracted by the audio decoder 312 is sent to the CPU 321. Then, under the control of the CPU 321, processing using the metadata is performed as appropriate in various parts of the television receiver 300. For example, image data is acquired from a server on the network and multi-screen display is performed.

[0140] [Example configuration of HDMI transmitter and HDMI receiver] Figure 28 shows an example configuration of the HDMI transmitter (HDMI source) 208 of the set-top box 200 shown in Figure 25 and the HDMI receiver (HDMI sink) 317 of the television receiver 300 shown in Figure 27.

[0141] The HDMI transmitter 208 transmits differential signals corresponding to the pixel data of one uncompressed image on multiple channels to the HDMI receiver 317 in one direction during the active image section (hereinafter also referred to as the active video section as appropriate). Here, the active image section is the section from one vertical synchronization signal to the next, excluding the horizontal retrace section and the vertical retrace section. In addition, the HDMI transmitter 208 transmits differential signals corresponding to at least audio data, control data, and other auxiliary data associated with the image on multiple channels to the HDMI receiver 317 in one direction during the horizontal retrace section or the vertical retrace section.

[0142] The HDMI system, consisting of an HDMI transmitter 208 and an HDMI receiver 317, has the following transmission channels: Specifically, there are three TMDS channels #0 to #2, which are transmission channels for unidirectional serial transmission of pixel data and audio data from the HDMI transmitter 208 to the HDMI receiver 317, synchronized with the pixel clock. There is also a TMDS clock channel, which is a transmission channel for transmitting the pixel clock.

[0143] The HDMI transmission unit 208 has an HDMI transmitter 81. The transmitter 81 converts, for example, uncompressed image pixel data into corresponding differential signals and transmits them serially in one direction to the HDMI receiver unit 317 connected via the HDMI cable 400, using three TMDS channels, which are multiple channels: #0, #1, and #2.

[0144] Furthermore, the transmitter 81 converts audio data associated with the uncompressed image, as well as necessary control data and other auxiliary data, into corresponding differential signals and transmits them serially in one direction to the HDMI receiver 317 via three TMDS channels #0, #1, and #2.

[0145] Furthermore, the transmitter 81 transmits a pixel clock synchronized with the pixel data transmitted on the three TMDS channels #0, #1, and #2 to the HDMI receiver 317 connected via the HDMI cable 400 on the TMDS clock channel. Here, on one TMDS channel #i (i=0,1,2), 10 bits of pixel data are transmitted during one clock cycle of the pixel clock.

[0146] The HDMI receiver 317 receives differential signals corresponding to pixel data transmitted unidirectionally from the HDMI transmitter 208 on multiple channels during the active video section. Furthermore, the HDMI receiver 317 also receives differential signals corresponding to audio data and control data transmitted unidirectionally from the HDMI transmitter 208 on multiple channels during the horizontal or vertical retrace section.

[0147] In other words, the HDMI receiving unit 317 has an HDMI receiver 82. This HDMI receiver 82 receives differential signals corresponding to pixel data and differential signals corresponding to audio data and control data, which are transmitted unidirectionally from the HDMI transmitting unit 208 on TMDS channels #0, #1, and #2. In this case, it receives the signals in synchronization with the pixel clock transmitted from the HDMI transmitting unit 208 on the TMDS clock channel.

[0148] In addition to the TMDS channels #0 to #2 and the TMDS clock channel mentioned above, the HDMI system also has transmission channels called DDC (Display Data Channel) 83 and CEC line 84. DDC 83 consists of two signal lines (not shown) included in the HDMI cable 400. DDC 83 is used by the HDMI transmitter 208 to read E-EDID (Enhanced Extended Display Identification Data) from the HDMI receiver 317.

[0149] In addition to the HDMI receiver 81, the HDMI receiver 317 has an EDID ROM (Read Only Memory) 85 that stores E-EDID, which is performance information related to its own performance (configuration / capability). The HDMI transmitter 208 reads the E-EDID from the HDMI receiver 317, which is connected via the HDMI cable 400, via the DDC 83, for example, in response to a request from the CPU 211 (see Figure 20).

[0150] The HDMI transmitter 208 sends the read E-EDID to the CPU 211. The CPU 211 stores this E-EDID in the flash ROM 212 or DRAM 213.

[0151] The CEC line 84 consists of a single signal line (not shown) included in the HDMI cable 400 and is used for bidirectional communication of control data between the HDMI transmitter 208 and the HDMI receiver 317. This CEC line 84 constitutes the control data line.

[0152] The HDMI cable 400 also includes a line (HPD line) 86 connected to a pin called HPD (Hot Plug Detect). The source device can use this line 86 to detect the connection of the sink device. This HPD line 86 is also used as a HEAC- line, which constitutes a bidirectional communication path. The HDMI cable 400 also includes a power line 87 used to supply power from the source device to the sink device. Furthermore, the HDMI cable 400 includes a utility line 88. This utility line 88 is also used as a HEAC+ line, which constitutes a bidirectional communication path.

[0153] Figure 29 shows the various transmission data intervals when image data with dimensions of 1920 pixels x 1080 lines is transmitted in TMDS channels #0, #1, and #2. In the video field transmitted through the three HDMI TMDS channels #0, #1, and #2, there are three types of intervals depending on the type of data being transmitted: the video data period 17, the data island period 18, and the control period 19.

[0154] Here, the video field section is the section from the rising edge (Active Edge) of one vertical synchronization signal to the rising edge of the next vertical synchronization signal, and is divided into a horizontal blanking period 15, a vertical blanking period 16, and an active video section 14, which is the section obtained by subtracting the horizontal blanking period and the vertical blanking period from the video field section.

[0155] The video data section 17 is assigned to the active pixel section 14. In this video data section 17, data for 1920 pixels × 1080 lines of active pixels, which constitute one uncompressed image data frame, is transmitted. The data island section 18 and the control section 19 are assigned to the horizontal retrace period 15 and the vertical retrace period 16. In these data island section 18 and the control section 19, auxiliary data is transmitted.

[0156] Specifically, the data island section 18 is allocated to a portion of the horizontal retrace period 15 and the vertical retrace period 16. In this data island section 18, auxiliary data that is not related to control, such as voice data packets, is transmitted. The control section 19 is allocated to the remaining portion of the horizontal retrace period 15 and the vertical retrace period 16. In this control section 19, auxiliary data that is related to control, such as vertical synchronization signals and horizontal synchronization signals, control packets, etc., is transmitted.

[0157] Next, with reference to Figure 30, a specific example of processing using metadata in the television receiver 300 will be explained. The television receiver 300 acquires, for example, the initial server URL, network service identification information, target file name, session start / end commands, media recording / playback commands, etc., as metadata. As mentioned above, metadata is described as access information for connecting to a predetermined network service, but here we will assume that other necessary information is also included in the metadata.

[0158] The television receiver 300, which is a network client, accesses the primary server using the initial server URL. The television receiver 300 then obtains information from the primary server, such as the streaming server URL, the target file name, the mime type indicating the file type, and media playback time information.

[0159] The television receiver 300 then accesses the streaming server using the streaming server URL. The television receiver 300 then specifies the target file name. If receiving a multicast service, the program service is identified using network identification information and service identification information.

[0160] The television receiver 300 then starts or ends a session with the streaming server using session start / end commands. Furthermore, while the session with the streaming server is ongoing, the television receiver 300 acquires media data from the streaming server using media record / playback commands.

[0161] In the example shown in Figure 30, the primary server and the streaming server exist separately. However, these servers may be configured as a single unit.

[0162] Figure 31 shows an example of screen display transitions when the television receiver 300 accesses a network service based on metadata. Figure 31(a) shows the state where no image is displayed on the display panel 311. Figure 31(b) shows the state where broadcast reception has started and the main content related to this broadcast reception is displayed in full screen on the display panel 311.

[0163] Figure 31(c) shows a state where metadata-based service access is available and a session has been initiated between the television receiver 300 and the server. In this case, the main content related to broadcast reception changes from full-screen display to partial-screen display.

[0164] Figure 31(d) shows the state where media playback from the server is performed, and network service content 1 is displayed on the display panel 311 in parallel with the display of the main content. Figure 31(e) shows the state where media playback from the server is performed, and network service content 2 is displayed on the display panel 311 in parallel with the display of network service content 1, superimposed on the display of the main content.

[0165] Figure 31(f) shows the state after playback of service content from the internet has finished and the session between the television receiver 300 and the server has ended. In this case, the display panel 311 returns to a state where the main content related to broadcast reception is displayed in full screen.

[0166] The television receiver 300 shown in Figure 27 is equipped with a speaker 315, and as shown in Figure 32, the audio data obtained by the audio decoder 312 is supplied to the speaker 315 via the audio processing circuit 313 and the audio amplification circuit 314, and the sound is output from this speaker 315.

[0167] However, as shown in Figure 33, the television receiver 300 does not have a speaker, and a configuration is also conceivable in which the audio stream obtained by the DASH / MP4 analysis unit 307 or the HDMI receiver unit 317 is supplied to the external speaker system 350 from the interface unit 331. The interface unit 331 is, for example, a digital interface such as HDMI (High-Definition Multimedia Interface), SPDIF (Sony Philips Digital Interface), or MHL (Mobile High-definition Link).

[0168] In this case, the audio stream is decoded by the audio decoder 351a of the external speaker system 350, and sound is output from the external speaker system 350. Furthermore, even if the television receiver 300 is equipped with a speaker 315 (see Figure 32), a configuration in which the audio stream is supplied from the interface unit 331 to the external speaker system 350 (see Figure 33) is also conceivable.

[0169] As described above, in the transmission and reception systems 10 and 10' shown in Figures 3(a) and 3(b), the service transmission system 100 inserts identification information into the MPD file indicating that metadata has been inserted into the audio stream. Therefore, the receiving side (set-top box 200, television receiver 300) can easily recognize that metadata has been inserted into the audio stream.

[0170] Furthermore, in the transmission / reception system 10 shown in Figure 3(a), the set-top box 200 transmits the audio stream with the inserted metadata, along with identification information indicating that metadata has been inserted into the audio stream, to the television receiver 300 via HDMI. Therefore, the television receiver 300 can easily recognize that metadata has been inserted into the audio stream, and based on this recognition, it can extract the metadata inserted into the audio stream, thereby efficiently and reliably acquiring and utilizing the metadata.

[0171] Furthermore, in the transmission / reception system 10' shown in Figure 3(b), the television receiver 300 extracts metadata from the audio stream based on the identification information inserted into the MPD file and uses it for processing. Therefore, metadata inserted into the audio stream can be acquired efficiently and reliably, and processing using the metadata can be performed appropriately.

[0172] <2. Variant> In the above-described embodiment, an example was shown where the transmission / reception system 10,10' handles DASH / MP4, but an example handling MPEG2-TS can be considered in the same way.

[0173] [Configuration of the transmission / reception system] Figure 34 shows an example configuration of a transmission / reception system that handles MPEG2-TS. The transmission / reception system 10A in Figure 34(a) includes a service transmission system 100A, a set-top box (STB) 200A, and a television receiver (TV) 300A. The set-top box 200A and the television receiver 300A are connected via an HDMI (High Definition Multimedia Interface) cable 400. The transmission / reception system 10A' in Figure 3(b) also includes a service transmission system 100A and a television receiver (TV) 300A.

[0174] The service transmission system 100A transmits the MPEG2-TS transport stream TS via an RF transmission path or a communication network transmission path. The service transmission system 100A inserts metadata into the audio stream. This metadata could include, for example, access information for connecting to a predetermined network service, or predetermined content information. Here, as in the embodiment described above, access information for connecting to a predetermined network service is inserted.

[0175] The service transmission system 100A inserts identification information into the container layer indicating that metadata has been inserted into the audio stream. The service transmission system 100A inserts this identification information, for example, as a descriptor into the audio elementary stream loop under the Program Map Table (PMT).

[0176] The set-top box 200A receives a transport stream TS from the service transmission system 100A via an RF transmission path or a communication network transmission path. This transport stream TS includes a video stream and an audio stream, with metadata inserted into the audio stream.

[0177] The set-top box 200A transmits the audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to the television receiver 300A via the HDMI cable 400.

[0178] Here, the set-top box 200A transmits the audio stream and identification information to the television receiver 300A by inserting the audio stream and identification information into the blanking period of the image data obtained by decoding the video stream and transmitting this image data to the television receiver 300A. The set-top box 200A inserts this identification information into, for example, an Audio InfoFrame packet (see Figure 26).

[0179] In the transmission / reception system 10A shown in Figure 34(a), the television receiver 300A receives an audio stream from the set-top box 200A via the HDMI cable 400, along with identification information indicating that metadata has been inserted into the audio stream. In other words, the television receiver 300A receives the audio stream and image data from the set-top box 200A, in which the identification information has been inserted during the blanking period.

[0180] The television receiver 300A then decodes the audio stream based on the identification information, extracts metadata, and processes this metadata. In this case, the television receiver 300A accesses a predetermined server on the network based on predetermined network service information as metadata.

[0181] Furthermore, in the transmission / reception system 10A' shown in Figure 34(b), the television receiver 300A receives a transport stream TS sent from the service transmission system 100A via an RF transmission path or a communication network transmission path. Access information for connecting to a predetermined network service is inserted as metadata into the audio stream contained in this transport stream TS. In addition, identification information indicating that metadata has been inserted into the audio stream is inserted into the container layer.

[0182] The television receiver 300A then decodes the audio stream based on the identification information, extracts metadata, and processes this metadata. In this case, the television receiver 300A accesses a predetermined server on the network based on predetermined network service information as metadata.

[0183] [TS generation unit of the service transmission system] Figure 35 shows an example configuration of the TS generation unit 110A included in the service transmission system 100A. In Figure 35, parts corresponding to those in Figure 8 are denoted by the same reference numerals. The TS generation unit 110A includes a control unit 111, a video encoder 112, an audio encoder 113, and a TS formatter 114A.

[0184] The control unit 111 is equipped with a CPU 111a and controls each part of the TS generation unit 110A. The video encoder 112 encodes the image data SV using MPEG2, H.264 / AVC, H.265 / HEVC, etc., to generate a video stream (video elementary stream). The image data SV is, for example, image data played back from a recording medium such as an HDD, or live image data obtained from a video camera.

[0185] The audio encoder 113 encodes the audio data SA using a compression format such as AAC, AC3, AC4, or MPEGH (3D audio) to generate an audio stream (audio elementary stream). The audio data SA is the audio data corresponding to the image data SV mentioned above, and can be audio data played back from a recording medium such as an HDD, or live audio data obtained from a microphone.

[0186] The audio encoder 113 has an audio encoding block section 113a and an audio framing section 113b. The audio encoding block section 113a generates encoding blocks, and the audio framing section 113b performs framing. In this case, the encoding blocks and framing will differ depending on the compression format.

[0187] The audio encoder 113 inserts metadata MD into the audio stream under the control of the control unit 111. This metadata MD could include, for example, access information for connecting to a predetermined network service or predetermined content information. Here, as in the embodiment described above, access information for connecting to a predetermined network service is inserted.

[0188] This metadata MD is inserted into the user data area of ​​the audio stream. Although a detailed explanation is omitted, the insertion of metadata MD in each compression format is performed in the same way as in the DASH / MP4 generation unit 110 in the above embodiment, and "SDO_payload()" is inserted as the metadata MD (see Figures 8-24).

[0189] The TS formatter 114A takes the video stream output from the video encoder 112 and the audio stream output from the audio encoder 113, converts them into PES packets, then converts them into transport packets and multiplexes them to obtain a transport stream TS as a multiplexed stream.

[0190] Furthermore, the TS formatter 114A inserts identification information under the program map table (PMT) to indicate that metadata MD has been inserted into the audio stream. This identification information is inserted using the audio user data descriptor (audio_userdata_descriptor). Details of this descriptor will be described later.

[0191] The operation of the TS generation unit 110A shown in Figure 35 will be briefly explained. The image data SV is supplied to the video encoder 112. The video encoder 112 encodes the image data SV using H.264 / AVC, H.265 / HEVC, etc., and generates a video stream containing encoded video data.

[0192] The audio data SA is also supplied to the audio encoder 113. The audio encoder 113 encodes the audio data SA using AAC, AC3, AC4, MPEGH (3D audio), etc., to generate an audio stream.

[0193] At this time, the control unit 111 supplies metadata MD to the audio encoder 113, along with size information for embedding this metadata MD into the user data area. The audio encoder 113 then embeds the metadata MD into the user data area of ​​the audio stream.

[0194] The video stream generated by the video encoder 112 is supplied to the TS formatter 114A. Similarly, the audio stream generated by the audio encoder 113, with metadata MD embedded in the user data area, is also supplied to the TS formatter 114A.

[0195] In this TS formatter 114A, the streams supplied from each encoder are packetized and multiplexed to obtain a transport stream TS as transmission data. In addition, in this TS formatter 114A, identification information indicating that metadata MD has been inserted into the audio stream is inserted under the program map table (PMT).

[0196] [Details of Audio User Data Descriptor] Figure 36 shows an example of the structure (syntax) of an audio user data descriptor (audio_userdata_descriptor). Figure 37 shows the content (semantics) of the main information in that example structure.

[0197] The 8-bit field "descriptor_tag" indicates the descriptor type. In this case, it indicates an audio user data descriptor. The 8-bit field "descriptor_length" indicates the descriptor length (size), showing the number of bytes remaining in the descriptor.

[0198] The 8-bit field of "audio_codec_type" indicates the audio encoding scheme (compression format). For example, "1" indicates "MPEGH", "2" indicates "AAC", "3" indicates "AC3", and "4" indicates "AC4". Adding this information makes it easy for the receiving end to understand the encoding scheme of the audio data in the audio stream.

[0199] The 3-bit field "metadata_type" indicates the type of metadata. For example, "1" indicates that the "userdata()" field will contain ATSC's "SDO_payload()", which has access information for connecting to a specified network service. Adding this information allows the receiving end to easily understand the type of metadata, i.e., what kind of metadata it is, and to decide, for example, whether or not to retrieve it.

[0200] The single-bit flag information in "coordinated_control_flag" indicates whether the metadata is inserted only into the audio stream. For example, "1" indicates that it is also inserted into the streams of other components, while "0" indicates that it is inserted only into the audio stream. Adding this information makes it easy for the receiving end to determine whether the metadata is inserted only into the audio stream.

[0201] The 3-bit field "frequency_type" indicates the type of metadata insertion frequency for the audio stream. For example, "1" indicates that one user data (metadata) is inserted into each audio access unit. "2" indicates that multiple user data (metadata) are inserted into a single audio access unit. Furthermore, "3" indicates that for each group containing random access points, at least one user data (metadata) is inserted into the first audio access unit. Adding this information makes it easy for the receiving end to understand the metadata insertion frequency for the audio stream.

[0202] [Transport Stream TS Configuration] Figure 38 shows an example of a transport stream (TS) configuration. In this configuration, there are PES packets for the video stream identified by PID1 ("video PES") and PES packets for the audio stream identified by PID2 ("audio PES"). A PES packet consists of a PES header (PES_header) and a PES payload (PES_payload). The PES header contains the DTS and PTS timestamps. The PES payload of the audio stream PES packet contains a user data area that includes metadata.

[0203] Furthermore, the transport stream TS contains a Program Map Table (PMT) as Program Specific Information (PSI). PSI is information that indicates which program each elementary stream included in the transport stream belongs to. The PMT contains a program loop that describes information related to the entire program.

[0204] Furthermore, the PMT has elementary stream loops that contain information related to each elementary stream. In this example configuration, there is a video elementary stream loop (video ES loop) corresponding to the video stream, as well as an audio elementary stream loop (audio ES loop) corresponding to the audio stream.

[0205] The video elementary stream loop (video ES loop) contains information such as the stream type and PID (packet identifier) ​​corresponding to the video stream, as well as a descriptor that describes information related to that video stream. The value of "Stream_type" for this video stream is set to "0x24", and the PID information is said to indicate PID1, which is assigned to the PES packet "video PES" of the video stream, as described above. One of the descriptors is the HEVC descriptor.

[0206] Furthermore, the audio elementary stream loop (audio ES loop) contains information such as the stream type and PID (packet identifier) ​​corresponding to the audio stream, as well as a descriptor that describes information related to that audio stream. The value of "Stream_type" for this audio stream is set to "0x11", and the PID information is said to indicate PID2, which is assigned to the audio stream's PES packet "audio PES" as described above. One of the descriptors is the audio user data descriptor (audio_userdata_descriptor) mentioned above.

[0207] [Example of a set-top box configuration] Figure 39 shows an example configuration of the set-top box 200A. In Figure 39, parts corresponding to those in Figure 25 are denoted by the same reference numerals. The receiving unit 204A receives the transport stream TS sent from the service transmission system 100A via an RF transmission line or a communication network transmission line.

[0208] The TS analysis unit 205A extracts video stream packets from the transport stream TS and sends them to the video decoder 206. The video decoder 206 reconstructs the video stream from the video packets extracted by the demultiplexer 205 and performs decoding to obtain uncompressed image data. The TS analysis unit 205A also extracts audio stream packets from the transport stream TS and reconstructs the audio stream. The audio framing unit 207 performs framing on the audio stream thus reconstructed.

[0209] In addition, it is possible to decode the audio stream transferred from the TS analysis unit 205A to the audio framing unit 207 using an audio decoder (not shown) and output the audio in parallel.

[0210] Furthermore, the TS analysis unit 205A extracts various descriptors from the transport stream TS and transmits them to the CPU 211. Here, the descriptors include the audio user data descriptor (see Figure 36), which is identification information indicating that metadata has been inserted into the audio stream.

[0211] Although a detailed explanation is omitted, the set-top box 200A shown in Figure 39 is otherwise configured and operates in the same way as the set-top box 200 shown in Figure 25.

[0212] [Example of a TV receiver configuration] Figure 40 shows an example configuration of the television receiver 300A. In Figure 40, parts corresponding to those in Figure 27 are denoted by the same reference numerals. The receiving unit 306A receives the transport stream TS sent from the service transmission system 100A via an RF transmission line or a communication network transmission line.

[0213] The TS analysis unit 307A extracts video stream packets from the transport stream TS and sends them to the video decoder 308. The video decoder 308 reconstructs the video stream from the video packets extracted by the demultiplexer 205 and performs decoding to obtain uncompressed image data. The TS analysis unit 307A also extracts audio stream packets from the transport stream TS and reconstructs the audio stream.

[0214] Furthermore, the TS analysis unit 307A extracts audio stream packets from the transport stream TS and reconstructs the audio stream. The TS analysis unit 307A also extracts various descriptors from the transport stream TS and transmits them to the CPU 321. Here, these descriptors include the audio user data descriptor (see Figure 36), which serves as identification information indicating that metadata has been inserted into the audio stream.

[0215] Although a detailed explanation is omitted, the television receiver 300A shown in Figure 40 is otherwise configured and operates in the same way as the television receiver 300 shown in Figure 27.

[0216] As described above, in the image display systems 10A and 10A' shown in Figures 34(a) and (b), the service transmission system 100A inserts metadata into the audio stream and also inserts identification information indicating that metadata has been inserted into the audio stream into the container layer. Therefore, the receiving side (set-top box 200A, television receiver 300A) can easily recognize that metadata has been inserted into the audio stream.

[0217] Furthermore, in the image display system 10A shown in Figure 34(a), the set-top box 200A transmits the audio stream with the inserted metadata, along with identification information indicating that metadata has been inserted into the audio stream, to the television receiver 300A via HDMI. Therefore, the television receiver 300A can easily recognize that metadata has been inserted into the audio stream, and based on this recognition, it can extract the metadata inserted into the audio stream, thereby efficiently and reliably acquiring and utilizing the metadata.

[0218] Furthermore, in the image display system 10A' shown in Figure 34(b), the television receiver 300A extracts metadata from the audio stream based on identification information received along with the audio stream and uses it for processing. Therefore, metadata inserted into the audio stream can be acquired efficiently and reliably, and processing using the metadata can be performed appropriately.

[0219] Furthermore, in the above-described embodiment, the set-top box 200 is configured to transmit image data and audio streams to the television receiver 300. However, a configuration in which the transmission is performed to a monitor device or projector instead of the television receiver 300 is also conceivable. Alternatively, a configuration in which a recorder with receiving capabilities, a personal computer, or the like is used instead of the set-top box 200 is also conceivable.

[0220] Furthermore, in the above-described embodiment, the set-top box 200 and the television receiver 300 are connected by an HDMI cable 400. However, it goes without saying that this invention can be similarly applied when these are connected by a wired connection using a digital interface similar to HDMI, or even when they are connected wirelessly.

[0221] Furthermore, this technology can also be configured as follows: (1) A transmission unit that transmits a metadata file containing metadata for receiving an audio stream into which metadata has been inserted, The system includes an information insertion unit that inserts identification information into the metafile indicating that the above metadata has been inserted into the above audio stream. Transmitter. (2) The metadata mentioned above is access information for connecting to a specified network service. The transmitting device described in (1) above. (3) The metadata above is a character code that indicates URI information. The transmitting device described in (2) above. (4) The above metafile is an MPD file. A transmitting device as described in any of (1) to (3) above. (5) The above information insertion unit is, Use "Supplementary Descriptor" to insert the above identification information into the above metafile. The transmitting device described in (4) above. (6) The above-mentioned transmission unit is: The above metadata is transmitted via an RF transmission path or a communication network transmission path. A transmitting device as described in any of (1) to (5) above. (7) The above-mentioned transmission unit is: Further send a container in a predetermined format containing the audio stream with the above metadata inserted. A transmitting device as described in any of (1) to (6) above. (8) The above container is MP4. The transmitting device described in (7) above. (9) A transmission step in which the transmitting unit transmits a metadata file containing metadata for the receiving device to acquire an audio stream into which metadata has been inserted, The process includes an information insertion step of inserting identification information into the metafile to indicate that the above metadata has been inserted into the above audio stream. Sending method. (10) A receiving unit that receives a metadata file having metadata for obtaining an audio stream into which metadata has been inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. The system further includes a transmitting unit that transmits the above audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to an external device via a predetermined transmission path. Receiving device. (11) The metadata above is access information for connecting to a specified network service. The receiving device described in (10) above. (12) The above metafile is an MPD file, The above metadata file contains the above identification information inserted by the "Supplementary Descriptor". The receiving device described in (10) or (11) above. (13) The above-mentioned transmission unit is The above audio stream and identification information are inserted during the blanking period of the image data, and the image data is transmitted to the above external device, thereby transmitting the above audio stream and identification information to the above external device. A receiving device as described in any of (10) to (12) above. (14) The specified transmission path is an HDMI cable. A receiving device as described in any of (10) to (13) above. (15) The receiving unit has a receiving step of receiving a metafile having metadata for obtaining an audio stream into which metadata has been inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. The method further includes a transmission step of transmitting the above audio stream, along with identification information indicating that metadata has been inserted into the audio stream, to an external device via a predetermined transmission path. Reception method. (16) A receiving unit that receives a metadata file having metadata for obtaining an audio stream into which metadata has been inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. A metadata extraction unit decodes the audio stream based on the above identification information and extracts the above metadata, The system further comprises a processing unit that performs processing using the above metadata. Receiving device. (17) The above metafile is an MPD file, The above metadata file contains the above identification information inserted by the "Supplementary Descriptor". The receiving device described in (16) above. (18) The above metadata is access information for connecting to a specified network service, The above-mentioned processing unit is, Based on the above network access information, access a designated server on the network. The receiving device described in (16) or (17) above. (19) The receiving unit has a receiving step of receiving a metafile having metadata for obtaining an audio stream into which metadata has been inserted, The above metadata file contains identification information indicating that the above metadata has been inserted into the above audio stream. A metadata extraction step is performed to decode the audio stream based on the above identification information and extract the above metadata. The process further includes a processing step that performs processing using the above metadata. Reception method. (20) A stream generation unit that generates an audio stream into which metadata including network access information is inserted, The system includes a transmitting unit that transmits a container in a predetermined format having the above-mentioned audio stream. Transmitter.

[0222] The main feature of this technology is that when metadata is inserted into the audio stream during DASH / MP4 distribution, identification information indicating that metadata has been inserted into the audio stream is inserted into the MPD file, making it easy for the receiving end to recognize that metadata has been inserted into the audio stream (see Figures 3 and 4). [Explanation of Symbols]

[0223] 10,10',10A,10A'...Transmit / receive system 14. Effective pixel section 15. Horizontal retrace period 16. Vertical retrace period 17. Video data section 18. Data Island Section 19. Control Section 30A, 30B...MPEG-DASH based streaming system 31... DASH Stream File Server 32... DASH MPD Server 33, 33-1 to 33-N... Receiving System 34... CDN 35, 35-1 to 35-M... Receiving System 36... Broadcasting Transmission System 81... HDMI Transmitter 82... HDMI Receiver 83... DDC 84... CEC Line 85... EDID ROM 100, 100A... Service Transmission System 110... DASH / MP4 Generation Unit 110A... TS Generation Unit 111... Control Unit 111a... CPU 112... Video Encoder 113... Audio Encoder 113a... Audio Encoding Block Unit 113b... Audio Framing Unit 114... DASH / MP4 Formatter 114A... TS Formatter 200, 200A... Set-Top Box (STB) 204, 204A... Receiver 205... DASH / MP4 Analysis Unit 205A... TS Analysis Unit 206... Video Decoder 207... Audio Framing Unit 208... HDMI Transmitter 209... HDMI Terminal 211... CPU211 212... Flash ROM 213... DRAM 214... Internal Bus 215... Remote Control Receiver 216... Remote Control Transmitter 300, 300A... TV receiver 306, 306A... Receiver 307...DASH / MP4 analysis section 307A...TS analysis department 308...Video Decoder 309...Video processing circuit 310... Panel drive circuit 311... Display Panel 312...Audio Decoder 313...Audio processing circuit 314...Audio Amplification Circuit 315...Speaker 316...HDMI terminal 317···HDMI receiver 318...Communication Interface 321···CPU 322...Flash ROM 323···DRAM 324...Internal bus 325... Remote control receiver 326... Remote control transmitter 350...External speaker system 400...HDMI cable< / baseurl> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / baseurl>

Claims

1. A receiving unit that receives an audio stream, A decoding unit performs a decoding process on the aforementioned audio stream to obtain audio data, The system includes an audio processing unit that processes the aforementioned audio data and outputs it to an output unit, The codec of the aforementioned audio stream is MPEG-H 3D audio. The aforementioned audio stream consists of packets composed of a header and a payload. The header of the aforementioned packet contains information about the packet type. The payload of the packet includes "SYNC", which corresponds to the synchronization start code as defined by the packet type information contained in the packet header, and "Frame", which is the actual data of the 3D audio transmission data, or "Config", which indicates the configuration of "Frame". The aforementioned "Frame" includes object encoding data that constitutes the transmission data of 3D audio, The object encoding data consists of SCE (Single Channel Element) encoding sample data and metadata for mapping the encoding sample data to a speaker located at an arbitrary position and rendering it. The aforementioned metadata is included as an extension element (Ext_element), Information processing device.

2. When the value of the extension element (Ext_element) is 0, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_FILL. The information processing apparatus according to claim 1.

3. When the value of the extension element (Ext_element) is 1, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_MPEGS. The information processing apparatus according to claim 1.

4. When the value of the extension element (Ext_element) is 2, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_SAOC. The information processing apparatus according to claim 1.

5. When the value of the extension element (Ext_element) is 3, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_AUDIOPREROLL. The information processing apparatus according to claim 1.

6. When the value of the extension element (Ext_element) is 4, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_UNI_DRC. The information processing apparatus according to claim 1.

7. When the value of the extension element (Ext_element) is 5, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_OBJ_METADATA. The information processing apparatus according to claim 1.

8. When the value of the extension element (Ext_element) is 6, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_SAOC_3D. The information processing apparatus according to claim 1.

9. When the value of the extension element (Ext_element) is 7, the corresponding type of the extension element (ExElementType) is ID_EXT_ELE_HOA. The information processing apparatus according to claim 1.

10. If the value of the aforementioned extension element (Ext_element) is 128 or greater, it is possible to extend to formats other than MPEG. The information processing apparatus according to claim 1.

11. The aforementioned packet is an MPEG Audio Stream Packet. The information processing apparatus according to claim 1.

12. The packet header includes, in addition to the packet type information, the packet label and packet length information. The information processing apparatus according to claim 1.

13. The aforementioned audio stream is included in the MPEG2-TS transport stream. The information processing apparatus according to claim 1.

14. The audio stream further includes identification information indicating that the audio stream contains the metadata. The information processing apparatus according to claim 1.

15. The procedure for receiving an audio stream, A procedure for obtaining audio data by performing a decoding process on the aforementioned audio stream, The procedure includes processing the aforementioned audio data and outputting it to the output unit, The codec of the aforementioned audio stream is MPEG-H 3D audio. The aforementioned audio stream consists of packets composed of a header and a payload. The header of the aforementioned packet contains information about the packet type. The payload of the packet includes "SYNC", which corresponds to the synchronization start code as defined by the packet type information contained in the packet header, and "Frame", which is the actual data of the 3D audio transmission data, or "Config", which indicates the configuration of "Frame". The aforementioned "Frame" includes object encoding data that constitutes the transmission data of 3D audio, The object encoding data consists of SCE (Single Channel Element) encoding sample data and metadata for mapping the encoding sample data to a speaker located at an arbitrary position and rendering it. The aforementioned metadata is included as an extension element (Ext_element), Information processing methods.

16. The procedure for receiving an audio stream, A procedure for obtaining audio data by performing a decoding process on the aforementioned audio stream, The procedure includes processing the aforementioned audio data and outputting it to the output unit, The codec of the aforementioned audio stream is MPEG-H 3D audio. The aforementioned audio stream consists of packets composed of a header and a payload. The header of the aforementioned packet contains information about the packet type. The payload of the packet includes "SYNC", which corresponds to the synchronization start code as defined by the packet type information contained in the packet header, and "Frame", which is the actual data of the 3D audio transmission data, or "Config", which indicates the configuration of "Frame". The aforementioned "Frame" includes object encoding data that constitutes the transmission data of 3D audio, The object encoding data consists of SCE (Single Channel Element) encoding sample data and metadata for mapping the encoding sample data to a speaker located at an arbitrary position and rendering it. The aforementioned metadata is included as an extension element (Ext_element), A program that causes a computer to execute an information processing method.

Citation Information

Patent Citations

  • Transmitter, transmission method, receiver, reception method and transmission / reception system

    JP2012010311A

  • System and method for adaptive audio signal generation, coding, and rendering

    JP2014522155A

  • Transmitting device, receiving device, and receiving method

    JP2024050685A

  • Playback device, playback method, and program

    WO2013065566A1

  • Systems, methods, apparatus, and computer-readable media for three-dimensional audio coding using basis function coefficients

    WO2014014757A1