Information processing device, information processing method, and program

JP7899406B2Active Publication Date: 2026-08-03SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2025-06-11
Publication Date
2026-08-03

Smart Images

  • Figure 0007899406000001
    Figure 0007899406000001
  • Figure 0007899406000002
    Figure 0007899406000002
  • Figure 0007899406000003
    Figure 0007899406000003
Patent Text Reader

Abstract

To reduce processing load on a receiving side when transmitting multiple types of audio data.SOLUTION: A meta file containing meta information for acquiring a predetermined number of audio streams, each containing encoded data from multiple groups, by a receiving device is transmitted. Attribute information indicating respective attributes of the encoded data of the multiple groups is inserted into this meta file. For example, stream correspondence information, which indicates which audio stream includes each of the encoded data of the multiple group, is further inserted into the meta file.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, as a stereophonic (3D) audio technology, a technology has been proposed in which encoded sample data is mapped to speakers existing at arbitrary positions based on metadata and then rendered (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is conceivable to transmit object-encoded data consisting of encoded sample data and metadata together with channel-encoded data such as 5.1-channel and 7.1-channel data, enabling acoustic reproduction with enhanced presence on the receiving side.

[0005] An object of this technology is to reduce the processing load on the receiving side when transmitting multiple types of encoded data.

Means for Solving the Problems

[0006] The concept of this technology is a transmission unit that transmits a meta file having meta information for acquiring, by a receiving device, a predetermined number of audio streams including a plurality of groups of encoded data, and an information insertion unit that inserts, into the meta file, attribute information indicating attributes of each of the plurality of groups of encoded data. This is in the transmission device.

[0007] In this technology, the transmitting unit transmits a metadata file containing metadata for a predetermined number of audio streams, each containing encoded data from multiple groups, to be acquired by a receiving device. For example, the encoded data from multiple groups may include either channel-encoded data or object-encoded data, or both.

[0008] The information insertion unit inserts attribute information into the metafile, indicating the attributes of each of the multiple groups of encoded data. For example, the metafile may be an MPD (Media Presentation Description) file. In this case, for example, the information insertion unit may use a “Supplementary Descriptor” to insert the attribute information into the metafile.

[0009] Furthermore, for example, the transmitting unit may transmit a metafile through an RF transmission path or a communication network transmission path. Alternatively, for example, the transmitting unit may further transmit a container of a predetermined format having a predetermined number of audio streams containing a plurality of groups of encoded data. For example, the container may be MP4. In this invention, MP4 refers to the ISO base media file format (ISOBMFF) (ISO / IEC 14496-12:2012).

[0010] In this technology, attribute information indicating the attributes of each of the encoded data groups is inserted into a metadata file containing metadata for acquiring a predetermined number of audio streams containing encoded data from multiple groups at the receiving device. Therefore, the receiving device can easily recognize the attributes of each of the encoded data groups before decoding the data, and can selectively decode and use only the encoded data from the necessary groups, thereby reducing the processing load.

[0011] Furthermore, in this technology, for example, the information insertion unit may further insert stream correspondence relationship information into the metafile, indicating which audio stream each of the encoded data from multiple groups is contained in. In this case, for example, the stream correspondence relationship information may be information indicating the correspondence between a group identifier that identifies each of the multiple groups of encoded data and an identifier that identifies each of a predetermined number of audio streams. In this case, the receiving side can easily recognize the audio stream containing the encoded data of the required group, thereby reducing the processing load.

[0012] Furthermore, other concepts of this technology include: The receiving unit includes a receiving unit that receives a metadata file containing metadata for acquiring a predetermined number of audio streams, including encoded data from multiple groups, at a receiving device. The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. The system further includes a processing unit that processes the predetermined number of audio streams based on the attribute information described above. It is located in the receiving device.

[0013] In this technology, a metafile is received by the receiving unit. This metafile contains metadata for the receiving device to acquire a predetermined number of audio streams, each containing encoded data from multiple groups. For example, the encoded data from multiple groups may include either channel-encoded data or object-encoded data, or both. Attribute information indicating the attributes of each of the encoded data from multiple groups is inserted into the metafile. The processing unit processes a predetermined number of audio streams based on this attribute information.

[0014] In this technology, a predetermined number of audio streams are processed based on attribute information indicating the attributes of each of the multiple groups of encoded data inserted into the metafile. Therefore, only the encoded data of the necessary groups can be selectively decoded and used, thereby reducing the processing load.

[0015] In this technology, for example, the metafile may also include stream correspondence relationship information indicating which audio stream each of several groups of encoded data is contained within. The processing unit may then process a predetermined number of audio streams based on this stream correspondence relationship information, in addition to the attribute information. In this case, the audio stream containing the encoded data of the required group can be easily identified, reducing the processing load.

[0016] Furthermore, in this technology, for example, the processing unit may selectively decode audio streams containing encoded data of groups having attributes that match speaker configuration and user selection information, based on attribute information and stream correspondence relationship information.

[0017] Furthermore, other concepts of this technology include: The receiving unit includes a receiving unit that receives a metadata file containing metadata for acquiring a predetermined number of audio streams, including encoded data from multiple groups, at a receiving device. The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. A processing unit that selectively obtains encoded data from a predetermined number of audio streams based on the attribute information, and reconstructs an audio stream containing the encoded data of the predetermined group, The system further comprises a stream transmission unit that transmits the reconstructed audio stream to an external device. It is located in the receiving device.

[0018] In this technology, a metafile is received by a receiving unit. This metafile has meta-information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups in a receiving device. Attribute information indicating each attribute of the encoded data of the plurality of groups is inserted into the metafile.

[0019] A processing unit selectively acquires encoded data of a predetermined group based on the attribute information from a predetermined number of audio streams, and an audio stream including the encoded data of the predetermined group is reconstructed. Then, the reconstructed audio stream is transmitted to an external device by a stream transmission unit.

[0020] Thus, in this technology, based on the attribute information indicating each attribute of the encoded data of the plurality of groups inserted into the metafile, the encoded data of a predetermined group is selectively acquired from a predetermined number of audio streams, and the audio stream to be transmitted to the external device is reconstructed. It is possible to easily acquire the encoded data of the necessary group, and the processing load can be reduced.

[0021] In this technology, for example, stream correspondence relationship information indicating which audio stream each of the encoded data of the plurality of groups is included in is further inserted into the metafile, and the processing unit may be configured to selectively acquire the encoded data of a predetermined group from a predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information. In this case, it is possible to easily recognize the audio stream including the encoded data of the predetermined group, and the processing load can be reduced.

Advantages of the Invention

[0022] According to this technology, it is possible to reduce the processing load on the receiving side when transmitting a plurality of types of encoded data. Note that the effects described in this specification are merely examples and are not limited, and there may be additional effects.

Brief Description of the Drawings

[0023] [Figure 1] It is a block diagram showing a configuration example of a stream distribution system based on MPEG - DASH. [Figure 2] It is a diagram showing an example of the relationship of each structure hierarchically arranged in the MPD file. [Figure 3] It is a block diagram showing a configuration example of a transmission - reception system as an embodiment. [Figure 4] It is a diagram showing the structure of an audio frame (1024 samples) in the transmission data of 3D audio. [Figure 5] It is a diagram showing a configuration example of the transmission data of 3D audio. [Figure 6] It is a diagram schematically showing a configuration example of an audio frame in the case of transmitting the transmission data of 3D audio with 1 track (1 audio stream) and in the case of transmitting with multiple tracks (multiple audio streams). [Figure 7] In the configuration example of the transmission data of 3D audio, it is a diagram showing an example of group division in the case of transmitting with 4 tracks. [Figure 8] It is a diagram showing the correspondence relationship between groups and tracks, etc. in the group division example (4 - division). [Figure 9] In the configuration example of the transmission data of 3D audio, it is a diagram showing an example of group division in the case of transmitting with 2 tracks. [Figure 10] It is a diagram showing the correspondence relationship between groups and tracks, etc. in the group division example (2 - division). [Figure 11] It is a diagram showing an example of MPD file description. [Figure 12] It is a diagram showing another example of MPD file description. [Figure 13] It is a diagram showing an example of the definition of "schemeIdUri" by "SupplementaryDescriptor". [Figure 14] 「 <baseurl>This is a diagram to explain the actual media file at the location indicated by "[ ]". [Figure 15] This diagram illustrates the correspondence between track identifiers (track IDs) and level identifiers (level IDs) within the "moov" box. [Figure 16] This diagram shows examples of transmissions from each box in a broadcasting system. [Figure 17] This block diagram shows an example configuration of the DASH / MP4 generation unit included in the service transmission system. [Figure 18] This block shows an example configuration of a service receiver. [Figure 19] This flowchart shows an example of the CPU's audio decoding control process in a service receiver. [Figure 20] This block shows another example configuration for the service receiver. [Modes for carrying out the invention]

[0024] The following describes embodiments for carrying out the invention. The description will be given in the following order. 1. Embodiment 2. Variations

[0025] <1. Embodiment> [Overview of MPEG-DASH-based streaming systems] First, we will describe an overview of an MPEG-DASH-based streaming system to which this technology can be applied.

[0026] Figure 1(a) shows an example configuration of an MPEG-DASH-based stream distribution system 30A. In this example configuration, media streams and MPD files are transmitted through a communication network transmission path. This stream distribution system 30A consists of a DASH stream file server 31 and a DASH MPD server 32, with N service receivers 33-1, 33-2, ..., 33-N connected via a CDN (Content Delivery Network) 34.

[0027] The DASH stream file server 31 generates a DASH-compliant stream segment (hereinafter referred to as "DASH segment" as appropriate) based on the media data (video data, audio data, subtitle data, etc.) of the specified content, and sends the segment in response to an HTTP request from a service receiver. This DASH stream file server 31 may be a server dedicated to streaming, or it may also be used as a web server.

[0028] Furthermore, the DASH stream file server 31 responds to requests for predetermined stream segments sent from service receivers 33 (33-1, 33-2, ..., 33-N) via the CDN 34, and transmits those stream segments to the requesting receiver via the CDN 34. In this case, the service receiver 33 refers to the rate values ​​listed in the MPD (Media Presentation Description) file and selects the stream with the optimal rate according to the network environment where the client is located, and makes the request.

[0029] The DASH MPD server 32 is a server that generates MPD files for obtaining DASH segments generated by the DASH stream file server 31. It generates MPD files based on content metadata from the content management server (not shown) and the address (url) of the segment generated by the DASH stream file server 31. Note that the DASH stream file server 31 and the DASH MPD server 32 may be physically the same.

[0030] In the MPD format, each stream, such as video and audio, uses an element called "Representation" to describe its attributes. For example, an MPD file contains separate representations for each of several video data streams with different rates, each describing its rate. The service receiver 33 uses these rate values ​​as a reference to select the optimal stream according to the network environment in which the service receiver 33 is located, as described above.

[0031] Figure 1(b) shows an example configuration of an MPEG-DASH-based stream distribution system 30B. In this example configuration, media streams and MPD files are transmitted via an RF transmission path. This stream distribution system 30B consists of a broadcast transmission system 36 to which a DASH stream file server 31 and a DASH MPD server 32 are connected, and M service receivers 35-1, 35-2, ..., 35-M.

[0032] In this stream distribution system 30B, the broadcast transmission system 36 transmits the DASH-compliant stream segment (DASH segment) generated by the DASH stream file server 31 and the MPD file generated by the DASH MPD server 32 on the broadcast wave.

[0033] Figure 2 shows an example of the hierarchical relationships between the various structures in an MPD file. As shown in Figure 2(a), the Media Presentation as a whole MPD file contains multiple periods separated by time intervals. For example, the first period starts at 0 seconds, the next period starts at 100 seconds, and so on.

[0034] As shown in Figure 2(b), a period has multiple representations. These multiple representations include groups of representations related to stream attributes, such as media streams with the same content but different rates, which are grouped by an AdaptationSet.

[0035] As shown in Figure 2(c), the representation includes SegmentInfo. As shown in Figure 2(d), this SegmentInfo contains an Initialization Segment and multiple Media Segments, each containing information for a segment further divided by periods. The Media Segments contain information such as addresses (urls) for actually retrieving segment data such as video and audio.

[0036] Furthermore, it is possible to freely switch between streams among multiple representations grouped in an adaptation set. This allows the service receiver to select the optimal rate stream depending on the network environment, enabling uninterrupted delivery.

[0037] [Example of a transmission / reception system configuration] Figure 3 shows an example configuration of a transmission / reception system 10 as an embodiment. This transmission / reception system 10 consists of a service transmission system 100 and a service receiver 200. In this transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31 and DASH MPD server 32 of the stream distribution system 30A shown in Figure 1(a) above. In this transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31, DASH MPD server 32 and broadcast transmission system 36 of the stream distribution system 30B shown in Figure 1(b) above.

[0038] Furthermore, in this transmission / reception system 10, the service receiver 200 corresponds to the service receivers 33 (33-1, 33-2, ..., 33-N) of the stream distribution system 30A shown in Figure 1(a) above. Also, in this transmission / reception system 10, the service receiver 200 corresponds to the service receivers 35 (35-1, 35-2, ..., 35-M) of the stream distribution system 30B shown in Figure 1(b) above.

[0039] The service transmission system 100 transmits DASH / MP4, that is, MP4 files containing MPD files as metafiles and media streams (media segments) such as video and audio, via an RF transmission path (see Figure 1(b)) or a communication network transmission path (see Figure 1(a)).

[0040] Figure 4 shows the structure of an audio frame (1024 samples) in the 3D audio (MPEGH) transmission data handled in this embodiment. This audio frame consists of multiple MPEG audio stream packets. Each MPEG audio stream packet consists of a header and a payload.

[0041] The header contains information such as the packet type, packet label, and packet length. The payload contains the information defined by the packet type in the header. This payload information includes "SYNC" information, which corresponds to the synchronization start code, "Frame" information, which is the actual data of the 3D audio transmission data, and "Config" information, which indicates the structure of this "Frame" information.

[0042] The "Frame" information includes channel encoding data and object encoding data that constitute the transmission data for 3D audio. Here, the channel encoding data consists of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), and LFE (Low Frequency Element). The object encoding data consists of encoded sample data for SCE (Single Channel Element) and metadata for mapping it to speakers located at arbitrary positions and rendering it. This metadata is included as an extension element (Ext_element).

[0043] Figure 5 shows an example of the structure of 3D audio transmission data. In this example, it consists of one channel-coded data and two object-coded data. The one channel-coded data is 5.1 channel-coded data (CD) and consists of coded sample data for SCE1, CPE1.1, CPE1.2, and LFE1.

[0044] The two object encoding data sets are the encoding data for the Immersive Audio Object (IAO) and the Speech Dialog Object (SDO). The Immersive Audio Object encoding data is object encoding data for immersive sound and consists of encoded sample data SCE2 and metadata EXE_El (Object metadata)2 for mapping it to a speaker located at an arbitrary position and rendering it.

[0045] Speech dialog object encoding data is object encoding data for a speech language. In this example, there is speech dialog object encoding data corresponding to the first and second languages. The speech dialog object encoding data corresponding to the first language consists of encoded sample data SCE3 and metadata EXE_El (Object metadata)3 for mapping it to a speaker located at an arbitrary position and rendering it. The speech dialog object encoding data corresponding to the second language consists of encoded sample data SCE4 and metadata EXE_El (Object metadata)4 for mapping it to a speaker located at an arbitrary position and rendering it.

[0046] Encoded data is distinguished by the concept of "groups" based on its type. In the example shown, 5.1 channel encoded channel data is designated as Group 1, immersive audio object encoded data as Group 2, speech dialogue object encoded data for the first language as Group 3, and speech dialogue object encoded data for the second language as Group 4.

[0047] Furthermore, at the receiving end, selectable groups are registered in a switch group (SW Group) and encoded. In the illustrated example, groups 3 and 4 are registered in switch group 1 (SW Group 1). Groups can also be bundled together to form a preset group, enabling playback according to the use case. In the illustrated example, groups 1, 2, and 3 are bundled together to form preset group 1, and groups 1, 2, and 4 are bundled together to form preset group 2.

[0048] Returning to Figure 3, the service transmission system 100 transmits the 3D audio transmission data, which includes encoded data from multiple groups, either as a single audio stream on one track, or as multiple audio streams on multiple tracks, as described above.

[0049] Figure 6(a) schematically shows an example of the audio frame configuration when transmitting 3D audio transmission data in one track (one audio stream) as shown in Figure 5. In this case, Audio track 1 includes channel coding data (CD), immersive audio object coding data (IAO), and speech dialogue object coding data (SDO), along with "SYNC" and "Config" information.

[0050] Figure 6(b) schematically shows an example of the audio frame configuration when transmitting with multiple tracks (multiple audio streams), in this case three tracks, as in the example of 3D audio transmission data configuration in Figure 5. In this case, Audio track 1 contains channel coding data (CD) along with "SYNC" information and "Config" information. Audio track 2 contains immersive audio object coding data (IAO) along with "SYNC" information and "Config" information. Furthermore, Audio track 3 contains speech dialogue object coding data (SDO) along with "SYNC" information and "Config" information.

[0051] Figure 7 shows an example of group division when transmitting 3D audio transmission data using four tracks, as in the example configuration of 3D audio transmission data in Figure 5. In this case, audio track 1 contains channel-coded data (CD) distinguished as group 1. Audio track 2 contains immersive audio object-coded data (IAO) distinguished as group 2. Audio track 3 contains speech dialogue object-coded data (SDO) of the first language distinguished as group 3. Furthermore, audio track 4 contains speech dialogue object-coded data (SDO) of the second language distinguished as group 4.

[0052] Figure 8 shows the correspondence between groups and audio tracks in the group division example (4 divisions) of Figure 7. Here, the group ID is an identifier for identifying a group. The attribute indicates the attribute of the encoded data for each group. The switch group ID is an identifier for identifying a switching group. The preset group ID is an identifier for identifying a preset group. The track ID is an identifier for identifying an audio track.

[0053] The diagram shows that the encoded data belonging to group 1 is channel encoded data, does not constitute a switch group, and is included in audio track 1. The diagram also shows that the encoded data belonging to group 2 is object encoded data for immersive sound (immersive audio object encoded data), does not constitute a switch group, and is included in audio track 2.

[0054] Furthermore, the diagram shows that the encoded data belonging to group 3 is object encoded data for the speech language of the first language (speech dialogue object encoded data), constitutes switch group 1, and is included in audio track 3. Furthermore, the diagram shows that the encoded data belonging to group 4 is object encoded data for the speech language of the second language (speech dialogue object encoded data), constitutes switch group 1, and is included in audio track 4.

[0055] Furthermore, the diagram shows that preset group 1 includes group 1, group 2, and group 3. Additionally, the diagram shows that preset group 2 includes group 1, group 2, and group 4.

[0056] Figure 9 shows an example of group division when transmitting 3D audio transmission data using two tracks, as in the example configuration of 3D audio transmission data in Figure 5. In this case, audio track 1 contains channel-coded data (CD) distinguished as group 1 and immersive audio object-coded data (IAO) distinguished as group 2. Audio track 2 contains speech dialogue object-coded data (SDO) of the first language distinguished as group 3 and speech dialogue object-coded data (SDO) of the second language distinguished as group 4.

[0057] Figure 10 shows the correspondence between groups and substreams in the group division example (2 divisions) of Figure 9. The illustrated correspondence indicates that the encoded data belonging to Group 1 is channel encoded data, does not constitute a switch group, and is included in Audio Track 1. The illustrated correspondence also indicates that the encoded data belonging to Group 2 is object encoded data for immersive sound (immersive audio object encoded data), does not constitute a switch group, and is included in Audio Track 1.

[0058] Furthermore, the diagram shows that the encoded data belonging to group 3 is object encoded data for the speech language of the first language (speech dialogue object encoded data), constitutes switch group 1, and is included in audio track 2. Furthermore, the diagram shows that the encoded data belonging to group 4 is object encoded data for the speech language of the second language (speech dialogue object encoded data), constitutes switch group 1, and is included in audio track 2.

[0059] Furthermore, the diagram shows that preset group 1 includes group 1, group 2, and group 3. Additionally, the diagram shows that preset group 2 includes group 1, group 2, and group 4.

[0060] Returning to Figure 3, the service transmission system 100 inserts attribute information into the MPD file that indicates the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data. The service transmission system 100 also inserts stream correspondence information into the MPD file that indicates which audio track (audio stream) each of these multiple groups of encoded data is included in. In this embodiment, this stream correspondence information is, for example, information that indicates the correspondence between a group ID and a track ID.

[0061] The service transmission system 100 inserts this attribute information and stream correspondence information into the MPD file. In this embodiment, the "SupplementaryDescriptor" makes it possible to newly define "schemeIdUri" as a broadcast or other application, separate from the predefined definitions in the conventional standard. The service transmission system 100 uses the "SupplementaryDescriptor" to insert this attribute information and stream correspondence information into the MPD file.

[0062] Figure 11 shows an example of an MPD file description corresponding to the group division example (4 divisions) in Figure 7. Figure 12 shows an example of an MPD file description corresponding to the group division example (2 divisions) in Figure 9. Here, for the sake of simplicity, only information about the audio stream is described, but in reality, information about other media streams such as video streams is also described. Figure 13 shows an example of the definition of "schemeIdUri" by "SupplementaryDescriptor".

[0063] First, let's explain the example MPD file description in Figure 11. <adaptationset mimetype=""audio / mp4”" group=""1”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 1.

[0064] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)". As shown in Figure 13, "schemeIdUri="urn:brdcst:codecType"" indicates the type of codec. In this case, it is "mpegh".

[0065] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group1” / ">The description " indicates that the audio stream contains encoded data for group 1 "group1". As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:groupId"" indicates the group identifier.

[0066] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""channeldata” / ">The description "" indicates that the encoded data for group 1 "group1" is channel encoded data "channeldata". As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:attribute"" indicates the attribute of the encoded data for the group in question.

[0067] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">The description "group1" indicates that the encoded data for group 1 does not belong to any switch group. As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:switchGroupId"" indicates the identifier of the switch group to which the group belongs. For example, when "value" is "0", it indicates that the data does not belong to any switch group. When "value" is anything other than "0", it indicates that the data belongs to a switch group.

[0068] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description " indicates that the encoded data of group 1 "group1" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description "" indicates that the encoded data of group 1 "group1" belongs to preset group 2 "preset2". As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:presetGroupId"" indicates the identifier of the preset group to which the group belongs.

[0069] " <representation id=""1”" bandwidth=""128000”">The description indicates that within the adaptation set of group 1, there is an audio stream with a bitrate of 128kbps containing the encoded data of group 1 "group1" as a representation identified by "Representation id="1"". <baseurl> audio / jp1 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp1 / 128.mp4".

[0070] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level1” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 1, "level1". As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:levelId" indicates the level identifier corresponding to the track identifier that transmits the audio stream containing the encoded data for the group. The mapping between the track identifier (track ID) and the level identifier (level ID) is described in the "moov" box, for example, as will be explained later.

[0071] Also," <adaptationset mimetype=""audio / mp4”" group=""2”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 2.

[0072] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group2” / ">The description indicates that the audio stream contains encoded data for group 2, "group2".

[0073] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectSound” / ">The description indicates that the encoded data for group 2, "group2," is object-encoded data for immersive sound, "objectSound." <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">The description indicates that the encoded data for group 2 "group2" does not belong to any switch group.

[0074] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description indicates that the encoded data for group 2 "group2" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description indicates that the encoded data for group 2 "group2" belongs to preset group 2 "preset2".

[0075] " <representation id=""2”" bandwidth=""128000”">The description indicates that within the adaptation set of group 2, there is an audio stream with a bitrate of 128kbps containing encoded data for group 2 "group2" as a representation identified by "Representation id="2"". <baseurl> audio / jp2 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp2 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level2” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 2.

[0076] Also," <adaptationset mimetype=""audio / mp4”" group=""3”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 3.

[0077] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group3” / ">The description indicates that the audio stream contains encoded data for group 3, "group3". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang1” / ">The description indicates that the encoded data for group 3, "group3", is the object encoded data "objectLang1" for the speech language of the first language.

[0078] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">The description indicates that the encoded data for group 3 belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description indicates that the encoded data for group 3 "group3" belongs to preset group 1 "preset1".

[0079] " <representation id=""3”" bandwidth=""128000”">The description indicates that within the adaptation set of group 3, there is an audio stream with a bitrate of 128kbps containing encoded data for group 3 "group3" as a representation identified by "Representation id="3"". <baseurl> audio / jp3 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp3 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level3” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 3.

[0080] Also," <adaptationset mimetype=""audio / mp4”" group=""4”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 4.

[0081] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group4” / ">The description indicates that the audio stream contains encoded data for group 4, "group4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang2” / ">The description indicates that the encoded data for group 4, "group4", is object-encoded data "objectLang2" for the speech language of the second language.

[0082] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">The description indicates that the encoded data for group 4 "group4" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description indicates that the encoded data for group 4 "group4" belongs to preset group 2 "preset2".

[0083] " <representation id=""4”" bandwidth=""128000”">The description indicates that within the adaptation set of group 4, there is an audio stream with a bitrate of 128kbps containing encoded data for group 4 "group4" as a representation identified by "Representation id="4"". <baseurl> audio / jp4 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp4 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level4” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 4.

[0084] Next, we will explain the example of an MPD file description in Figure 12. <adaptationset mimetype=""audio / mp4”" group=""1”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)".

[0085] " <representation id=""1”" bandwidth=""128000”">The description indicates that within the adaptation set of Group 1, there is an audio stream with a bitrate of 128kbps, identified as "Representation id="1"". <baseurl> audio / jp1 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp1 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level1” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 1.

[0086] " <subrepresentation id=""11”" subgroupset=""1”">The description indicates that within the representation identified by "Representation id="1"", there exists a subrepresentation identified by "SubRepresentation id="11"", to which subgroup set 1 has been assigned.

[0087] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group1” / ">The description " indicates that the audio stream contains encoded data for group 1 "group1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""channeldata” / ">The description indicates that the encoded data for group 1 "group1" is channel encoded data "channeldata".

[0088] <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">The description "" indicates that the encoded data of group 1 "group1" does not belong to any switch group. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description " indicates that the encoded data of group 1 "group1" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description indicates that the encoded data for group 1 "group1" belongs to preset group 2 "preset2".

[0089] " <subrepresentation id=""12”" subgroupset=""2”">The description indicates that within the representation identified by "Representation id="1"", there exists a subrepresentation identified by "SubRepresentation id="12"", to which subgroup set 2 has been assigned.

[0090] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group2” / ">The description indicates that the audio stream contains encoded data for group 2, "group2". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectSound” / ">The description indicates that the encoded data for group 2, "group2," is object-encoded data for immersive sound, "objectSound."

[0091] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">The description indicates that the encoded data for group 2 "group2" does not belong to any switch group. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description indicates that the encoded data for group 2 "group2" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description indicates that the encoded data for group 2 "group2" belongs to preset group 2 "preset2".

[0092] Also," <adaptationset mimetype=""audio / mp4”" group=""2”">The description indicates that an AdaptationSet exists for the audio stream, that the audio stream is supplied in an MP4 file structure, and that it is assigned to group 2. <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">The description indicates that the audio stream's codec is "MPEGH (3D audio)".

[0093] " <representation id=""2”" bandwidth=""128000”">The description indicates that within the adaptation set of Group 1, there is an audio stream with a bitrate of 128kbps, identified as "Representation id="2"". <baseurl> audio / jp2 / 128.mp4< / baseurl> The description indicates that the location of the audio stream is "audio / jp2 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level2” / ">The description indicates that the audio stream will be transmitted on a track corresponding to level 2.

[0094] " <subrepresentation id=""21”" subgroupset=""3”">The description indicates that within the representation identified by "Representation id="2"", there exists a subrepresentation identified by "SubRepresentation id="21"", to which subgroup set 3 has been assigned.

[0095] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group3” / ">The description indicates that the audio stream contains encoded data for group 3, "group3". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang1” / ">The description indicates that the encoded data for group 3, "group3", is the object encoded data "objectLang1" for the speech language of the first language.

[0096] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">The description indicates that the encoded data for group 3 belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">The description indicates that the encoded data for group 3 "group3" belongs to preset group 1 "preset1".

[0097] " <subrepresentation id=""22”" subgroupset=""4”">The description indicates that within the representation identified as "Representation id="2"", there exists a subrepresentation identified as "SubRepresentation id="22"", to which subgroup set 4 has been assigned.

[0098] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group4” / ">The description indicates that the audio stream contains encoded data for group 4, "group4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang2” / ">The description indicates that the encoded data for group 4, "group4", is object-encoded data "objectLang2" for the speech language of the second language.

[0099] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">The description indicates that the encoded data for group 4 "group4" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">The description indicates that the encoded data for group 4 "group4" belongs to preset group 2 "preset2".

[0100] Here, <baseurl>This section describes the media file entity at the location indicated by ", i.e., the file that is containerized in each audio track. In the case of Non-Fragmented MP4, for example, it may be defined as "url 1" as shown in Figure 14(a). In this case, the "ftyp" box, which describes the file type, is placed first. This "ftyp" box indicates that it is an unfragmented MP4 file. Next, the "moov" box and the "mdat" box are placed. The "moov" box contains all metadata, such as header information and content metadata for each track, time information, etc. The "mdat" box contains the media data itself.

[0101] In the case of Fragmented MP4, for example, it may be defined as "url 2" as shown in Figure 14(b). In this case, a "styp" box describing the segment type is placed first. Next, a "sidx" box describing the segment index is placed. Following that, a predetermined number of movie fragments are placed. Here, a movie fragment consists of a "moof" box containing control information and an "mdat" box containing the media data itself. Since the "mdat" box of a single movie fragment contains a fragment obtained by fragmenting the transmission media, the control information placed in the box is control information related to that fragment. "styp", "sidx", "moof", and "mdat" are the units that make up a segment.

[0102] Furthermore, the combination of "url 1" and "url 2" mentioned above is also possible. In this case, for example, "url 1" can be used as the initialization segment, and "url 1" and "url 2" can be combined into a single MP4 service. Alternatively, "url 1" and "url 2" can be combined into one and defined as "url 3," as shown in Figure 14(c).

[0103] As mentioned above, the "moov" box contains the mapping between track IDs and level IDs. As shown in Figure 15(a), the "ftyp" box and the "moov" box constitute the Initialization segment. Inside the "moov" box is the "mvex" box, which in turn contains the "leva" box.

[0104] As shown in Figure 15(b), the mapping between track identifiers (track IDs) and level identifiers (level IDs) is defined in this "leva" box. In the example shown, "level0" is mapped to "track0", "level1" is mapped to "track1", and "level2" is mapped to "track2".

[0105] Figure 16(a) shows an example of transmission for each box in a broadcast system. One segment consists of an initialization segment (is) at the beginning, followed by "styp", then a "sidx" box, and then a predetermined number of movie fragments (consisting of "moof" boxes and "mdat" boxes). The example shown illustrates the case where the predetermined number is 1.

[0106] As described above, the "moov" box that makes up the initialization segment (is) contains the mapping between track identifiers (track IDs) and level identifiers (level IDs). Also, as shown in Figure 16(b), the "sidx" box contains the level information for each track, and the range information for each track is registered therein. That is, playback time information and track start position information on the file are registered corresponding to each level. On the receiving side, with regard to audio, it is possible to selectively extract the audio stream of the desired audio track based on this range information.

[0107] Returning to Figure 3, the service receiver 200 receives DASH / MP4, that is, an MPD file as a metafile, and an MP4 containing media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission path or a communication network transmission path.

[0108] As mentioned above, MP4 has a predetermined number of audio tracks (audio streams) that contain encoded data for multiple groups, in addition to the video stream, which constitute the 3D audio transmission data. The MPD file contains attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, as well as stream correspondence information indicating which audio track (audio stream) each of these multiple groups of encoded data belongs to.

[0109] The service receiver 200 selectively decodes audio streams containing encoded data of groups having attributes that match the speaker configuration and user selection information, based on attribute information and stream correspondence information, to obtain 3D audio output.

[0110] [DASH / MP4 generation unit of the service transmission system] Figure 17 shows an example configuration of the DASH / MP4 generation unit 110 included in the service transmission system 100. This DASH / MP4 generation unit 110 includes a control unit 111, a video encoder 112, an audio encoder 113, and a DASH / MP4 formatter 114.

[0111] The video encoder 112 receives video data SV as input, and applies encoding such as MPEG2, H.264 / AVC, or H.265 / HEVC to this video data SV to generate a video stream (video elementary stream). The audio encoder 113 receives audio data SA, along with channel data, and object data for immersive audio and speech dialogue.

[0112] The audio encoder 113 applies MPEGH encoding to the audio data SA to obtain 3D audio transmission data. This 3D audio transmission data includes channel-coded data (CD), immersive audio object-coded data (IAO), and speech dialogue object-coded data (SDO), as shown in Figure 5. The audio encoder 113 generates one or more audio streams (audio elementary streams) containing encoded data from multiple groups, in this case four groups (see Figures 6(a) and (b)).

[0113] The DASH / MP4 formatter 114 generates an MP4 file containing media streams (media segments) such as video and audio, based on the video stream generated by the video encoder 112 and a predetermined number of audio streams generated by the audio encoder 113. Here, each video and audio stream is stored in the MP4 as a separate track.

[0114] Furthermore, the DASH / MP4 formatter 114 generates an MPD file using content metadata, segment URL information, etc. In this embodiment, the DASH / MP4 formatter 114 inserts attribute information into the MPD file that indicates the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, and also inserts stream correspondence information that indicates which audio track (audio stream) each of these multiple groups of encoded data is included in (see Figures 11 and 12).

[0115] The operation of the DASH / MP4 generation unit 110 shown in Figure 17 will be briefly explained. Video data SV is supplied to the video encoder 112. The video encoder 112 encodes the video data SV using H.264 / AVC, H.265 / HEVC, etc., and generates a video stream containing the encoded video data. This video stream is supplied to the DASH / MP4 formatter 114.

[0116] Audio data SA is supplied to the audio encoder 113. This audio data SA includes channel data and object data for immersive audio and speech dialogue. The audio encoder 113 applies MPEGH encoding to the audio data SA to obtain 3D audio transmission data.

[0117] The transmission data for this 3D audio includes channel coding data (CD), as well as immersive audio object coding data (IAO) and speech dialogue object coding data (SDO) (see Figure 5). The audio encoder 113 then generates one or more audio streams containing the coded data for four groups (see Figures 6(a) and 6(b)). These audio streams are supplied to the DASH / MP4 formatter 114.

[0118] The DASH / MP4 formatter 114 generates an MP4 file containing media streams (media segments) such as video and audio content, based on the video stream generated by the video encoder 112 and a predetermined number of audio streams generated by the audio encoder 113. Here, each video and audio stream is stored in the MP4 as a separate track.

[0119] Furthermore, the DASH / MP4 formatter 114 generates an MPD file using content metadata and segment URL information. This MPD file contains attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, as well as stream correspondence information indicating which audio track (audio stream) each of these multiple groups of encoded data belongs to.

[0120] [Example of service receiver configuration] Figure 18 shows an example configuration of the service receiver 200. This service receiver 200 includes a receiving unit 201, a DASH / MP4 analysis unit 202, a video decoder 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. The service receiver 200 also includes container buffers 211-1 to 211-N, a combiner 212, a 3D audio decoder 213, an audio output processing circuit 214, and a speaker system 215. Furthermore, the service receiver 200 includes a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiver 225, and a remote control transmitter 226.

[0121] The CPU 221 controls the operation of each part of the service receiver 200. The flash ROM 222 stores the control software and data. The DRAM 223 constitutes the work area of ​​the CPU 221. The CPU 221 loads the software and data read from the flash ROM 222 onto the DRAM 223, starts the software, and controls each part of the service receiver 200.

[0122] The remote control receiver 225 receives the remote control signal (remote control code) transmitted from the remote control transmitter 226 and supplies it to the CPU 221. The CPU 221 controls each part of the service receiver 200 based on this remote control code. The CPU 221, flash ROM 222, and DRAM 223 are connected to the internal bus 224.

[0123] The receiving unit 201 receives DASH / MP4, that is, an MPD file as a metafile, and an MP4 containing media streams (media segments) such as video and audio, which are sent from the service transmission system 100 via an RF transmission path or a communication network transmission path.

[0124] MP4 files contain a predetermined number of audio tracks (audio streams) in addition to the video stream, which include encoded data for multiple groups that constitute the 3D audio transmission data. Furthermore, MPD files contain attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, as well as stream correspondence information indicating which audio track (audio stream) each of these multiple groups of encoded data belongs to.

[0125] The DASH / MP4 analysis unit 202 analyzes the MPD file and MP4 received by the receiving unit 201. The DASH / MP4 analysis unit 202 extracts the video stream from the MP4 and sends it to the video decoder 203. The video decoder 203 performs a decoding process on the video stream to obtain uncompressed video data.

[0126] The video processing circuit 204 performs scaling and image quality adjustment on the video data obtained from the video decoder 203 to obtain video data for display. The panel driving circuit 205 drives the display panel 206 based on the video data for display obtained from the video processing circuit 204. The display panel 206 is composed of, for example, an LCD (Liquid Crystal Display) or an organic EL display (organic electroluminescence display).

[0127] Furthermore, the DASH / MP4 analysis unit 202 extracts MPD information contained in the MPD file and sends it to the CPU 221. The CPU 221 controls the acquisition process of video and audio streams based on this MPD information. The DASH / MP4 analysis unit 202 also extracts metadata from the MP4, such as header information for each track, metadata descriptions of the content, and time information, and sends it to the CPU 221.

[0128] CPU21 recognizes the audio track (audio stream) containing the encoded data of a group with attributes that match the speaker configuration and listener (user) selection information, based on attribute information indicating the attributes of the encoded data of each group contained in the MPD file, and stream correspondence information indicating which audio track (audio stream) each group belongs to.

[0129] Furthermore, under the control of the CPU 221, the DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams contained in the MP4 file, which include encoded data of a group having attributes that match the speaker configuration and listener (user) selection information, by referring to the level ID, and therefore the track ID.

[0130] Container buffers 211-1 to 211-N each capture the audio stream extracted by the DASH / MP4 analysis unit 202. The number N of container buffers 211-1 to 211-N is considered to be a necessary and sufficient number, but in actual operation, the number of buffers used will be equal to the number of audio streams extracted by the DASH / MP4 analysis unit 202.

[0131] The combiner 212 reads each audio stream from the container buffers 211-1 to 211-N, each containing the audio stream extracted by the DASH / MP4 analysis unit 202, and supplies it to the 3D audio decoder 213 as encoded data for groups with attributes that match the speaker configuration and viewer (user) selection information.

[0132] The 3D audio decoder 213 decodes the encoded data supplied from the combiner 212 to obtain audio data for driving each speaker of the speaker system 215. Here, the encoded data to be decoded can be in three cases: containing only channel encoded data, containing only object encoded data, or containing both channel encoded data and object encoded data.

[0133] When decoding channel-encoded data, the 3D audio decoder 213 performs downmixing and upmixing to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When decoding object-encoded data, the 3D audio decoder 213 calculates speaker rendering (mixing ratio for each speaker) based on object information (metadata), and mixes the object's audio data into audio data for driving each speaker according to the calculation result.

[0134] The audio output processing circuit 214 performs necessary processing, such as D / A conversion and amplification, on the audio data obtained by the 3D audio decoder 213 to drive each speaker, and supplies it to the speaker system 215. The speaker system 215 has multiple channels, such as 2-channel, 5.1-channel, 7.1-channel, or 22.2-channel speakers.

[0135] The operation of the service receiver 200 shown in Figure 18 will be briefly explained. The receiving unit 201 receives DASH / MP4, which is an MP4 file containing an MPD file as a metafile and media streams (media segments) such as video and audio, from the service transmission system 100 via an RF transmission line or a communication network transmission line. The MPD file and MP4 file received in this way are supplied to the DASH / MP4 analysis unit 202.

[0136] The DASH / MP4 analysis unit 202 analyzes the MPD file and MP4 received by the receiving unit 201. The DASH / MP4 analysis unit 202 then extracts the video stream from the MP4 and sends it to the video decoder 203. The video decoder 203 decodes the video stream to obtain uncompressed video data. This video data is then supplied to the video processing circuit 204.

[0137] In the video processing circuit 204, scaling and image quality adjustment processing are performed on the video data obtained from the video decoder 203 to obtain video data for display. This video data for display is supplied to the panel driving circuit 205. The panel driving circuit 205 drives the display panel 206 based on the video data for display. As a result, the display panel 206 displays an image corresponding to the video data for display.

[0138] Furthermore, the DASH / MP4 analysis unit 202 extracts MPD information contained in the MPD file and sends it to the CPU 221. The DASH / MP4 analysis unit 202 also extracts metadata from the MP4, such as header information and content metadata for each track, and time information, and sends it to the CPU 221. The CPU 221 recognizes audio tracks (audio streams) that contain encoded data for groups with attributes that match the speaker configuration and viewer (user) selection information, based on the attribute information and stream correspondence information contained in the MPD file.

[0139] Furthermore, under the control of the CPU 221, the DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams contained in the MP4 file, which include encoded data of a group having attributes that match the speaker configuration and viewer (user) selection information, by referring to the track ID.

[0140] The audio stream extracted by the DASH / MP4 analysis unit 202 is loaded into the corresponding container buffer from container buffers 211-1 to 211-N. The combiner 212 reads the audio stream frame by frame from each container buffer into which the audio stream has been loaded, and supplies it to the 3D audio decoder 213 as encoded data of groups with attributes that match the speaker configuration and listener selection information. The 3D audio decoder 213 decodes the encoded data supplied from the combiner 212 to obtain audio data for driving each speaker of the speaker system 215.

[0141] When channel-encoded data is decoded, downmixing or upmixing is performed to match the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When object-encoded data is decoded, speaker rendering (mixing ratio for each speaker) is calculated based on object information (metadata), and the object's audio data is mixed with the audio data for driving each speaker according to the calculation result.

[0142] The audio data obtained by the 3D audio decoder 213 to drive each speaker is supplied to the audio output processing circuit 214. In this audio output processing circuit 214, necessary processing such as D / A conversion and amplification is performed on the audio data to drive each speaker. The processed audio data is then supplied to the speaker system 215. As a result, the speaker system 215 provides an audio output corresponding to the display image on the display panel 206.

[0143] Figure 19 shows an example of the audio decoding control processing of the CPU 221 in the service receiver 200 shown in Figure 18. In step ST1, the CPU 221 starts processing. Then, in step ST2, the CPU 221 detects the receiver speaker configuration, that is, the speaker configuration of the speaker system 215. Next, in step ST3, the CPU 221 obtains selection information regarding the audio output by the viewer (user).

[0144] Next, in step ST4, the CPU 221 reads the information related to each audio stream of the MPD information, namely "groupID", "attribute", "switchGroupID", "presetGroupID", and "levelID". Then, in step ST5, the CPU 221 recognizes the track ID of the audio track to which the encoded data group having attributes that match the speaker configuration and listener selection information belongs.

[0145] Next, in step ST6, the CPU 221 selects each audio track based on the recognition results and imports the stored audio streams into a container buffer. Then, in step ST7, the CPU 221 reads the audio stream from the container buffer frame by frame and supplies the necessary group encoding data to the 3D audio decoder 213.

[0146] Next, in step ST8, the CPU 221 decides whether or not to decode the object-encoded data. If the object-encoded data is decoded, in step ST9, the CPU 221 calculates the speaker rendering (mixing ratio to each speaker) based on the object information (metadata), using the azimuth (direction information) and elevation (elevation angle information). After that, the CPU 221 proceeds to step ST10. If the object-encoded data is not decoded in step ST8, the CPU 221 immediately proceeds to step ST10.

[0147] In step ST10, the CPU 221 determines whether or not to decode the channel-encoded data. If the channel-encoded data is decoded, in step ST11, the CPU 221 performs downmixing or upmixing to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. After that, the CPU 221 proceeds to step ST12. If the object-encoded data is not decoded in step ST10, the CPU 221 immediately proceeds to step ST12.

[0148] In step ST12, when the CPU 221 decodes the object-encoded data, it mixes the object's audio data with audio data to drive each speaker according to the calculation result in step ST9, and then performs dynamic range control. After that, the CPU 21 terminates processing in step ST13. If the object-encoded data is not to be decoded, the CPU 221 skips step ST12.

[0149] As described above, in the transmission / reception system 10 shown in Figure 3, the service transmission system 100 inserts attribute information into the MPD file that indicates the attributes of each of the multiple groups of encoded data contained in a predetermined number of audio streams. Therefore, the receiving side can easily recognize the attributes of each of the multiple groups of encoded data before decoding the encoded data, and can selectively decode and use only the encoded data of the necessary groups, thereby reducing the processing load.

[0150] Furthermore, in the transmission / reception system 10 shown in Figure 3, the service transmission system 100 inserts stream correspondence information into the MPD file that indicates which audio track (audio stream) each of the encoded data from multiple groups is contained in. As a result, the receiving side can easily recognize the audio track (audio stream) containing the encoded data from the required group, thereby reducing the processing load.

[0151] <2. Variant> In the above-described embodiment, the service receiver 200 is configured to selectively extract from multiple audio streams transmitted from the service transmission system 100 the audio streams containing encoded data of groups having attributes that match the speaker configuration and viewer selection information, and to perform a decoding process to obtain a predetermined number of audio data for driving the speakers.

[0152] However, as a service receiver, it is also conceivable to selectively extract one or more audio streams from multiple audio streams transmitted from the service transmission system 100 that have encoded data for groups with attributes that match the speaker configuration and listener selection information, reconstruct the audio streams that have encoded data for groups with attributes that match the speaker configuration and listener selection information, and distribute the reconstructed audio streams to devices connected to the local area network (including DLNA devices).

[0153] Figure 20 shows an example configuration of a service receiver 200A that distributes the reconstructed audio stream to devices connected to the local network, as described above. In Figure 20, parts corresponding to those in Figure 18 are denoted by the same reference numerals, and their detailed descriptions are omitted where appropriate.

[0154] Under the control of the CPU 221, the DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams in the MP4 file, which include encoded data for groups with attributes that match the speaker configuration and listener (user) selection information, by referring to the level ID, and therefore the track ID.

[0155] The audio stream extracted by the DASH / MP4 analysis unit 202 is loaded into the corresponding container buffer from container buffers 211-1 to 211-N. The combiner 212 reads the audio stream frame by frame from each container buffer into which the audio stream has been loaded and supplies it to the stream reconstruction unit 231.

[0156] In the stream reconstruction unit 231, encoded data of a predetermined group having attributes that match the speaker configuration and viewer selection information is selectively acquired, and an audio stream containing the encoded data of this predetermined group is reconstructed. This reconstructed audio stream is supplied to the distribution interface 232. From this distribution interface 232, it is distributed (transmitted) to the devices 300 connected to the local area network.

[0157] This on-premises network connectivity includes Ethernet connections and wireless connections such as "Wi-Fi" or "Bluetooth." "Wi-Fi" and "Bluetooth" are registered trademarks.

[0158] Furthermore, device 300 includes surround speakers, a second display, and an audio output device attached to the network terminal. Device 300, which receives the reconstructed audio stream, performs decoding processing similar to that of the 3D audio decoder 213 in the service receiver 200 in Figure 18 to obtain audio data for driving a predetermined number of speakers.

[0159] Furthermore, as a service receiver, a configuration is also conceivable in which the reconstructed audio stream described above is transmitted to a device connected via a digital interface such as "HDMI (High-Definition Multimedia Interface)", "MHL (Mobile High Definition Link)", or "DisplayPort". Note that "HDMI" and "MHL" are registered trademarks.

[0160] Furthermore, in the embodiments described above, an example was shown in which attribute information of the encoded data of each group is transmitted with an "attribute" field (see Figures 11 to 13). However, this technology also includes a method in which a special meaning is defined for the value of the Group ID itself between the transceiver and receiver, so that the type (attribute) of the encoded data can be recognized by recognizing a specific Group ID. In this case, the Group ID functions not only as a group identifier but also as attribute information of the encoded data of that group, and the "attribute" field becomes unnecessary.

[0161] Furthermore, the above-described embodiment showed an example in which the encoded data of multiple groups includes both channel-encoded data and object-encoded data (see Figure 5). However, this technology can be similarly applied when the encoded data of multiple groups includes only channel-encoded data or only object-encoded data.

[0162] Furthermore, this technology can also be configured as follows: (1) A transmission unit that transmits a metadata file containing metadata for acquiring a predetermined number of audio streams, including encoded data from multiple groups, at a receiving device, The metafile includes an information insertion unit that inserts attribute information indicating the attributes of each of the multiple groups of encoded data. Transmitter. (2) The above information insertion unit is, The above metafile will further include stream correspondence information indicating which audio stream each of the above groups of encoded data belongs to. The transmitting device described in (1) above. (3) The above stream compatibility information is This information shows the correspondence between a group identifier that identifies each of the above-mentioned multiple groups of encoded data and an identifier that identifies each of the above-mentioned predetermined number of audio streams. The transmitting device described in (2) above. (4) The above metafile is an MPD file. A transmitting device as described in any of (1) to (3) above. (5) The above information insertion unit is, Use "Supplementary Descriptor" to insert the above attribute information into the above metafile. The transmitting device described in (4) above. (6) The above-mentioned transmission unit is: The above metadata is transmitted via an RF transmission path or a communication network transmission path. A transmitting device as described in any of (1) to (5) above. (7) The above-mentioned transmitting unit is Further transmit a container in a predetermined format having a predetermined number of audio streams, including encoded data from the above-mentioned multiple groups. A transmitting device as described in any of (1) to (6) above. (8) The above container is MP4. The transmitting device described in (7) above. (9) The coded data of the above-mentioned multiple groups includes either or both channel coded data and object coded data. A transmitting device as described in any of (1) to (8) above. (10) A transmission step in which the transmitting unit transmits a metadata file having metadata for receiving a predetermined number of audio streams containing encoded data of multiple groups to be acquired by a receiving device, The metafile includes an information insertion step of inserting attribute information that indicates the attributes of each of the multiple groups of encoded data. Sending method. (11) The receiving unit has a metadata file that has metadata for acquiring a predetermined number of audio streams containing encoded data of multiple groups at a receiving device, The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. The system further includes a processing unit that processes the predetermined number of audio streams based on the attribute information described above. Receiving device. (12) The metafile further contains stream correspondence information indicating which audio stream each of the above groups of encoded data is included in: The above processing unit is, In addition to the attribute information described above, the predetermined number of audio streams are processed based on the stream correspondence information described above. The receiving device described in (11) above. (13) The above processing unit is Based on the attribute information and stream correspondence information described above, the system selectively decodes audio streams containing encoded data of groups with attributes that match the speaker configuration and user selection information. The receiving device described in (12) above. (14) The coded data of the above-mentioned multiple groups includes either or both channel coded data and object coded data. A receiving device as described in any of (11) to (13) above. (15) The receiving unit has a receiving step in which it receives a metadata file having metadata for acquiring a predetermined number of audio streams containing encoded data of multiple groups at a receiving device, The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. The process further includes a processing step of processing the predetermined number of audio streams based on the attribute information described above. Reception method. (16) A receiving unit that receives a metadata file having metadata for acquiring a predetermined number of audio streams containing encoded data of multiple groups, The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. A processing unit that selectively obtains encoded data from a predetermined number of audio streams based on the attribute information, and reconstructs an audio stream containing the encoded data of the predetermined group, The system further comprises a stream transmission unit that transmits the reconstructed audio stream to an external device. Receiving device. (17) The above metafile further contains stream correspondence information indicating which audio stream each of the above groups of encoded data is included in: The above processing unit is, In addition to the attribute information described above, based on the stream correspondence information described above, the encoded data of a predetermined group is selectively obtained from a predetermined number of audio streams. The receiving device described in (16) above. (18) The receiving unit has a receiving step in which it receives a metadata file having metadata for acquiring a predetermined number of audio streams containing encoded data of multiple groups at a receiving device, The above metafile contains attribute information indicating the attributes of each of the multiple groups of encoded data mentioned above. A processing step of selectively obtaining encoded data of a predetermined group from the above predetermined number of audio streams based on the above attribute information, and reconstructing an audio stream containing the encoded data of the predetermined group, The system further comprises a stream transmission step of transmitting the reconstructed audio stream to an external device. Reception method.

[0163] The main feature of this technology is that it reduces the processing load on the receiving end by inserting attribute information that shows the attributes of each of the encoded data of multiple groups contained in a predetermined number of audio streams, as well as stream correspondence information that shows which audio track (audio stream) each of the encoded data of multiple groups is contained in, into the MPD file (see Figures 11, 12, and 17). [Explanation of symbols]

[0164] 10. Transmit / Receive System 30A, 30B...MPEG-DASH based streaming system 31. DASH Stream File Server 32. DASH MPD Server 33, 33-1~33-N) ···Service receiver 34..CDN 35, 35-1~35-M)...Service Receiver 36. Broadcast transmission system 100...Service transmission system 110...DASH / MP4 generation section 112... Video Encoder 113... Audio Encoder 114···DASH / MP4 Formatter 200... Service Receiver 201... Receiver 202...DASH / MP4 analysis section 203...Video Decoder 204...Video processing circuit 205... Panel drive circuit 206...Display Panel 211-1~211-N···Container buffer 212... Combiner 213...3D Audio Decoder 214...Audio output processing circuit 215...Speaker System 221...CPU 222...Flash ROM 223···DRAM 224...Internal bus 225... Remote control receiver 226... Remote control transmitter 231...Stream Reconstruction Unit 232...Distribution Interface 300 devices< / baseurl> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / baseurl>

Claims

1. An acquisition unit that acquires a predetermined number of audio streams, The system includes a decoder that performs decoding on a predetermined number of audio streams and obtains audio data for driving each speaker of the speaker system, The aforementioned audio stream is associated with a data type-specific group ID and a preset group ID corresponding to multiple different group IDs, and furthermore, Among the aforementioned audio streams, multiple audio streams belonging to multiple selectable groups are associated with a switch group ID. The acquisition unit is, Select and retrieve a predetermined number of audio streams corresponding to the same preset group ID. Information processing device.

2. The system further comprises a combiner for integrating the predetermined number of audio streams. The information processing apparatus according to claim 1.

3. The codec of the aforementioned audio stream is MPEG-H 3D Audio. The information processing apparatus according to claim 1.

4. The aforementioned audio stream consists of MPEG audio stream packets. The information processing apparatus according to claim 1.

5. The aforementioned MPEG audio stream packet consists of a header and a payload. The information processing apparatus according to claim 4.

6. The header includes information such as the packet type, packet label, and packet length. The information processing apparatus according to claim 5.

7. The payload includes "SYNC" information corresponding to the synchronization start code, "Frame" information which is the actual data of the audio stream, and "Config" information which indicates the configuration of the "Frame" information. The information processing apparatus according to claim 5.

8. The "Frame" information includes channel-encoded data or object-encoded data that constitutes the audio stream. The information processing apparatus according to claim 7.

9. The object encoding data consists of SCE (Single Channel Element) encoding sample data and metadata for mapping the encoding sample data to a speaker located at an arbitrary position and rendering it. The information processing apparatus according to claim 8.

10. The aforementioned metadata is included as an extension element (Ext_element), The information processing apparatus according to claim 9.

11. The object-encoded data includes speech dialogue object-encoded data, which is object-encoded data for a speech language. The information processing apparatus according to claim 8.

12. The audio stream includes a group ID for identifying a group of the object-encoded data. The information processing apparatus according to claim 8.

13. The audio stream includes attribute information indicating the attributes of the group of multiple object-encoded data. The information processing apparatus according to claim 12.

14. Multiple object-encoded data are registered in a switch group (SW Group) and encoded. The information processing apparatus according to claim 8.

15. Multiple object-encoded data are associated with a switch group ID, which is an identifier for identifying the switch group. The information processing apparatus according to claim 14.

16. Multiple object-encoded data include a preset group in which multiple groups are bundled together. The information processing apparatus according to claim 12.

17. Multiple object-encoded data are associated with a preset group ID, which is an identifier for identifying the preset group. The information processing apparatus according to claim 16.

18. The multiple object-encoded data are associated with stream correspondence information indicating which audio stream each of the multiple object-encoded data is included in. The information processing apparatus according to claim 8.

19. The procedure for the acquisition unit to acquire a predetermined number of audio streams, The decoder has a procedure for decoding a predetermined number of audio streams and obtaining audio data for driving each speaker of the speaker system. The aforementioned audio stream is associated with a data type-specific group ID and a preset group ID corresponding to multiple different group IDs, and furthermore, Among the aforementioned audio streams, multiple audio streams belonging to multiple selectable groups are associated with a switch group ID. The acquisition unit is, Select and retrieve a predetermined number of audio streams corresponding to the same preset group ID. Information processing methods.

20. Computers, Means for acquiring a predetermined number of audio streams, The predetermined number of audio streams are decoded and used as a means to acquire audio data for driving each speaker of the speaker system. The aforementioned audio stream is associated with a data type-specific group ID and a preset group ID corresponding to multiple different group IDs, and furthermore, Among the aforementioned audio streams, multiple audio streams belonging to multiple selectable groups are associated with a switch group ID. The aforementioned computer, During the process of acquiring the aforementioned multiple audio streams, Select and retrieve a predetermined number of audio streams corresponding to the same preset group ID. program.