Information processing device, information processing method and program
By using a metafile with attribute and stream correspondence information to manage encoded audio data, the processing load on receiving devices is reduced, enabling efficient selective decoding and use of necessary audio data.
Patent Information
- Application Number
- JP2025097692
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-09-12
- Filing Date
- 2025-06-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2035-09-07
AI Technical Summary
Existing technologies face high processing loads when transmitting multiple types of encoded audio data, such as channel-coded and object-coded data, which can overwhelm the receiving devices.
A metafile is used to transmit meta information for acquiring audio streams, with attribute information and stream correspondence information inserted to facilitate selective decoding and processing of necessary groups of encoded data, reducing the processing load on receiving devices.
This approach allows for efficient recognition and selective decoding of required audio data, thereby reducing the processing load on receiving devices.
Smart Images

Figure 2025124902000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] BACKGROUND ART Conventionally, a technique has been proposed as a stereoscopic (3D) audio technique in which encoded sample data is mapped to speakers located at arbitrary positions based on metadata and then rendered (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2014-520491 Summary of the Invention [Problem to be solved by the invention]
[0004] It is conceivable that object-coded data consisting of coded sample data and metadata can be transmitted together with channel-coded data such as 5.1 channels or 7.1 channels, thereby enabling audio reproduction with enhanced realism on the receiving side.
[0005] An object of the present technology is to reduce the processing load on the receiving side when multiple types of encoded data are transmitted. [Means for solving the problem]
[0006] The concept of this technology is: a transmitter for transmitting a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device; an information inserting unit that inserts attribute information indicating attributes of each of the plurality of groups of coded data into the metafile; Located in the transmitting device.
[0007] In the present technology, a transmitter transmits a metafile having meta information for allowing a receiving device to acquire a predetermined number of audio streams including multiple groups of encoded data, where the multiple groups of encoded data may include either or both of channel-encoded data and object-encoded data.
[0008] The information inserting unit inserts attribute information indicating attributes of each of the multiple groups of encoded data into the metafile. For example, the metafile may be an MPD (Media Presentation Description) file. In this case, for example, the information inserting unit may insert the attribute information into the metafile using a "Supplementary Descriptor."
[0009] Furthermore, for example, the transmitter may be configured to transmit the metafile via an RF transmission path or a communication network transmission path. Furthermore, for example, the transmitter may be configured to further transmit a container in a predetermined format having a predetermined number of audio streams including the encoded data of the multiple groups. For example, the container is MP4. In this invention report, MP4 refers to the ISO base media file format (ISOBMFF) (ISO / IEC 14496-12:2012).
[0010] In this way, in the present technology, attribute information indicating the attributes of each of the coded data of the multiple groups is inserted into a metafile having meta information for acquiring a predetermined number of audio streams including coded data of the multiple groups at a receiving device. As a result, the receiving device can easily recognize the attributes of each of the coded data of the multiple groups before decoding the coded data, and can selectively decode and use only the coded data of the necessary groups, thereby reducing the processing load.
[0011] In the present technology, for example, the information insertion unit may further insert stream correspondence information into the metafile, the stream correspondence information indicating which audio streams each contain encoded data of the plurality of groups. In this case, for example, the stream correspondence information may be information indicating a correspondence between a group identifier that identifies each of the encoded data of the plurality of groups and an identifier that identifies each of the predetermined number of audio streams. In this case, the receiving side can easily recognize the audio stream that contains the encoded data of the required group, thereby reducing the processing load.
[0012] Another concept of the present technology is a receiving unit for receiving a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device; attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; The audio signal processing device further includes a processor for processing the predetermined number of audio streams based on the attribute information. It is in the receiving device.
[0013] In the present technology, a receiving unit receives a metafile. The metafile has metainformation for acquiring a predetermined number of audio streams including multiple groups of encoded data in a receiving device. For example, the multiple groups of encoded data may include either or both of channel-encoded data and object-encoded data. Attribute information indicating attributes of each of the multiple groups of encoded data is inserted into the metafile. A processing unit processes the predetermined number of audio streams based on the attribute information.
[0014] In this way, with this technology, a predetermined number of audio streams are processed based on attribute information indicating the attributes of each of the multiple groups of encoded data inserted into the metafile, which allows selective decoding and use of only the encoded data of the necessary groups, thereby reducing the processing load.
[0015] In the present technology, for example, stream correspondence information indicating which audio streams contain each of the multiple groups of encoded data may be further inserted into the metafile, and the processing unit may process a predetermined number of audio streams based on the stream correspondence information in addition to the attribute information. In this case, it is possible to easily identify the audio stream containing the encoded data of the required group, thereby reducing the processing load.
[0016] Furthermore, in the present technology, for example, the processing unit may be configured to selectively perform decoding processing on an audio stream including encoded data of a group having attributes that match the speaker configuration and user selection information, based on the attribute information and the stream correspondence information.
[0017] Furthermore, another concept of the present technology is a receiving unit for receiving a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device; attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; a processing unit that selectively acquires a predetermined group of coded data from the predetermined number of audio streams based on the attribute information, and reconstructs an audio stream including the coded data of the predetermined group; a stream sending unit that sends the reconstructed audio stream to an external device. It is in the receiving device.
[0018] In the present technology, a receiving unit receives a metafile, which has meta information for allowing a receiving device to acquire a predetermined number of audio streams including multiple groups of encoded data, and attribute information indicating attributes of each of the multiple groups of encoded data is inserted into the metafile.
[0019] The processing unit selectively acquires a predetermined group of encoded data from a predetermined number of audio streams based on the attribute information, reconstructs an audio stream including the encoded data of the predetermined group, and transmits the reconstructed audio stream to an external device by the stream transmission unit.
[0020] In this way, with this technology, a predetermined group of coded data is selectively acquired from a predetermined number of audio streams based on attribute information indicating the attributes of each of the multiple groups of coded data inserted into a metafile, and an audio stream to be transmitted to an external device is reconstructed. This makes it possible to easily acquire the coded data of the necessary groups, thereby reducing the processing load.
[0021] In the present technology, for example, the metafile may further include stream correspondence information indicating which audio streams each contain encoded data of a plurality of groups, and the processing unit may selectively acquire encoded data of a predetermined group from a predetermined number of audio streams based on the stream correspondence information in addition to the attribute information. In this case, audio streams containing encoded data of a predetermined group can be easily recognized, thereby reducing the processing load. [Effects of the Invention]
[0022] According to the present technology, it is possible to reduce the processing load on the receiving side when transmitting multiple types of encoded data. Note that the effects described in this specification are merely examples and are not limiting, and additional effects may also be provided. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of an MPEG-DASH-based stream distribution system. [Figure 2] FIG. 2 is a diagram showing an example of the relationship between structures hierarchically arranged in an MPD file. [Figure 3] 1 is a block diagram showing an example of the configuration of a transmission / reception system according to an embodiment; [Figure 4] FIG. 1 is a diagram showing the structure of an audio frame (1024 samples) in 3D audio transmission data. [Figure 5] FIG. 10 is a diagram illustrating an example of the configuration of 3D audio transmission data. [Figure 6] 10A and 10B are diagrams illustrating an example of the configuration of an audio frame when 3D audio transmission data is transmitted in one track (one audio stream) and when 3D audio transmission data is transmitted in multiple tracks (multiple audio streams). [Figure 7] FIG. 10 is a diagram showing an example of group division when transmitting on four tracks in an example configuration of 3D audio transmission data. [Figure 8] FIG. 10 is a diagram showing the correspondence between groups and tracks in an example of group division (divided into four). [Figure 9] FIG. 10 is a diagram showing an example of group division when transmitting on two tracks in an example configuration of 3D audio transmission data. [Figure 10] FIG. 10 is a diagram showing the correspondence between groups and tracks in an example of group division (division into two). [Figure 11] FIG. 10 is a diagram illustrating an example of an MPD file description. [Figure 12] 10A and 10B are diagrams illustrating examples of descriptions in an MPD file and other files. [Figure 13] FIG. 10 is a diagram showing an example of a definition of "schemeIdUri" by "SupplementaryDescriptor." [Figure 14] " <baseurl>This is a diagram for explaining the media file entity at the location indicated by ". [Figure 15] FIG. 10 is a diagram for explaining the description of the correspondence between track identifiers (track IDs) and level identifiers (level IDs) in the "moov" box. [Figure 16] FIG. 10 is a diagram showing an example of transmission from each box in the case of a broadcasting system. [Figure 17] 10 is a block diagram showing an example of the configuration of a DASH / MP4 generation unit included in the service transmission system. FIG. [Figure 18] FIG. 2 is a block diagram illustrating an example of the configuration of a service receiver. [Figure 19] 10 is a flowchart showing an example of an audio decode control process of a CPU in the service receiver. [Figure 20] FIG. 10 is a block diagram showing another example of the configuration of a service receiver. DETAILED DESCRIPTION OF THE INVENTION
[0024] Hereinafter, modes for carrying out the invention (hereinafter referred to as "embodiments") will be described. The description will be made in the following order. 1. Embodiment 2. Variations
[0025] <1. Embodiment> [Overview of MPEG-DASH-based streaming system] First, an outline of an MPEG-DASH-based stream distribution system to which the present technology can be applied will be described.
[0026] 1(a) shows an example of the configuration of an MPEG-DASH-based stream distribution system 30A. In this example, a media stream and an MPD file are transmitted over a communication network transmission path. In this stream distribution system 30A, N service receivers 33-1, 33-2, . . . , 33-N are connected to a DASH stream file server 31 and a DASH MPD server 32 via a CDN (Content Delivery Network) 34.
[0027] The DASH stream file server 31 generates stream segments conforming to the DASH specification (hereinafter referred to as "DASH segments") based on media data of predetermined content (video data, audio data, subtitle data, etc.) and transmits the segments in response to HTTP requests from service receivers. This DASH stream file server 31 may be a server dedicated to streaming, or may also serve as a web server.
[0028] Furthermore, in response to a request for a segment of a predetermined stream sent from a service receiver 33 (33-1, 33-2, . . . , 33-N) via the CDN 34, the DASH stream file server 31 transmits the segment of the stream to the requesting receiver via the CDN 34. In this case, the service receiver 33 refers to the rate value described in the MPD (Media Presentation Description) file, selects the stream with the optimal rate depending on the state of the network environment in which the client is located, and makes the request.
[0029] The DASH MPD server 32 is a server that generates an MPD file for acquiring DASH segments generated in the DASH stream file server 31. The MPD file is generated based on content metadata from a content management server (not shown) and addresses (URLs) of segments generated in the DASH stream file server 31. Note that the DASH stream file server 31 and the DASH MPD server 32 may be physically the same server.
[0030] In the MPD format, attributes of each stream, such as video and audio, are described using an element called a representation. For example, an MPD file describes the rate of each of multiple video data streams with different rates using separate representations. The service receiver 33 can refer to the rate values and select the optimal stream depending on the state of the network environment in which the service receiver 33 is located, as described above.
[0031] 1(b) shows an example of the configuration of an MPEG-DASH-based stream distribution system 30B. In this example, a media stream and an MPD file are transmitted over an RF transmission path. This stream distribution system 30B is composed of a broadcast transmission system 36 connected to a DASH stream file server 31 and a DASH MPD server 32, and M service receivers 35-1, 35-2, . . . , 35-M.
[0032] In the case of this stream distribution system 30B, the broadcast transmission system 36 transmits DASH-specification stream segments (DASH segments) generated by the DASH stream file server 31 and MPD files generated by the DASH MPD server 32 over broadcast waves.
[0033] Figure 2 shows an example of the relationship between the various structures arranged hierarchically in an MPD file. As shown in Figure 2(a), the media presentation as an entire MPD file contains multiple periods separated by time intervals. For example, the first period starts at 0 seconds, the next period starts at 100 seconds, and so on.
[0034] As shown in Figure 2(b), a period has multiple representations, which are grouped into adaptation sets, each representing a media stream with the same content but different stream attributes, such as different rates.
[0035] As shown in Figure 2(c), the representation contains segment info (SegmentInfo). As shown in Figure 2(d), this segment info contains an initialization segment and multiple media segments in which information for each segment, which are further divided into smaller periods, is described. The media segment contains information such as the address (URL) for actually obtaining segment data such as video and audio.
[0036] Furthermore, streams can be freely switched between multiple representations grouped in an adaptation set, allowing the optimal stream rate to be selected depending on the state of the network environment where the service receiver is located, enabling uninterrupted distribution.
[0037] [Example of transmission and reception system configuration] Fig. 3 shows an example of the configuration of a transmission / reception system 10 according to an embodiment. The transmission / reception system 10 is made up of a service transmission system 100 and a service receiver 200. In the transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31 and the DASH MPD server 32 of the stream delivery system 30A shown in Fig. 1(a) above. In the transmission / reception system 10, the service transmission system 100 corresponds to the DASH stream file server 31, the DASH MPD server 32, and the broadcast transmission system 36 of the stream delivery system 30B shown in Fig. 1(b) above.
[0038] In addition, in this transmission / reception system 10, the service receiver 200 corresponds to the service receivers 33 (33-1, 33-2, . . . , 33-N) of the stream distribution system 30A shown in Fig. 1(a) above. In addition, in this transmission / reception system 10, the service receiver 200 corresponds to the service receivers 35 (35-1, 35-2, . . . , 35-M) of the stream distribution system 30B shown in Fig. 1(b) above.
[0039] The service transmission system 100 transmits DASH / MP4, i.e., an MP4 containing an MPD file as a metafile and media streams (media segments) such as video and audio, via an RF transmission path (see Figure 1(b)) or a communication network transmission path (see Figure 1(a)).
[0040] Figure 4 shows the structure of an audio frame (1024 samples) in the 3D audio (MPEGH) transmission data handled in this embodiment. This audio frame is made up of multiple MPEG Audio Stream Packets. Each MPEG Audio Stream Packet is made up of a header and a payload.
[0041] The header contains information such as packet type, packet label, and packet length. The payload contains information defined by the packet type in the header. This payload information contains "SYNC" information, which is equivalent to a synchronization start code, "Frame" information, which is the actual data for 3D audio transmission, and "Config" information, which indicates the configuration of this "Frame" information.
[0042] The "Frame" information includes channel-encoded data and object-encoded data that make up the 3D audio transmission data. Here, the channel-encoded data is made up of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), and LFE (Low Frequency Element). The object-encoded data is made up of encoded sample data of SCE (Single Channel Element) and metadata for mapping and rendering it to speakers located at any position. This metadata is included as an extension element (Ext_element).
[0043] Figure 5 shows an example of the structure of 3D audio transmission data. This example consists of one channel coded data and two object coded data. The one channel coded data is 5.1 channel coded data (CD) and consists of coded sample data for SCE1, CPE1.1, CPE1.2, and LFE1.
[0044] The two object coded data are coded data for an immersive audio object (IAO) and a speech dialog object (SDO). The immersive audio object coded data is object coded data for immersive sound, and consists of coded sample data SCE2 and metadata EXE_El (Object metadata)2 for mapping the coded sample data to speakers located at any position and rendering it.
[0045] The speech dialogue object coded data is object coded data for a speech language. In this example, there are speech dialogue object coded data corresponding to each of the first and second languages. The speech dialogue object coded data corresponding to the first language consists of coded sample data SCE3 and metadata EXE_El (Object metadata)3 for mapping the coded sample data to a speaker located at an arbitrary position and rendering it. The speech dialogue object coded data corresponding to the second language consists of coded sample data SCE4 and metadata EXE_El (Object metadata)4 for mapping the coded sample data to a speaker located at an arbitrary position and rendering it.
[0046] The coded data are classified into groups according to their types. In the illustrated example, coded channel data of 5.1 channels are grouped as Group 1, coded immersive audio object data are grouped as Group 2, coded speech dialogue object data of a first language are grouped as Group 3, and coded speech dialogue object data of a second language are grouped as Group 4.
[0047] Furthermore, on the receiving side, selectable groups are registered in a switch group (SW Group) and encoded. In the illustrated example, groups 3 and 4 are registered in switch group 1 (SW Group 1). Groups are also bundled together to form preset groups, enabling playback according to the use case. In the illustrated example, groups 1, 2, and 3 are bundled together to form preset group 1, and groups 1, 2, and 4 are bundled together to form preset group 2.
[0048] Returning to FIG. 3, the service transmission system 100 transmits 3D audio transmission data including multiple groups of encoded data as described above in one track as one audio stream, or in multiple tracks as multiple audio streams.
[0049] Fig. 6(a) shows a schematic example of an audio frame configuration when transmitting one track (one audio stream) in the example configuration of 3D audio transmission data in Fig. 5. In this case, audio track 1 includes channel coded data (CD), immersive audio object coded data (IAO), and speech dialogue object coded data (SDO) along with "SYNC" information and "Config" information.
[0050] Fig. 6(b) shows a schematic diagram of an example of an audio frame configuration when transmitting multiple tracks (multiple audio streams), three tracks in this case, in the example configuration of 3D audio transmission data in Fig. 5. In this case, audio track 1 contains channel-encoded data (CD) along with "SYNC" information and "Config" information. Audio track 2 contains immersive audio object-encoded data (IAO) along with "SYNC" information and "Config" information. Audio track 3 contains speech dialogue object-encoded data (SDO) along with "SYNC" information and "Config" information.
[0051] Fig. 7 shows an example of group division when transmitting on four tracks in the example configuration of 3D audio transmission data in Fig. 5. In this case, audio track 1 contains channel-encoded data (CD) classified as group 1. Audio track 2 contains immersive audio object-encoded data (IAO) classified as group 2. Audio track 3 contains speech dialogue object-encoded data (SDO) in a first language classified as group 3. Audio track 4 contains speech dialogue object-encoded data (SDO) in a second language classified as group 4.
[0052] FIG. 8 shows the correspondence between groups and audio tracks in the group division example (divided into four) of FIG. 7. Here, group ID is an identifier for identifying a group. Attribute indicates the attribute of the coded data of each group. Switch Group ID is an identifier for identifying a switching group. Preset Group ID is an identifier for identifying a preset group. Track ID is an identifier for identifying an audio track.
[0053] The illustrated correspondence relationship indicates that the coded data belonging to group 1 is channel coded data, does not constitute a switch group, and is included in audio track 1. The illustrated correspondence relationship also indicates that the coded data belonging to group 2 is object coded data for immersive sound (immersive audio object coded data), does not constitute a switch group, and is included in audio track 2.
[0054] The illustrated correspondence relationship also indicates that the coded data belonging to group 3 is object coded data for a first speech language (speech dialogue object coded data), constitutes switch group 1, and is included in audio track 3. The illustrated correspondence relationship also indicates that the coded data belonging to group 4 is object coded data for a second speech language (speech dialogue object coded data), constitutes switch group 1, and is included in audio track 4.
[0055] The illustrated correspondence relationship also indicates that preset group 1 includes group 1, group 2, and group 3. The illustrated correspondence relationship also indicates that preset group 2 includes group 1, group 2, and group 4.
[0056] Fig. 9 shows an example of group division when transmitting on two tracks in the example configuration of 3D audio transmission data in Fig. 5. In this case, audio track 1 includes channel coded data (CD) classified as group 1 and immersive audio object coded data (IAO) classified as group 2. Audio track 2 also includes speech dialogue object coded data (SDO) in a first language classified as group 3 and speech dialogue object coded data (SDO) in a second language classified as group 4.
[0057] Fig. 10 shows the correspondence between groups and substreams in the group division example (two division) of Fig. 9. The correspondence shown in the figure indicates that the coded data belonging to group 1 is channel coded data, does not constitute a switch group, and is included in audio track 1. The correspondence shown in the figure also indicates that the coded data belonging to group 2 is object coded data for immersive sound (immersive audio object coded data), does not constitute a switch group, and is included in audio track 1.
[0058] The illustrated correspondence relationship also indicates that the coded data belonging to group 3 is object coded data for a first speech language (speech dialogue object coded data), constitutes switch group 1, and is included in audio track 2. The illustrated correspondence relationship also indicates that the coded data belonging to group 4 is object coded data for a second speech language (speech dialogue object coded data), constitutes switch group 1, and is included in audio track 2.
[0059] The illustrated correspondence relationship also indicates that preset group 1 includes group 1, group 2, and group 3. The illustrated correspondence relationship also indicates that preset group 2 includes group 1, group 2, and group 4.
[0060] Returning to Fig. 3, the service transmission system 100 inserts attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data into the MPD file. The service transmission system 100 also inserts stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is included. In this embodiment, the stream correspondence information is, for example, information indicating the correspondence between group IDs and track IDs.
[0061] The service transmission system 100 inserts this attribute information and stream correspondence information into the MPD file. "SupplementaryDescriptor" makes it possible to newly define "schemeIdUri" as a broadcast or other application, separate from the existing definitions in conventional standards. In this embodiment, the service transmission system 100 uses "SupplementaryDescriptor" to insert this attribute information and stream correspondence information into the MPD file.
[0062] Fig. 11 shows an example of an MPD file description corresponding to the group division example (divided into four) in Fig. 7. Fig. 12 shows an example of an MPD file description corresponding to the group division example (divided into two) in Fig. 9. Here, for simplicity of explanation, an example is shown in which only information about the audio stream is described, but in reality, information about other media streams such as the video stream is also described. Fig. 13 is a diagram showing an example of a definition of "schemeIdUri" by "SupplementaryDescriptor".
[0063] First, we will explain the example of MPD file description in Figure 11. <adaptationset mimetype=""audio / mp4”" group=""1”">" indicates that an adaptation set (AdaptationSet) for the audio stream exists, that the audio stream is supplied in an MP4 file structure, and that group 1 is assigned to the audio stream.
[0064] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This description indicates that the codec of the audio stream is "MPEGH (3D audio)". As shown in Fig. 13, "schemeIdUri="urn:brdcst:codecType"" indicates the type of codec. In this case, it is "mpegh".
[0065] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group1” / ">" indicates that the audio stream contains encoded data of group 1 "group1". As shown in FIG. 13, "schemeIdUri="urn:brdcst:3dAudio:groupId"" indicates the identifier of the group.
[0066] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""channeldata” / ">" indicates that the coded data of group 1 "group1" is channel coded data "channeldata". As shown in FIG. 13, "schemeIdUri="urn:brdcst:3dAudio:attribute"" indicates the attribute of the coded data of the corresponding group.
[0067] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">" indicates that the coded data of group 1 "group1" does not belong to any switch group. As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:switchGroupId"" indicates the identifier of the switch group to which the group belongs. For example, when "value" is "0", it indicates that it does not belong to any switch group. When "value" is other than "0", it indicates the switch group to which it belongs.
[0068] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 1 "group1" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the encoded data of group 1 "group1" belongs to preset group 2 "preset2". As shown in FIG. 13, "schemeIdUri="urn:brdcst:3dAudio:presetGroupId"" indicates the identifier of the preset group to which the group belongs.
[0069] " <representation id=""1”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps that contains the coded data of group 1 "group1" as a representation identified by "Representation id="1"" in the adaptation set of group 1. And, " <baseurl> audio / jp1 / 128.mp4< / baseurl> " indicates the location of the audio stream as "audio / jp1 / 128.mp4".
[0070] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level1” / ">" indicates that the audio stream will be transmitted on a track corresponding to level 1 "level1". As shown in Figure 13, "schemeIdUri="urn:brdcst:3dAudio:levelId" indicates the level identifier corresponding to the identifier of the track that transmits the audio stream containing the encoded data of the corresponding group. Note that the correspondence between the track identifier (track ID) and the level identifier (level ID) is described in, for example, the "moov" box, as will be described later.
[0071] Also," <adaptationset mimetype=""audio / mp4”" group=""2”">" indicates that an adaptation set (AdaptationSet) for the audio stream exists, that the audio stream is supplied in an MP4 file structure, and that group 2 is assigned to the audio stream.
[0072] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This indicates that the audio stream codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group2” / ">" indicates that the audio stream contains coded data of group 2 "group2".
[0073] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectSound” / ">" indicates that the coded data of group 2 "group2" is object coded data "objectSound" for immersive sound. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">" indicates that the coded data of group 2 "group2" does not belong to any switch group.
[0074] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 2 "group2" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the coded data of group 2 "group2" belongs to preset group 2 "preset2".
[0075] " <representation id=""2”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps that contains the coded data of group 2 "group2" as a representation identified by "Representation id="2"" in the adaptation set of group 2. <baseurl> audio / jp2 / 128.mp4< / baseurl> " indicates the location of the audio stream as "audio / jp2 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level2” / ">" indicates that the audio stream is transmitted on a track that supports level 2.
[0076] Also," <adaptationset mimetype=""audio / mp4”" group=""3”">" indicates that an adaptation set (AdaptationSet) for the audio stream exists, that the audio stream is supplied in an MP4 file structure, and that group 3 is assigned to it.
[0077] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This indicates that the audio stream codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group3” / ">" indicates that the audio stream contains coded data for group 3. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang1” / ">" indicates that the coded data of group 3 "group3" is object coded data "objectLang1" for the first speech language.
[0078] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">" indicates that the coded data of group 3 "group3" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 3 "group3" belongs to preset group 1 "preset1".
[0079] " <representation id=""3”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps that contains the coded data of group 3 "group3" as a representation identified by "Representation id="3"" in the adaptation set of group 3. And, " <baseurl> audio / jp3 / 128.mp4< / baseurl> " indicates the location of the audio stream as "audio / jp3 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level3” / ">" indicates that the audio stream is transmitted on a track that supports level 3.
[0080] Also," <adaptationset mimetype=""audio / mp4”" group=""4”">" indicates that an adaptation set (AdaptationSet) for the audio stream exists, that the audio stream is supplied in an MP4 file structure, and that group 4 is assigned to the audio stream.
[0081] " <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This indicates that the audio stream codec is "MPEGH (3D audio)". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group4” / ">" indicates that the audio stream contains coded data for group 4. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang2” / ">" indicates that the coded data of group 4 "group4" is object coded data "objectLang2" for the second speech language.
[0082] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">" indicates that the coded data of group 4 "group4" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the coded data of group 4 "group4" belongs to preset group 2 "preset2".
[0083] " <representation id=""4”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps that contains the coded data of group 4 "group4" as a representation identified by "Representation id="4"" in the adaptation set of group 4. And, " <baseurl> audio / jp4 / 128.mp4< / baseurl> " indicates the location of the audio stream as "audio / jp4 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level4” / ">" indicates that the audio stream will be transmitted on a track that supports level 4.
[0084] Next, an example of an MPD file description in FIG. 12 will be described. <adaptationset mimetype=""audio / mp4”" group=""1”">" indicates that an adaptation set (AdaptationSet) exists for the audio stream, that the audio stream is provided in an MP4 file structure, and that group 1 is assigned to the audio stream. <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This description indicates that the codec of the audio stream is "MPEGH (3D audio)".
[0085] " <representation id=""1”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps in the adaptation set of group 1 as a representation identified by "Representation id="1"". And, " <baseurl> audio / jp1 / 128.mp4< / baseurl> " indicates that the location of the audio stream is "audio / jp1 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level1” / ">" indicates that the audio stream is transmitted on a track corresponding to level 1 "level1".
[0086] " <subrepresentation id=""11”" subgroupset=""1”">" indicates that the representation identified by "Representation id="1"" contains a sub-representation identified by "SubRepresentation id="11"", and that subgroup set 1 is assigned to it.
[0087] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group1” / ">" indicates that the audio stream contains coded data for group 1 "group1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""channeldata” / ">" indicates that the coded data of group 1 "group1" is channel coded data "channeldata".
[0088] <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">" indicates that the coded data of group 1 "group1" does not belong to any switch group. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 1 "group1" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the coded data of group 1 "group1" belongs to preset group 2 "preset2".
[0089] " <subrepresentation id=""12”" subgroupset=""2”">" indicates that the representation identified by "Representation id="1"" contains a sub-representation identified by "SubRepresentation id="12"", and that subgroup set 2 is assigned to it.
[0090] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group2” / ">" indicates that the audio stream contains coded data for group 2 "group2". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectSound” / ">" indicates that the coded data of group 2 "group2" is coded object data "objectSound" for immersive sound.
[0091] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""0” / ">" indicates that the coded data of group 2 "group2" does not belong to any switch group. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 2 "group2" belongs to preset group 1 "preset1". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the coded data of group 2 "group2" belongs to preset group 2 "preset2".
[0092] Also," <adaptationset mimetype=""audio / mp4”" group=""2”">" indicates that an adaptation set (AdaptationSet) exists for the audio stream, that the audio stream is provided in an MP4 file structure, and that group 2 is assigned to it. <supplementarydescriptor schemeiduri=""urn:brdcst:codecType”" value=""mpegh” / ">This description indicates that the codec of the audio stream is "MPEGH (3D audio)".
[0093] " <representation id=""2”" bandwidth=""128000”">" indicates that there is an audio stream with a bit rate of 128 kbps in the adaptation set of group 1 as a representation identified by "Representation id="2"". And, " <baseurl> audio / jp2 / 128.mp4< / baseurl> " indicates the location of the audio stream as "audio / jp2 / 128.mp4". <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:levelId”" value=""level2” / ">" indicates that the audio stream is transmitted on a track that supports level 2.
[0094] " <subrepresentation id=""21”" subgroupset=""3”">" indicates that the representation identified by "Representation id="2"" contains a sub-representation identified by "SubRepresentation id="21"", and that subgroup set 3 is assigned to it.
[0095] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group3” / ">" indicates that the audio stream contains coded data for group 3. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang1” / ">" indicates that the coded data of group 3 "group3" is object coded data "objectLang1" for the first speech language.
[0096] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">" indicates that the coded data of group 3 "group3" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset1” / ">" indicates that the coded data of group 3 "group3" belongs to preset group 1 "preset1".
[0097] " <subrepresentation id=""22”" subgroupset=""4”">" indicates that the representation identified by "Representation id="2"" contains a sub-representation identified by "SubRepresentation id="22"", and that subgroup set 4 is assigned to it.
[0098] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:groupId”" value=""group4” / ">" indicates that the audio stream contains coded data for group 4. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:attribute”" value=""objectLang2” / ">" indicates that the coded data of group 4 "group4" is object coded data "objectLang2" for the second speech language.
[0099] " <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:switchGroupId”" value=""1” / ">" indicates that the coded data of group 4 "group4" belongs to switch group 1. <supplementarydescriptor schemeiduri=""urn:brdcst:3dAudio:presetGroupId”" value=""preset2” / ">" indicates that the coded data of group 4 "group4" belongs to preset group 2 "preset2".
[0100] Here, " <baseurl>", that is, the file contained in each audio track. In the case of non-fragmented MP4, for example, it may be defined as "url 1" as shown in FIG. 14(a). In this case, an "ftyp" box describing the file type is placed first. This "ftyp" box indicates that it is a non-fragmented MP4 file. Next, a "moov" box and an "mdat" box are placed. The "moov" box contains all metadata, such as header information for each track, meta descriptions of the content, and time information. The "mdat" box contains the media data itself.
[0101] In the case of Fragmented MP4, for example, it may be defined as "url 2", as shown in Figure 14(b). In this case, a "styp" box describing the segment type is placed first. Next is a "sidx" box describing the segment index. Following that, a predetermined number of movie fragments are placed. Here, a movie fragment consists of a "moof" box containing control information and an "mdat" box containing the media data itself. The "mdat" box of one movie fragment contains a fragment obtained by fragmenting the transmission media, so the control information placed in the box is control information related to that fragment. "styp", "sidx", "moof", and "mdat" are the units that make up a segment.
[0102] Also, a combination of the above-mentioned "url 1" and "url 2" is possible. In this case, for example, "url 1" can be an initialization segment, and "url 1" and "url 2" can be defined as one service MP4. Alternatively, "url 1" and "url 2" can be combined into one and defined as "url 3" as shown in Figure 14(c).
[0103] As mentioned above, the "moov" box describes the correspondence between track identifiers (track IDs) and level identifiers (level IDs). As shown in FIG. 15(a), the "ftyp" box and the "moov" box make up an initialization segment. The "moov" box contains a "mvex" box, which in turn contains a "leva" box.
[0104] As shown in Figure 15(b), the "leva" box defines the correspondence between track identifiers (track IDs) and level identifiers (level IDs). In the example shown, "level0" is associated with "track0", "level1" is associated with "track1", and "level2" is associated with "track2".
[0105] Figure 16(a) shows an example of how each box is transmitted in a broadcast system. One segment consists of an initialisation segment (is), followed by a "styp" and an "sidx" box, followed by a predetermined number of movie fragments (consisting of "moof" and "mdat" boxes). The example shown shows the case where the predetermined number is 1.
[0106] As described above, the "moov" box that constitutes the initialization segment (is) describes the correspondence between track identifiers (track IDs) and level identifiers (level IDs). Also, as shown in Figure 16(b), the "sidx" box indicates each track by its level, and registers range information for each track. That is, playback time information and track start position information in the file are registered corresponding to each level. On the receiving side, with regard to audio, it is possible to selectively extract the audio stream of the desired audio track based on this range information.
[0107] Returning to Figure 3, the service receiver 200 receives DASH / MP4, i.e., an MP4 containing an MPD file as a metafile and media streams (media segments) such as video and audio, sent from the service transmission system 100 via an RF transmission path or a communication network transmission path.
[0108] As described above, in addition to a video stream, MP4 has a predetermined number of audio tracks (audio streams) containing multiple groups of encoded data that make up the 3D audio transmission data. An MPD file also contains attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, as well as stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is included.
[0109] Based on the attribute information and stream correspondence information, the service receiver 200 selectively performs decoding processing on audio streams containing encoded data of groups having attributes that match the speaker configuration and user selection information, to obtain 3D audio output.
[0110] [DASH / MP4 generation part of the service transmission system] 17 shows an example of the configuration of the DASH / MP4 generation unit 110 included in the service transmission system 100. The DASH / MP4 generation unit 110 includes a control unit 111, a video encoder 112, an audio encoder 113, and a DASH / MP4 formatter 114.
[0111] The video encoder 112 receives video data SV and encodes the video data SV using MPEG2, H.264 / AVC, H.265 / HEVC, etc. to generate a video stream (video elementary stream). The audio encoder 113 receives immersive audio and speech dialogue object data as audio data SA, along with channel data.
[0112] The audio encoder 113 performs MPEGH encoding on the audio data SA to obtain 3D audio transmission data. This 3D audio transmission data includes channel-encoded data (CD), immersive audio object-encoded data (IAO), and speech dialogue object-encoded data (SDO), as shown in Fig. 5. The audio encoder 113 generates one or more audio streams (audio elementary streams) containing encoded data for multiple groups, four groups in this case (see Figs. 6(a) and (b)).
[0113] The DASH / MP4 formatter 114 generates an MP4 file containing media streams (media segments) such as video and audio content, based on the video stream generated by the video encoder 112 and a predetermined number of audio streams generated by the audio encoder 113. Here, each video and audio stream is stored as a separate track in the MP4 file.
[0114] Furthermore, the DASH / MP4 formatter 114 generates an MPD file using content metadata, segment URL information, etc. In this embodiment, the DASH / MP4 formatter 114 inserts into this MPD file attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, and also inserts stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is included (see FIGS. 11 and 12).
[0115] The operation of the DASH / MP4 generation unit 110 shown in Fig. 17 will be briefly described. Video data SV is supplied to a video encoder 112. This video encoder 112 encodes the video data SV using H.264 / AVC, H.265 / HEVC, or the like, and generates a video stream including encoded video data. This video stream is supplied to a DASH / MP4 formatter 114.
[0116] The audio data SA is supplied to an audio encoder 113. This audio data SA includes channel data and object data for immersive audio and speech dialogue. The audio encoder 113 performs MPEGH encoding on the audio data SA to obtain 3D audio transmission data.
[0117] This 3D audio transmission data includes immersive audio object coded data (IAO) and speech dialogue object coded data (SDO) in addition to channel coded data (CD) (see FIG. 5). The audio encoder 113 then generates one or more audio streams containing four groups of coded data (see FIGS. 6(a) and 6(b)). This audio stream is supplied to a DASH / MP4 formatter 114.
[0118] The DASH / MP4 formatter 114 generates an MP4 file including media streams (media segments) such as video and audio content based on the video stream generated by the video encoder 112 and a predetermined number of audio streams generated by the audio encoder 113. Here, each video and audio stream is stored as a separate track in the MP4 file.
[0119] Furthermore, the DASH / MP4 formatter 114 generates an MPD file using content metadata, segment URL information, etc. Attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data is inserted into this MPD file, as well as stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is included.
[0120] [Example of service receiver configuration] 18 shows an example configuration of a service receiver 200. The service receiver 200 includes a receiving unit 201, a DASH / MP4 analysis unit 202, a video decoder 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. The service receiver 200 also includes container buffers 211-1 to 211-N, a combiner 212, a 3D audio decoder 213, an audio output processing circuit 214, and a speaker system 215. The service receiver 200 also includes a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiving unit 225, and a remote control transmitter 226.
[0121] The CPU 221 controls the operation of each part of the service receiver 200. The flash ROM 222 stores control software and data. The DRAM 223 constitutes a work area for the CPU 221. The CPU 221 loads software and data read from the flash ROM 222 onto the DRAM 223, starts up the software, and controls each part of the service receiver 200.
[0122] The remote control receiving unit 225 receives a remote control signal (remote control code) transmitted from the remote control transmitter 226 and supplies it to the CPU 221. The CPU 221 controls each unit of the service receiver 200 based on the remote control code. The CPU 221, the flash ROM 222, and the DRAM 223 are connected to an internal bus 224.
[0123] The receiving unit 201 receives DASH / MP4, i.e., MP4 containing an MPD file as a metafile and media streams (media segments) such as video and audio, sent from the service transmission system 100 via an RF transmission path or a communication network transmission path.
[0124] In addition to a video stream, MP4 has a predetermined number of audio tracks (audio streams) containing multiple groups of encoded data that make up the 3D audio transmission data. Also, an MPD file contains attribute information indicating the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data, as well as stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is included.
[0125] The DASH / MP4 parser 202 analyzes the MPD file and MP4 received by the receiver 201. The DASH / MP4 parser 202 extracts a video stream from the MP4 and sends it to the video decoder 203. The video decoder 203 performs a decoding process on the video stream to obtain uncompressed video data.
[0126] Video processing circuit 204 performs scaling processing, image quality adjustment processing, etc. on the video data obtained by video decoder 203 to obtain video data for display. Panel drive circuit 205 drives display panel 206 based on the video data for display obtained by video processing circuit 204. Display panel 206 is configured, for example, by an LCD (Liquid Crystal Display), an organic EL display (organic electroluminescence display), or the like.
[0127] The DASH / MP4 analyzer 202 also extracts MPD information included in the MPD file and sends it to the CPU 221. The CPU 221 controls the acquisition process of video and audio streams based on this MPD information. The DASH / MP4 analyzer 202 also extracts metadata from the MP4, such as header information for each track, meta descriptions of the content, and time information, and sends it to the CPU 221.
[0128] CPU21 recognizes an audio track (audio stream) containing encoded data of a group having attributes that match the speaker configuration and viewer (user) selection information, based on attribute information contained in the MPD file that indicates the attributes of the encoded data of each group, and stream correspondence information that indicates in which audio track (audio stream) each group is included.
[0129] In addition, under the control of CPU 221, DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams contained in MP4, the audio streams including encoded data of a group having attributes that match the speaker configuration and viewer (user) selection information, by referring to the level ID and therefore the track ID.
[0130] The container buffers 211-1 to 211-N each take in one of the audio streams extracted by the DASH / MP4 analysis unit 202. Here, the number N of the container buffers 211-1 to 211-N is set to a necessary and sufficient number, but in actual operation, the number of container buffers used is the same as the number of audio streams extracted by the DASH / MP4 analysis unit 202.
[0131] The combiner 212 reads the audio stream for each audio frame from the container buffers 211-1 to 211-N into which each audio stream extracted by the DASH / MP4 analysis unit 202 has been taken, and supplies the audio stream to the 3D audio decoder 213 as coded data of a group having attributes that match the speaker configuration and viewer (user) selection information.
[0132] The 3D audio decoder 213 performs a decoding process on the coded data supplied from the combiner 212 to obtain audio data for driving each speaker of the speaker system 215. Here, the coded data to be decoded can be in three cases: when it includes only channel coded data, when it includes only object coded data, or when it includes both channel coded data and object coded data.
[0133] When decoding channel-encoded data, the 3D audio decoder 213 performs downmixing or upmixing to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When decoding object-encoded data, the 3D audio decoder 213 calculates speaker rendering (mixing ratio for each speaker) based on object information (metadata), and mixes the audio data of the object into audio data for driving each speaker according to the calculation result.
[0134] The audio output processing circuit 214 performs necessary processing such as D / A conversion and amplification on the audio data for driving each speaker obtained by the 3D audio decoder 213, and supplies the audio data to the speaker system 215. The speaker system 215 includes multiple channels, such as 2 channels, 5.1 channels, 7.1 channels, 22.2 channels, etc.
[0135] The operation of the service receiver 200 shown in Fig. 18 will be briefly described. The receiving unit 201 receives DASH / MP4, that is, MP4 including an MPD file as a metafile and media streams (media segments) such as video and audio, sent from the service transmission system 100 via an RF transmission path or a communication network transmission path. The MPD file and MP4 received in this manner are supplied to the DASH / MP4 analyzing unit 202.
[0136] The DASH / MP4 analysis unit 202 analyzes the MPD file and MP4 received by the receiving unit 201. The DASH / MP4 analysis unit 202 then extracts a video stream from the MP4 and sends it to the video decoder 203. The video decoder 203 decodes the video stream to obtain uncompressed video data. This video data is then supplied to the video processing circuit 204.
[0137] The video processing circuit 204 performs scaling, image quality adjustment, and the like on the video data obtained by the video decoder 203 to obtain video data for display. This video data for display is supplied to a panel driving circuit 205. The panel driving circuit 205 drives a display panel 206 based on the video data for display. As a result, an image corresponding to the video data for display is displayed on the display panel 206.
[0138] Furthermore, the DASH / MP4 analysis unit 202 extracts MPD information contained in the MPD file and sends it to the CPU 221. The DASH / MP4 analysis unit 202 also extracts metadata from the MP4, such as header information for each track, meta descriptions of the content, and time information, and sends it to the CPU 221. The CPU 221 recognizes audio tracks (audio streams) that include coded data of groups with attributes that match the speaker configuration and viewer (user) selection information, based on attribute information, stream correspondence information, and the like, that are contained in the MPD file.
[0139] In addition, under the control of the CPU 221, the DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams contained in the MP4, the audio streams including encoded data of a group having attributes that match the speaker configuration and viewer (user) selection information, by referring to the track ID.
[0140] The audio stream extracted by the DASH / MP4 analyzer 202 is loaded into a corresponding one of the container buffers 211-1 to 211-N. The combiner 212 reads the audio stream for each audio frame from each container buffer into which the audio stream has been loaded, and supplies the audio stream to the 3D audio decoder 213 as coded data of a group having attributes that match the speaker configuration and viewer selection information. The 3D audio decoder 213 decodes the coded data supplied from the combiner 212, and obtains audio data for driving each speaker of the speaker system 215.
[0141] When the channel-encoded data is decoded, it is downmixed or upmixed to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When the object-encoded data is decoded, speaker rendering (mixing ratio for each speaker) is calculated based on the object information (metadata), and the audio data of the object is mixed into audio data for driving each speaker according to the calculation result.
[0142] The audio data for driving each speaker obtained by the 3D audio decoder 213 is supplied to an audio output processing circuit 214. In this audio output processing circuit 214, the audio data for driving each speaker is subjected to necessary processing such as D / A conversion and amplification. The processed audio data is then supplied to a speaker system 215. As a result, an acoustic output corresponding to the image displayed on the display panel 206 is obtained from the speaker system 215.
[0143] Fig. 19 shows an example of audio decode control processing by the CPU 221 in the service receiver 200 shown in Fig. 18. The CPU 221 starts processing in step ST1. Then, in step ST2, the CPU 221 detects the receiver speaker configuration, that is, the speaker configuration of the speaker system 215. Next, in step ST3, the CPU 221 obtains selection information regarding audio output by the viewer (user).
[0144] Next, in step ST4, the CPU 221 reads information related to each audio stream in the MPD information, namely, "groupID," "attribute," "switchGroupID," "presetGroupID," and "levelID." Then, in step ST5, the CPU 221 recognizes the track ID of the audio track to which belongs an encoded data group having attributes that match the speaker configuration and viewer selection information.
[0145] Next, in step ST6, the CPU 221 selects each audio track based on the recognition result and imports the stored audio stream into the container buffer. Then, in step ST7, the CPU 221 reads the audio stream for each audio frame from the container buffer and supplies the encoded data of the required group to the 3D audio decoder 213.
[0146] Next, in step ST8, the CPU 221 determines whether or not to decode the coded object data. If the coded object data is to be decoded, in step ST9 the CPU 221 calculates speaker rendering (mixing ratio for each speaker) using azimuth (direction information) and elevation (elevation angle information) based on the object information (metadata). Thereafter, the CPU 221 proceeds to step ST10. Note that if the coded object data is not to be decoded in step ST8, the CPU 221 immediately proceeds to step ST10.
[0147] In step ST10, the CPU 221 determines whether or not to decode the channel-encoded data. If the channel-encoded data is to be decoded, in step ST11 the CPU 221 performs downmixing or upmixing processing for the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. Thereafter, the CPU 221 proceeds to step ST12. Note that if the object-encoded data is not to be decoded in step ST10, the CPU 221 immediately proceeds to step ST12.
[0148] In step ST12, when decoding the coded object data, the CPU 221 mixes the audio data of the object with the audio data for driving each speaker according to the calculation result of step ST9, and then performs dynamic range control. Thereafter, the CPU 21 ends the process in step ST13. Note that if the coded object data is not to be decoded, the CPU 221 skips step ST12.
[0149] 3, the service transmission system 100 inserts attribute information indicating the attributes of each of the multiple groups of encoded data included in a predetermined number of audio streams into an MPD file. This allows the receiving side to easily recognize the attributes of each of the multiple groups of encoded data before decoding the encoded data, and to selectively decode and use only the encoded data of the necessary groups, thereby reducing the processing load.
[0150] 3, the service transmission system 100 inserts stream correspondence information into the MPD file, the stream correspondence information indicating which audio track (audio stream) contains each of the encoded data of multiple groups. This allows the receiving side to easily recognize the audio track (audio stream) containing the encoded data of the required group, thereby reducing the processing load.
[0151] <2. Modifications> In the above-described embodiment, the service receiver 200 is configured to selectively extract, from the multiple audio streams transmitted from the service transmission system 100, audio streams containing coded data of groups having attributes that match the speaker configuration and viewer selection information, and perform decoding processing to obtain audio data for driving a predetermined number of speakers.
[0152] However, as a service receiver, it is also possible to selectively extract one or more audio streams having encoded data of a group with attributes that match the speaker configuration and viewer selection information from multiple audio streams transmitted from the service transmission system 100, reconstruct an audio stream having encoded data of a group with attributes that match the speaker configuration and viewer selection information, and distribute the reconstructed audio stream to devices (including DLNA devices) connected to the local network.
[0153] Fig. 20 shows an example of the configuration of a service receiver 200A that distributes the reconstructed audio stream to devices connected to an in-house network as described above. In Fig. 20, parts corresponding to those in Fig. 18 are given the same reference numerals, and detailed descriptions thereof will be omitted as appropriate.
[0154] Under the control of CPU 221, the DASH / MP4 analysis unit 202 selectively extracts one or more audio streams from a predetermined number of audio streams contained in the MP4, the audio streams containing encoded data of a group having attributes that match the speaker configuration and viewer (user) selection information, by referring to the level ID (level ID) and therefore the track ID (track ID).
[0155] The audio stream extracted by the DASH / MP4 analysis unit 202 is loaded into a corresponding one of the container buffers 211-1 to 211-N. The combiner 212 reads the audio stream for each audio frame from each container buffer into which the audio stream has been loaded, and supplies the read audio stream to the stream reconstruction unit 231.
[0156] The stream reconstruction unit 231 selectively acquires a predetermined group of coded data having attributes that match the speaker configuration and viewer selection information, and reconstructs an audio stream having the coded data of the predetermined group. This reconstructed audio stream is supplied to a distribution interface 232. Then, the audio stream is distributed (transmitted) from the distribution interface 232 to devices 300 connected to the local network.
[0157] This local network connection may include an Ethernet connection, or a wireless connection such as "WiFi" or "Bluetooth." Note that "WiFi" and "Bluetooth" are registered trademarks.
[0158] The device 300 also includes surround speakers, a second display, and an audio output device attached to a network terminal. The device 300 receiving the reconstructed audio stream performs a decoding process similar to that of the 3D audio decoder 213 in the service receiver 200 in Fig. 18 to obtain audio data for driving a predetermined number of speakers.
[0159] The service receiver may also be configured to transmit the reconstructed audio stream to a device connected via a digital interface such as "HDMI (High-Definition Multimedia Interface)", "MHL (Mobile High-definition Link)", or "DisplayPort". Note that "HDMI" and "MHL" are registered trademarks.
[0160] In the above-described embodiment, an example was shown in which attribute information of encoded data of each group was transmitted by providing an "attribute" field (see FIGS. 11 to 13). However, the present technology also includes a method in which a special meaning is defined for the value of a group ID (GroupID) itself between a transmitter and a receiver, so that the type (attribute) of encoded data can be recognized by recognizing a specific group ID. In this case, the group ID functions as an identifier for the group as well as attribute information for the encoded data of that group, making the "attribute" field unnecessary.
[0161] In the above-described embodiment, an example has been shown in which the coded data of the multiple groups includes both channel-coded data and object-coded data (see FIG. 5 ). However, the present technology can also be applied to cases in which the coded data of the multiple groups includes only channel-coded data or only object-coded data.
[0162] The present technology can also be configured as follows. (1) a transmitter for transmitting a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device; an information inserting unit that inserts attribute information indicating attributes of each of the plurality of groups of coded data into the metafile; Transmitting device. (2) The information insertion unit is Stream correspondence relationship information indicating which audio stream contains each of the plurality of groups of coded data is further inserted into the metafile. The transmitting device according to (1) above. (3) The above stream correspondence information is information indicating a correspondence between a group identifier for identifying each of the plurality of groups of coded data and an identifier for identifying each of the predetermined number of audio streams. The transmitting device according to (2) above. (4) The metafile is an MPD file. The transmitting device according to any one of (1) to (3). (5) The information insertion unit is The attribute information is inserted into the metafile using "Supplementary Descriptor". The transmitting device according to (4) above. (6) The transmitting unit is The metafile is transmitted via an RF transmission path or a communication network transmission path. The transmitting device according to any one of (1) to (5). (7) The transmitting unit is and transmitting a container in a predetermined format having a predetermined number of audio streams including the encoded data of the plurality of groups. The transmitting device according to any one of (1) to (6). (8) The container is MP4. The transmitting device according to (7) above. (9) The plurality of groups of coded data include either or both of channel coded data and object coded data. The transmitting device according to any one of (1) to (8). (10) a transmitting step of transmitting, by a transmitting unit, a metafile having meta information for acquiring, at a receiving device, a predetermined number of audio streams including encoded data of a plurality of groups; an information inserting step of inserting attribute information indicating attributes of each of the plurality of groups of coded data into the metafile. Sending method. (11) A receiving unit for receiving a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups by a receiving device, attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; The audio signal processing device further includes a processor for processing the predetermined number of audio streams based on the attribute information. Receiving device. (12) The metafile further includes stream correspondence information indicating which audio stream contains each of the plurality of groups of coded data; The processing unit The predetermined number of audio streams are processed based on the attribute information as well as the stream correspondence information. The receiving device according to (11) above. (13) The processing unit Selectively performing decoding processing on audio streams including encoded data of a group having attributes that match the speaker configuration and user selection information based on the attribute information and the stream correspondence relationship information. The receiving device according to (12) above. (14) The plurality of groups of coded data include either or both of channel coded data and object coded data. The receiving device according to any one of (11) to (13). (15) A receiving step of receiving, by a receiving unit, a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device, attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; The method further includes a processing step of processing the predetermined number of audio streams based on the attribute information. Receiving method. (16) A receiving unit for receiving a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device, attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; a processing unit that selectively acquires a predetermined group of coded data from the predetermined number of audio streams based on the attribute information, and reconstructs an audio stream including the coded data of the predetermined group; a stream sending unit that sends the reconstructed audio stream to an external device. Receiving device. (17) The metafile further includes stream correspondence information indicating which audio stream contains each of the plurality of groups of coded data; The processing unit Selectively acquiring the predetermined group of coded data from the predetermined number of audio streams based on the attribute information and the stream correspondence information. The receiving device according to (16) above. (18) A receiving step of receiving, by a receiving unit, a metafile having meta information for acquiring a predetermined number of audio streams including encoded data of a plurality of groups at a receiving device, attribute information indicating attributes of each of the plurality of groups of coded data is inserted into the metafile; a processing step of selectively acquiring a predetermined group of coded data from the predetermined number of audio streams based on the attribute information, and reconstructing an audio stream including the coded data of the predetermined group; and a stream transmitting step of transmitting the reconstructed audio stream to an external device. Receiving method.
[0163] The main feature of this technology is that it makes it possible to reduce the processing load on the receiving side by inserting into the MPD file attribute information indicating the attributes of each of multiple groups of encoded data contained in a predetermined number of audio streams, and stream correspondence information indicating in which audio track (audio stream) each of the multiple groups of encoded data is contained (see Figures 11, 12, and 17). [Explanation of symbols]
[0164] 10. Transmitting and receiving system 30A, 30B: MPEG-DASH-based streaming system 31···DASH Stream File Server 32 DASH MPD Server 33, 33-1 to 33-N) Service receiver 34···CDN 35, 35-1 to 35-M) Service receiver 36 Broadcasting System 100···Service Transmission System 110...DASH / MP4 generation section 112...Video Encoder 113 Audio Encoder 114···DASH / MP4 Formatter 200···Service receiver 201 Receiving unit 202...DASH / MP4 analysis section 203 Video decoder 204...Video processing circuit 205 Panel drive circuit 206···Display panel 211-1 to 211-N Container buffer 212 Combiner 213···3D Audio Decoder 214...Audio output processing circuit 215···Speaker System 221 CPU 222···Flash ROM 223 DRAM 224 Internal Bus 225 Remote control receiver 226···Remote control transmitter 231 Stream reconstruction unit 232···Distribution Interface 300 devices< / baseurl> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / subrepresentation> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / supplementarydescriptor> < / representation> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / supplementarydescriptor> < / adaptationset> < / baseurl>
Claims
1. an acquisition unit for acquiring a plurality of audio streams; a decoder that performs a decoding process on the plurality of audio streams to obtain audio data for driving each speaker of a speaker system. Information processing device.
2. a combiner for combining the plurality of audio streams; The information processing device according to claim 1 .
3. The codec of the audio stream is MPEG-H 3D Audio; The information processing device according to claim 1 .
4. The audio stream is composed of MPEG Audio Stream Packets. The information processing device according to claim 1 .
5. The MPEG Audio Stream Packet is composed of a header and a payload. The information processing device according to claim 4 .
6. The header includes information such as a packet type, a packet label, and a packet length. The information processing device according to claim 5 .
7. The payload includes "SYNC" information corresponding to a synchronization start code, "Frame" information which is the actual data of the audio stream, and "Config" information which indicates the configuration of the "Frame" information. The information processing device according to claim 5 .
8. The "Frame" information includes channel-encoded data or object-encoded data that constitutes the audio stream. The information processing device according to claim 7 .
9. The object coded data is composed of coded sample data of SCE (Single Channel Element) and metadata for mapping the coded sample data to speakers located at arbitrary positions and rendering the data. The information processing device according to claim 8 .
10. The metadata is included as an extension element (Ext_element), The information processing device according to claim 9 .
11. the object coding data includes speech dialogue object coding data, which is object coding data for a speech language; The information processing device according to claim 8 .
12. the audio stream includes a group ID for identifying a group of the plurality of coded object data; The information processing device according to claim 8 .
13. the audio stream includes attribute information indicating attributes of the group of the plurality of coded object data; The information processing device according to claim 12.
14. The plurality of encoded object data are registered in a switch group (SW Group) and encoded. The information processing device according to claim 8 .
15. The plurality of coded object data include a switch group ID which is an identifier for identifying the switch group. The information processing device according to claim 14.
16. The plurality of encoded object data includes a preset group in which a plurality of the groups are bundled together. The information processing device according to claim 12.
17. The plurality of pieces of coded object data include a preset group ID which is an identifier for identifying the preset group.
17. The information processing device according to claim 16.
18. the plurality of coded object data includes stream correspondence information indicating in which of the audio streams each of the plurality of coded object data is included; The information processing device according to claim 8 .
19. an acquisition unit acquiring a plurality of audio streams; a decoder performing a decoding process on the plurality of audio streams to obtain audio data for driving each speaker of a speaker system. Information processing methods.
20. Computer, a means for obtaining multiple audio streams; The audio data processing unit performs a decoding process on the plurality of audio streams and functions as a means for obtaining audio data for driving each speaker of a speaker system. program.
Citation Information
Patent Citations
Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information
JP2009278381A
Data generation device and data generation method, data processing device and data processing method
JP2012033243A
Systems and tools for improved 3D audio creation and expression
JP2014520491A