Transmission device, transmission method, reception device, and reception method

By inserting attribute and stream mapping information into the layers of containers and audio streams, the problem of heavy processing load on the receiving side when sending multiple audio data items is solved, achieving more efficient processing.

CN113921019BActive Publication Date: 2025-11-07SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111172292.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-09-30
Filing Date
2015-09-16
Publication Date
2025-11-07
Estimated Expiration
2035-10-16

AI Technical Summary

Technical Problem

When sending multiple audio data items, the processing load on the receiving side is heavy, and existing technologies are unable to effectively reduce it.

Method used

By inserting information representing the corresponding attributes of encoded data items for multiple groups into the layers of the container and the audio stream, including attribute information and stream correspondence information, the processing load on the receiving side is reduced.

Benefits of technology

On the receiving side, it is easier to identify and selectively decode the necessary encoded data items, reducing processing load and improving processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113921019B_ABST
    Figure CN113921019B_ABST
Patent Text Reader

Abstract

The present application provides a transmitting apparatus, a transmitting method, a receiving apparatus and a receiving method. The transmitting apparatus includes a transmitting unit configured to transmit a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items; and an information inserting unit configured to insert attribute information indicating respective attributes of the plurality of groups of encoded data items into a layer of the container and / or a layer of the audio streams.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese national phase application of international application No. PCT / JP2015 / 076259, filed on September 16, 2015, with the title of "Transmission device, transmission method, reception device, and reception method", the entry date into the Chinese national phase of which is March 23, 2017, the application number of which is 201580051430.9, and the title of which is "Transmission device, transmission method, reception device, and reception method". TECHNICAL FIELD

[0002] The present technology relates to a transmission device, a transmission method, a reception device, and a reception method, and more particularly, to a transmission device and the like for transmitting a plurality of audio data items. BACKGROUND

[0003] In the related art, as a three-dimensional (3D) acoustic technology, a technology has been proposed that maps and renders encoded sample data items based on metadata items to speakers at any position (for example, refer to Patent Literature 1).

[0004] LIST OF CITATIONS

[0005] PATENT LITERATURE

[0006] Patent Literature 1: Translation of PCT International Application Publication No. 2014-520491. SUMMARY

[0007] TECHNICAL PROBLEM

[0008] It is conceivable that, by transmitting object encoded data items including encoded sample data items and metadata items together with channel encoded data items of 5.1 channels, 7.1 channels, or the like, sound with enhanced realism can be reproduced at the reception side.

[0009] An object of the present technology is to reduce the processing load at the reception side in a case where a plurality of audio data items are transmitted.

[0010] SOLUTION TO PROBLEM

[0011] The concept of the present technology is a transmission device including a transmission unit that transmits a container of a predetermined format having a predetermined number of audio streams including a plurality of groups of encoded data items, and an information insertion unit that inserts information indicating respective attributes of the plurality of groups of encoded data items into a layer of the container and / or a layer of the audio streams.

[0012] In the present technology, the transmission unit transmits a container of a predetermined format having a predetermined number of audio streams including a plurality of groups of encoded data items. For example, the plurality of groups of encoded data items can include one or both of channel encoded data items and object encoded data items.

[0013] The attribute information indicating the respective attributes of the encoded data items of the plurality of groups is inserted by the information insertion unit into a layer of a container and / or a layer of an audio stream. For example, the container can be a transport stream (MPEG2-TS) applicable in a digital broadcasting standard. Further, for example, the container can be an MP4 format or other format for Internet transmission.

[0014] Thus, in the present technology, the attribute information indicating the respective attributes of the encoded data items of the plurality of groups included in the predetermined number of audio streams is inserted into a layer of a container and / or a layer of an audio stream. Therefore, at the reception side, the respective attributes of the encoded data items of the plurality of groups can be easily recognized before the encoded data items are decoded, and only the encoded data items of the necessary groups can be selectively decoded and used, whereby the processing load can be reduced.

[0015] Further, in the present technology, for example, the information insertion unit can also insert stream correspondence relationship information into a layer of a container and / or a layer of an audio stream, the stream correspondence relationship information indicating in which audio stream the encoded data items of the plurality of groups are respectively included. In this way, the audio stream including the encoded data items of the necessary groups can be easily recognized at the reception side, and thus the processing load can be reduced.

[0016] In this case, for example, the container can be an MPEG2-TS, and in a case where the attribute information and the stream identifier information are inserted into the container, the information insertion unit can insert the attribute information and the stream identifier information into an audio elementary stream loop corresponding to at least one or a plurality of audio streams among the predetermined number of audio streams present under a program map table.

[0017] Further, in this case, for example, in a case where the attribute information and the stream correspondence relationship information are inserted into the audio stream, the information insertion unit can insert the attribute information and the stream correspondence relationship information into a PES payload of a PES packet in at least one or a plurality of audio streams among the predetermined number of audio streams.

[0018] For example, the stream correspondence relationship information can be information indicating a correspondence relationship between a group identifier identifying each of the encoded data items of the plurality of groups and a stream identifier identifying each of the predetermined number of audio streams. In this case, for example, the information insertion unit can insert stream identifier information into a layer of a container and / or a layer of an audio stream, the stream identifier information indicating the stream identifier of each of the predetermined number of audio streams.

[0019] For example, the container can be an MPEG2-TS, and in a case where the stream identifier information is inserted into the container, the information insertion unit can insert the stream identifier information into an audio elementary stream loop corresponding to each of the predetermined number of audio streams present under the program map table. Also, for example, in a case where the stream identifier information is inserted into the audio stream, the information insertion unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.

[0020] Also, for example, the stream correspondence relationship information can be information indicating a correspondence relationship between a group identifier identifying each of the plurality of groups of encoded data items and a packet identifier added when each of the predetermined number of audio streams is packetized. Also, for example, the stream correspondence relationship information can be information indicating a correspondence relationship between a group identifier identifying each of the plurality of groups of encoded data items and type information indicating a stream type of each of the predetermined number of audio streams.

[0021] Further, another concept of the present technology is a receiving apparatus including: a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, attribute information indicating respective attributes of the plurality of groups of encoded data items being inserted into a layer of the container and / or a layer of the audio streams; and a processing unit that processes the predetermined number of audio streams included in the received container based on the attribute information.

[0022] In the present technology, the receiving unit receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items. For example, the plurality of groups of encoded data items can include one or both of channel encoded data items and object encoded data items. Attribute information indicating respective attributes of the plurality of groups of encoded data items is inserted into a layer of the container and / or a layer of the audio streams. The processing unit processes the predetermined number of audio streams included in the received container based on the attribute information.

[0023] Thus, in the present technology, the predetermined number of audio streams contained in the received container is processed based on attribute information indicating respective attributes of the plurality of groups of encoded data items inserted into a layer of the container and / or a layer of the audio streams. Therefore, it is possible to selectively decode and use only necessary groups of encoded data items, and thus it is possible to reduce a processing load.

[0024] Further, in the present technology, for example, stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is contained is further inserted into a layer of the container and / or a layer of the audio streams. The processing unit can process the predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information. In this case, it is possible to easily identify an audio stream containing a necessary group of encoded data items, and thus it is possible to reduce a processing load.

[0025] In addition, in the present technology, for example, based on the attribute information and the stream correspondence relationship information, the processing unit can perform a selective decoding process on an audio stream including an encoded data item of a group having an attribute suitable for a speaker configuration and user selection information.

[0026] Further, another concept of the present technology is a reception device including: a reception unit that receives a container of a predetermined format having a predetermined number of audio streams including encoded data items of a plurality of groups; attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio streams; a processing unit that selectively acquires encoded data items of a predetermined group from the predetermined number of audio streams included in the received container based on the attribute information, and reconfigures an audio stream including the encoded data items of the predetermined group; and a stream transmission unit that transmits the audio stream reconfigured by the processing unit to an external device.

[0027] In the present technology, the reception unit receives a container of a predetermined format having a predetermined number of audio streams including encoded data items of a plurality of groups. Attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio streams. The processing unit selectively acquires encoded data items of a predetermined group from the predetermined number of audio streams included in the received container based on the attribute information, and reconfigures an audio stream including the encoded data items of the predetermined group. The stream transmission unit transmits the audio stream reconfigured by the processing unit to an external device.

[0028] Thus, in the present technology, based on the attribute information indicating respective attributes of the encoded data items of the plurality of groups inserted into a layer of the container and / or a layer of the audio streams, encoded data items of a predetermined group are selectively acquired from the predetermined number of audio streams, and an audio stream to be transmitted to an external device is reconfigured. It is possible to easily acquire encoded data items of a necessary group, and thus it is possible to reduce a processing load.

[0029] In addition, in the present technology, for example, stream correspondence relationship information indicating in which audio stream each of the encoded data items of the plurality of groups is included is further inserted into a layer of the container and / or a layer of the audio streams. The processing unit can selectively acquire encoded data items of a predetermined group from the predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information. In this case, it is possible to easily identify an audio stream including encoded data items of a predetermined group, and thus it is possible to reduce a processing load.

[0030] Advantages of the Invention

[0031] According to this technology, the processing load at the receiving end can be reduced when transmitting multiple audio data items. It should be noted that the effects described in this specification are illustrative rather than limiting, and may have additional effects. Attached Figure Description

[0032] [ Figure 1 [ ] is a block diagram illustrating a configuration example of a transmitting / receiving system as an implementation.

[0033] [ Figure 2 [Illustration] is a diagram showing the structure of an audio frame (1024 samples) in a 3D audio transmission data item.

[0034] [ Figure 3 [This is a diagram showing a configuration example of a 3D audio transmission data item.]

[0035] [ Figure 4 [Illustrated diagram showing a configuration example of an audio frame in the case of sending 3D audio data items through one or more streams.]

[0036] [ Figure 5 [Illustration] is a diagram illustrating an example of group partitioning in the case of sending 3D audio data items via two streams.

[0037] [ Figure 6 [] is a diagram showing the correspondence between groups and flows in a group partitioning instance (two partitions).

[0038] [ Figure 7 [Illustration] is a diagram illustrating an example of group partitioning in the case of sending 3D audio data items via two streams.

[0039] [ Figure 8 [] is a diagram showing the correspondence between groups and flows in a group partitioning instance (two partitions).

[0040] [ Figure 9 [ ] is a block diagram showing a configuration instance of a stream generation unit contained in a service sender.

[0041] [ Figure 10 [] is a diagram showing a configuration instance of a 3D audio stream configuration descriptor.

[0042] [ Figure 11 This displays the main information in a configuration instance of the 3D audio stream configuration descriptor.

[0043] [ Figure 12 [] is a diagram showing the types of content defined in "contentKind".

[0044] [ Figure 13FIG. 1 is a diagram showing a configuration example of a 3D audio stream ID descriptor and contents of main information in the configuration example.

[0045] [ Figure 14 FIG. 2 is a diagram showing a configuration example of a transport stream.

[0046] [ Figure 15 FIG. 3 is a block diagram showing a configuration example of a service receiver.

[0047] [ Figure 16 FIG. 4 is a diagram showing an example of a received audio stream.

[0048] [ Figure 17 FIG. 5 is a diagram schematically showing a decoding process in a case where descriptor information is not present within an audio stream.

[0049] [ Figure 18 FIG. 6 is a diagram showing a configuration example of an audio access unit (audio frame) of an audio stream in a case where descriptor information is not present within the audio stream.

[0050] [ Figure 19 FIG. 7 is a diagram schematically showing a decoding process in a case where descriptor information is present within an audio stream.

[0051] [ Figure 20 FIG. 8 is a diagram showing a configuration example of an audio access unit (audio frame) of an audio stream in a case where descriptor information is present within the audio stream.

[0052] [ Figure 21 FIG. 9 is a diagram showing another configuration example of an audio access unit (audio frame) of an audio stream in a case where descriptor information is present within the audio stream.

[0053] [ Figure 22 FIG. 10 is a flowchart (1 / 2) showing an example of an audio decoding control process of a CPU in a service receiver.

[0054] [ Figure 23 FIG. 11 is a flowchart (2 / 2) showing an example of an audio decoding control process of a CPU in a service receiver.

[0055] [ Figure 24 FIG. 12 is a block diagram showing another configuration example of a service receiver. DETAILED DESCRIPTION

[0056] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The present specification will proceed in the following order.

[0057] 1. Embodiment

[0058] 2. Modified Example

[0059] <1. Embodiment>

[0060] [Configuration example of transmission / reception system]

[0061] Figure 1 A configuration example of a transmission / reception system 10 as an embodiment is shown. The transmission / reception system 10 includes a service transmitter 100 and a service receiver 200. The service transmitter 100 transmits a transport stream TS through a broadcast wave or a network packet. The transport stream TS has a video stream and an audio stream of a predetermined number (i.e., one or more) of encoded data items including a plurality of groups.

[0062] Figure 2 A structure of an audio frame (1024 samples) in a 3D audio transmission data item handled in the present embodiment is shown. The audio frame includes a plurality of MPEG audio stream packets. Each MPEG audio stream packet includes a header and a payload.

[0063] The header has information on a packet type, a packet label, a packet length, and the like. Information set by the packet type in the header is set on the payload. In the payload information, there are "SYNC" information corresponding to a synchronization start code, "Frame" which is actual data of a 3D audio transmission data item, and "Config" indicating a configuration of the "Frame".

[0064] The "Frame" includes a channel encoded data item and an object encoded data item configuring the 3D audio transmission data item. Here, the channel encoded data item includes encoded sample data items such as SCE (Single Channel Element), CPE (Channel Pair Element), LFE (Low Frequency Element), and the like. Further, the object encoded data item includes an encoded sample data item of SCE (Single Channel Element), and a metadata item for mapping the encoded sample data item of SCE to a speaker existing at any position and rendering the encoded sample data item of SCE. The metadata item is included as an Ext_element.

[0065] Figure 3 A configuration example of a 3D audio transmission data item is shown. In this example, the 3D audio transmission data item includes one channel encoded data item and two object encoded data items. The channel encoded data item is a 5.1 channel channel encoded data item (CD), and includes each encoded sample data item of SCE1, CPE1.1, CPE1.2, and LFE1.

[0066] Two object encoded data items are an immersive audio object (IAO) and a speech dialogue object (SDO) encoded data item. The immersive audio object encoded data item is an object encoded data item for immersive sound, and includes an encoded sample data item SCE2 and a metadata item EXE_E1 (object metadata) 2 for mapping the encoded sample data item SCE2 to a speaker existing at an arbitrary position and rendering the encoded sample data item SCE2.

[0067] The speech dialogue object encoded data item is an object encoded data item for speech language. In this example, there are speech dialogue object encoded data items corresponding to each of a first language and a second language. The speech dialogue object encoded data item corresponding to the first language includes an encoded sample data item SCE3 and a metadata item EXE_E1 (object metadata) 3 for mapping the encoded sample data item SCE3 to a speaker existing at an arbitrary position and rendering the encoded sample data item SCE3. In addition, the speech dialogue object encoded data item corresponding to the second language includes an encoded sample data item SCE4 and a metadata item EXE_E1 (object metadata) 4 for mapping the encoded sample data item SCE4 to a speaker existing at an arbitrary position and rendering the encoded sample data item SCE4.

[0068] The encoded data items are classified by the concept of groups (Group) based on the type. In the example shown, the encoded channel data items of 5.1 channels are classified as Group 1, the immersive audio object encoded data items are classified as Group 2, the speech dialogue object encoded data items according to the first language are classified as Group 3, and the speech dialogue object encoded data items according to the second language are classified as Group 4.

[0069] In addition, the groups selected among the groups at the receiving side are registered as a switching group (SW group), and are encoded. Furthermore, the groups are bundled as preset groups, so that reproduction corresponding to use cases can be performed. In the example shown, Group 1, Group 2, and Group 3 are bundled as a preset group 1, and Group 1, Group 2, and Group 4 are bundled as a preset group 2.

[0070] Returning to Figure 1 The service transmitter 100 transmits a 3D audio transmission data item including the encoded data items of the plurality of groups by one stream or a plurality of streams (streams), as described above.

[0071] Figure 4 (a) schematically shows a configuration example in the case where the 3D audio transmission data item in Figure 3 is transmitted by one stream (main stream). In this case, one stream includes the channel encoded data item (CD), the immersive audio object encoded data item (IAO), and the speech dialogue object encoded data item (SDO), and "SYNC" and "Config".

[0072] Figure 4 (b) shows a configuration example in a case where the 3D audio transmission data items in Figure 3 are transmitted through a plurality of streams. In this case, the main stream includes the channel coded data items (CD) and the immersive audio object coded data items (IAO) along with "SYNC" and "Config". Further, the sub stream includes the speech dialogue object coded data items (SDO) along with "SYNC" and "Config".

[0073] Figure 5 shows a group division example in a case where the 3D audio transmission data items in Figure 3 are transmitted through two kinds of streams. In this case, the main stream includes the channel coded data items (CD) classified into group 1 and the immersive audio object coded data items (IAO) classified into group 2. Further, the sub stream includes the speech dialogue object coded data items (SDO) according to the first language classified into group 3, and the speech dialogue object coded data items (SDO) according to the second language classified into group 4.

[0074] Figure 6 shows a diagram of the correspondence relation between the groups and the streams in the group division example (two divisions) in Figure 5 . Here, the group ID is an identifier for identifying a group. The attribute shows the attribute of the coded data item of each group. The switch group ID is an identifier for identifying a switch group. The preset group ID is an identifier for identifying a preset group. The stream ID (sub stream ID) is an identifier for identifying a sub stream. The kind shows the kind of the content of each group.

[0075] The shown correspondence relation shows that the coded data items belonging to group 1 are channel coded data items, do not constitute a switch group, and are included in stream 1. Further, the shown correspondence relation shows that the coded data items belonging to group 2 are object coded data items for immersive sound (immersive audio object coded data items), do not constitute a switch group, and are included in stream 1.

[0076] Further, the shown correspondence relation shows that the coded data items belonging to group 3 are object coded data items for speech language according to the first language (speech dialogue object coded data items), constitute a switch group 1, and are included in stream 2. Further, the shown correspondence relation shows that the coded data items belonging to group 4 are object coded data items for speech language according to the second language (speech dialogue object coded data items), constitute a switch group 1, and are included in stream 2.

[0077] Further, the illustrated correspondence shows that the preset group 1 includes the group 1, the group 2, and the group 3. Further, the illustrated correspondence shows that the preset group 2 includes the group 1, the group 2, and the group 4.

[0078] Figure 7 An example of the group division in the case where the 3D audio transmission data item is transmitted through two streams is illustrated. In this case, the main stream includes a channel coded data item (CD) classified as the group 1 and an immersive audio object coded data item (IAO) classified as the group 2.

[0079] Further, the main stream includes an SAOC (Spatial Audio Object Coding) object coded data item classified as the group 5 and an HOA (Higher Order Ambisonics) object coded data item classified as the group 6. The SAOC object coded data item is a data item that utilizes the characteristics of the object data item, and performs higher compression of the object coding. The purpose of the HOA object coded data item is to reproduce the sound direction from the sound incoming direction of the microphone to the auditory position by a technology that processes the 3D sound as the entire sound field.

[0080] The sub stream includes a speech dialogue object coded data item (SDO) according to a first language classified as the group 3, and a speech dialogue object coded data item (SDO) according to a second language classified as the group 4. Further, the sub stream includes a first audio description coded data item classified as the group 7 and a second audio description coded data item classified as the group 8. The audio description coded data item is used to explain the content (mainly video) in sound, and is transmitted separately from the general sound, for people with mainly visual disabilities.

[0081] Figure 8 An example of the group division in the case where the 3D audio transmission data item is transmitted through two streams is illustrated. In this case, the main stream includes a channel coded data item (CD) classified as the group 1 and an immersive audio object coded data item (IAO) classified as the group 2. Figure 7 The illustrated correspondence shows that the encoded data item belonging to the group 1 is a channel coded data item, does not constitute a switching group, and is included in the stream 1. Further, the illustrated correspondence shows that the encoded data item belonging to the group 2 is an object coded data item for immersive sound (immersive audio object coded data item), does not constitute a switching group, and is included in the stream 1.

[0082] Further, the illustrated correspondence shows that the encoded data item belonging to the group 3 is an object coded data item for a speech language according to a first language (speech dialogue object coded data item), constitutes a switching group 1, and is included in the stream 2. Further, the illustrated correspondence shows that the encoded data item belonging to the group 4 is an object coded data item for a speech language according to a second language (speech dialogue object coded data item), constitutes a switching group 1, and is included in the stream 2.

[0083] Further, the illustrated correspondence shows that the encoded data items belonging to group 5 are SAOC object encoded data items, constitute a switching group 2, and are included in stream 1. Further, the illustrated correspondence shows that the encoded data items belonging to group 6 are HAO object encoded data items, constitute a switching group 2, and are included in stream 1.

[0084] Further, the illustrated correspondence shows that the encoded data items belonging to group 5 are first audio description object encoded data items, constitute a switching group 3, and are included in stream 2. Further, the illustrated correspondence shows that the encoded data items belonging to group 8 are second audio description object encoded data items, constitute a switching group 3, and are included in stream 2.

[0085] Further, the illustrated correspondence shows that preset group 1 includes group 1, group 2, group 3, and group 7. Further, the illustrated correspondence shows that preset group 2 includes group 1, group 2, group 4, and group 8.

[0086] Returning to Figure 1 The service transmitter 100 inserts attribute information indicating respective attributes of the encoded data items of the plurality of groups included in the 3D audio transmission data item into a layer of the container. Further, the service transmitter 100 inserts stream correspondence information indicating in which audio stream each of the encoded data items of the plurality of groups is respectively included into a layer of the container. In this embodiment, for example, the stream correspondence information is regarded as information indicating a correspondence between a group ID and a stream identifier.

[0087] The service transmitter 100 inserts the attribute information and the stream correspondence information as descriptors into an audio elementary stream loop corresponding to one or more of a predetermined number of audio streams existing, for example, under a program map table (PMT: Program Map Table).

[0088] Further, the service transmitter 100 inserts stream identifier information indicating respective stream identifiers of the predetermined number of audio streams into a layer of the container. The service transmitter 100, for example, inserts the stream identifier information as descriptors into an audio elementary stream loop corresponding to the respective predetermined number of audio streams existing under a program map table (PMT: Program Map Table).

[0089] Further, the service transmitter 100 inserts attribute information indicating respective attributes of the encoded data items of the plurality of groups included in the 3D audio transmission data item into a layer of the audio stream. Further, the service transmitter 100 inserts stream correspondence information indicating in which audio stream each of the encoded data items of the plurality of groups is respectively included into a layer of the audio stream. The service transmitter 100, for example, inserts the attribute information and the stream correspondence information into a PES payload of a PES packet of one or more of the predetermined number of audio streams.

[0090] Further, the service transmitter 100 inserts stream identifier information indicating respective stream identifiers of the predetermined number of audio streams into the layer of the audio streams. The service transmitter 100, for example, inserts the stream identifier information into the PES payload of the respective PES packets of the predetermined number of audio streams.

[0091] By inserting "Desc", i.e., descriptor information, between "SYNC" and "Config", the service transmitter 100 inserts information into the layer of the audio streams as shown in (a) and (b). Figure 4

[0092] As described above, although the present embodiment shows that each piece of information (attribute information, stream correspondence relationship information, stream identifier information) is inserted into both the layer of the container and the layer of the audio streams, it is envisaged that each piece of information is inserted into only the layer of the container or only the layer of the audio streams.

[0093] The service receiver 200 receives the transport stream TS transmitted from the service transmitter 100 through a broadcast wave or a network packet. As described above, the transport stream TS includes the predetermined number of audio streams in addition to the video stream, and the audio stream includes a plurality of groups of encoded data items configuring 3D audio transmission data items.

[0094] Attribute information indicating respective attributes of the plurality of groups of encoded data items included in the 3D audio transmission data items is inserted into the layer of the container and / or the layer of the audio streams, and stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included is inserted.

[0095] Based on the attribute information and the stream correspondence relationship information, the service receiver 200 performs selective decoding processing on the audio stream including the encoded data items of the groups having attributes suitable for the speaker configuration and the user selection information, and acquires an audio output of the 3D audio.

[0096] [Stream generating unit of service transmitter]

[0097] Figure 9 A configuration example of a stream generating unit 110 included in the service transmitter 100 is shown. The stream generating unit 110 includes a video encoder 112, an audio encoder 113, and a multiplexer 114. Here, the audio transmission data item includes one encoded channel data item and two object encoded data items as shown in (a) and (b) in Figure 3

[0098] The video encoder 112 inputs video data SV, encodes the video data SV, and generates a video stream (video elementary stream). The audio encoder 113 inputs immersive audio and a voice dialogue object data item as audio data items SA together with a channel data item.

[0099] ​​The audio encoder 113 encodes the audio data item SA and acquires a 3D audio transmission data item. As shown in Figure 3 The 3D audio transmission data item includes a channel encoded data item (CD), an immersive audio object encoded data item (IAO), and a speech dialogue object encoded data item (SDO).

[0100] The audio encoder 113 generates one or more audio streams (audio elementary streams) including a plurality of groups (four groups here) of encoded data items (see Figure 4 (a), (b)). At this time, as described above, the audio encoder 113 inserts descriptor information ("Desc") including attribute information, stream correspondence relationship information, and stream identifier information between "SYNC" and "Config".

[0101] The multiplexer 114 PES-packetizes the video stream output from the video encoder 112 and the predetermined number of audio streams output from the audio encoder 113, further transport stream-packetizes the audio streams so as to be multiplexed, and acquires a transport stream TS as a multiplexed stream.

[0102] Further, the multiplexer 114 inserts attribute information indicating respective attributes of the encoded data items of the plurality of groups and stream correspondence relationship information indicating in which audio stream each of the encoded data items in the plurality of groups is included, under a program map table (PMT). The multiplexer 114 inserts the information using a 3D audio stream configuration descriptor (3Daudio_stream_config_descriptor) into an audio elementary stream loop corresponding to at least one or a plurality of audio streams among the predetermined number of audio streams. Details of the descriptor will be described later.

[0103] Further, the multiplexer 114 inserts stream identifier information indicating respective stream identifiers of the predetermined number of audio streams, under the program map table (PMT). The multiplexer 114 inserts the information using a 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) into an audio elementary stream loop corresponding to the respective predetermined number of audio streams. Details of the descriptor will be described later.

[0104] The operation of the stream generation unit 110 shown in Figure 9 will be briefly described. The video data item is supplied to the video encoder 112. In the video encoder 112, the video data item SV is encoded, and a video stream including the encoded video data item is generated. The video stream is supplied to the multiplexer 114.

[0105] The audio data item SA is supplied to the audio encoder 113. The audio data item SA includes a channel data item, an immersive audio, and an object data item of a speech dialogue. In the audio encoder 113, the audio data item SA is encoded, and a 3D audio transmission data item is obtained.

[0106] The 3D audio transmission data item includes, in addition to the channel encoded data item (CD), an immersive audio object encoded data item (IAO) and a speech dialogue object encoded data item (SDO) (see Figure 3 ). In the audio encoder 113, one or more audio streams including four groups of encoded data items are generated (see Figure 4 (a), (b)).

[0107] At this time, as described above, the audio encoder 113 inserts descriptor information ("Desc") including attribute information, stream correspondence relationship information, and stream identifier information between "SYNC" and "Config".

[0108] The video stream generated in the video encoder 112 is supplied to the multiplexer 114. In addition, the audio stream generated in the audio encoder 113 is supplied to the multiplexer 114. In the multiplexer 114, the stream supplied from each encoder passes through PES packetization and transport packetization so as to be multiplexed, and a transport stream TS as a multiplexed stream is acquired.

[0109] Further, in the multiplexer 114, for example, a 3D audio stream configuration descriptor is inserted into an audio elementary stream loop corresponding to at least one or more of a predetermined number of audio streams. The descriptor includes attribute information indicating respective attributes of the plurality of groups of encoded data items, and stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included.

[0110] Further, in the multiplexer 114, a 3D audio stream ID descriptor is inserted into an audio elementary stream loop corresponding to the respective predetermined number of audio. The descriptor includes stream identifier information indicating respective stream identifiers of the predetermined number of audio streams.

[0111] [Details of 3D audio stream configuration descriptor]

[0112] Figure 10 An example of the structure (syntax) of the 3D audio stream configuration descriptor is shown. Further, Figure 11 The contents (semantics) of the main information in the configuration example are shown.

[0113] The 8-bit field "descriptor_tag" indicates the descriptor type. Here, it shows that it is a 3D audio stream configuration descriptor. The 8-bit field "descriptor_length" indicates the descriptor length (size), and shows the number of subsequent bytes as the descriptor length.

[0114] 8-bit field "NumOfGroups, N" indicates the number of groups. 8-bit field "NumOfPresetGroups, P" indicates the number of preset groups. For the number of groups, 8-bit field "group ID", 8-bit field "attribute of group ID", 8-bit field "SwitchGroupID" and 8-bit field "audio_streamID" are repeated.

[0115] Field "group ID" indicates the identifier of a group. Field "attribute of group ID" indicates the relevant attribute of the encoded data item of a group. Field "SwitchGroupID" is the identifier indicating the switch group to which the relevant group belongs. "0" indicates that it does not belong to any switch group. Non-"0" indicates the switch group to which it belongs. 8-bit field "contentKind" indicates the kind of the content of a group. "audio_streamID" is the identifier indicating the audio stream including the relevant group. Figure 12 The kind of the content defined in "contentKind" is shown.

[0116] For the number of preset groups, 8-bit field "presetGroupID" and 8-bit field "NumOfGroups_in_preset, R" are repeated. Field "presetGroupID" is the identifier indicating the bundle of preset groups. Field "NumOfGroups_in_preset, R" indicates the number of groups belonging to a preset group. 8-bit field "group ID" is repeated for each preset group (for the number of groups belonging thereto) and shows the groups belonging to a preset group. The descriptor can be set under an extended descriptor.

[0117] [Details of 3D audio stream ID descriptor]

[0118] Figure 13 (a) A configuration example (syntax) of 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) is shown. Figure 13 (b) The content (semantics) of the main information in the configuration example is shown.

[0119] 8-bit field "descriptor_tag" indicates the descriptor type. Here, it shows that it is a 3D audio stream ID descriptor. 8-bit field "descriptor_length" indicates the descriptor length (size) and indicates the number of subsequent bytes as the descriptor length. 8-bit field "audio_streamID" indicates the identifier of an audio stream. The descriptor can be set under an extended descriptor.

[0120] [Configuration of transport stream TS]

[0121] Figure 14 A configuration example of a transport stream TS is shown. The configuration example corresponds to a case where 3D audio transmission data items are transmitted by two streams (see Fig. 1). In the configuration example, there is a video stream PES packet "video PES" identified by PID 1. In addition, in the configuration example, there are two audio stream PES packets "audio PES" identified by PIDs 2 and 3, respectively. The PES packet includes a PES header and a PES payload. DTS and PTS time stamps are inserted into the PES header. At the time of multiplexing, the PIDs 2 and 3 time stamps are matched to provide precision, whereby synchronization therebetween can be ensured throughout the system. Figure 5 ). In the configuration example, there is a video stream PES packet "video PES" identified by PID 1. In addition, in the configuration example, there are two audio stream PES packets "audio PES" identified by PIDs 2 and 3, respectively. The PES packet includes a PES header and a PES payload. DTS and PTS time stamps are inserted into the PES header. At the time of multiplexing, the PIDs 2 and 3 time stamps are matched to provide precision, whereby synchronization therebetween can be ensured throughout the system.

[0122] Here, the audio stream PES packet "audio PES" identified by PID 2 includes a channel coded data item (CD) classified into group 1 and an immersive audio object coded data item (IAO) classified into group 2. In addition, the audio stream PES packet "audio PES" identified by PID 3 includes a speech dialogue object coded data item (SDO) according to a first language classified into group 3, and a speech dialogue object coded data item (SDO) according to a second language classified into group 4.

[0123] In addition, the transport stream TS includes a PMT (Program Map Table) as PSI (Program Specific Information). The PSI is information describing to which program each elementary stream included in the transport stream belongs. A program loop in which information about the entire program is described exists in the PMT.

[0124] In addition, an elementary stream loop having information about each elementary stream exists in the PMT. In the configuration example, there is a video elementary stream loop (video ES loop) corresponding to the video stream, and there are audio elementary stream loops (audio ES loops) corresponding to the two audio streams.

[0125] In the video elementary stream loop (video ES loop), information about the stream type, PID (Packet Identifier), etc. is set corresponding to the video stream, and a descriptor describing information about the video stream is also set. The value of the video stream "Stream_type" is set to "0x24", and the PID information indicates PID 1 added to the video stream PES packet "video PES", as described above. As one of the descriptors, an HEVC descriptor is set.

[0126] At each audio base stream loop (audio ES loop), information such as stream type, PID (group identifier), etc., is set for the corresponding audio stream, and descriptors describing information related to the audio stream are also set. PID2 is the main audio stream, and the value of "Stream_type" is set to "0x2C". The PID information indicates the PID added to the audio stream PES packet "audio PES", as described above. Additionally, PID3 is a sub-audio stream, and the value of "Stream_type" is set to "0x2D". The PID information indicates the PID added to the audio stream PES packet "audio PES", as described above.

[0127] In addition, at each Audio ES loop, set both the 3D audio stream configuration descriptor and the 3D audio stream ID descriptor mentioned above.

[0128] Additionally, descriptor information is inserted into the PES payload of each audio base stream's PES packet. This descriptor information is the "Desc" inserted between "SYNC" and "Config" as described above (see...). Figure 4 Assuming the information included in the 3D audio stream configuration descriptor is represented as D1, and the information included in the 3D audio stream ID descriptor is represented as D2, then the descriptor information includes "D1+D2" information.

[0129] [Service Receiver Configuration Example]

[0130] Figure 15 An example configuration of the service receiver 200 is shown. The service receiver 200 includes a receiving unit 201, a demultiplexer 202, a video decoder 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. Furthermore, the service receiver 200 includes multiplexing buffers 211-1 to 211-N, a combiner 212, a 3D audio decoder 213, a sound output processing circuit 214, and a speaker system 215. Additionally, the service receiver 200 includes a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiving unit 225, and a remote control transmitter 226.

[0131] CPU 221 controls the operation of each unit in service receiver 200. Flash ROM 222 stores control software and saves data. DRAM 223 configures the working area of ​​CPU 221. CPU 221 decompresses software or data read from flash ROM 222 onto DRAM 223 to start the software and control each unit in service receiver 200.

[0132] The remote control receiving unit 225 receives a remote control signal (remote control code) transmitted from the remote control transmitter 226, and supplies it to the CPU 221. The CPU 221 controls each unit in the service receiver 200 based on the remote control code. The CPU 221, the flash ROM 222, and the DRAM 223 are connected to the internal bus 224.

[0133] The receiving unit 201 receives a transport stream TS transmitted from the service transmitter 100 on a broadcast wave or a network packet. The transport stream TS includes a predetermined number of audio streams in addition to a video stream, which includes a plurality of groups of encoded data items configured 3D audio transmission data items.

[0134] Figure 16 An example of the received audio stream is shown. Figure 16 (a) An example of one stream (main stream) is shown. The stream includes channel encoded data items (CD), immersive audio object encoded data items (IAO), and speech dialogue object encoded data items (SDO) along with "SYNC" and "Config". The stream is identified by PID 2.

[0135] In addition, between "SYNC" and "Config", descriptor information ("Desc") is included. Attribute information indicating respective attributes of the plurality of groups of encoded data items, stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included, and stream identifier information indicating a self stream identifier are inserted into the descriptor information.

[0136] Figure 16 (b) An example of two streams is shown. A main stream identified by PID 2 includes channel encoded data items (CD) and immersive audio object encoded data items (IAO) along with "SYNC" and "Config". In addition, a sub stream identified by PID 3 includes speech dialogue object encoded data items (SDO) along with "SYNC" and "Config".

[0137] In addition, each stream includes descriptor information ("Desc") between "SYNC" and "Config". Attribute information indicating respective attributes of the plurality of groups of encoded data items, stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included, and stream identifier information indicating a self stream identifier are inserted into the descriptor information.

[0138] The demultiplexer 202 extracts a video stream packet from the transport stream TS, and transmits it to the video decoder 203. The video decoder 203 reconfigures a video stream of the video packet extracted by the demultiplexer 202, and performs a decoding process to obtain an uncompressed video data item.

[0139] The video processing circuit 204 performs scaling processing, image quality adjustment processing, and the like on the video data item obtained at the video decoder 203, thereby obtaining a video data item for display. The panel drive circuit 205 drives the display panel 206 based on the image data item for display obtained at the video processing circuit 204. The display panel 206 includes, for example, an LCD (Liquid Crystal Display), an organic EL display (Organic Electroluminescence Display), or the like.

[0140] In addition, the demultiplexer 202 extracts various information such as descriptor information from the transport stream TS, and sends it to the CPU 221. The various information also includes the above-described information on the 3D audio stream configuration descriptor (3Daudio_stream_config_descriptor) and the 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) (see FIG. 6). Figure 14 ).

[0141] Based on the attribute information indicating the attribute of the encoded data item of each group included in the descriptor information and the stream relationship information indicating in which video stream each group is included, the CPU 221 identifies the video stream including the encoded data item of the group having an attribute suitable for the speaker configuration and the viewer and audience (user) selection information.

[0142] In addition, the demultiplexer 202 selectively extracts one or more audio stream packets including the encoded data item of the group having an attribute suitable for the speaker configuration and the viewer and audience (user) selection information from among the predetermined number of audio streams possessed by the transport stream TS under the control of the CPU 221 by means of a PID filter.

[0143] The multiplexing buffers 211-1 to 211-N each collect each audio stream taken out at the demultiplexer 202. Here, the N multiplexing buffers 211-1 to 211-N are necessary and sufficient. In actual operation, a plurality of audio streams taken out at the demultiplexer 202 will be used.

[0144] The combiner 212 reads the audio stream every audio frame from the multiplexing buffer receiving each audio stream taken out at the demultiplexer 202 among the multiplexing buffers 211-1 to 211-N, and sends it to the 3D audio decoder 213.

[0145] When the audio stream supplied from combiner 212 includes descriptor information (“Desc”), 3D audio decoder 213 sends the descriptor information to CPU 221. Under the control of CPU 221, 3D audio decoder 213 selectively retrieves encoded data items with groups of attributes suitable for speaker configuration and viewer / audience (user) selection information, performs decoding processing, and obtains audio data items for driving each speaker of speaker system 215.

[0146] Here, the encoded data items that have undergone decoding can have three modes: only channel encoded data items, only object encoded data items, or both channel encoded data items and object encoded data items.

[0147] When decoding channel encoded data items, the 3D audio decoder 213 performs downmixing or upmixing processing on the speaker configuration of the speaker system 215 to obtain audio data items for driving each speaker. Furthermore, when decoding object encoded data items, the 3D audio decoder 213 calculates speaker rendering (mixing ratio for each speaker) based on object information (metadata items), and mixes the object audio data items into the audio data items according to the calculation results to drive each speaker.

[0148] The sound output processing circuit 214 performs necessary processing, such as D / A conversion and amplification, on the audio data items obtained at the 3D audio decoder 213 for driving each speaker, and supplies them to the speaker system 215. The speaker system 215 includes multiple speakers with multiple channels, such as 2-channel, 5.1-channel, 7.1-channel, or 22.2-channel.

[0149] A brief description Figure 15 The operation of the service receiver 200 is shown. The receiving device 201 receives a transport stream TS transmitted by the service transmitter 100 on a broadcast wave or network packet. In addition to the video stream, the transport stream TS includes a predetermined number of audio streams, which comprise multiple groups of encoded data items configured as 3D audio transmission data items. The transport stream TS is supplied to the demultiplexer 202.

[0150] In demultiplexer 202, video stream packets are extracted from the transport stream TS, and these packets are supplied to video decoder 203. In video decoder 203, the video stream is reconfigured using the video packets extracted at demultiplexer 202, decoding is performed, and uncompressed video data items are obtained. These video data items are then supplied to video processing circuitry 204.

[0151] The video processing circuit 204 performs scaling processing, image quality adjustment processing, and the like on the video data items obtained at the video decoder 203, thereby obtaining video data items for display. The video data items for display are supplied to the panel drive circuit 205. The panel drive circuit 205 drives the display panel 206 based on the image data items for display. In this way, an image corresponding to the image data items for display is displayed on the display panel 206.

[0152] In addition, the demultiplexer 202 extracts various information such as descriptor information from the transport stream TS, which is delivered to the CPU 221. The various information also includes information on a 3D audio stream configuration descriptor and a 3D audio stream ID descriptor (see FIG. 6). Figure 14 Based on the attribute information and stream relationship information contained in the descriptor information, the CPU 221 identifies a video stream including encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience (user) selection information.

[0153] Further, the demultiplexer 202 selectively extracts one or more audio stream packets including encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience selection information from among a predetermined number of audio streams possessed by the transport stream TS under the control of the CPU 221 by a PID filter.

[0154] The audio stream taken out at the demultiplexer 202 is received into a corresponding one of the multiplex buffers 211-1 to 211-N. In the combiner 212, an audio stream is read out from each audio frame of each multiplex buffer that receives an audio stream, and the audio stream is supplied to the 3D audio decoder 213.

[0155] In a case where the audio stream supplied from the combiner 212 includes descriptor information ("Desc"), the descriptor information is extracted and sent to the CPU 221 in the 3D audio decoder 213. The 3D audio decoder 213 selectively extracts encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience (user) selection information under the control of the CPU 221, performs decoding processing, and obtains audio data items for driving each speaker of the speaker system 215.

[0156] Here, when decoding a channel encoded data item, downmix or upmix processing is performed on a speaker configuration of the speaker system 215, and an audio data item for driving each speaker is obtained. Further, when decoding an object encoded data item, speaker rendering (mixing ratio for each speaker) is calculated based on object information (metadata item), and an object audio data item is mixed into an audio data item in accordance with the calculation result so as to drive each speaker.

[0157] The audio data items for driving each speaker acquired at the 3D audio decoder 213 are supplied to the sound output processing circuit 214. The sound output processing circuit 214 performs necessary processing such as D / A conversion, amplification, and the like on the audio data items for driving each speaker. The processed audio data items are supplied to the speaker system 215. In this way, the audio output corresponding to the display image of the display panel 206 is obtained from the speaker system 215.

[0158] Figure 17 The decoding process in the case where there is no descriptor information within the audio stream is schematically shown. The transport stream TS that is a multiplexed stream is input to the demultiplexer 202. In the demultiplexer 202, the system layer is analyzed, and the descriptor information 1 (information on the 3D audio stream configuration descriptor or the 3D audio stream ID descriptor) is supplied to the CPU 221.

[0159] In the CPU 221, the audio stream including the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience (user) selection information is identified based on the descriptor information 1. In the demultiplexer 202, the selection between the streams is performed under the control of the CPU 221.

[0160] In other words, in the demultiplexer 202, the PID filter selectively takes out one or a plurality of audio stream packets from among the predetermined number of audio streams of the transport stream TS, the one or a plurality of audio stream packets including the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience selection information. The audio stream thus taken out is received into the multiplex buffer 211 (211-1 to 211-N).

[0161] The 3D audio decoder 213 performs packet type analysis on each audio stream received in the multiplex buffer 211. Then, in the demultiplexer 202, under the control of the CPU 221, the selection within the stream is performed based on the above-described descriptor information 1.

[0162] Specifically, the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience (user) selection information are selectively taken out from each audio stream as a decoding object, and a decoding process and a mix rendering process are applied thereto, thereby obtaining the audio data items (uncompressed audio) for driving each speaker.

[0163] Figure 18 An example of the configuration of the audio access unit (audio frame) of the audio stream in the case where there is no descriptor information within the audio stream is shown. Here, an example of two streams is shown.

[0164] With respect to the audio stream identified by PID2, the information "FrWork#ch=2, #obj=1" included in the "Config" indicates that there is a "frame" including a channel coded data item and an object coded data item in two channels. The information "GroupID[0]=1, GroupID[1]=2" registered in this order within the "AudioSceneInfo()" included in the "Config" indicates that a "frame" having a coded data item of group 1 and a "frame" having a coded data item of group 2 are set in this order. Note that the value of the packet label (PL) is considered to be the same in the "Config" and each "frame" corresponding thereto.

[0165] Here, the "frame" having a coded data item of group 1 includes a coded sample data item of a CPE (channel pair element). Further, the "frame" having a coded data item of group 2 includes a "frame" having a metadata item as an extension element (Ext_element), and a "frame" having a coded sample data item of a SCE (single channel element).

[0166] With respect to the audio stream identified by PID3, the information "FrWork#ch=0, #obj=2" included in the "Config" indicates that there is a "frame" including two object coded data items. The information "GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=1" registered in this order within the "AudioSceneInfo()" included in the "Config" indicates that a "frame" having a coded data item of group 3 and a "frame" having a coded data item of group 4 are set in this order and these group configuration switching group 1. Note that the value of the packet label (PL) is considered to be the same in the "Config" and each "frame" corresponding thereto.

[0167] Here, the "frame" having a coded data item of group 3 includes a "frame" having a metadata item as an extension element (Ext_element) and a "frame" having a coded sample data item of a SCE (single channel element). Similarly, the "frame" having a coded data item of group 4 includes a "frame" having a metadata item as an extension element (Ext_element) and a "frame" having a coded sample data item of a SCE (single channel element).

[0168] Figure 19 A diagram schematically showing a decoding process in a case where descriptor information is present within an audio stream. A transport stream TS which is a multiplexed stream is input to a demultiplexer 202. In the demultiplexer 202, the system layer is analyzed, and descriptor information 1 (information on a 3D audio stream configuration descriptor or a 3D audio stream ID descriptor) is supplied to a CPU 221.

[0169] In the CPU 221, the audio stream including the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience (user) selection information is identified based on the descriptor information 1. The selection among the streams is performed under the control of the CPU 221 in the demultiplexer 202.

[0170] In other words, the demultiplexer 202 selectively extracts one or more audio stream packets from among the predetermined number of audio streams possessed by the transport stream TS through the PID filter, the extracted one or more audio stream packets including the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience selection information. The thus extracted audio stream is received into the multiplex buffer 211 (211-1 to 211-N).

[0171] The 3D audio decoder 213 performs packet type analysis on each audio stream received in the multiplex buffer 211, and transmits the descriptor information 2 present within the audio stream to the CPU 221. The presence of the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience (user) selection information is identified based on the descriptor information 2. Then, in the demultiplexer 202, the selection among the streams is performed based on the descriptor information 2 under the control of the CPU 221.

[0172] Specifically, the encoded data items of the group having the attribute suitable for the speaker configuration and the viewer and audience (user) selection information are selectively extracted from each audio stream as a decoding object, and a decoding process and a mix rendering process are applied thereto, thereby obtaining the audio data items (uncompressed audio) for driving each speaker.

[0173] Figure 20 An example of the configuration of the audio access unit (audio frame) of the audio stream in the case where the descriptor information is present within the audio stream is shown. Here, an example of two streams is shown. Figure 20 Similar to Figure 18 , except that "Desc", i.e., descriptor information, is inserted between "SYNC" and "Config".

[0174] With regard to the audio stream identified by PID2, the information "GroupID[0] = 1, channel data" contained in "Desc" indicates that the encoded data items of group 1 are channel encoded data items. The information "GroupID[1] = 2, object sound" contained in "Desc" indicates that the encoded data items of group 2 are object encoded data items for immersive sound. Further, the information of "Stream_ID" indicates the stream identifier of the audio stream.

[0175] As for the audio stream identified by the PID 3, the information "GroupID[2]=3, object langl" contained in the "Desc" indicates that the encoded data item of the group 3 is an object encoded data item according to the voice language of the first language. The information "GroupID[3]=4, object lang2" contained in the "Desc" indicates that the encoded data item of the group 4 is an object encoded data item according to the voice language of the second language. Further, the information "SW_GRPID[0]=l" contained in the "Desc" indicates that the groups 3 and 4 configure the switch group 1. Further, the information of the "StreamJD" indicates the stream identifier of the audio stream.

[0176] Figure 21 A configuration example of the audio access unit (audio frame) of the audio stream in the case where the descriptor information exists within the audio stream is shown. Here, an example of one stream is shown.

[0177] The information "FrWork#ch=2, #obj=3" contained in the "Config" indicates that there is a "frame" including the channel encoded data item included in two channels and three object encoded data items. The information "GroupID[0]=l, GroupID[l]=2, GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=l" of the "GroupID" sequentially registered in the "AudioSceneInfo()" contained in the "Config" indicates that the "frame" having the encoded data item of the group 1 and the "frame" having the encoded data item of the group 2, the "frame" having the encoded data item of the group 3 and the "frame" having the encoded data item of the group 4 are set in this order, and these group 3 and group 4 configure the switch group 1. Note that the value of the packet label (PL) is considered to be the same in the "Config" and each "frame" corresponding thereto.

[0178] Here, the "frame" having the encoded data item of the group 1 includes the encoded sample data item of the CPE (channel pair element). Further, the "frame" having the encoded data items of the groups 2 to 4 includes the "frame" having the metadata item as the extension element (Ext_element), and the "frame" having the encoded sample data item of the SCE (single channel element).

[0179] The information "GroupID[0]=l, channel data" contained in the "Desc" indicates that the encoded data item of the group 1 is a channel encoded data item. The information "GroupID[l]=2, object sound" contained in the "Desc" indicates that the encoded data item of the group 2 is an object encoded data item for immersive sound.

[0180] The information "GroupID[2]=3, object langl" included in the "Desc" indicates that the encoded data items of group 3 are object encoded data items according to the voice language of the first language. The information "GroupID[3]=4, object lang2" included in the "Desc" indicates that the encoded data items of group 4 are object encoded data items according to the voice language of the second language. In addition, the information "SW_GRPID[0]=l" included in the "Desc" indicates that groups 3 and 4 constitute a switch group 1. Further, the information "StreamJD" indicates the stream identifier of the audio stream.

[0181] Figure 22 and Figure 23 the flowchart shown in Figure 15 the CPU 221 in the service receiver 200 shown in Fig. 2. The CPU 221 starts the process at step STl. Then, the CPU 221 detects the receiver speaker configuration, i.e., the speaker configuration of the speaker system 215, at step ST2. Next, the CPU 221 acquires selection information on the audio output by the viewer and the audience (user) at step ST3.

[0182] Next, at step ST4, the CPU 221 reads the descriptor information on the main stream within the PMT, selects the audio stream belonging to the group having attributes suitable for the speaker configuration as well as the viewer and audience selection information, and receives it into a buffer. Then, at step ST5, the CPU 221 checks whether a descriptor type packet exists in the audio stream.

[0183] Next, at step ST6, the CPU 221 determines whether a descriptor type packet exists. If it exists, at step ST7, the CPU 221 reads the descriptor information of the relevant packet, detects the information of "groupID", "attribute", "switchGroupID", and "presetGroupID", and then proceeds to the process at step ST9. On the other hand, if it does not exist, the CPU 221 detects the information of "groupID", "attribute", "switchGroupID", and "presetGroupID" from the descriptor information of the PMT at step ST8, and then proceeds to step ST9. Note that step ST8 can not be executed, and the entire audio stream to be processed can be decoded.

[0184] In step ST9, the CPU 221 determines whether or not to decode the object encoded data item. If decoding, the CPU 221 decodes the object encoded data item in step ST10, and then proceeds to the processing in step ST11. On the other hand, if not decoding, the CPU 221 immediately proceeds to the processing in step ST11.

[0185] In step ST11, the CPU 221 determines whether or not to decode the channel encoded data item. If decoding, the CPU 221 decodes the channel encoded data item as necessary, performs downmixing or upmixing processing for the speaker configuration of the speaker system 215, and obtains the audio data item for driving each speaker in step ST12. Thereafter, the CPU 221 proceeds to the processing in step ST13. On the other hand, if not decoding, the CPU 221 immediately proceeds to the processing in step ST13.

[0186] In step ST13, in the case where the CPU 221 decodes the object encoded data item, based on the information, it mixes or calculates the speaker rendering with the channel data item. In the speaker rendering calculation, the speaker rendering (mixing ratio for each speaker) is calculated by the azimuth angle (azimuth angle information) and the elevation angle (elevation angle information). According to the calculation result, the object audio data item is mixed with the channel data for driving each speaker.

[0187] Next, the CPU 221 performs dynamic range control of the audio data item for driving each speaker, and outputs in step ST14. Thereafter, the CPU 221 ends the processing in step ST15.

[0188] As described above, in the transmission / reception system 10 shown in Figure 1 In the transmission / reception system 10 shown in

[0189] In the transmission / reception system 10 shown in Figure 1 In the transmission / reception system 10 shown in

[0190] <2. Variants>

[0191] In the above-described embodiments, the service receiver 200 selectively takes out one or more audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information from among a plurality of audio streams transmitted via the service transmitter 100, performs a decoding process, and obtains a predetermined number of audio data items for driving the speakers.

[0192] However, it is conceivable that the service receiver selectively takes out one or more audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information from among a plurality of audio streams transmitted via the service transmitter 100, reconfigures the audio streams including the encoded data items having the group of attributes suitable for the speaker configuration and the viewer and audience selection information, and transmits the reconfigured audio streams to devices connected to the in-house network (including also DLNA devices).

[0193] Figure 24 A configuration example of the service receiver 200A that transmits the reconfigured audio streams to the devices connected to the in-house network is shown, as described above. The components corresponding to those in Figure 15 Figure 24 The components in

[0194] The demultiplexer 202 selectively takes out packets of one or more audio streams having a predetermined number of audio streams in the transport stream TS under the control of the CPU 221 by a PID filter, the taken-out audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information.

[0195] The audio streams taken out by the demultiplexer 202 are received into the respective ones of the multiplex buffers 211-1 to 211-N. In the combiner 212, the audio streams are read out from each audio frame of each multiplex buffer that receives the audio, and the audio streams are supplied to the stream reconfiguration unit 231.

[0196] In the stream reconfiguration unit 231, descriptor information ("Desc") is extracted and transmitted to the CPU 221 in the case where the descriptor information is contained in the audio streams supplied from the combiner 212. In the stream reconfiguration unit 231, the encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience (user) selection information are selectively acquired under the control of the CPU 221, and the audio streams having the encoded data items are reconfigured. The reconfigured audio streams are supplied to the transmission interface 232. Then, they are transmitted (sent) from the transmission interface 232 to the devices 300 connected to the in-house network.

[0197] ​The indoor network connection includes an Ethernet connection and a wireless connection of "WiFi" or "Bluetooth". "WiFi" and "Bluetooth" are registered trademarks.

[0198] In addition, the device 300 includes a surround speaker, a second display, and an audio output device attached to the network terminal. The device 200 to which the reconfigured audio stream is transmitted performs a decoding process similar to the 3D audio decoder 213 in the service receiver 200 in Figure 15 and obtains an audio data item for driving a predetermined number of speakers.

[0199] Furthermore, as the service receiver, it is conceivable to send the above-described reconfigured audio stream to a device connected with a digital interface such as "HDMI (High-Definition Multimedia Interface)", "MHL (Mobile High-definition Link)", and "DisplayPort". "HDMI" and "MHL" are registered trademarks.

[0200] In addition, in the above-described embodiment, the stream correspondence relationship information inserted into the layer or the like of the container is information indicating a correspondence relationship between a group ID and a substream ID. Specifically, the substream ID is used to associate a group with an audio stream. However, it is conceivable to use a packet identifier (PID: packet ID) or a stream type in order to associate the group with the audio stream. In the case of using the stream type, the stream type of each audio stream should be changed.

[0201] In addition, the above-described embodiment shows an example in which attribute information of the encoded data item of each group is sent by providing an "attribute_of_groupID" field (see Figure 10 ). However, the present technology also includes a method that can identify the type (attribute) of the encoded data item if a specific group ID is identified by defining a specific meaning in the value of the group ID (GroupID) itself between the transmitter and the receiver. In this case, the group ID serves as an identifier of the group, and also serves as attribute information of the encoded data item of the group, so that the field "attribute_of_groupID" becomes unnecessary.

[0202] In addition, the above-described embodiment shows an example in which the encoded data items of a plurality of groups include both a channel encoded data item and an object encoded data item (see Figure 3 ). However, the present technology can also be similarly applied to a case in which the encoded data items of a plurality of groups include only a channel encoded data item or only an object encoded data item.

[0203] In addition, the above-described embodiments show an example in which the container is a transport stream (MPEG-2 TS). However, the present technology can also be similarly applied to a system in which a stream is delivered by a container of a format such as MP4. For example, the system includes an MPEG-DASH basic stream delivery system, or a transmission / reception system that handles an MMT (MPEG Media Transport) structure transmission stream.

[0204] The present technology can also have the following configuration.

[0205] 1. A transmission apparatus comprising:

[0206] a transmission unit that transmits a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items; and

[0207] an information insertion unit that inserts attribute information indicating respective attributes of the plurality of groups of encoded data items into a layer of the container and / or a layer of the audio streams.

[0208] (2) The transmission apparatus according to (1), wherein

[0209] the information insertion unit further inserts stream correspondence relationship information into the layer of the container and / or the layer of the audio streams, the stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included.

[0210] (3) The transmission apparatus according to (2), wherein

[0211] the stream correspondence relationship information is information indicating a correspondence relationship between a group identifier identifying each of the plurality of groups of encoded data items and a stream identifier identifying each of the predetermined number of audio streams.

[0212] (4) The transmission apparatus according to (3), wherein

[0213] the information insertion unit further inserts stream identifier information into a layer of the container and / or a layer of the audio streams, the stream identifier information indicating the stream identifier of each of the predetermined number of audio streams.

[0214] (5) The transmission apparatus according to (4), wherein

[0215] the container is an MPEG2-TS, and

[0216] in a case where the stream identifier information is inserted into the container, the information insertion unit inserts the stream identifier information into an audio elementary stream loop corresponding to each of the predetermined number of audio streams present under a program map table.

[0217] (6) The transmission apparatus according to (4) or (5) above, wherein

[0218] In a case where the stream identifier information is inserted into the audio stream, the information insertion unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.

[0219] (7) The transmission apparatus according to (2), wherein

[0220] The stream correspondence relationship information is information indicating a correspondence relationship between a group identifier identifying each of the encoded data items of the plurality of groups and a packet identifier added in a case where each of the predetermined number of audio streams is packetized.

[0221] (8) The transmission apparatus according to (2), wherein

[0222] The stream correspondence relationship information is information indicating a correspondence relationship between a group identifier identifying each of the encoded data items of the plurality of groups and type information indicating a stream type of each of the predetermined number of audio streams.

[0223] (9) The transmission apparatus according to any one of (2) to (8) above, wherein

[0224] The container is MPEG2-TS, and

[0225] In a case where the attribute information and the stream correspondence relationship information are inserted into the container, the information insertion unit inserts the attribute information and the stream correspondence relationship information into the audio elementary stream loop corresponding to at least one or a plurality of audio streams among the predetermined number of audio streams existing below the program map table.

[0226] (10) The transmission apparatus according to any one of (2) to (8) above, wherein

[0227] In a case where the attribute information and the stream correspondence relationship are inserted into the audio stream, the information insertion unit inserts the attribute information and the stream correspondence relationship information into a PES payload of a PES packet in at least one or a plurality of audio streams among the predetermined number of audio streams.

[0228] (11) The transmission apparatus according to any one of (1) to (10) above, wherein

[0229] The encoded data items of the plurality of groups include one or both of a channel encoded data item and an object encoded data item.

[0230] (12) A transmission method comprising:

[0231] a sending step of sending, by a sending unit, a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items; and

[0232] an information inserting step of inserting, into a layer of the container and / or a layer of the audio streams, attribute information indicating respective attributes of the encoded data items of the plurality of groups.

[0233] (13) A receiving apparatus comprising:

[0234] a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items,

[0235] attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio streams; and

[0236] a processing unit that processes the predetermined number of audio streams contained in the received container based on the attribute information.

[0237] (14) The receiving apparatus according to the above (13), wherein,

[0238] stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is contained is further inserted into a layer of the container and / or a layer of the audio streams, and

[0239] the processing unit processes the predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information.

[0240] (15) The receiving apparatus according to the above (14), wherein,

[0241] based on the attribute information and the stream correspondence relationship information, the processing unit performs selective decoding processing on an audio stream containing encoded data items of a group having attributes suitable for a speaker configuration and user selection information.

[0242] (16) The receiving apparatus according to any one of the above (13) to (15), wherein,

[0243] the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items.

[0244] (17) A receiving method comprising:

[0245] a receiving step of receiving, by a receiving unit, a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items,

[0246] attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream; and

[0247] a processing step of processing the predetermined number of audio streams included in the received container based on the attribute information.

[0248] (18) A reception apparatus comprising:

[0249] a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams including encoded data items of a plurality of groups,

[0250] attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream; and

[0251] a processing unit that selectively acquires encoded data items of a predetermined group from the predetermined number of audio streams contained in the received container based on the attribute information, and reconfigures an audio stream including the encoded data items of the predetermined group; and

[0252] a stream transmitting unit that transmits the audio stream reconfigured by the processing unit to an external device.

[0253] (19) The reception apparatus according to the above (18), wherein

[0254] stream correspondence relationship information indicating in which audio stream the encoded data items of the plurality of groups are respectively contained is further inserted into a layer of the container and / or a layer of the audio stream, and

[0255] the processing unit selectively acquires the encoded data items of the predetermined group from the predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information.

[0256] (20) A reception method comprising:

[0257] a receiving step of receiving, by a receiving unit, a container of a predetermined format having a predetermined number of audio streams including encoded data items of a plurality of groups,

[0258] attribute information indicating respective attributes of the encoded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream; and

[0259] a processing step of processing the predetermined number of audio streams included in the received container based on the attribute information.

[0260] A stream transmission step transmits the audio stream reconfigured in the processing step to an external device.

[0261] The main feature of the present technology is that stream correspondence relationship information is inserted into a layer of a container and / or a layer of an audio stream, the stream correspondence relationship information indicating which audio stream includes each attribute information, the attribute information indicating a plurality of groups of encoded data items included in a predetermined number of audio streams and respective attributes of the plurality of groups of encoded data items, so that a processing load at a reception side can be reduced (see Figure 14 ).

[0262] Symbol explanation

[0263] 10 Transmission / reception system

[0264] 100 Service transmitter

[0265] 100 Stream generating unit

[0266] 112 Video encoder

[0267] 113 Audio encoder

[0268] 114 Multiplexer

[0269] 200, 200A Service receiver

[0270] 201 Reception unit

[0271] 202 Demultiplexer

[0272] 203 Video decoder

[0273] 204 Video processing circuit

[0274] 205 Panel drive circuit

[0275] 206 Display panel

[0276] 211-1 to 211-N Multiplexing buffer

[0277] 212 Combiner

[0278] 213 3D audio decoder

[0279] 214 Sound output processing circuit

[0280] 215 Speaker system

[0281] 221 CPU

[0282] 222 Flash ROM

[0283] 223 DRAM

[0284] 224 Internal bus

[0285] 225 remote control receiving unit

[0286] 226 remote control transmitter

[0287] 231 flow reconfiguration unit

[0288] 232 transmission interface

[0289] 300 device

Claims

1. A transmitting apparatus comprising: a transmitting unit configured to transmit a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements; and an information inserting unit configured to insert, into a layer of the container and / or a layer of an audio stream, attribute information representing respective attributes of the plurality of groups of encoded data items and stream correspondence information representing a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a stream identifier identifying each of the predetermined number of audio streams.

2. The transmitting apparatus according to claim 1, wherein the stream correspondence information represents in which audio stream the encoded data items of the plurality of groups are respectively included.

3. The transmitting apparatus according to claim 2, wherein the stream correspondence information includes stream correspondence information representing an identifier of a group groupID, an identifier of a switch group switchGroupID, and a kind of content of a group contentKind.

4. The transmitting apparatus according to claim 3, wherein the information inserting unit further inserts, into the layer of the container and / or the layer of the audio stream, stream identifier information representing the stream identifier of each of the predetermined number of audio streams.

5. The transmitting apparatus according to claim 4, wherein the container is an MPEG2-TS, and in a case where the stream identifier information is inserted into the container, the information inserting unit inserts the stream identifier information into an audio elementary stream loop corresponding to each of the predetermined number of audio streams existing under a program map table.

6. The transmitting apparatus according to claim 4, wherein in a case where the stream identifier information is inserted into the audio stream, the information inserting unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.

7. The transmitting apparatus according to claim 2, wherein the container is an MPEG2-TS, and in a case where the attribute information and the stream correspondence information are inserted into the container, the information inserting unit inserts the attribute information and the stream correspondence information into an audio elementary stream loop corresponding to at least one or a plurality of audio streams of the predetermined number of audio streams existing under a program map table.

8. The transmitting apparatus according to claim 2, wherein in a case where the attribute information and the stream correspondence information are inserted into the audio stream, the information inserting unit inserts the attribute information and the stream correspondence information into a PES payload of a PES packet in at least one or a plurality of audio streams of the predetermined number of audio streams.

9. A transmitting method comprising: a sending step of sending, by a sending unit, a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements; and an information inserting step of inserting, into a layer of the container and / or a layer of an audio stream, attribute information indicating respective attributes of the encoded data items of the plurality of groups and stream correspondence information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a stream identifier identifying each of the predetermined number of audio streams.

10. A receiving apparatus comprising: a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information indicating respective attributes of the encoded data items of the plurality of groups and stream correspondence information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a stream identifier identifying each of the predetermined number of audio streams are inserted into a layer of the container and / or a layer of an audio stream; and a processing unit that processes the predetermined number of audio streams contained in the received container based on the attribute information.

11. The receiving apparatus according to claim 10, wherein the stream correspondence information indicates in which audio stream the encoded data items of the plurality of groups are respectively contained, the stream correspondence information including an identifier of a group groupID, an identifier of a switch group switchGroupID, and a kind of content of a group contentKind, and the processing unit processes the predetermined number of audio streams based on the stream correspondence information in addition to the attribute information.

12. The receiving apparatus according to claim 11, wherein based on the attribute information and the stream correspondence information, the processing unit performs a selective decoding process on an audio stream containing encoded data items of a group having attributes suitable for a speaker configuration and user selection information.

13. A receiving method comprising: a receiving step of receiving, by a receiving unit, a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information indicating respective attributes of the encoded data items of the plurality of groups and stream correspondence information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a stream identifier identifying each of the predetermined number of audio streams are inserted into a layer of the container and / or a layer of an audio stream; and a processing step of processing the predetermined number of audio streams included in the received container based on the attribute information.

14. A reception apparatus comprising: a reception unit that receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information representing respective attributes of the plurality of groups of encoded data items and stream correspondence information representing a correspondence between a group identifier identifying each of the plurality of groups of encoded data items and a stream identifier identifying each of the predetermined number of audio streams are inserted into a layer of the container and / or a layer of an audio stream; a processing unit that selectively acquires a predetermined group of encoded data items from the predetermined number of audio streams included in the received container based on the attribute information, and reconfigures an audio stream including the predetermined group of encoded data items; and a stream transmission unit that transmits the audio stream reconfigured by the processing unit to an external device.

15. The reception apparatus according to claim 14, wherein the stream correspondence information represents in which audio stream the plurality of groups of encoded data items are respectively included, the stream correspondence information including an identifier of a group, group ID, an identifier of a switch group, switchGroupID, and a kind of content of a group, contentKind, and the processing unit selectively acquires the predetermined group of encoded data items from the predetermined number of audio streams based on the stream correspondence information in addition to the attribute information.

16. A reception method comprising: a reception step of receiving, by a reception unit, a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items, wherein the plurality of groups of encoded data items include one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information representing respective attributes of the plurality of groups of encoded data items and stream correspondence information representing a correspondence between a group identifier identifying each of the plurality of groups of encoded data items and a stream identifier identifying each of the predetermined number of audio streams are inserted into a layer of the container and / or a layer of an audio stream; a processing step of selectively acquiring a predetermined group of encoded data items from the predetermined number of audio streams included in the received container based on the attribute information, and reconfiguring an audio stream including the predetermined group of encoded data items; and a stream transmission step of transmitting the audio stream reconfigured in the processing step to an external device.

Citation Information

Patent Citations

  • System and tools for enhanced 3D audio authoring and rendering

    CN103650535A

  • Transmission device, transmission method, receiving device and receiving method

    CN103843330A

  • Information processing device, information processing method, program, and data structure

    CN1926872A