Transmission device, transmission method, reception device, and reception method
By inserting attribute and stream correspondence information into the container and audio stream layers, the problem of heavy processing load on the receiving side when sending multiple audio data items is solved, and more efficient processing is achieved.
Patent Information
- Application Number
- CN202111173401.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-09-30
- Filing Date
- 2015-09-16
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2035-09-16
AI Technical Summary
When sending multiple audio data items, the processing load on the receiving side is heavy, which is difficult to effectively reduce with existing technologies.
By inserting information indicating attributes of coded data items, including attribute information and stream correspondence information, into the layers of the container and audio stream, the receiving side is allowed to selectively decode and process necessary coded data items.
The processing load on the receiving side is reduced and the processing efficiency is improved.
Smart Images

Figure CN113921020B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese national phase application with an international filing date of September 16, 2015, an international application number of PCT / JP2015 / 076259, and an invention title of “Transmitting Device, Transmitting Method, Receiving Device and Receiving Method”. The national phase entry date of the Chinese national phase application is March 23, 2017, the application number is 201580051430.9, and the invention title of “Transmitting Device, Transmitting Method, Receiving Device and Receiving Method”. Technical Field
[0002] The present technology relates to a transmitting device, a transmitting method, a receiving device, and a receiving method, and more particularly, to a transmitting device and the like for transmitting a variety of audio data items. Background Art
[0003] In related art, as a three-dimensional (3D) acoustic technology, a technology has been proposed in which encoded sample data items are mapped and rendered to speakers at any positions based on metadata items (for example, refer to Patent Document 1).
[0004] Reference List
[0005] Patent Literature
[0006] Patent Document 1: Translation of PCT International Application Publication No. 2014-520491. Summary of the Invention
[0007] Technical issues
[0008] It is conceivable that by transmitting object coded data items including coded sample data items and metadata items together with channel coded data items of 5.1 channels, 7.1 channels, etc., sound with enhanced realism can be reproduced on the receiving side.
[0009] An object of the present technology is to reduce the processing load on the receiving side in the case where a variety of audio data items are transmitted.
[0010] Solution to the problem
[0011] The concept of the present technology is a sending device including: a sending unit that sends a container of a predetermined format having a predetermined number of audio streams, the audio stream including a plurality of groups of encoded data items; and an information inserting unit that inserts information indicating respective attributes of the plurality of groups of encoded data items into a layer of the container and / or a layer of the audio stream.
[0012] In the present technology, a transmitting unit transmits a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of sets of encoded data items. For example, the plurality of sets of encoded data items may include one or both of channel encoded data items and object encoded data items.
[0013] The information insertion unit inserts attribute information indicating corresponding attributes of the plurality of groups of encoded data items into a container layer and / or an audio stream layer. For example, the container may be a transport stream (MPEG2-TS) applicable to digital broadcasting standards. Furthermore, the container may be in MP4 format or other formats for Internet transmission, for example.
[0014] Therefore, in this technology, attribute information indicating the corresponding attributes of multiple groups of encoded data items included in a predetermined number of audio streams is inserted into the container layer and / or the audio stream layer. Therefore, on the receiving side, before decoding the encoded data items, the corresponding attributes of the multiple groups of encoded data items can be easily identified, and only the necessary groups of encoded data items can be selectively decoded and used, thereby reducing the processing load.
[0015] Furthermore, in this technology, for example, the information insertion unit can also insert stream correspondence information into the container layer and / or the audio stream layer, indicating which audio stream each of the plurality of groups of coded data items is contained in. This makes it easier for the receiving end to identify the audio stream containing the necessary group of coded data items, thereby reducing the processing load.
[0016] In this case, for example, the container may be MPEG2-TS, and in the case where attribute information and stream identifier information are inserted into the container, the information insertion unit may insert the attribute information and stream identifier information into an audio elementary stream loop corresponding to at least one or more audio streams among a predetermined number of audio streams existing under a program map table.
[0017] In addition, in this case, for example, in the case where attribute information and stream correspondence information are inserted into the audio stream, the information insertion unit can insert the attribute information and stream correspondence information into the PES payload of the PES packet in at least one or more audio streams among a predetermined number of audio streams.
[0018] For example, the stream correspondence information may be information indicating a correspondence between a group identifier identifying each of a plurality of groups of encoded data items and a stream identifier identifying each of a predetermined number of audio streams. In this case, for example, the information insertion unit may insert stream identifier information indicating a stream identifier for each of the predetermined number of audio streams into a container layer and / or an audio stream layer.
[0019] For example, the container may be MPEG2-TS, and in the case where the stream identifier information is inserted into the container, the information insertion unit may insert the stream identifier information into an audio elementary stream loop corresponding to each of a predetermined number of audio streams existing under the program map table. Also, for example, in the case where the stream identifier information is inserted into the audio stream, the information insertion unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.
[0020] Furthermore, for example, the stream correspondence information may be information indicating a correspondence between a group identifier identifying each of a plurality of groups of coded data items and a packet identifier added when each of a predetermined number of audio streams is packetized. Furthermore, for example, the stream correspondence information may be information indicating a correspondence between a group identifier identifying each of a plurality of groups of coded data items and type information indicating a stream type of each of a predetermined number of audio streams.
[0021] In addition, other concepts of the present technology are a receiving device including: a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, attribute information representing corresponding attributes of the plurality of groups of encoded data items being inserted into a layer of the container and / or a layer of the audio stream; and a processing unit that processes the predetermined number of audio streams included in the received container based on the attribute information.
[0022] In the present technology, a receiving unit receives a container in a predetermined format containing a predetermined number of audio streams, the predetermined number of audio streams comprising a plurality of groups of encoded data items. For example, the plurality of groups of encoded data items may include one or both of channel encoded data items and object encoded data items. Attribute information indicating corresponding attributes of the plurality of groups of encoded data items is inserted into a layer of the container and / or a layer of the audio stream. A processing unit processes the predetermined number of audio streams included in the received container based on the attribute information.
[0023] Therefore, in this technology, a predetermined number of audio streams contained in a received container are processed based on attribute information indicating the respective attributes of a plurality of groups of coded data items inserted into the container layers and / or audio stream layers. Therefore, only the necessary groups of coded data items can be selectively decoded and used, thereby reducing the processing load.
[0024] Furthermore, in this technology, for example, stream correspondence information indicating which audio streams contain each of multiple groups of coded data items is further inserted into the container layer and / or the audio stream layer. In addition to attribute information, the processing unit can also process a predetermined number of audio streams based on this stream correspondence information. In this case, the audio stream containing the necessary group of coded data items can be easily identified, thereby reducing the processing load.
[0025] In addition, in the present technology, for example, based on the attribute information and the stream correspondence information, the processing unit can perform selective decoding processing on the audio stream including the encoded data items having a set of attributes applicable to the speaker configuration and the user selection information.
[0026] In addition, other concepts of the present technology are a receiving device, including: a receiving unit, which receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including multiple groups of encoded data items; attribute information representing corresponding attributes of the multiple groups of encoded data items is inserted into a layer of the container and / or a layer of the audio stream; a processing unit, which selectively obtains a predetermined group of encoded data items from a predetermined number of audio streams included in the received container based on the attribute information, and reconfigures the audio stream including the predetermined group of encoded data items; and a stream sending unit, which sends the audio stream reconfigured by the processing unit to an external device.
[0027] In this technology, a receiving unit receives a container in a predetermined format having a predetermined number of audio streams including a plurality of groups of coded data items. Attribute information indicating the respective attributes of the plurality of groups of coded data items is inserted into a layer of the container and / or a layer of the audio stream. A processing unit selectively extracts a predetermined group of coded data items from the predetermined number of audio streams included in the received container based on the attribute information and reconfigures the audio stream including the predetermined group of coded data items. A stream transmitting unit transmits the audio stream reconfigured by the processing unit to an external device.
[0028] Therefore, in this technology, based on attribute information indicating the respective attributes of a plurality of groups of coded data items inserted into a container layer and / or an audio stream layer, a predetermined group of coded data items is selectively acquired from a predetermined number of audio streams, and the audio stream to be transmitted to an external device is reconfigured. The necessary group of coded data items can be readily acquired, thereby reducing the processing load.
[0029] Furthermore, in this technology, for example, stream correspondence information indicating which audio streams contain each of multiple groups of coded data items is further inserted into the container layer and / or the audio stream layer. In addition to attribute information, the processing unit can selectively retrieve a predetermined group of coded data items from a predetermined number of audio streams based on this stream correspondence information. In this case, the audio stream containing the predetermined group of coded data items can be easily identified, thereby reducing the processing load.
[0030] Beneficial effects of the present invention
[0031] According to the present technology, when transmitting a variety of audio data items, the processing load at the receiving side can be reduced. It should be noted that the effects described in this specification are merely illustrative and not restrictive, and there may be additional effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] [ Figure 1 ] is a block diagram showing a configuration example of a sending / receiving system as an embodiment.
[0033] [ Figure 2 ] is a diagram showing the structure of an audio frame (1024 samples) in a 3D audio transmission data item.
[0034] [ Figure 3 ] is a diagram showing a configuration example of 3D audio transmission data items.
[0035] [ Figure 4 ] is a diagram schematically showing a configuration example of an audio frame in the case where 3D audio transmission data items are transmitted through one stream and a plurality of streams.
[0036] [ Figure 5 ] is a diagram showing an example of group division in the case of transmitting 3D audio transmission data items through two streams.
[0037] [ Figure 6 ] is a diagram showing the correspondence between groups and flows in a group division example (two divisions), etc.
[0038] [ Figure 7 ] is a diagram showing an example of group division in the case of transmitting 3D audio transmission data items through two streams.
[0039] [ Figure 8 ] is a diagram showing the correspondence between groups and flows in a group division example (two divisions), etc.
[0040] [ Figure 9 ] is a block diagram showing a configuration example of a stream generation unit included in a service transmitter.
[0041] [ Figure 10 ] is a diagram showing a configuration example of a 3D audio stream configuration descriptor.
[0042] [ Figure 11 ] shows the contents of main information in the configuration instance of the 3D audio stream configuration descriptor.
[0043] [ Figure 12 ] is a diagram showing the kind of content defined in "contentKind".
[0044] [ Figure 13] is a diagram showing a configuration example of the contents of a 3D audio stream ID descriptor and main information in the configuration example.
[0045] [ Figure 14 ] is a diagram showing an example of the configuration of a transport stream.
[0046] [ Figure 15 ] is a block diagram showing a configuration example of a service receiver.
[0047] [ Figure 16 ] is a diagram showing an example of a received audio stream.
[0048] [ Figure 17 ] is a diagram schematically illustrating a decoding process when descriptor information does not exist in an audio stream.
[0049] [ Figure 18 ] is a diagram showing an example of the configuration of an audio access unit (audio frame) of an audio stream when descriptor information does not exist within the audio stream.
[0050] [ Figure 19 ] is a diagram schematically illustrating a decoding process when descriptor information exists within an audio stream.
[0051] [ Figure 20 ] is a diagram showing a configuration example of an audio access unit (audio frame) of an audio stream in the case where descriptor information exists within the audio stream.
[0052] [ Figure 21 ] is a diagram showing another configuration example of an audio access unit (audio frame) of an audio stream in the case where descriptor information exists within the audio stream.
[0053] [ Figure 22 ] is a flowchart (1 / 2) showing an example of audio decoding control processing by the CPU in the service receiver.
[0054] [ Figure 23 ] is a flowchart (2 / 2) showing an example of audio decoding control processing by the CPU in the service receiver.
[0055] [ Figure 24 ] is a block diagram showing other configuration examples of a service receiver. DETAILED DESCRIPTION
[0056] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. This specification will proceed in the following order.
[0057] 1. Implementation Method
[0058] 2. Modifications
[0059] <1. Implementation Method>
[0060] [Configuration example of transmission / reception system]
[0061] Figure 1 The configuration example of the transmission / reception system 10 as an embodiment is shown. The transmission / reception system 10 includes a service transmitter 100 and a service receiver 200. The service transmitter 100 transmits a transport stream TS via broadcast waves or network packets. The transport stream TS has a video stream and a predetermined number (i.e., one or more) of audio streams including a plurality of sets of encoded data items.
[0062] Figure 2 The structure of an audio frame (1024 samples) in the 3D audio transmission data item processed in this embodiment is shown. An audio frame includes multiple MPEG audio stream packets. Each MPEG audio stream packet includes a header and a payload.
[0063] The header contains information such as the packet type (Packet Type), packet label (Packet Label), and packet length (Packet Length). Information defined by the packet type in the header is included in the payload. The payload contains "SYNC" information corresponding to the synchronization start code, "Frame" which is the actual data item for 3D audio transmission, and "Config" indicating the configuration of the "Frame."
[0064] A "frame" includes channel coded data items and object coded data items that configure 3D audio transmission data items. Here, the channel coded data items include coded sample data items such as SCE (mono channel element), CPE (channel pair element), LFE (low frequency element), etc. In addition, the object coded data items include coded sample data items of SCE (mono channel element) and metadata items for mapping the coded sample data items of SCE to speakers existing at any position and rendering the coded sample data items of SCE. The metadata items are included as extension elements (Ext_element).
[0065] Figure 3 An example configuration of a 3D audio transmission data item is shown. In this example, the 3D audio transmission data item includes one channel encoding data item and two object encoding data items. The channel encoding data item is a 5.1-channel channel encoding data item (CD) and includes each encoding sample data item of SCE1, CPE1.1, CPE1.2, and LFE1.
[0066] The two object coded data items are coded data items for an immersive audio object (IAO) and a speech dialog object (SDO). The immersive audio object coded data item is an object coded data item for immersive sound and includes a coded sample data item SCE2 and a metadata item EXE_E1 (object metadata) 2 for mapping the coded sample data item SCE2 to speakers at arbitrary positions and rendering the coded sample data item SCE2.
[0067] The speech dialogue object coded data items are object coded data items for speech languages. In this example, there are speech dialogue object coded data items corresponding to each of the first and second languages. The speech dialogue object coded data items corresponding to the first language include coded sample data items SCE3 and metadata items EXE_E1 (object metadata) 3 for mapping coded sample data items SCE3 to speakers located at any location and rendering coded sample data items SCE3. Furthermore, the speech dialogue object coded data items corresponding to the second language include coded sample data items SCE4 and metadata items EXE_E1 (object metadata) 4 for mapping coded sample data items SCE4 to speakers located at any location and rendering coded sample data items SCE4.
[0068] Based on the type, the encoded data items are categorized into groups. In the example shown, 5.1-channel encoded channel data items are categorized into Group 1, immersive audio object encoded data items are categorized into Group 2, first language voice dialogue object encoded data items are categorized into Group 3, and second language voice dialogue object encoded data items are categorized into Group 4.
[0069] The selection between groups on the receiving side is registered as a switch group (SW group) and encoded. Furthermore, groups are bundled as preset groups (preset groups), enabling playback tailored to the use case. In the example shown, groups 1, 2, and 3 are bundled into preset group 1, and groups 1, 2, and 4 are bundled into preset group 2.
[0070] return Figure 1 , the service transmitter 100 transmits 3D audio transmission data items including a plurality of groups of encoded data items through one stream or a plurality of streams (multiple streams), as described above.
[0071] Figure 4 (a) Schematically shows the transmission of Figure 3 This is a configuration example of the case of 3D audio transmission data items in . In this case, one stream includes channel coding data items (CD), immersive audio object coding data items (IAO), and voice dialogue object coding data items (SDO), as well as "SYNC" and "Config".
[0072] Figure 4 (b) schematically illustrates the transmission of Figure 3 3D audio transmission data items in the configuration example. In this case, the main stream includes channel coding data items (CD) and immersive audio object coding data items (IAO) along with "SYNC" and "Config". In addition, the substream includes voice dialogue object coding data items (SDO) along with "SYNC" and "Config".
[0073] Figure 5 Shown in two streams sent Figure 3 : This figure shows an example of group division in the case of 3D audio transmission data items in [1]. In this case, the main stream includes channel coded data items (CD) classified as group 1 and immersive audio object coded data items (IAO) classified as group 2. In addition, the substream includes speech dialogue object coded data items (SDO) according to the first language classified as group 3, and speech dialogue object coded data items (SDO) according to the second language classified as group 4.
[0074] Figure 6 Shown in Figure 5 This figure shows the correspondence between groups and streams in the group division example (two divisions) in [1]. Here, the group ID (groupID) is an identifier for identifying a group. The attribute (attribute) shows the attributes of the encoded data item of each group. The switch group ID (switchGroupID) is an identifier for identifying a switch group. The preset group ID (presetGroupID) is an identifier for identifying a preset group. The stream ID (sub Stream ID) is an identifier for identifying a substream. The kind (Kind) shows the kind of content of each group.
[0075] The illustrated correspondence relationship shows that the encoded data items belonging to group 1 are channel encoded data items, do not constitute a switching group, and are included in stream 1. In addition, the illustrated correspondence relationship shows that the encoded data items belonging to group 2 are object encoded data items for immersive sound (immersive audio object encoded data items), do not constitute a switching group, and are included in stream 1.
[0076] In addition, the illustrated correspondence relationship shows that the coded data items belonging to group 3 are object coded data items (speech dialogue object coded data items) for a speech language according to the first language, constitute switching group 1, and are included in stream 2. In addition, the illustrated correspondence relationship shows that the coded data items belonging to group 4 are object coded data items (speech dialogue object coded data items) for a speech language according to the second language, constitute switching group 1, and are included in stream 2.
[0077] Further, the illustrated correspondence shows that the preset group 1 includes the group 1, the group 2, and the group 3. Further, the illustrated correspondence shows that the preset group 2 includes the group 1, the group 2, and the group 4.
[0078] Figure 7 An example of the group division in the case where the 3D audio transmission data item is transmitted through two streams is illustrated. In this case, the main stream includes a channel coded data item (CD) classified as the group 1 and an immersive audio object coded data item (IAO) classified as the group 2.
[0079] Further, the main stream includes an SAOC (Spatial Audio Object Coding) object coded data item classified as the group 5 and an HOA (Higher Order Ambisonics) object coded data item classified as the group 6. The SAOC object coded data item is a data item that utilizes the characteristics of the object data item, and performs higher compression of the object coding. The purpose of the HOA object coded data item is to reproduce the sound direction from the sound incoming direction of the microphone to the auditory position by a technology that processes the 3D sound as the entire sound field.
[0080] The sub stream includes a speech dialogue object coded data item (SDO) according to a first language classified as the group 3, and a speech dialogue object coded data item (SDO) according to a second language classified as the group 4. Further, the sub stream includes a first audio description coded data item classified as the group 7 and a second audio description coded data item classified as the group 8. The audio description coded data item is used to explain the content (mainly video) in sound, and is transmitted separately from the general sound, for people with mainly visual disabilities.
[0081] Figure 8 An example of the group division in the case where the 3D audio transmission data item is transmitted through two streams is illustrated. In this case, the main stream includes a channel coded data item (CD) classified as the group 1 and an immersive audio object coded data item (IAO) classified as the group 2. Figure 7 The illustrated correspondence shows that the encoded data item belonging to the group 1 is a channel coded data item, does not constitute a switching group, and is included in the stream 1. Further, the illustrated correspondence shows that the encoded data item belonging to the group 2 is an object coded data item for immersive sound (immersive audio object coded data item), does not constitute a switching group, and is included in the stream 1.
[0082] Further, the illustrated correspondence shows that the encoded data item belonging to the group 3 is an object coded data item for a speech language according to a first language (speech dialogue object coded data item), constitutes a switching group 1, and is included in the stream 2. Further, the illustrated correspondence shows that the encoded data item belonging to the group 4 is an object coded data item for a speech language according to a second language (speech dialogue object coded data item), constitutes a switching group 1, and is included in the stream 2.
[0083] In addition, the illustrated correspondence shows that the coded data items belonging to group 5 are SAOC object coded data items, constitute switching group 2, and are included in stream 1. In addition, the illustrated correspondence shows that the coded data items belonging to group 6 are HAO object coded data items, constitute switching group 2, and are included in stream 1.
[0084] Furthermore, the illustrated correspondence shows that the coded data items belonging to group 5 are first audio description object coded data items, constitute switching group 3, and are included in stream 2. Furthermore, the illustrated correspondence shows that the coded data items belonging to group 8 are second audio description object coded data items, constitute switching group 3, and are included in stream 2.
[0085] In addition, the illustrated correspondence shows that the preset group 1 includes group 1, group 2, group 3, and group 7. In addition, the illustrated correspondence shows that the preset group 2 includes group 1, group 2, group 4, and group 8.
[0086] return Figure 1 The service transmitter 100 inserts attribute information indicating the corresponding attributes of the coded data items of the multiple groups included in the 3D audio transmission data item into the container layer. Furthermore, the service transmitter 100 inserts stream correspondence information indicating which audio streams the coded data items of the multiple groups are included in into the container layer. In this embodiment, the stream correspondence information is considered to indicate, for example, the correspondence between the group ID and the stream identifier.
[0087] The service transmitter 100 inserts attribute information and stream correspondence information as descriptors into an audio elementary stream loop corresponding to one or more audio streams among a predetermined number of audio streams existing under a program map table (PMT: Program Map Table), for example.
[0088] In addition, the service transmitter 100 inserts stream identifier information indicating the stream identifiers of the predetermined number of audio streams into the container layer. For example, the service transmitter 100 inserts the stream identifier information as a descriptor into the audio elementary stream loop corresponding to the predetermined number of audio streams in the program map table (PMT).
[0089] Furthermore, the service transmitter 100 inserts attribute information indicating the respective attributes of the plurality of groups of encoded data items included in the 3D audio transmission data items into the audio stream layer. Furthermore, the service transmitter 100 inserts stream correspondence information indicating which audio streams the plurality of groups of encoded data items are included in into the audio stream layer. For example, the service transmitter 100 inserts the attribute information and stream correspondence information into the PES payloads of PES packets of one or more audio streams from a predetermined number of audio streams.
[0090] In addition, the service transmitter 100 inserts stream identifier information indicating the respective stream identifiers of the predetermined number of audio streams into the layer of the audio stream. The service transmitter 100 inserts the stream identifier information into the PES payloads of the respective PES packets of the predetermined number of audio streams, for example.
[0091] By inserting "Desc", ie, descriptor information, between "SYNC" and "Config", the service transmitter 100 inserts information into the layer of the audio stream, such as Figure 4 As shown in (a) and (b).
[0092] As described above, although this embodiment shows that each information (attribute information, stream correspondence information, stream identifier information) is inserted into both the container layer and the audio stream layer, it is assumed that each information is inserted only into the container layer or only into the audio stream layer.
[0093] The service receiver 200 receives the transport stream TS transmitted from the service transmitter 100 via broadcast waves or network packets. As described above, the transport stream TS includes a predetermined number of audio streams including a plurality of sets of encoded data items configuring 3D audio transmission data items in addition to the video stream.
[0094] Attribute information indicating respective attributes of the plurality of groups of encoded data items included in the 3D audio transmission data item is inserted into a layer of a container and / or a layer of an audio stream, and stream correspondence information indicating in which audio stream the encoded data items in the plurality of groups are respectively included is inserted.
[0095] Based on the attribute information and the stream correspondence information, the service receiver 200 performs a selective decoding process on the audio stream including the encoded data items having a group of attributes suitable for the speaker configuration and the user selection information, and obtains an audio output of 3D audio.
[0096] [Stream generation unit of service sender]
[0097] Figure 9 FIG1 shows a configuration example of the stream generation unit 110 included in the service transmitter 100. The stream generation unit 110 includes a video encoder 112, an audio encoder 113, and a multiplexer 114. Here, the audio transmission data item includes one coded channel data item and two object coded data items, as shown in FIG1 . Figure 3 As shown in .
[0098] The video encoder 112 inputs video data SV, encodes the video data SV, and generates a video stream (video elementary stream). The audio encoder 113 inputs immersive audio and voice dialogue object data items as audio data items SA along with channel data items.
[0099] The audio encoder 113 encodes the audio data item SA and obtains a 3D audio transmission data item. Figure 3 As shown, the 3D audio transmission data item includes a channel coding data item (CD), an immersive audio object coding data item (IAO), and a voice dialogue object coding data item (SDO).
[0100] The audio encoder 113 generates one or more audio streams (audio elementary streams) including a plurality of groups (here, four groups) of encoded data items (see Figure 4 (a), (b)) At this time, as described above, the audio encoder 113 inserts descriptor information ("Desc") including attribute information, stream correspondence information, and stream identifier information between "SYNC" and "Config".
[0101] The multiplexer 114 PES-packetizes the video stream output from the video encoder 112 and a predetermined number of audio streams output from the audio encoder 113 , further PES-packetizes the audio stream for multiplexing, and acquires a transport stream TS as a multiplexed stream.
[0102] Furthermore, the multiplexer 114 inserts attribute information indicating the respective attributes of the coded data items of the plurality of groups and stream correspondence information indicating which audio streams the coded data items in the plurality of groups are included in, under the program map table (PMT). The multiplexer 114 uses a 3D audio stream configuration descriptor (3Daudio_stream_config_descriptor) to insert information into the audio elementary stream loop corresponding to at least one or more audio streams among a predetermined number of audio streams. The details of the descriptor will be described later.
[0103] In addition, the multiplexer 114 inserts stream identifier information indicating the stream identifiers of the predetermined number of audio streams under the program map table (PMT). The multiplexer 114 uses a 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) to insert information into the audio elementary stream loop corresponding to the predetermined number of audio streams. The details of the descriptor will be described later.
[0104] Will briefly describe Figure 9 The operation of the stream generation unit 110 is shown. The video data item is supplied to the video encoder 112. In the video encoder 112, the video data item SV is encoded, and a video stream including the encoded video data item is generated. The video stream is supplied to the multiplexer 114.
[0105] The audio data item SA is supplied to the audio encoder 113. The audio data item SA includes channel data items, object data items of immersive audio and voice dialogue. In the audio encoder 113, the audio data item SA is encoded, and a 3D audio transmission data item is obtained.
[0106] In addition to the channel coded data item (CD), the 3D audio transmission data item also includes an immersive audio object coded data item (IAO) and a voice dialogue object coded data item (SDO) (see Figure 3 In the audio encoder 113, one or more audio streams including four sets of coded data items are generated (see Figure 4 (a), (b)).
[0107] At this time, as described above, the audio encoder 113 inserts descriptor information ("Desc") including attribute information, stream correspondence information, and stream identifier information between "SYNC" and "Config."
[0108] The video stream generated by the video encoder 112 is supplied to the multiplexer 114. In addition, the audio stream generated by the audio encoder 113 is supplied to the multiplexer 114. In the multiplexer 114, the stream supplied from each encoder is subjected to PES packetization and transport packetization for multiplexing, and a transport stream TS as a multiplexed stream is obtained.
[0109] Furthermore, in the multiplexer 114, for example, a 3D audio stream configuration descriptor is inserted into an audio elementary stream loop corresponding to at least one or more of the predetermined number of audio streams. The descriptor includes attribute information indicating respective attributes of the plurality of groups of coded data items, and stream correspondence information indicating in which audio stream each of the plurality of groups of coded data items is contained.
[0110] In addition, in the multiplexer 114, a 3D audio stream ID descriptor is inserted into the audio elementary stream loop corresponding to the respective predetermined number of audios. The descriptor includes stream identifier information indicating the respective stream identifiers of the predetermined number of audio streams.
[0111] [Details of 3D audio stream configuration descriptor]
[0112] Figure 10 The structure example (syntax) of the 3D audio stream configuration descriptor is shown. Figure 11 Shows the contents (semantics) of the main information in the configuration instance.
[0113] The 8-bit field "descriptor_tag" indicates the descriptor type. Here, it shows that it is a 3D audio stream configuration descriptor. The 8-bit field "descriptor_length" indicates the descriptor length (size) and shows the number of bytes that follow as the descriptor length.
[0114] 8-bit field "NumOfGroups, N" indicates the number of groups. 8-bit field "NumOfPresetGroups, P" indicates the number of preset groups. For the number of groups, 8-bit field "group ID", 8-bit field "attribute of group ID", 8-bit field "SwitchGroupID" and 8-bit field "audio_streamID" are repeated.
[0115] Field "group ID" indicates the identifier of a group. Field "attribute of group ID" indicates the relevant attribute of the encoded data item of a group. Field "SwitchGroupID" is the identifier indicating the switch group to which the relevant group belongs. "0" indicates that it does not belong to any switch group. Non-"0" indicates the switch group to which it belongs. 8-bit field "contentKind" indicates the kind of the content of a group. "audio_streamID" is the identifier indicating the audio stream including the relevant group. Figure 12 The kind of the content defined in "contentKind" is shown.
[0116] For the number of preset groups, 8-bit field "presetGroupID" and 8-bit field "NumOfGroups_in_preset, R" are repeated. Field "presetGroupID" is the identifier indicating the bundle of preset groups. Field "NumOfGroups_in_preset, R" indicates the number of groups belonging to a preset group. 8-bit field "group ID" is repeated for each preset group (for the number of groups belonging thereto), and groups belonging to a preset group are shown. A descriptor can be set under an extended descriptor.
[0117] [Details of 3D audio stream ID descriptor]
[0118] Figure 13 (a) A configuration example (syntax) of 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) is shown. Figure 13 (b) The content (semantics) of the main information in the configuration example is shown.
[0119] 8-bit field "descriptor_tag" indicates the descriptor type. Here, it shows that it is a 3D audio stream ID descriptor. 8-bit field "descriptor_length" indicates the descriptor length (size), and the number of subsequent bytes is indicated as the descriptor length. 8-bit field "audio_streamID" indicates the identifier of an audio stream. A descriptor can be set under an extended descriptor.
[0120] [Configuration of transport stream TS]
[0121] Figure 14 The configuration example of the transport stream TS is shown. This configuration example corresponds to the case where 3D audio transmission data items are transmitted through two streams (see Figure 5 ). In this configuration example, there is a video stream PES packet "video PES" identified by PID1. In addition, in this configuration example, there are two audio stream PES packets "audio PES" identified by PID2 and PID3, respectively. The PES packet includes a PES header (PES_header) and a PES payload (PES_payload). DTS and PTS timestamps are inserted into the PES header. During multiplexing, the PID2 and PID3 timestamps are matched to provide accuracy, thereby ensuring synchronization between them throughout the system.
[0122] Here, the audio stream PES packet "Audio PES" identified by PID2 includes channel coded data items (CD) classified as group 1 and immersive audio object coded data items (IAO) classified as group 2. In addition, the audio stream PES packet "Audio PES" identified by PID3 includes speech dialog object coded data items (SDO) according to the first language classified as group 3, and speech dialog object coded data items (SDO) according to the second language classified as group 4.
[0123] The transport stream TS also includes a PMT (Program Map Table) as PSI (Program Specific Information). PSI describes which programs each elementary stream included in the transport stream belongs to. The PMT contains a program loop that describes information about the entire program.
[0124] In addition, an elementary stream loop having information about each elementary stream exists in the PMT. In this configuration example, there is a video elementary stream loop (video ES loop) corresponding to the video stream, and there is an audio elementary stream loop (audio ES loop) corresponding to the two audio streams.
[0125] In the video elementary stream loop (video ES loop), information about the stream type, PID (packet identifier), etc. is set for the video stream, and descriptors describing information related to the video stream are also set. The value of the video stream "Stream_type" is set to "0x24", and the PID information indicates PID1 added to the video stream PES packet "Video PES" as described above. As one of the descriptors, the HEVC descriptor is set.
[0126] At each audio elementary stream loop (audio ES loop), information about the stream type, PID (packet identifier), and so on is set for the audio stream, and a descriptor describing information related to the audio stream is also set. PID2 is the main audio stream, and the value of "Stream_type" is set to "0x2C." The PID information indicates the PID added to the audio stream PES packet "Audio PES," as described above. In addition, PID3 is the sub-audio stream, and the value of "Stream_type" is set to "0x2D." The PID information indicates the PID added to the audio stream PES packet "Audio PES," as described above.
[0127] In addition, at each audio elementary stream loop (Audio ES loop), both the above-mentioned 3D audio stream configuration descriptor and 3D audio stream ID descriptor are set.
[0128] In addition, the descriptor information is inserted into the PES payload of the PES packet of each audio elementary stream. The descriptor information is "Desc" inserted between "SYNC" and "Config" as described above (see Figure 4 ). Assuming that information included in the 3D audio stream configuration descriptor is represented as D1, and information included in the 3D audio stream ID descriptor is represented as D2, the descriptor information includes "D1+D2" information.
[0129] [Configuration example of service receiver]
[0130] Figure 15 1 shows an example of a configuration of a service receiver 200. The service receiver 200 includes a receiving unit 201, a demultiplexer 202, a video decoder 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. In addition, the service receiver 200 includes multiplexing buffers 211-1 to 211-N, a combiner 212, a 3D audio decoder 213, a sound output processing circuit 214, and a speaker system 215. In addition, the service receiver 200 includes a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiving unit 225, and a remote control transmitter 226.
[0131] The CPU 221 controls the operation of each unit in the service receiver 200. The flash ROM 222 stores control software and saves data. The DRAM 223 configures a work area for the CPU 221. The CPU 221 decompresses software or data read from the flash ROM 222 onto the DRAM 223 to start the software and controls each unit in the service receiver 200.
[0132] The remote control receiving unit 225 receives a remote control signal (remote control code) transmitted from the remote control transmitter 226, and supplies it to the CPU 221. The CPU 221 controls each unit in the service receiver 200 based on the remote control code. The CPU 221, the flash ROM 222, and the DRAM 223 are connected to the internal bus 224.
[0133] The receiving unit 201 receives a transport stream TS transmitted from the service transmitter 100 on a broadcast wave or a network packet. The transport stream TS includes a predetermined number of audio streams in addition to a video stream, which includes a plurality of groups of encoded data items configured 3D audio transmission data items.
[0134] Figure 16 An example of the received audio stream is shown. Figure 16 (a) An example of one stream (main stream) is shown. The stream includes channel encoded data items (CD), immersive audio object encoded data items (IAO), and speech dialogue object encoded data items (SDO) along with "SYNC" and "Config". The stream is identified by PID 2.
[0135] In addition, between "SYNC" and "Config", descriptor information ("Desc") is included. Attribute information indicating respective attributes of the plurality of groups of encoded data items, stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included, and stream identifier information indicating a self stream identifier are inserted into the descriptor information.
[0136] Figure 16 (b) An example of two streams is shown. A main stream identified by PID 2 includes channel encoded data items (CD) and immersive audio object encoded data items (IAO) along with "SYNC" and "Config". In addition, a sub stream identified by PID 3 includes speech dialogue object encoded data items (SDO) along with "SYNC" and "Config".
[0137] In addition, each stream includes descriptor information ("Desc") between "SYNC" and "Config". Attribute information indicating respective attributes of the plurality of groups of encoded data items, stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included, and stream identifier information indicating a self stream identifier are inserted into the descriptor information.
[0138] The demultiplexer 202 extracts a video stream packet from the transport stream TS, and transmits it to the video decoder 203. The video decoder 203 reconfigures a video stream of the video packet extracted by the demultiplexer 202, and performs a decoding process to obtain an uncompressed video data item.
[0139] The video processing circuit 204 performs scaling processing, image quality adjustment processing, and the like on the video data item obtained at the video decoder 203, thereby obtaining a video data item for display. The panel drive circuit 205 drives the display panel 206 based on the image data item for display obtained at the video processing circuit 204. The display panel 206 includes, for example, an LCD (Liquid Crystal Display), an organic EL display (Organic Electroluminescence Display), or the like.
[0140] In addition, the demultiplexer 202 extracts various information such as descriptor information from the transport stream TS, and sends it to the CPU 221. The various information also includes the above-described information on the 3D audio stream configuration descriptor (3Daudio_stream_config_descriptor) and the 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) (see FIG. 6). Figure 14 ).
[0141] Based on the attribute information indicating the attribute of the encoded data item of each group included in the descriptor information and the stream relationship information indicating in which video stream each group is included, the CPU 221 identifies the video stream including the encoded data item of the group having an attribute suitable for the speaker configuration and the viewer and audience (user) selection information.
[0142] In addition, the demultiplexer 202 selectively extracts one or more audio stream packets including the encoded data item of the group having an attribute suitable for the speaker configuration and the viewer and audience (user) selection information from among the predetermined number of audio streams possessed by the transport stream TS under the control of the CPU 221 by means of a PID filter.
[0143] The multiplexing buffers 211-1 to 211-N each collect each audio stream taken out at the demultiplexer 202. Here, the N multiplexing buffers 211-1 to 211-N are necessary and sufficient. In actual operation, a plurality of audio streams taken out at the demultiplexer 202 will be used.
[0144] The combiner 212 reads the audio stream every audio frame from the multiplexing buffer receiving each audio stream taken out at the demultiplexer 202 among the multiplexing buffers 211-1 to 211-N, and sends it to the 3D audio decoder 213.
[0145] When the audio stream supplied from the combiner 212 includes descriptor information ("Desc"), the 3D audio decoder 213 transmits the descriptor information to the CPU 221. Under the control of the CPU 221, the 3D audio decoder 213 selectively extracts a set of encoded data items having attributes suitable for the speaker configuration and viewer and audience (user) selection information, performs decoding processing, and obtains audio data items for driving each speaker of the speaker system 215.
[0146] Here, the coded data items to which the decoding process is applied may have three modes: including only channel coded data items, including only object coded data items, or including both channel coded data items and object coded data items.
[0147] When decoding channel coded data items, the 3D audio decoder 213 performs downmixing or upmixing processing on the speaker configuration of the speaker system 215, obtaining audio data items for driving each speaker. In addition, when decoding object coded data items, the 3D audio decoder 213 calculates speaker rendering (mixing ratio for each speaker) based on object information (metadata items), and mixes the object audio data items into the audio data items based on the calculation results to drive each speaker.
[0148] The sound output processing circuit 214 performs necessary processing such as D / A conversion, amplification, etc. on the audio data items for driving each speaker obtained at the 3D audio decoder 213, and supplies them to the speaker system 215. The speaker system 215 includes a plurality of speakers having multiple channels, for example, 2 channels, 5.1 channels, 7.1 channels, or 22.2 channels.
[0149] Will briefly describe Figure 15 The operation of the service receiver 200 is shown. The receiving device 201 receives the transport stream TS transmitted by the service transmitter 100 via broadcast waves or network packets. In addition to the video stream, the transport stream TS includes a predetermined number of audio streams, each of which includes multiple sets of encoded data items, which constitute 3D audio transmission data items. The transport stream TS is supplied to the demultiplexer 202.
[0150] In the demultiplexer 202, video stream packets are extracted from the transport stream TS, and the video stream packets are supplied to the video decoder 203. In the video decoder 203, the video stream is reconfigured from the video packets extracted at the demultiplexer 202, and decoding processing is performed to obtain uncompressed video data items. The video data items are supplied to the video processing circuit 204.
[0151] The video processing circuit 204 performs scaling processing, image quality adjustment processing, and the like on the video data items obtained at the video decoder 203, thereby obtaining video data items for display. The video data items for display are supplied to the panel drive circuit 205. The panel drive circuit 205 drives the display panel 206 based on the image data items for display. In this way, an image corresponding to the image data items for display is displayed on the display panel 206.
[0152] In addition, the demultiplexer 202 extracts various information such as descriptor information from the transport stream TS, which is delivered to the CPU 221. The various information also includes information on a 3D audio stream configuration descriptor and a 3D audio stream ID descriptor (see FIG. 6). Figure 14 Based on the attribute information and stream relationship information contained in the descriptor information, the CPU 221 identifies a video stream including encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience (user) selection information.
[0153] Further, the demultiplexer 202 selectively extracts one or more audio stream packets including encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience selection information from among a predetermined number of audio streams possessed by the transport stream TS under the control of the CPU 221 by a PID filter.
[0154] The audio stream taken out at the demultiplexer 202 is received into a corresponding one of the multiplex buffers 211-1 to 211-N. In the combiner 212, an audio stream is read out from each audio frame of each multiplex buffer that receives an audio stream, and the audio stream is supplied to the 3D audio decoder 213.
[0155] In a case where the audio stream supplied from the combiner 212 includes descriptor information ("Desc"), the descriptor information is extracted and sent to the CPU 221 in the 3D audio decoder 213. The 3D audio decoder 213 selectively extracts encoded data items of a group having attributes suitable for a speaker configuration and viewer and audience (user) selection information under the control of the CPU 221, performs decoding processing, and obtains audio data items for driving each speaker of the speaker system 215.
[0156] Here, when decoding a channel encoded data item, downmix or upmix processing is performed on a speaker configuration of the speaker system 215, and an audio data item for driving each speaker is obtained. Further, when decoding an object encoded data item, speaker rendering (mixing ratio for each speaker) is calculated based on object information (metadata item), and an object audio data item is mixed into an audio data item in accordance with the calculation result so as to drive each speaker.
[0157] The audio data items for driving each speaker acquired at the 3D audio decoder 213 are supplied to the sound output processing circuit 214. The sound output processing circuit 214 performs necessary processing such as D / A conversion and amplification on the audio data items for driving each speaker. The processed audio data items are supplied to the speaker system 215. In this way, audio output corresponding to the display image of the display panel 206 is obtained from the speaker system 215.
[0158] Figure 17 The decoding process in the case where descriptor information does not exist in the audio stream is schematically shown. The transport stream TS as a multiplexed stream is input to the demultiplexer 202. In the demultiplexer 202, the system layer is analyzed and descriptor information 1 (information about the 3D audio stream configuration descriptor or the 3D audio stream ID descriptor) is supplied to the CPU 221.
[0159] In the CPU 221, an audio stream including a group of coded data items having attributes suitable for the speaker configuration and viewer and audience (user) selection information is identified based on the descriptor information 1. In the demultiplexer 202, selection between streams is performed under the control of the CPU 221.
[0160] In other words, in the demultiplexer 202, the PID filter selectively extracts one or more audio stream packets from a predetermined number of audio streams in the transport stream TS. The extracted one or more audio stream packets include a set of encoded data items having attributes suitable for the speaker configuration and the viewer and audience selection information. The audio streams thus extracted are received into the multiplexing buffer 211 (211-1 to 211-N).
[0161] The 3D audio decoder 213 performs packet type analysis on each audio stream received at the multiplexing buffer 211. Then, in the demultiplexer 202, under the control of the CPU 221, selection within the stream is performed based on the above-mentioned descriptor information 1.
[0162] Specifically, a group of encoded data items having attributes applicable to the speaker configuration and viewer and audience (user) selection information is selectively taken out from each audio stream as a decoding object, and decoding processing and mixed rendering processing are applied thereto to obtain audio data items (uncompressed audio) for driving each speaker.
[0163] Figure 18 An example of the configuration of an audio access unit (audio frame) of an audio stream is shown in the case where no descriptor information exists in the audio stream. Here, an example of two streams is shown.
[0164] With respect to the audio stream identified by PID2, the information "FrWork#ch=2, #obj=1" included in the "Config" indicates that there is a "frame" including a channel coded data item and an object coded data item in two channels. The information "GroupID[0]=1, GroupID[1]=2" registered in this order within the "AudioSceneInfo()" included in the "Config" indicates that a "frame" having a coded data item of group 1 and a "frame" having a coded data item of group 2 are set in this order. Note that the value of the packet label (PL) is considered to be the same in the "Config" and each "frame" corresponding thereto.
[0165] Here, the "frame" having a coded data item of group 1 includes a coded sample data item of a CPE (channel pair element). Further, the "frame" having a coded data item of group 2 includes a "frame" having a metadata item as an extension element (Ext_element), and a "frame" having a coded sample data item of a SCE (single channel element).
[0166] With respect to the audio stream identified by PID3, the information "FrWork#ch=0, #obj=2" included in the "Config" indicates that there is a "frame" including two object coded data items. The information "GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=1" registered in this order within the "AudioSceneInfo()" included in the "Config" indicates that a "frame" having a coded data item of group 3 and a "frame" having a coded data item of group 4 are set in this order and these group configuration switching group 1. Note that the value of the packet label (PL) is considered to be the same in the "Config" and each "frame" corresponding thereto.
[0167] Here, the "frame" having a coded data item of group 3 includes a "frame" having a metadata item as an extension element (Ext_element) and a "frame" having a coded sample data item of a SCE (single channel element). Similarly, the "frame" having a coded data item of group 4 includes a "frame" having a metadata item as an extension element (Ext_element) and a "frame" having a coded sample data item of a SCE (single channel element).
[0168] Figure 19 A diagram schematically showing a decoding process in a case where descriptor information is present within an audio stream. A transport stream TS which is a multiplexed stream is input to a demultiplexer 202. In the demultiplexer 202, the system layer is analyzed, and descriptor information 1 (information on a 3D audio stream configuration descriptor or a 3D audio stream ID descriptor) is supplied to a CPU 221.
[0169] In the CPU 221, an audio stream including a group of coded data items having attributes suitable for the speaker configuration and viewer and audience (user) selection information is identified based on the descriptor information 1. In the demultiplexer 202, selection between streams is performed under the control of the CPU 221.
[0170] In other words, the demultiplexer 202 selectively extracts one or more audio stream packets from a predetermined number of audio streams in the transport stream TS through a PID filter. Each of the extracted audio stream packets includes a set of coded data items having attributes suitable for the speaker configuration and viewer and audience selection information. The audio streams thus extracted are received into the multiplexing buffer 211 (211-1 to 211-N).
[0171] The 3D audio decoder 213 performs packet type analysis on each audio stream received in the multiplexing buffer 211 and sends the descriptor information 2 present in the audio stream to the CPU 221. Based on the descriptor information 2, the presence of a group of encoded data items having attributes suitable for speaker configuration and viewer and audience (user) selection information is identified. Then, in the demultiplexer 202, under the control of the CPU 221, selection is performed within the stream based on the descriptor information 2.
[0172] Specifically, a group of encoded data items having attributes applicable to the speaker configuration and viewer and audience (user) selection information is selectively taken out from each audio stream as a decoding object, and decoding processing and mixed rendering processing are applied thereto to obtain audio data items (uncompressed audio) for driving each speaker.
[0173] Figure 20 An example of the configuration of an audio access unit (audio frame) of an audio stream is shown in the case where descriptor information exists in the audio stream. Here, an example of two streams is shown. Figure 20 Similar to Figure 18 , except that "Desc", the descriptor information, is inserted between "SYNC" and "Config".
[0174] Regarding the audio stream identified by PID2, the information "GroupID[0]=1, channeldata" included in "Desc" indicates that the encoded data items of group 1 are channel encoded data items. The information "GroupID[1]=2, object sound" included in "Desc" indicates that the encoded data items of group 2 are object encoded data items for immersive sound. In addition, the information "Stream_ID" indicates the stream identifier of the audio stream.
[0175] Regarding the audio stream identified by PID3, the information "GroupID[2]=3, object lang1" included in "Desc" indicates that the coded data items of group 3 are object coded data items according to the speech language of the first language. The information "GroupID[3]=4, object lang2" included in "Desc" indicates that the coded data items of group 4 are object coded data items according to the speech language of the second language. In addition, the information "SW_GRPID[0]=1" included in "Desc" indicates that groups 3 and 4 configure switch group 1. In addition, the information "Stream_ID" indicates the stream identifier of the audio stream.
[0176] Figure 21 An example of the configuration of an audio access unit (audio frame) of an audio stream is shown in the case where descriptor information exists in the audio stream. Here, an example of one stream is shown.
[0177] The information of "FrWork#ch=2, #obj=3" included in "Config" indicates that there are "frames" of channel coded data items and three object coded data items included in two channels. The information of "GroupID[0]=1, GroupID[1]=2, GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=1" sequentially registered in "AudioSceneInfo()" included in "Config" indicates that "frames" with coded data items of group 1 and "frames" with coded data items of group 2, "frames" with coded data items of group 3 and "frames" with coded data items of group 4 are set in this order, and these group 3 and group 4 configure switching group 1. Note that the value of the packet tag (PL) is considered to be the same in "Config" and each "frame" corresponding thereto.
[0178] Here, the "frame" having the coded data items of Group 1 includes coded sample data items of CPE (Channel Pair Element). In addition, the "frame" having the coded data items of Groups 2 to 4 includes "frames" having metadata items as extension elements (Ext_element) and "frames" having coded sample data items of SCE (Mono Channel Element).
[0179] The information "GroupID[0]=1, channeldata" included in "Desc" indicates that the coded data items of group 1 are channel coded data items. The information "GroupID[1]=2, object sound" included in "Desc" indicates that the coded data items of group 2 are object coded data items for immersive sound.
[0180] The information "GroupID[2]=3, object lang1" included in "Desc" indicates that the coded data items of group 3 are object coded data items according to the second language. The information "GroupID[3]=4, object lang2" included in "Desc" indicates that the coded data items of group 4 are object coded data items according to the second language. In addition, the information "SW_GRPID[0]=1" included in "Desc" indicates that groups 3 and 4 constitute switch group 1. Furthermore, the information "Stream_ID" indicates the stream identifier of the audio stream.
[0181] Figure 22 and Figure 23 The flowchart in Figure 15 1 shows an example of audio decoding control processing by the CPU 221 in the service receiver 200. The CPU 221 starts processing in step ST1. Then, the CPU 221 detects the receiver speaker configuration, that is, the speaker configuration of the speaker system 215 in step ST2. Next, the CPU 221 obtains selection information about the audio output by the viewer and the audience (user) in step ST3.
[0182] Next, in step ST4, the CPU 221 reads the descriptor information about the main stream in the PMT, selects the audio stream to which the group having the attributes suitable for the speaker configuration and the viewer and audience selection information belongs, and receives it into the buffer. Then, in step ST5, the CPU 221 checks whether a descriptor type packet exists in the audio stream.
[0183] Next, in step ST6, the CPU 221 determines whether a descriptor type packet exists. If so, in step ST7, the CPU 221 reads the descriptor information of the relevant packet, detects the information of "groupID", "attribute", "switchGroupID", and "presetGroupID", and then proceeds to the processing in step ST9. On the other hand, if it does not exist, the CPU 221 detects the information of "groupID", "attribute", "switchGroupID", and "presetGroupID" from the descriptor information of the PMT in step ST8, and then proceeds to step ST9. It should be noted that step ST8 does not need to be executed, and the entire audio stream can be decoded.
[0184] In step ST9, the CPU 221 determines whether to decode the target coded data item. If decoded, the CPU 221 decodes the target coded data item in step ST10 and then proceeds to the processing in step ST11. On the other hand, if not decoded, the CPU 221 immediately proceeds to the processing in step ST11.
[0185] In step ST11, the CPU 221 determines whether the channel coded data item is decoded. If it is decoded, the CPU 221 decodes the channel coded data item as needed in step ST12, performs downmixing or upmixing processing on the speaker configuration of the speaker system 215, and obtains audio data items for driving each speaker. Thereafter, the CPU 221 proceeds to the processing in step ST13. On the other hand, if it is not decoded, the CPU 221 immediately proceeds to the processing in step ST13.
[0186] In step ST13, when the CPU 221 decodes the target coded data item, it is mixed with the channel data item or a speaker rendering is calculated based on the information. In the speaker rendering calculation, the speaker rendering (mixing ratio of each speaker) is calculated using the azimuth (azimuth information) and elevation (elevation information). Based on the calculation result, the target audio data item is mixed with the channel data to drive each speaker.
[0187] Next, the CPU 221 performs dynamic range control of the audio data item for driving each speaker and outputs it in step ST 14. Thereafter, the CPU 221 ends the processing in step ST 15.
[0188] As mentioned above, in Figure 1 In the illustrated transmission / reception system 10, the service transmitter 100 inserts attribute information indicating respective attributes of a plurality of groups of coded data items included in a predetermined number of audio streams into a container layer and / or an audio stream layer. Therefore, on the receiving side, the respective attributes of the plurality of groups of coded data items can be easily identified before decoding the coded data items, and only the necessary groups of coded data items can be selectively decoded and used, thereby reducing the processing load.
[0189] exist Figure 1 In the illustrated transmission / reception system 10, the service transmitter 100 inserts stream correspondence information indicating which audio streams contain each of the plurality of groups of coded data items into the container layer and / or the audio stream layer. This makes it easy for the receiving end to identify the audio stream containing the necessary group of coded data items, thereby reducing the processing load.
[0190] <2. Modifications>
[0191] In the above-described embodiments, the service receiver 200 selectively takes out one or more audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information from among a plurality of audio streams transmitted via the service transmitter 100, performs a decoding process, and obtains a predetermined number of audio data items for driving the speakers.
[0192] However, it is conceivable that the service receiver selectively takes out one or more audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information from among a plurality of audio streams transmitted via the service transmitter 100, reconfigures the audio streams including the encoded data items having the group of attributes suitable for the speaker configuration and the viewer and audience selection information, and transmits the reconfigured audio streams to devices connected to the in-house network (including also DLNA devices).
[0193] Figure 24 A configuration example of the service receiver 200A that transmits the reconfigured audio streams to the devices connected to the in-house network is shown, as described above. The components corresponding to those in Figure 15 Figure 24 The components in
[0194] The demultiplexer 202 selectively takes out packets of one or more audio streams having a predetermined number of audio streams in the transport stream TS under the control of the CPU 221 by a PID filter, the taken-out audio streams including encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience selection information.
[0195] The audio streams taken out by the demultiplexer 202 are received into the respective ones of the multiplex buffers 211-1 to 211-N. In the combiner 212, the audio streams are read out from each audio frame of each multiplex buffer that receives the audio, and the audio streams are supplied to the stream reconfiguration unit 231.
[0196] In the stream reconfiguration unit 231, descriptor information ("Desc") is extracted and transmitted to the CPU 221 in the case where the descriptor information is contained in the audio streams supplied from the combiner 212. In the stream reconfiguration unit 231, the encoded data items having a group of attributes suitable for a speaker configuration and viewer and audience (user) selection information are selectively acquired under the control of the CPU 221, and the audio streams having the encoded data items are reconfigured. The reconfigured audio streams are supplied to the transmission interface 232. Then, they are transmitted (sent) from the transmission interface 232 to the devices 300 connected to the in-house network.
[0197] In-room network connections include Ethernet and wireless connections using "WiFi" or "Bluetooth." "WiFi" and "Bluetooth" are registered trademarks.
[0198] In addition, the device 300 includes surround speakers, a second display, and an audio output device attached to the network terminal. The device 200 to which the reconfigured audio stream is transmitted performs the same operation as in the Figure 15 The 3D audio decoder 213 in the service receiver 200 performs similar decoding processing and obtains audio data items for driving a predetermined number of speakers.
[0199] Furthermore, as a service receiver, it is conceivable to transmit the above-described reconfigured audio stream to a device connected with a digital interface such as "HDMI (High-Definition Multimedia Interface)", "MHL (Mobile High-Definition Link)" and "DisplayPort". "HDMI" and "MHL" are registered trademarks.
[0200] In the above-described embodiment, the stream correspondence information inserted into the container layer, etc., is information indicating the correspondence between the group ID and the substream ID. Specifically, the substream ID is used to associate the group with the audio stream. However, it is conceivable to use a packet identifier (PID: packet ID) or a stream type (stream_type) to associate the group with the audio stream. In the case of using a stream type, the stream type should be changed for each audio stream.
[0201] In addition, the above embodiment shows an example in which the attribute information of the coded data item of each group is transmitted by setting the "attribute_of_groupID" field (see Figure 10 ). However, the present technology also includes a method that can identify the type (attribute) of the encoded data item by defining a specific meaning in the value of the group ID (GroupID) between the transmitter and the receiver to identify the specific group ID. In this case, the group ID is used as an identifier of the group and also as attribute information of the encoded data item of the group, so that the field "attribute_of_groupID" becomes unnecessary.
[0202] In addition, the above embodiment shows an example in which a plurality of sets of coded data items include both channel coded data items and object coded data items (see Figure 3 ). However, the present technology can also be similarly applied to a case where a plurality of groups of encoded data items include only channel encoded data items or only object encoded data items.
[0203] Further, the above-described embodiments show an example in which the container is a transport stream (MPEG-2 TS). However, the present technology can also be similarly applied to a system in which a stream is delivered by a container of a format such as MP4. For example, the system includes an MPEG-DASH basic stream delivery system, or a transmission / reception system that handles an MMT (MPEG Media Transport) structure transmission stream.
[0204] The present technology can also have the following configuration.
[0205] 1. A transmission apparatus comprising:
[0206] a transmission unit that transmits a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items; and
[0207] an information insertion unit that inserts attribute information indicating respective attributes of the plurality of groups of encoded data items into a layer of the container and / or a layer of the audio streams.
[0208] (2) The transmission apparatus according to (1), wherein
[0209] the information insertion unit further inserts stream correspondence relationship information into the layer of the container and / or the layer of the audio streams, the stream correspondence relationship information indicating in which audio stream each of the plurality of groups of encoded data items is included.
[0210] (3) The transmission apparatus according to (2), wherein
[0211] the stream correspondence relationship information is information indicating a correspondence relationship between a group identifier identifying each of the plurality of groups of encoded data items and a stream identifier identifying each of the predetermined number of audio streams.
[0212] (4) The transmission apparatus according to (3), wherein
[0213] the information insertion unit further inserts stream identifier information into a layer of the container and / or a layer of the audio streams, the stream identifier information indicating the stream identifier of each of the predetermined number of audio streams.
[0214] (5) The transmission apparatus according to (4), wherein
[0215] the container is an MPEG2-TS, and
[0216] in a case where the stream identifier information is inserted into the container, the information insertion unit inserts the stream identifier information into an audio elementary stream loop corresponding to each of the predetermined number of audio streams present under a program map table.
[0217] (6) The transmitting device according to (4) or (5) above, wherein:
[0218] In a case where the stream identifier information is inserted into the audio stream, the information inserting unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.
[0219] (7) The transmitting device according to (2), wherein
[0220] The stream correspondence information is information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a packet identifier added in a case where each of the predetermined number of audio streams is packetized.
[0221] (8) The transmitting device according to (2), wherein
[0222] The stream correspondence information is information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and type information indicating a stream type of each of the predetermined number of audio streams.
[0223] (9) The transmitting device according to any one of (2) to (8) above, wherein:
[0224] The container is MPEG2-TS, and
[0225] In the case where the attribute information and the stream correspondence information are inserted into the container, the information insertion unit inserts the attribute information and the stream correspondence information into the audio elementary stream loop corresponding to at least one or more audio streams among the predetermined number of audio streams existing under the program map table.
[0226] (10) The transmitting device according to any one of (2) to (8) above, wherein:
[0227] When the attribute information and the stream correspondence information are inserted into the audio stream, the information inserting unit inserts the attribute information and the stream correspondence information into the PES payload of the PES packet in at least one or more audio streams among the predetermined number of audio streams.
[0228] (11) The transmitting device according to (1) to (10) above, wherein
[0229] The plurality of groups of the encoded data items include one or both of channel encoded data items and object encoded data items.
[0230] (12) A sending method, comprising:
[0231] a sending step in which a sending unit sends a container in a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items; and
[0232] An information inserting step of inserting attribute information indicating respective attributes of the plurality of groups of the encoded data items into the layer of the container and / or the layer of the audio stream.
[0233] (13) A receiving device comprising:
[0234] a receiving unit receiving a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items,
[0235] Attribute information indicating respective attributes of the plurality of groups of the encoded data items is inserted into a layer of the container and / or a layer of the audio stream; and
[0236] A processing unit processes the predetermined number of audio streams contained in the received container based on the attribute information.
[0237] (14) The receiving device according to (13) above, wherein
[0238] Stream correspondence information indicating in which audio stream the coded data items of the plurality of groups are respectively contained is further inserted into the layer of the container and / or the layer of the audio stream, and
[0239] The processing unit processes the predetermined number of audio streams based on the stream correspondence information in addition to the attribute information.
[0240] (15) The receiving device according to (14) above, wherein
[0241] The processing unit performs selective decoding processing on an audio stream including encoded data items having a group of attributes applicable to a speaker configuration and user selection information, based on the attribute information and the stream correspondence relationship information.
[0242] (16) The receiving device according to any one of (13) to (15) above, wherein
[0243] The plurality of groups of the encoded data items include one or both of channel encoded data items and object encoded data items.
[0244] (17) A receiving method, comprising:
[0245] a receiving step in which a receiving unit receives a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items,
[0246] Attribute information indicating respective attributes of the plurality of groups of the encoded data items is inserted into a layer of the container and / or a layer of the audio stream; and
[0247] A processing step of processing the predetermined number of audio streams included in the received container based on the attribute information.
[0248] (18) A receiving device comprising:
[0249] a receiving unit that receives a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items,
[0250] Attribute information representing respective attributes of the coded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream;
[0251] a processing unit that selectively acquires a predetermined group of coded data items from the predetermined number of audio streams contained in the received container based on the attribute information, and reconfigures an audio stream including the predetermined group of coded data items; and
[0252] A stream sending unit sends the audio stream reconfigured by the processing unit to an external device.
[0253] (19) The receiving device according to (18) above, wherein
[0254] Stream correspondence information indicating in which audio stream the coded data items of the plurality of groups are respectively contained is further inserted into the layer of the container and / or the layer of the audio stream, and
[0255] The processing unit selectively acquires the predetermined group of encoded data items from the predetermined number of audio streams based on the stream correspondence information in addition to the attribute information.
[0256] (20) A receiving method, comprising:
[0257] a receiving step in which a receiving unit receives a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items;
[0258] Attribute information representing respective attributes of the coded data items of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream;
[0259] a processing step of selectively acquiring a predetermined group of coded data items from the predetermined number of audio streams included in the received container based on the attribute information, and reconfiguring the audio stream including the predetermined group of coded data items; and
[0260] A stream sending step of sending the audio stream reconfigured in the processing step to an external device.
[0261] The main feature of the present technology is that stream correspondence information is inserted into a layer of a container and / or a layer of an audio stream, the stream correspondence information indicates which audio stream includes each attribute information, and the attribute information indicates a plurality of groups of coded data items included in a predetermined number of audio streams and corresponding attributes of the plurality of groups of coded data items, thereby making it possible to reduce the processing load on the receiving side (see Figure 14 ).
[0262] Explanation of symbols
[0263] 10 Transmit / Receive System
[0264] 100 Service Sender
[0265] 100 Stream Generation Units
[0266] 112 Video Encoder
[0267] 113 Audio Codec
[0268] 114 Multiplexer
[0269] 200, 200A service receiver
[0270] 201 Receiving Unit
[0271] 202 Demultiplexer
[0272] 203 Video Decoder
[0273] 204 video processing circuit
[0274] 205 panel drive circuit
[0275] 206 display panel
[0276] 211-1 to 211-N multiplexing buffer
[0277] 212 Combiner
[0278] 213 3D Audio Decoder
[0279] 214 Sound output processing circuit
[0280] 215 speaker system
[0281] 221 CPU
[0282] 222 Flash ROM
[0283] 223 DRAM
[0284] 224 internal bus
[0285] 225 Remote Control Receiver Unit
[0286] 226 Remote Control Transmitter
[0287] 231 Stream Reconfiguration Unit
[0288] 232 Send interface
[0289] 300 devices.
Claims
1. A sending device, comprising: a transmitting unit configured to transmit a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of encoded data items, the plurality of groups of encoded data items including one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements; as well as An information insertion unit is configured to insert attribute information indicating respective attributes of the plurality of groups of coded data items into a layer of the container to select an audio stream of a group having attributes applicable to the selection information, and into a layer of the audio stream to selectively decode, at a receiving device, the coded data items of a necessary group within the selected audio stream.
2. The transmitting device according to claim 1, wherein The information inserting unit further inserts stream correspondence information into the layer of the container and / or the layer of the audio stream, the stream correspondence information indicating in which audio stream the encoded data items of the plurality of groups are respectively included.
3. The transmitting device according to claim 2, wherein: The stream correspondence information is information representing the correspondence between a group identifier identifying each of the encoded data items of the multiple groups and a stream identifier identifying each of the predetermined number of audio streams, and the stream correspondence information includes an identifier groupID representing the group, an identifier switchGroupID representing the switching group, and contentKind representing the type of content of the group. The transmitting device according to claim 3 , wherein: The information inserting unit further inserts stream identifier information into the layer of the container and / or the layer of the audio stream, the stream identifier information indicating the stream identifier of each of the predetermined number of audio streams.
5. The transmitting device according to claim 4, wherein: The container is MPEG2-TS, and In a case where the stream identifier information is inserted into the container, the information inserting unit inserts the stream identifier information into an audio elementary stream loop corresponding to each of the predetermined number of audio streams existing under a program map table. The transmitting device according to claim 4 , wherein: In a case where the stream identifier information is inserted into the audio stream, the information inserting unit inserts the stream identifier information into a PES payload of a PES packet of each of the predetermined number of audio streams.
7. The transmitting device according to claim 2, wherein: The stream correspondence information is information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and a packet identifier added in a case where each of the predetermined number of audio streams is packetized.
8. The transmitting device according to claim 2, wherein: The stream correspondence information is information indicating a correspondence between a group identifier identifying each of the encoded data items of the plurality of groups and type information indicating a stream type of each of the predetermined number of audio streams.
9. The transmitting device according to claim 2, wherein: The container is MPEG2-TS, and In the case where the attribute information and the stream correspondence information are inserted into the container, the information insertion unit inserts the attribute information and the stream correspondence information into the audio elementary stream loop corresponding to at least one or more audio streams among the predetermined number of audio streams existing under the program map table.
10. The transmitting device according to claim 2, wherein: When the attribute information and the stream correspondence information are inserted into the audio stream, the information inserting unit inserts the attribute information and the stream correspondence information into the PES payload of the PES packet in at least one or more audio streams among the predetermined number of audio streams.
11. A sending method, comprising: a sending step in which a sending unit sends a container in a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of coded data items, the plurality of groups of coded data items including one or both of channel coded data items and object coded data items, the object coded data items including metadata items as extension elements; as well as An information inserting step of inserting attribute information indicating corresponding attributes of the coded data items of the plurality of groups into a layer of the container to select an audio stream of a group having attributes suitable for the selection information, and into a layer of the audio stream to selectively decode the coded data items of a necessary group within the selected audio stream at a receiving device.
12. The sending method according to claim 11, wherein: The information inserting step further inserts stream correspondence information into the layer of the container and / or the layer of the audio stream, the stream correspondence information including an identifier groupID indicating a group, an identifier switchGroupID indicating a switch group, and contentKind indicating a type of content of the group.
13. A receiving device, comprising: a receiving unit configured to receive a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, the plurality of groups of encoded data items including one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information indicating respective attributes of the plurality of groups of the coded data items is inserted into a layer of the container to select an audio stream of a group having attributes applicable to the selection information, and is inserted into a layer of an audio stream to selectively decode a necessary group of coded data items within the selected audio stream at the receiving device; as well as A processing unit processes the predetermined number of audio streams contained in the received container based on the attribute information. The receiving device according to claim 13 , wherein: Stream correspondence information indicating in which audio stream the coded data items of the plurality of groups are respectively contained is further inserted into the layer of the container and / or the layer of the audio stream, and The processing unit processes the predetermined number of audio streams based on the stream correspondence information in addition to the attribute information.
15. The receiving device according to claim 14, wherein: Based on the attribute information and the stream correspondence information, the processing unit performs selective decoding processing on an audio stream containing a group of encoded data items having attributes suitable for speaker configuration and user selection information, and the stream correspondence information includes an identifier groupID representing a group, an identifier switchGroupID representing a switching group, and contentKind representing the type of content of the group.
16. A receiving method, comprising: a receiving step in which a receiving unit receives a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of coded data items, the plurality of groups of coded data items including one or both of channel coded data items and object coded data items, the object coded data items including metadata items as extension elements, Attribute information indicating respective attributes of the plurality of groups of coded data items is inserted into a layer of the container to select an audio stream of a group having attributes applicable to the selection information, and is inserted into a layer of the audio stream to selectively decode a necessary group of coded data items within the selected audio stream at a receiving device; as well as A processing step of processing the predetermined number of audio streams included in the received container based on the attribute information.
17. The receiving method according to claim 16, wherein: Stream correspondence information is inserted into the layer of the container and / or the layer of the audio stream, the stream correspondence information including an identifier groupID indicating a group, an identifier switchGroupID indicating a switch group, and contentKind indicating a kind of content of the group.
18. A receiving device, comprising: a receiving unit configured to receive a container of a predetermined format having a predetermined number of audio streams, the predetermined number of audio streams including a plurality of groups of encoded data items, the plurality of groups of encoded data items including one or both of channel encoded data items and object encoded data items, the object encoded data items including metadata items as extension elements, attribute information indicating respective attributes of the plurality of groups of the coded data items is inserted into a layer of the container to select an audio stream of a group having attributes applicable to the selection information, and is inserted into a layer of an audio stream to selectively decode a necessary group of coded data items within the selected audio stream at the receiving device; a processing unit that selectively acquires a predetermined group of coded data items from the predetermined number of audio streams contained in the received container based on the attribute information, and reconfigures the audio stream including the predetermined group of coded data items; as well as A stream sending unit sends the audio stream reconfigured by the processing unit to an external device.
19. The receiving device according to claim 18, wherein Stream correspondence information is inserted into the layer of the container and / or the layer of the audio stream, the stream correspondence information including an identifier groupID indicating a group, an identifier switchGroupID indicating a switch group, and contentKind indicating a type of content of the group.
20. A receiving method, comprising: a receiving step in which a receiving unit receives a container of a predetermined format having a predetermined number of audio streams, the audio streams including a plurality of groups of coded data items, the plurality of groups of coded data items including one or both of channel coded data items and object coded data items, the object coded data items including metadata items as extension elements, Attribute information indicating respective attributes of the plurality of groups of coded data items is inserted into a layer of the container to select an audio stream of a group having attributes applicable to the selection information, and is inserted into a layer of the audio stream to selectively decode a necessary group of coded data items within the selected audio stream at a receiving device; a processing step of selectively acquiring a predetermined group of coded data items from the predetermined number of audio streams included in the received container based on the attribute information, and reconfiguring the audio stream including the predetermined group of coded data items; as well as A stream sending step of sending the audio stream reconfigured in the processing step to an external device.
Citation Information
Patent Citations
System and tools for enhanced 3D audio authoring and rendering
CN103650535A
Transmission device, transmission method, receiving device and receiving method
CN103843330A
Information processing device, information processing method, program, and data structure
CN1926872A