Information processing device, information processing method, and program

JP7899285B2Active Publication Date: 2026-08-03SONY GROUP CORP
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0025】 本技術によれば、複数種類のオーディデータを送信する場合にあって受信側の処理負荷を軽減することが可能となる。なお、本明細書に記載された効果はあくまで例示であって限定されるものではなく、また付加的な効果があってもよい。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899285000001
    Figure 0007899285000001
  • Figure 0007899285000002
    Figure 0007899285000002
  • Figure 0007899285000003
    Figure 0007899285000003
Patent Text Reader

Abstract

To reduce a processing load of a reception side when transmitting a plurality of types of audio data.SOLUTION: A container of a prescribed format having a prescribed number of audio streams including encoded data of a plurality of groups is transmitted. For example, either or both of channel encoded data and object encoded data are included in encoded data of the plurality of groups. Attribute information indicating respective attributes of encoded data of the plurality of groups is inserted into a layer of the container and / or a layer of the audio stream. For example, stream correspondence information indicating which audio stream contains which encoded data of the plurality of groups is further inserted.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, as a three-dimensional (3D) audio technology, a technology has been proposed in which encoded sample data is mapped to speakers existing at arbitrary positions based on metadata and rendered (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is conceivable to transmit object-encoded data composed of encoded sample data and metadata together with channel-encoded data such as 5.1-channel and 7.1-channel data, and enable acoustic playback with enhanced presence on the receiving side.

[0005] An object of this technology is to reduce the processing load on the receiving side when transmitting multiple types of audio data.

Means for Solving the Problems

[0006] The concept of this technology is a transmission unit that transmits a container in a predetermined format having a predetermined number of audio streams including a plurality of groups of encoded data, and an information insertion unit that inserts attribute information indicating the respective attributes of the plurality of groups of encoded data into the layer of the container and / or the layer of the audio stream. It is in a transmission device.

[0007] In this technology, the transmitting unit transmits a container in a predetermined format having a predetermined number of audio streams, each containing encoded data from multiple groups. For example, the encoded data from multiple groups may include either channel-encoded data or object-encoded data, or both.

[0008] The information insertion unit inserts attribute information indicating the attributes of each of the multiple groups of encoded data into the container layer and / or the audio stream layer. For example, the container may be a transport stream (MPEG-2 TS) adopted in digital broadcasting standards. Alternatively, the container may be MP4, used for internet distribution, or a container in another format.

[0009] In this technology, attribute information indicating the attributes of each of the multiple groups of encoded data contained in a predetermined number of audio streams is inserted into the container layer and / or the audio stream layer. Therefore, the receiving side can easily recognize the attributes of each of the multiple groups of encoded data before decoding the encoded data, and selectively decode and use only the encoded data of the necessary groups, thereby reducing the processing load.

[0010] Furthermore, in this technology, for example, the information insertion unit may further insert stream correspondence information into the container layer and / or the audio stream layer, indicating which audio stream each group of encoded data is contained in. This allows the receiving side to easily recognize the audio stream containing the required group of encoded data, thereby reducing the processing load.

[0011] In this case, for example, if the container is MPEG2-TS, and the information insertion unit inserts attribute information and stream correspondence information into the container, it may insert this attribute information and stream correspondence information into an audio elementary stream loop corresponding to at least one audio stream from a predetermined number of audio streams under the program map table.

[0012] In this case, for example, when the information insertion unit inserts attribute information and stream correspondence information into the audio stream layer, it may insert this attribute information and stream correspondence information into the PES payload of the PES packet of at least one audio stream out of a predetermined number of audio streams.

[0013] For example, the stream correspondence information may be information that shows the correspondence between a group identifier that identifies each of the encoded data of multiple groups and a stream identifier that identifies each of a predetermined number of audio streams. In this case, for example, the information insertion unit may further insert stream identifier information that shows each of the stream identifiers of a predetermined number of audio streams into the container layer and / or the audio stream layer.

[0014] For example, if the container is MPEG2-TS and the information insertion unit inserts stream identifier information into the container, this stream identifier information may be inserted into the audio elementary stream loop corresponding to each of the predetermined number of audio streams under the program map table. Alternatively, for example, if the information insertion unit inserts stream identifier information into an audio stream, this stream identifier information may be inserted into the PES payload of each PES packet of the predetermined number of audio streams.

[0015] Furthermore, for example, the stream correspondence information may be information that shows the correspondence between a group identifier that identifies each of the encoded data of multiple groups and a packet identifier that is assigned when each of a predetermined number of audio streams is packetized. Alternatively, for example, the stream correspondence information may be information that shows the correspondence between a group identifier that identifies each of the encoded data of multiple groups and type information that indicates the stream type of each of a predetermined number of audio streams.

[0016] Furthermore, other concepts of this technology include: The receiving unit includes a container in a predetermined format having a predetermined number of audio streams containing encoded data from multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. The processing unit further comprises processing the predetermined number of audio streams contained in the received container based on the attribute information. It is located in the receiving device.

[0017] In this technology, the receiving unit receives a container in a predetermined format having a predetermined number of audio streams containing encoded data from multiple groups. For example, the encoded data from multiple groups may include either channel-encoded data or object-encoded data, or both. Attribute information indicating the respective attributes of the encoded data from multiple groups is inserted into the container layer and / or the audio stream layer. The processing unit processes the predetermined number of audio streams contained in the received container based on their attribute information.

[0018] In this technology, the processing of a predetermined number of audio streams in a received container is performed based on attribute information indicating the attributes of each of the multiple groups of encoded data inserted into the container layer and / or the audio stream layer. Therefore, only the encoded data of the necessary groups can be selectively decoded and used, thereby reducing the processing load.

[0019] In this technology, for example, the container layer and / or the audio stream layer may also have stream correspondence relationship information inserted that indicates which audio stream each of several groups of encoded data is contained in. The processing unit may process a predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information. In this case, the audio stream containing the encoded data of the required group can be easily identified, and the processing load can be reduced.

[0020] Furthermore, in this technology, for example, the processing unit may selectively decode audio streams containing encoded data of groups having attributes that match speaker configuration and user selection information, based on attribute information and stream correspondence relationship information.

[0021] Furthermore, other concepts of this technology include: The receiving unit includes a container in a predetermined format having a predetermined number of audio streams containing encoded data from multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. A processing unit that selectively obtains encoded data of a predetermined group from the predetermined number of audio streams held by the received container based on the attribute information, and reconstructs an audio stream containing the encoded data of the predetermined group, It further includes a stream transmission unit that transmits the audio stream reconfigured by the above processing unit to an external device. It is in the receiving device.

[0022] In this technology, a container in a predetermined format having a predetermined number of audio streams including encoded data of a plurality of groups is received by the receiving unit. Attribute information indicating each attribute of the encoded data of the plurality of groups is inserted into the layer of the container and / or the layer of the audio stream. By the processing unit, encoded data of a predetermined group is selectively acquired from the predetermined number of audio streams based on the attribute information, and an audio stream including the encoded data of the predetermined group is reconfigured. Then, the reconfigured audio stream is transmitted to an external device by the stream transmission unit.

[0023] Thus, in this technology, based on the attribute information indicating each attribute of the encoded data of the plurality of groups inserted into the layer of the container and / or the layer of the audio stream, encoded data of a predetermined group is selectively acquired from the predetermined number of audio streams, and the audio stream to be transmitted to the external device is reconfigured. It is possible to easily acquire the encoded data of the necessary group, and it is possible to reduce the processing load.

[0024] In this technology, for example, stream correspondence relationship information indicating which audio stream each of the encoded data of the plurality of groups is included in is further inserted into the layer of the container and / or the layer of the audio stream, and the processing unit may be configured to selectively acquire the encoded data of a predetermined group from the predetermined number of audio streams based on the stream correspondence relationship information in addition to the attribute information. In this case, it is possible to easily recognize the audio stream including the encoded data of the predetermined group, and it is possible to reduce the processing load.

Effects of the Invention

[0025] This technology makes it possible to reduce the processing load on the receiving end when transmitting multiple types of audio data. The effects described herein are merely illustrative and not limited to those described herein, and additional effects may also exist. [Brief explanation of the drawing]

[0026] [Figure 1] This block diagram shows an example configuration of a transmission and reception system as an embodiment. [Figure 2] This diagram shows the structure of an audio frame (1024 samples) in 3D audio transmission data. [Figure 3] This figure shows an example of the data structure for 3D audio transmission. [Figure 4] This diagram schematically shows examples of audio frame configurations when transmitting 3D audio data in a single stream and when transmitting it in multiple streams. [Figure 5] This figure shows an example of group division when transmitting 3D audio data using two streams. [Figure 6] This diagram shows the correspondence between groups and streams in an example of group division (2 divisions). [Figure 7] This figure shows an example of group division when transmitting 3D audio data using two streams. [Figure 8] This diagram shows the correspondence between groups and streams in an example of group division (2 divisions). [Figure 9] This block diagram shows an example configuration of the stream generation unit included in the service transmitter. [Figure 10] This diagram shows an example of the structure of a 3D audio stream config descriptor. [Figure 11] This shows the key information content in an example of the structure of a 3D audio stream config descriptor. [Figure 12] This diagram shows the types of content defined in "contentKind". [Figure 13] This diagram shows an example of a 3D audio stream ID descriptor structure and the content of the main information within that structure. [Figure 14] This figure shows an example of a transport stream configuration. [Figure 15] This block shows an example configuration of a service receiver. [Figure 16] This figure shows an example of a received audio stream. [Figure 17] This diagram schematically illustrates the decoding process when descriptor information is not present in the audio stream. [Figure 18] This figure shows an example of the configuration of an audio access unit (audio frame) in an audio stream when descriptor information is not present in the audio stream. [Figure 19] This diagram schematically illustrates the decoding process when descriptor information is present in the audio stream. [Figure 20] This figure shows an example of the configuration of an audio access unit (audio frame) in an audio stream when descriptor information is present in the audio stream. [Figure 21] This figure shows another example of the configuration of an audio access unit (audio frame) in an audio stream when descriptor information is present in the audio stream. [Figure 22] This flowchart (1 / 2) shows an example of the CPU's audio decoding control process in a service receiver. [Figure 23] This flowchart (2 / 2) shows an example of the CPU's audio decoding control process in a service receiver. [Figure 24] This block shows another example configuration for the service receiver. [Modes for carrying out the invention]

[0027] The following describes embodiments for carrying out the invention. The description will be given in the following order. 1. Embodiment 2. Variations

[0028] <1. Embodiment> [Example of a transmission / reception system configuration] Figure 1 shows an example configuration of a transmission / reception system 10 as an embodiment. This transmission / reception system 10 consists of a service transmitter 100 and a service receiver 200. The service transmitter 100 transmits a transport stream TS on a broadcast wave or network packet. This transport stream TS has a video stream and a predetermined number of encoded data in multiple groups, i.e., one or more audio streams.

[0029] Figure 2 shows the structure of an audio frame (1024 samples) in the 3D audio transmission data handled in this embodiment. This audio frame consists of multiple MPEG audio stream packets. Each MPEG audio stream packet consists of a header and a payload.

[0030] The header contains information such as the packet type, packet label, and packet length. The payload contains the information defined by the packet type in the header. This payload information includes "SYNC," which corresponds to the synchronization start code, "Frame," which is the actual data of the 3D audio transmission data, and "Config," which indicates the structure of this "Frame."

[0031] A “Frame” contains channel encoding data and object encoding data that make up the transmission data for 3D audio. Here, the channel encoding data consists of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), and LFE (Low Frequency Element). The object encoding data consists of encoded sample data for SCE (Single Channel Element) and metadata for mapping it to speakers located at arbitrary positions and rendering it. This metadata is included as an extension element (Ext_element).

[0032] Figure 3 shows an example of the structure of 3D audio transmission data. In this example, it consists of one channel-coded data and two object-coded data. The one channel-coded data is 5.1 channel-coded data (CD) and consists of coded sample data for SCE1, CPE1.1, CPE1.2, and LFE1.

[0033] The two object encoding data sets are the encoding data for the Immersive Audio Object (IAO) and the Speech Dialog Object (SDO). The Immersive Audio Object encoding data is object encoding data for immersive sound and consists of encoded sample data SCE2 and metadata EXE_El (Object metadata)2 for mapping it to a speaker located at an arbitrary position and rendering it.

[0034] Speech dialog object encoding data is object encoding data for a speech language. In this example, there is speech dialog object encoding data corresponding to the first and second languages. The speech dialog object encoding data corresponding to the first language consists of encoded sample data SCE3 and metadata EXE_El (Object metadata)3 for mapping it to a speaker located at an arbitrary position and rendering it. The speech dialog object encoding data corresponding to the second language consists of encoded sample data SCE4 and metadata EXE_El (Object metadata)4 for mapping it to a speaker located at an arbitrary position and rendering it.

[0035] Encoded data is distinguished by the concept of "groups" based on its type. In the example shown, 5.1 channel encoded channel data is group 1, immersive audio object encoded data is group 2, speech dialogue object encoded data for the first language is group 3, and speech dialogue object encoded data for the second language is group 4.

[0036] Furthermore, at the receiving end, selectable groups are registered in a switch group (SW Group) and encoded. Groups can also be combined into a preset group, enabling playback according to the use case. In the illustrated example, groups 1, 2, and 3 are combined into preset group 1, and groups 1, 2, and 4 are combined into preset group 2.

[0037] Returning to Figure 1, the service transmitter 100 transmits 3D audio transmission data, including encoded data from multiple groups, in one stream or multiple streams, as described above.

[0038] Figure 4(a) schematically shows an example of the audio frame configuration when transmitting 3D audio transmission data in one stream (main stream) as in the example configuration of 3D audio transmission data in Figure 3. In this case, this one stream includes channel coding data (CD), immersive audio object coding data (IAO), and speech dialogue object coding data (SDO), along with "SYNC" and "Config".

[0039] Figure 4(b) schematically shows an example of the audio frame configuration when transmitting multiple streams, specifically two streams, in the same configuration example as in Figure 3 for 3D audio transmission data. In this case, the main stream contains channel-coded data (CD) and immersive audio object-coded data (IAO), along with "SYNC" and "Config". The substream contains speech dialogue object-coded data (SDO), along with "SYNC" and "Config".

[0040] Figure 5 shows an example of group division when transmitting 3D audio transmission data in two streams, as in the example configuration of Figure 3. In this case, the main stream contains channel-coded data (CD) distinguished as Group 1 and immersive audio object-coded data (IAO) distinguished as Group 2. The substream contains speech dialogue object-coded data (SDO) of the first language distinguished as Group 3 and speech dialogue object-coded data (SDO) of the second language distinguished as Group 4.

[0041] Figure 6 shows the correspondence between groups and streams in the group division example (2 divisions) of Figure 5. Here, the group ID is an identifier for identifying a group. The attribute indicates the attribute of the encoded data for each group. The switch group ID is an identifier for identifying a switching group. The preset group ID is an identifier for identifying a preset group. The substream ID is an identifier for identifying a substream. Kind indicates the type of content for each group.

[0042] The diagram shows that the encoded data belonging to group 1 is channel encoded data, does not constitute a switch group, and is included in stream 1. The diagram also shows that the encoded data belonging to group 2 is object encoded data for immersive sound (immersive audio object encoded data), does not constitute a switch group, and is included in stream 1.

[0043] Furthermore, the diagram shows that the encoded data belonging to group 3 is object encoded data for the speech language of the first language (speech dialogue object encoded data), constitutes switch group 1, and is included in stream 2. Furthermore, the diagram shows that the encoded data belonging to group 4 is object encoded data for the speech language of the second language (speech dialogue object encoded data), constitutes switch group 1, and is included in stream 2.

[0044] Furthermore, the diagram shows that preset group 1 includes group 1, group 2, and group 3. Additionally, the diagram shows that preset group 2 includes group 1, group 2, and group 4.

[0045] Figure 7 shows an example of group division when transmitting 3D audio data using two streams. In this case, the main stream contains channel-coded data (CD) distinguished as Group 1 and immersive audio object-coded data (IAO) distinguished as Group 2.

[0046] Furthermore, this mainstream includes SAOC (Spatial Audio Object Coding) object coding data, which is distinguished as Group 5, and HOA (Higher Order Ambisonics) object coding data, which is distinguished as Group 6. SAOC object coding data is data that utilizes the characteristics of object data to achieve higher compression of object coding. HOA object coding data is a technology that treats 3D sound as a whole sound field, and is data that aims to reproduce the direction of sound at the auditory position from the direction of sound arrival at the microphone.

[0047] The substream includes speech dialogue object encoding data (SDO) for the first language, distinguished as Group 3, and speech dialogue object encoding data (SDO) for the second language, distinguished as Group 4. This substream also includes audio description encoding data for the first language, distinguished as Group 7, and audio description encoding data for the second language, distinguished as Group 8. Audio description encoding data is data used to provide audio commentary on content (primarily video), and is primarily intended for transmission separately from regular audio, especially for visually impaired users.

[0048] Figure 8 shows the correspondence between groups and substreams in the group division example (2 divisions) of Figure 7. The illustrated correspondence indicates that the encoded data belonging to Group 1 is channel encoded data, does not constitute a switch group, and is included in Stream 1. The illustrated correspondence also indicates that the encoded data belonging to Group 2 is object encoded data for immersive sound (immersive audio object encoded data), does not constitute a switch group, and is included in Stream 1.

[0049] Furthermore, the diagram shows that the encoded data belonging to group 3 is object encoded data for the speech language of the first language (speech dialogue object encoded data), constitutes switch group 1, and is included in stream 2. Furthermore, the diagram shows that the encoded data belonging to group 4 is object encoded data for the speech language of the second language (speech dialogue object encoded data), constitutes switch group 1, and is included in stream 2.

[0050] Furthermore, the diagram shows that the encoded data belonging to group 5 is SAOC object encoded data, constitutes switch group 2, and is included in stream 1. Also, the diagram shows that the encoded data belonging to group 6 is HOA object encoded data, constitutes switch group 2, and is included in stream 1.

[0051] Furthermore, the diagram shows that the encoded data belonging to group 7 is the object encoded data of the first audio description, constitutes switch group 3, and is included in stream 2. Also, the diagram shows that the encoded data belonging to group 8 is the object encoded data of the second audio description, constitutes switch group 3, and is included in stream 2.

[0052] Furthermore, the diagram shows that preset group 1 includes group 1, group 2, group 3, and group 7. Additionally, the diagram shows that preset group 2 includes group 1, group 2, group 4, and group 8.

[0053] Returning to Figure 1, the service transmitter 100 inserts attribute information into the container layer that indicates the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data. The service transmitter 100 also inserts stream correspondence information into the container layer that indicates which audio stream each of these multiple groups of encoded data is included in. In this embodiment, this stream correspondence information is, for example, information that indicates the correspondence between a group ID and a stream identifier.

[0054] The service transmitter 100 inserts this attribute information and stream correspondence information as descriptors into an audio elementary stream loop corresponding to one or more audio streams from a predetermined number of audio streams under the Program Map Table (PMT), for example.

[0055] Furthermore, the service transmitter 100 inserts stream identifier information into the container layer, indicating the stream identifier of each of a predetermined number of audio streams. The service transmitter 100 then inserts this stream identifier information as a descriptor into the audio elementary stream loop corresponding to each of the predetermined number of audio streams that exist under the Program Map Table (PMT), for example.

[0056] Furthermore, the service transmitter 100 inserts attribute information into the audio stream layer that indicates the attributes of each of the multiple groups of encoded data included in the 3D audio transmission data. The service transmitter 100 also inserts stream correspondence information into the audio stream layer that indicates which audio stream each of these multiple groups of encoded data is included in. The service transmitter 100 inserts this attribute information and stream correspondence information into the PES payload of a PES packet of one or more audio streams from a predetermined number of audio streams, for example.

[0057] Furthermore, the service transmitter 100 inserts stream identifier information, which indicates the stream identifier of each of a predetermined number of audio streams, into the audio stream layer. The service transmitter 100 then inserts this stream identifier information, for example, into the PES payload of each PES packet of the predetermined number of audio streams.

[0058] The service transmitter 100 inserts information into these audio stream layers by inserting "Desc," or descriptor information, between "SYNC" and "Config," as shown in Figures 4(a) and 4(b).

[0059] In this embodiment, we have shown an example in which the information (attribute information, stream correspondence information, and stream identifier information) is inserted into both the container layer and the audio stream layer as described above. However, it is also possible to insert the information into only the container layer or only the audio stream layer.

[0060] The service receiver 200 receives the transport stream TS transmitted from the service transmitter 100 on a broadcast wave or network packet. As described above, this transport stream TS has a predetermined number of audio streams, including a video stream and encoded data for multiple groups that constitute the transmission data for 3D audio.

[0061] Furthermore, attribute information indicating the attributes of each of the multiple groups of encoded data contained in the 3D audio transmission data is inserted into the container layer and / or the audio stream layer, along with stream correspondence information indicating which audio stream each of these multiple groups of encoded data is contained in.

[0062] The service receiver 200 selectively decodes audio streams containing encoded data of groups having attributes that match the speaker configuration and user selection information, based on attribute information and stream correspondence information, to obtain 3D audio output.

[0063] [Stream generation unit of the service transmitter] Figure 9 shows an example configuration of the stream generation unit 110 included in the service transmitter 100. This stream generation unit 110 includes a video encoder 112, an audio encoder 113, and a multiplexer 114. Here, we assume that the audio transmission data consists of one encoded channel data and two object encoded data, as shown in Figure 3.

[0064] The video encoder 112 receives video data SV as input, encodes this video data SV, and generates a video stream (video elementary stream). The audio encoder 113 receives immersive audio and speech dialogue object data along with channel data as audio data SA.

[0065] The audio encoder 113 encodes the audio data SA to obtain 3D audio transmission data. As shown in Figure 3, this 3D audio transmission data includes channel encoded data (CD), immersive audio object encoded data (IAO), and speech dialogue object encoded data (SDO).

[0066] The audio encoder 113 generates one or more audio streams (audio elementary streams) containing encoded data from multiple groups, in this case four groups (see Figures 4(a) and 4(b)). At this time, the audio encoder 113 inserts descriptor information ("Desc") containing attribute information, stream correspondence information, and stream identifier information between "SYNC" and "Config," as described above.

[0067] The multiplexer 114 takes the video stream output from the video encoder 112 and a predetermined number of audio streams output from the audio encoder 113, converts them into PES packets, further converts them into transport packets, and multiplexes them to obtain a transport stream TS as a multiplexed stream.

[0068] Furthermore, the multiplexer 114 inserts attribute information indicating the attributes of each of the multiple groups of encoded data, and stream correspondence information indicating which audio stream each of the multiple groups of encoded data is contained in, under the Program Map Table (PMT). The multiplexer 114 inserts this information into the audio elementary stream loop corresponding to at least one of the predetermined number of audio streams, using a 3D audio stream config descriptor (3Daudio_stream_config_descriptor). Details of this descriptor will be described later.

[0069] Furthermore, the multiplexer 114 inserts stream identifier information, indicating the stream identifier of each of a predetermined number of audio streams, under the Program Map Table (PMT). The multiplexer 114 then inserts this information into the corresponding audio elementary stream loop for each of the predetermined number of audio streams using a 3D audio stream ID descriptor (3Daudio_substreamID_descriptor). Details of this descriptor will be described later.

[0070] The operation of the stream generation unit 110 shown in Figure 9 will be briefly explained. Video data is supplied to the video encoder 112. The video encoder 112 encodes the video data SV and generates a video stream containing the encoded video data. This video stream is supplied to the multiplexer 114.

[0071] Audio data SA is supplied to the audio encoder 113. This audio data SA includes channel data and object data for immersive audio and speech dialogue. The audio encoder 113 encodes the audio data SA to obtain 3D audio transmission data.

[0072] The transmission data for this 3D audio includes channel coding data (CD), as well as immersive audio object coding data (IAO) and speech dialogue object coding data (SDO) (see Figure 3). The audio encoder 113 then generates one or more audio streams containing the coded data of four groups (see Figures 4(a) and 4(b)).

[0073] At this time, the audio encoder 113 inserts descriptor information ("Desc"), which includes the attribute information, stream correspondence information, and stream identifier information mentioned above, between "SYNC" and "Config".

[0074] The video stream generated by the video encoder 112 is supplied to the multiplexer 114. Similarly, the audio stream generated by the audio encoder 113 is supplied to the multiplexer 114. In the multiplexer 114, the streams supplied from each encoder are packetized into PES packets, then further packetized into transport packets and multiplexed to obtain a transport stream TS as a multiplexed stream.

[0075] Furthermore, in the multiplexer 114, a 3D audio stream configuration descriptor is inserted into an audio elementary stream loop that corresponds to at least one audio stream from a given set of audio streams. This descriptor contains attribute information that indicates the attributes of each of the multiple groups of encoded data, and stream correspondence information that indicates which audio stream each of the multiple groups of encoded data is included in.

[0076] Furthermore, in descriptor 114, a 3D audio stream ID descriptor is inserted into the audio elemental stream loop corresponding to each of the predetermined number of audio streams. This descriptor contains stream identifier information indicating the stream identifier of each of the predetermined number of audio streams.

[0077] [Details of the 3D Audio Stream Config Descriptor] Figure 10 shows an example of the structure (syntax) of a 3D audio stream config descriptor (3Daudio_stream_config_descriptor). Figure 11 shows the content (semantics) of the main information in that example structure.

[0078] The 8-bit field "descriptor_tag" indicates the descriptor type. In this case, it indicates a 3D audio stream configuration descriptor. The 8-bit field "descriptor_length" indicates the descriptor length (size), showing the number of bytes remaining as the descriptor length.

[0079] The 8-bit field "NumOfGroups, N" indicates the number of groups. The 8-bit field "NumOfPresetGroups, P" indicates the number of preset groups. The 8-bit fields "groupID", "attribute_of_groupID", "SwitchGroupID", and "audio_streamID" are repeated for each group.

[0080] The "groupID" field indicates the group identifier. The "attribute_of_groupID" field indicates the attribute of the encoded data for the group. The "SwitchGroupID" field is an identifier indicating which switch group the group belongs to. "0" indicates that the group does not belong to any switch group. Any value other than "0" indicates the switch group to which the group belongs. The 8-bit "contentKind" field indicates the type of content for the group. "audio_streamID" is an identifier indicating the audio stream to which the group is contained. Figure 12 shows the types of content defined in "contentKind".

[0081] Furthermore, the 8-bit fields of "presetGroupID" and "NumOfGroups_in_preset, R" are repeated for each preset group. The "presetGroupID" field is an identifier that indicates the bundle of preset groups. The "NumOfGroups_in_preset, R" field indicates the number of groups belonging to the preset group. Then, for each preset group, the 8-bit field of "groupID" is repeated for each group belonging to it, indicating the groups that belong to the preset group. This descriptor may be placed under an extended descriptor.

[0082] [Details of 3D Audio Stream ID Descriptor] Figure 13(a) shows an example of the structure (syntax) of a 3D audio stream ID descriptor (3Daudio_substreamID_descriptor). Figure 13(b) shows the content (semantics) of the main information in that example structure.

[0083] The 8-bit field "descriptor_tag" indicates the descriptor type. Here, it indicates a 3D audio stream ID descriptor. The 8-bit field "descriptor_length" indicates the length (size) of the descriptor, showing the number of subsequent bytes as the descriptor length. The 8-bit field "audio_streamID" indicates the identifier of the audio stream. This descriptor may be placed under an extended descriptor.

[0084] [Transport Stream TS Configuration] Figure 14 shows an example of a transport stream (TS) configuration. This configuration is suitable for transmitting 3D audio data in two streams (see Figure 5). In this configuration, there is a PES packet "video PES" for the video stream identified by PID1. In this configuration, there are also two PES packets "audio PES" for the audio streams identified by PID2 and PID3, respectively. A PES packet consists of a PES header (PES_header) and a PES payload (PES_payload). The PES header contains timestamps for the DTS and PTS. By accurately matching the timestamps of PID2 and PID3 during multiplexing, it is possible to ensure synchronization between the two streams throughout the entire system.

[0085] Here, the PES packet "audio PES" for the audio stream identified by PID2 contains channel-coded data (CD) distinguished as group 1 and immersive audio object-coded data (IAO) distinguished as group 2. Similarly, the PES packet "audio PES" for the audio stream identified by PID3 contains speech dialogue object-coded data (SDO) for the first language distinguished as group 3 and speech dialogue object-coded data (SDO) for the second language distinguished as group 4.

[0086] Furthermore, the transport stream TS contains a Program Map Table (PMT) as Program Specific Information (PSI). PSI is information that indicates which program each elementary stream included in the transport stream belongs to. The PMT contains a program loop that describes information related to the entire program.

[0087] Furthermore, the PMT has elementary stream loops that hold information related to each elementary stream. In this example configuration, there is a video elementary stream loop (video ES loop) corresponding to the video stream, as well as audio elementary stream loops (audio ES loops) corresponding to the two audio streams.

[0088] The video elementary stream loop (video ES loop) contains information such as the stream type and PID (packet identifier) ​​corresponding to the video stream, as well as a descriptor that describes information related to that video stream. The value of "Stream_type" for this video stream is set to "0x24", and the PID information is said to indicate PID1, which is assigned to the PES packet "video PES" of the video stream, as described above. One of the descriptors is the HEVC descriptor.

[0089] Each audio elementary stream loop (audio ES loop) contains information such as the stream type and PID (packet identifier) ​​corresponding to the audio stream, as well as a descriptor that describes information related to that audio stream. PID2 is the main audio stream, and the value of "Stream_type" is set to "0x2C". The PID information indicates the PID assigned to the audio stream's PES packet "audio PES" as described above. PID3 is a sub-audio stream, and the value of "Stream_type" is set to "0x2D". The PID information indicates the PID assigned to the audio stream's PES packet "audio PES" as described above.

[0090] Furthermore, each audio elemental stream loop (audio ES loop) contains both the 3D audio stream config descriptor and the 3D audio stream ID descriptor mentioned above.

[0091] Furthermore, descriptor information is inserted into the PES payload of each audio elemental stream's PES packet. This descriptor information is "Desc," which is inserted between "SYNC" and "Config" as described above (see Figure 4). If we denote the information contained in the 3D audio stream config descriptor as D1 and the information contained in the 3D audio stream ID descriptor as D2, then this descriptor information contains "D1 + D2" information.

[0092] [Example of service receiver configuration] Figure 15 shows an example configuration of a service receiver 200. This service receiver 200 includes a receiving unit 201, a demultiplexer 202, a video decoder 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. The service receiver 200 also includes multiplexing buffers 211-1 to 211-N, a combiner 212, a 3D audio decoder 213, an audio output processing circuit 214, and a speaker system 215. Furthermore, the service receiver 200 includes a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiver 225, and a remote control transmitter 226.

[0093] The CPU 221 controls the operation of each part of the service receiver 200. The flash ROM 222 stores the control software and data. The DRAM 223 constitutes the work area of ​​the CPU 221. The CPU 221 loads the software and data read from the flash ROM 222 onto the DRAM 223, starts the software, and controls each part of the service receiver 200.

[0094] The remote control receiver 225 receives the remote control signal (remote control code) transmitted from the remote control transmitter 226 and supplies it to the CPU 221. The CPU 221 controls each part of the service receiver 200 based on this remote control code. The CPU 221, flash ROM 222, and DRAM 223 are connected to the internal bus 224.

[0095] The receiving unit 201 receives the transport stream TS sent from the service transmitter 100 on a broadcast wave or network packet. In addition to the video stream, this transport stream TS has a predetermined number of audio streams, including encoded data for multiple groups that constitute the transmission data for 3D audio.

[0096] Figure 16 shows an example of a received audio stream. Figure 16(a) shows an example of one stream (main stream). This stream contains channel-coded data (CD), immersive audio object-coded data (IAO), and speech dialogue object-coded data (SDO), along with "SYNC" and "Config". This stream is identified by PID2.

[0097] Furthermore, descriptor information ("Desc") is included between "SYNC" and "Config". This descriptor information contains attribute information indicating the attributes of each of the multiple groups of encoded data, stream correspondence information indicating which audio stream each of the multiple groups of encoded data is included in, and stream identifier information indicating its own stream identifier.

[0098] Figure 16(b) shows an example of a two-stream configuration. The main stream, identified by PID2, contains channel-coded data (CD) and immersive audio object-coded data (IAO), along with "SYNC" and "Config". The substream, identified by PID3, contains speech dialogue object-coded data (SDO), along with "SYNC" and "Config".

[0099] Furthermore, each stream contains descriptor information ("Desc") between "SYNC" and "Config". This descriptor information includes attribute information indicating the attributes of each of the multiple groups of encoded data, stream correspondence information indicating which audio stream each of the multiple groups of encoded data is contained in, and stream identifier information indicating its own stream identifier.

[0100] The demultiplexer 202 extracts video stream packets from the transport stream TS and sends them to the video decoder 203. The video decoder 203 reconstructs the video stream from the video packets extracted by the demultiplexer 202 and performs decoding to obtain uncompressed video data.

[0101] The video processing circuit 204 performs scaling and image quality adjustment on the video data obtained from the video decoder 203 to obtain video data for display. The panel driving circuit 205 drives the display panel 206 based on the display image data obtained from the video processing circuit 204. The display panel 206 is composed of, for example, an LCD (Liquid Crystal Display) or an organic EL display (organic electroluminescence display).

[0102] Furthermore, the demultiplexer 202 extracts various information, such as descriptor information, from the transport stream TS and sends it to the CPU 221. This information includes the aforementioned 3D audio stream config descriptor (3Daudio_stream_config_descriptor) and 3D audio stream ID descriptor (3Daudio_substreamID_descriptor) information (see Figure 14).

[0103] Based on attribute information indicating the attributes of the encoded data for each group, and stream relationship information indicating which audio stream each group belongs to, the CPU221 recognizes the audio stream containing encoded data for groups with attributes that match the speaker configuration and listener (user) selection information.

[0104] Furthermore, under the control of the CPU 221, the demultiplexer 202 selectively extracts packets of one or more audio streams from a predetermined number of audio streams possessed by the transport stream TS, using a PID filter, which include encoded data from groups having attributes that match the speaker configuration and viewer (user) selection information.

[0105] The multiplexing buffers 211-1 to 211-N each capture the audio stream extracted by the demultiplexer 202. While the number of multiplexing buffers 211-1 to 211-N, N, is considered sufficient, in actual operation, the number of buffers used will be equal to the number of audio streams extracted by the demultiplexer 202.

[0106] The combiner 212 reads each audio stream from the multiplexing buffers 211-1 to 211-N, each of which contains the audio stream extracted by the demultiplexer 202, and sends it to the 3D audio decoder 213.

[0107] If the audio stream supplied by the combiner 212 contains descriptor information ("Desc"), the 3D audio decoder 213 sends that descriptor information to the CPU 221. Under the control of the CPU 221, the 3D audio decoder 213 selectively extracts encoded data from groups with attributes that match the speaker configuration and viewer (user) selection information, performs decoding, and obtains audio data to drive each speaker of the speaker system 215.

[0108] Here, the encoded data to be decoded can be in three cases: it contains only channel-encoded data, it contains only object-encoded data, or it contains both channel-encoded and object-encoded data.

[0109] When decoding channel-encoded data, the 3D audio decoder 213 performs downmixing and upmixing to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When decoding object-encoded data, the 3D audio decoder 213 calculates speaker rendering (mixing ratio for each speaker) based on object information (metadata), and mixes the object's audio data into audio data for driving each speaker according to the calculation result.

[0110] The audio output processing circuit 214 performs necessary processing, such as D / A conversion and amplification, on the audio data obtained by the 3D audio decoder 213 to drive each speaker, and supplies it to the speaker system 215. The speaker system 215 has multiple channels, such as 2-channel, 5.1-channel, 7.1-channel, or 22.2-channel speakers.

[0111] The operation of the service receiver 200 shown in Figure 15 will be briefly explained. The receiver 201 receives the transport stream TS that is sent from the service transmitter 100 on a broadcast wave or network packet. In addition to the video stream, this transport stream TS has a predetermined number of audio streams, which include encoded data for multiple groups that constitute the transmission data for 3D audio. This transport stream TS is supplied to the demultiplexer 202.

[0112] In the demultiplexer 202, video stream packets are extracted from the transport stream TS and supplied to the video decoder 203. In the video decoder 203, the video stream is reconstructed from the video packets extracted by the demultiplexer 202, and decoding is performed to obtain uncompressed video data. This video data is supplied to the video processing circuit 204.

[0113] In the video processing circuit 204, scaling and image quality adjustment processing are performed on the video data obtained from the video decoder 203 to obtain video data for display. This video data for display is supplied to the panel driving circuit 205. The panel driving circuit 205 drives the display panel 206 based on the video data for display. As a result, the display panel 206 displays an image corresponding to the video data for display.

[0114] Furthermore, the demultiplexer 202 extracts various information, such as descriptor information, from the transport stream TS and sends it to the CPU 221. This information includes 3D audio stream configuration descriptor and 3D audio stream ID descriptor information. Based on the attribute information and stream relationship information contained in this descriptor information, the CPU 221 recognizes audio streams that contain encoded data for groups with attributes that match the speaker configuration and viewer (user) selection information.

[0115] Furthermore, under the control of the CPU 221, the demultiplexer 202 selectively extracts packets of one or more audio streams from a predetermined number of audio streams possessed by the transport stream TS, which include encoded data for groups with attributes that match the speaker configuration and listener selection information, using a PID filter.

[0116] The audio stream extracted by the demultiplexer 202 is loaded into the corresponding multiplexing buffer from the multiplexing buffers 211-1 to 211-N. The combiner 212 reads the audio stream frame by frame from each multiplexing buffer into which the audio stream was loaded and supplies it to the 3D audio decoder 213.

[0117] In the 3D audio decoder 213, if the audio stream supplied from the combiner 212 contains descriptor information ("Desc"), that descriptor information is extracted and sent to the CPU 221. Under the control of the CPU 221, the 3D audio decoder 213 selectively extracts encoded data from groups of data that have attributes matching the speaker configuration and viewer (user) selection information, decodes them, and obtains audio data to drive each speaker of the speaker system 215.

[0118] When channel-encoded data is decoded, downmixing or upmixing is performed to match the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. When object-encoded data is decoded, speaker rendering (mixing ratio for each speaker) is calculated based on object information (metadata), and the object's audio data is mixed with the audio data for driving each speaker according to the calculation result.

[0119] The audio data obtained by the 3D audio decoder 213 to drive each speaker is supplied to the audio output processing circuit 214. In this audio output processing circuit 214, necessary processing such as D / A conversion and amplification is performed on the audio data to drive each speaker. The processed audio data is then supplied to the speaker system 215. As a result, the speaker system 215 provides an audio output corresponding to the display image on the display panel 206.

[0120] Figure 17 schematically illustrates the decoding process when descriptor information is not present in the audio stream. The transport stream TS, which is a multiplexed stream, is input to the demultiplexer 202. The demultiplexer 202 performs system layer analysis and supplies descriptor information 1 (information on the 3D audio stream configuration descriptor and the 3D audio stream ID descriptor) to the CPU 221.

[0121] Based on this descriptor information 1, CPU 221 recognizes audio streams containing encoded data for groups with attributes that match the speaker configuration and listener (user) selection information. Under the control of CPU 221, demultiplexer 202 performs selection between streams.

[0122] In other words, the demultiplexer 202 selectively extracts packets of one or more audio streams from a predetermined number of audio streams in the transport stream TS, which contain encoded data for groups with attributes that match the speaker configuration and listener selection information, using a PID filter. The extracted audio streams are then placed into the multiplexing buffer 211 (211-1 to 211-N).

[0123] The 3D audio decoder 213 analyzes the packet type of each audio stream captured in the multiplexing buffer 211. Then, the demultiplexer 202 performs selection within the stream under the control of the CPU 221 based on the descriptor information 1 described above.

[0124] In other words, encoded data of groups with attributes that match the speaker configuration and listener (user) selection information are selectively extracted from each audio stream as the target for decoding, subjected to decoding, mixing and rendering, etc., to obtain audio data (uncompressed audio) for driving each speaker.

[0125] Figure 18 shows an example of the configuration of audio access units (audio frames) in an audio stream when descriptor information is not present in the audio stream. This example shows two streams.

[0126] For audio streams identified by PID2, the information "FrWork #ch=2, #obj=1" in "Config" indicates the existence of a "Frame" containing two channel-encoded data and one object-encoded data. Furthermore, the information "GroupID[0]=1, GroupID[1]=2" registered sequentially within "AudioSceneInfo()" in "Config" indicates that the "Frame" with the encoded data for group 1 and the "Frame" with the encoded data for group 2 are arranged in that order. Note that the packet label (PL) value is the same for "Config" and each corresponding "Frame".

[0127] Here, the "Frame" containing encoded data for Group 1 includes encoded sample data for CPE (Channel Pair Element). The "Frame" containing encoded data for Group 2 consists of a "Frame" containing metadata as an extension element (Ext_element) and a "Frame" containing encoded sample data for SCE (Single Channel Element).

[0128] For audio streams identified by PID3, the information "FrWork #ch=0, #obj=2" in "Config" indicates the existence of a "Frame" with two object-encoded data. Furthermore, the information "GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=1" registered sequentially within "AudioSceneInfo()" in "Config" indicates that a "Frame" with encoding data for group 3 and a "Frame" with encoding data for group 4 are arranged in that order, and these groups constitute switch group 1. Note that the packet label (PL) value is the same in "Config" and for each corresponding "Frame".

[0129] Here, a "Frame" containing encoded data for Group 3 consists of a "Frame" containing metadata as an extension element (Ext_element) and a "Frame" containing encoded sample data for SCE (Single Channel Element). Similarly, a "Frame" containing encoded data for Group 4 consists of a "Frame" containing metadata as an extension element (Ext_element) and a "Frame" containing encoded sample data for SCE (Single Channel Element).

[0130] Figure 19 schematically illustrates the decoding process when descriptor information is present in the audio stream. The transport stream TS, which is a multiplexed stream, is input to the demultiplexer 202. The demultiplexer 202 performs system layer analysis and sends descriptor information 1 (information on the 3D audio stream configuration descriptor and the 3D audio stream ID descriptor) to the CPU 221.

[0131] Based on this descriptor information 1, CPU 221 recognizes audio streams containing encoded data for groups with attributes that match the speaker configuration and listener (user) selection information. Under the control of CPU 221, demultiplexer 202 performs selection between streams.

[0132] In other words, the demultiplexer 202 selectively extracts packets of one or more audio streams from a predetermined number of audio streams in the transport stream TS, which contain encoded data for groups with attributes that match the speaker configuration and listener selection information, using a PID filter. The extracted audio streams are then placed into the multiplexing buffer 211 (211-1 to 211-N).

[0133] In the 3D audio decoder 213, packet type analysis is performed on each audio stream taken into the multiplexing buffer 211, and descriptor information 2 present in the audio stream is sent to the CPU 221. Based on this descriptor information 2, the existence of encoded data for groups with attributes that match the speaker configuration and viewer (user) selection information is recognized. Then, in the demultiplexer 202, under the control of the CPU 221 based on this descriptor information 2, selection is made within the stream.

[0134] In other words, encoded data of groups with attributes that match the speaker configuration and listener (user) selection information are selectively extracted from each audio stream as the target for decoding, subjected to decoding, mixing and rendering, etc., to obtain audio data (uncompressed audio) for driving each speaker.

[0135] Figure 20 shows an example of the configuration of an audio access unit (audio frame) in an audio stream when descriptor information is present in the audio stream. This example shows two streams. Figure 20 is similar to Figure 18, except that "Desc," i.e., descriptor information, is inserted between "SYNC" and "Config."

[0136] For audio streams identified by PID2, the "GroupID[0]=1, channeldata" information in "Desc" indicates that the encoded data for group 1 is channel encoded data. Furthermore, the "GroupID[1]=2, object sound" information in "Desc" indicates that the encoded data for group 2 is object encoded data for immersive sound. Additionally, the "Stream_ID" information indicates the stream identifier for the audio stream.

[0137] For audio streams identified by PID3, the information "GroupID[2]=3, object lang1" in "Desc" indicates that the encoded data for group 3 is object encoded data for the speech language of the first language. Similarly, the information "GroupID[3]=4, object lang2" in "Desc" indicates that the encoded data for group 4 is object encoded data for the speech language of the second language. Furthermore, the information "SW_GRPID[0]=1" in "Desc" indicates that groups 3 and 4 constitute switch group 1. Finally, the "Stream_ID" information indicates the stream identifier of the audio stream.

[0138] Figure 21 shows an example of the configuration of an audio access unit (audio frame) in an audio stream when descriptor information is present in the audio stream. This example shows a single stream.

[0139] The information "FrWork #ch=2, #obj=3" in "Config" indicates the existence of a "Frame" containing two channel-encoded data and three object-encoded data. Furthermore, the information "GroupID[0]=1, GroupID[1]=2, GroupID[2]=3, GroupID[3]=4, SW_GRPID[0]=1" registered sequentially within "AudioSceneInfo()" in "Config" indicates that the "Frame" with the encoded data of group 1, the "Frame" with the encoded data of group 2, the "Frame" with the encoded data of group 3, and the "Frame" with the encoded data of group 4 are arranged in this order, and that groups 3 and 4 constitute switch group 1. Note that the packet label (PL) value is the same in "Config" and each corresponding "Frame".

[0140] Here, the "Frame" containing encoded data for Group 1 includes encoded sample data for CPE (Channel Pair Element). Furthermore, the "Frames" containing encoded data for Groups 2-4 each consist of a "Frame" containing metadata as an extension element (Ext_element) and a "Frame" containing encoded sample data for SCE (Single Channel Element).

[0141] The information "GroupID[0]=1, channeldata" in "Desc" indicates that the encoded data for group 1 is channel-encoded data. Additionally, the information "GroupID[1]=2, object sound" in "Desc" indicates that the encoded data for group 2 is object-encoded data for immersive sound.

[0142] Furthermore, the information "GroupID[2]=3, object lang1" contained in "Desc" indicates that the encoded data for group 3 is object encoded data for the speech language of the first language. Similarly, the information "GroupID[3]=4, object lang2" contained in "Desc" indicates that the encoded data for group 4 is object encoded data for the speech language of the second language. Additionally, the information "SW_GRPID[0]=1" contained in "Desc" indicates that groups 3 and 4 constitute switch group 1. Finally, the information "Stream_ID" indicates the stream identifier of the audio stream in question.

[0143] The flowcharts in Figures 22 and 23 show an example of the audio decoding control process performed by the CPU 221 in the service receiver 200 shown in Figure 15. In step ST1, the CPU 221 starts processing. Then, in step ST2, the CPU 221 detects the receiver speaker configuration, that is, the speaker configuration of the speaker system 215. Next, in step ST3, the CPU 221 obtains selection information regarding the audio output by the viewer (user).

[0144] Next, in step ST4, the CPU221 reads descriptor information for the mainstream in the PMT, selects the audio stream to which a group with attributes matching the speaker configuration and listener selection information belongs, and loads it into the buffer. Then, in step ST5, the CPU221 checks whether or not a descriptor type packet exists in the audio stream.

[0145] Next, in step ST6, CPU221 determines whether a descriptor-type packet exists. If it does, in step ST7, CPU221 reads the descriptor information of the packet, detects the "groupID", "attribute", "switchGroupID", and "presetGroupID" information, and then proceeds to step ST9. On the other hand, if it does not exist, in step ST8, CPU221 detects the "groupID", "attribute", "switchGroupID", and "presetGroupID" information from the PMT descriptor information, and then proceeds to step ST9. It is also possible to perform a process that decodes the entire target audio stream without executing step ST8.

[0146] In step ST9, the CPU 221 decides whether or not to decode the object-encoded data. If it decides to decode, the CPU 221 decodes the object-encoded data in step ST10 and then proceeds to the process in step ST11. On the other hand, if it decides not to decode, the CPU 221 immediately proceeds to the process in step ST11.

[0147] In step ST11, the CPU 221 decides whether or not to decode the channel-encoded data. If decoding is chosen, in step ST12, the CPU 221 decodes the channel-encoded data and, if necessary, performs downmixing or upmixing to the speaker configuration of the speaker system 215 to obtain audio data for driving each speaker. After that, the CPU 221 proceeds to step ST13. On the other hand, if decoding is not chosen, the CPU 221 immediately proceeds to step ST13.

[0148] In step ST13, if the CPU 221 decodes the object-encoded data, it calculates either mixing it with channel data or speaker rendering based on that information. In the speaker rendering calculation, the speaker rendering (mixing ratio to each speaker) is calculated using azimuth (direction information) and elevation (elevation angle information), and according to the calculation result, the object's audio data is mixed with channel data to drive each speaker.

[0149] Next, in step ST14, the CPU 221 controls the dynamic range of the audio data to drive each speaker and outputs it. After that, in step ST15, the CPU 221 terminates the process.

[0150] As described above, in the transmission / reception system 10 shown in Figure 1, the service transmitter 100 inserts attribute information indicating the attributes of each of the multiple groups of encoded data contained in a predetermined number of audio streams into the container layer and / or the audio stream layer. Therefore, the receiving side can easily recognize the attributes of each of the multiple groups of encoded data before decoding the encoded data, and can selectively decode and use only the encoded data of the necessary groups, thereby reducing the processing load.

[0151] Furthermore, in the transmission / reception system 10 shown in Figure 1, the service transmitter 100 inserts stream correspondence information into the container layer and / or audio stream layer, indicating which audio stream each group of encoded data is contained in. As a result, the receiving side can easily recognize the audio stream containing the encoded data of the required group, thereby reducing the processing load.

[0152] <2. Variant> In the above-described embodiment, the service receiver 200 is configured to selectively extract from multiple audio streams transmitted from the service transmitter 100 the audio stream containing encoded data of a group having attributes that match the speaker configuration and viewer selection information, and to perform a decoding process to obtain a predetermined number of audio data for driving the speakers.

[0153] However, as a service receiver, it is also conceivable to selectively extract one or more audio streams from multiple audio streams transmitted from the service transmitter 100 that have encoded data for groups with attributes that match the speaker configuration and listener selection information, reconstruct the audio streams that have encoded data for groups with attributes that match the speaker configuration and listener selection information, and distribute the reconstructed audio streams to devices connected to the local area network (including DLNA devices).

[0154] Figure 24 shows an example configuration of a service receiver 200A that distributes the reconstructed audio stream to devices connected to the local network, as described above. In Figure 24, parts corresponding to those in Figure 15 are denoted by the same reference numerals, and their detailed descriptions are omitted where appropriate.

[0155] In the demultiplexer 202, under the control of the CPU 221, packets of one or more audio streams containing encoded data of groups with attributes that match the speaker configuration and listener selection information are selectively extracted by a PID filter from a predetermined number of audio streams possessed by the transport stream TS.

[0156] The audio stream extracted by the demultiplexer 202 is loaded into the corresponding multiplexing buffer from the multiplexing buffers 211-1 to 211-N. The combiner 212 reads the audio stream frame by frame from each multiplexing buffer into which the audio stream was loaded and supplies it to the stream reconstruction unit 231.

[0157] In the stream reconstruction unit 231, if the audio stream supplied from the combiner 212 contains descriptor information ("Desc"), the descriptor information is extracted and sent to the CPU 221. Under the control of the CPU 221, the stream reconstruction unit 231 selectively acquires encoded data for groups with attributes that match the speaker configuration and viewer (user) selection information, and reconstructs an audio stream containing that encoded data. This reconstructed audio stream is supplied to the distribution interface 232. From this distribution interface 232, it is distributed (transmitted) to the devices 300 connected to the local area network.

[0158] This on-premises network connectivity includes Ethernet connections and wireless connections such as "Wi-Fi" or "Bluetooth." "Wi-Fi" and "Bluetooth" are registered trademarks.

[0159] Furthermore, device 300 includes surround speakers, a second display, and an audio output device attached to the network terminal. Device 200, which receives the reconstructed audio stream, performs decoding processing similar to that of the 3D audio decoder 213 in the service receiver 200 in Figure 15 to obtain audio data for driving a predetermined number of speakers.

[0160] Furthermore, as a service receiver, a configuration is also conceivable in which the reconstructed audio stream described above is transmitted to a device connected via a digital interface such as "HDMI (High-Definition Multimedia Interface)", "MHL (Mobile High Definition Link)", or "DisplayPort". Note that "HDMI" and "MHL" are registered trademarks.

[0161] Furthermore, in the above-described embodiment, the stream correspondence information inserted into the container layer, etc., was information indicating the correspondence between the group ID and the substream ID. In other words, the substream ID was used to associate the group with the audio stream. However, it is also conceivable to use a packet identifier (PID: Packet ID) or stream type (stream_type) to associate the group with the audio stream. Note that if the stream type is used, the stream type of each audio stream needs to be changed.

[0162] Furthermore, in the above-described embodiment, an example was shown in which attribute information of the encoded data of each group is transmitted with an "attribute_of_groupID" field (see Figure 10). However, this technology also includes a method in which a special meaning is defined for the value of the group ID itself between the transceiver and receiver, so that the type (attribute) of the encoded data can be recognized by recognizing a specific group ID. In this case, the group ID functions not only as a group identifier but also as attribute information of the encoded data of that group, and the "attribute_of_groupID" field becomes unnecessary.

[0163] Furthermore, the above-described embodiment shows an example in which the encoded data of multiple groups includes both channel-encoded data and object-encoded data (see Figure 3). However, this technology can be similarly applied when the encoded data of multiple groups includes only channel-encoded data or only object-encoded data.

[0164] Furthermore, the above-described embodiment showed an example where the container is a transport stream (MPEG-2 TS). However, this technology can also be applied to systems that deliver content in MP4 or other formats. For example, MPEG-DASH based stream delivery systems, or transmission and reception systems that handle MMT (MPEG Media Transport) structured transmission streams.

[0165] Furthermore, this technology can also be configured as follows: (1) A transmission unit that transmits a container in a predetermined format having a predetermined number of audio streams containing encoded data of multiple groups, The container layer and / or audio stream layer are provided with an information insertion unit that inserts attribute information indicating the attributes of each of the multiple groups of encoded data. Transmitter. (2) The above information insertion unit is, Further insert stream correspondence information into the container layer and / or audio stream layer, indicating which audio stream each of the multiple groups of encoded data is contained within. The transmitting device described in (1) above. (3) The above stream compatibility information is This information shows the correspondence between a group identifier that identifies each of the above-mentioned multiple groups of encoded data and a stream identifier that identifies each of the above-mentioned predetermined number of audio streams. The transmitting device described in (2) above. (4) The above information insertion unit is, Further insert stream identifier information, indicating the stream identifier of each of the predetermined number of audio streams, into the container layer and / or the audio stream layer. The transmitting device described in (3) above. (5) The above container is MPEG2-TS, The above information insertion section is, When inserting the above stream identifier information into the above container, the stream identifier information is inserted into the audio element stream loop corresponding to each of the predetermined number of audio streams located under the program map table. The transmitting device described in (4) above. (6) The above information insertion unit is, When inserting the above stream identifier information into the audio stream, the stream identifier information is inserted into the PES payload of each PES packet of the predetermined number of audio streams. The transmitting device described in (4) or (5) above. (7) The above stream correspondence information is This information shows the correspondence between a group identifier that identifies each of the above-mentioned multiple groups of encoded data and a packet identifier that is assigned when each of the above-mentioned predetermined number of audio streams is packetized. The transmitting device described in (2) above. (8) The above stream correspondence information is This information shows the correspondence between a group identifier that identifies each of the above-mentioned multiple groups of encoded data and type information that indicates the stream type of each of the above-mentioned predetermined number of audio streams. The transmitting device described in (2) above. (9) The above container is MPEG2-TS, The above information insertion section is, When inserting the above attribute information and stream correspondence information into the above container, the attribute information and stream correspondence information are inserted into an audio elementary stream loop corresponding to at least one of the predetermined number of audio streams located under the program map table. A transmitting device as described in any of (2) to (8) above. (10) The above information insertion unit is, When inserting the above attribute information and stream correspondence information into the audio stream, the attribute information and stream correspondence information are inserted into the PES payload of the PES packet of at least one of the predetermined number of audio streams. A transmitting device as described in any of (2) to (8) above. (11) The coded data of the above-mentioned multiple groups includes either or both channel coded data and object coded data. A transmitting device as described in any of (1) to (10) above. (12) A transmission step in which the transmission unit transmits a container of a predetermined format having a predetermined number of audio streams including encoded data of multiple groups, The container layer and / or audio stream layer have an information insertion step which inserts attribute information indicating the attributes of each of the multiple groups of encoded data. Sending method. (13) A receiving unit that receives a container in a predetermined format having a predetermined number of audio streams including encoded data of multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. The processing unit further comprises processing the predetermined number of audio streams contained in the received container based on the attribute information. Receiving device. (14) The container layer and / or the audio stream layer further insert stream correspondence information indicating which audio stream each of the above groups of encoded data is contained in, The above-mentioned processing unit is, In addition to the attribute information described above, the predetermined number of audio streams are processed based on the stream correspondence information described above. The receiving device described in (13) above. (15) The above processing unit is Based on the attribute information and stream correspondence information described above, the system selectively decodes audio streams containing encoded data of groups with attributes that match the speaker configuration and user selection information. The receiving device described in (14) above. (16) The coded data of the above-mentioned multiple groups includes either or both channel coded data and object coded data. A receiving device as described in any of (13) to (15) above. (17) The receiving unit has a receiving step of receiving a container of a predetermined format having a predetermined number of audio streams including encoded data of multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. The process further includes a processing step of processing the predetermined number of audio streams contained in the received container based on the attribute information. Reception method. (18) A receiving unit that receives a container in a predetermined format having a predetermined number of audio streams including encoded data of multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. A processing unit that selectively obtains encoded data of a predetermined group from the predetermined number of audio streams held by the received container based on the attribute information, and reconstructs an audio stream containing the encoded data of the predetermined group, The processing unit further comprises a stream transmission unit that transmits the reconstructed audio stream to an external device. Receiving device. (19) The container layer and / or the audio stream layer further insert stream correspondence information indicating which audio stream each of the above groups of encoded data is contained in, The above-mentioned processing unit is, In addition to the attribute information described above, based on the stream correspondence information described above, the encoded data of a predetermined group is selectively obtained from a predetermined number of audio streams. The receiving device described in (18) above. (20) The receiving unit has a receiving step of receiving a container of a predetermined format having a predetermined number of audio streams including encoded data of multiple groups, The container layer and / or audio stream layer have attribute information inserted that indicates the attributes of each of the multiple groups of encoded data. A processing step of selectively obtaining encoded data of a predetermined group from the predetermined number of audio streams held by the received container based on the attribute information, and reconstructing an audio stream containing the encoded data of the predetermined group, The process further includes a stream transmission step for transmitting the audio stream reconstructed in the above processing step to an external device. Reception method.

[0166] The main feature of this technology is that it reduces the processing load on the receiving side by inserting attribute information indicating the attributes of each of the encoded data in multiple groups contained in a predetermined number of audio streams, as well as stream correspondence information indicating which audio stream each of the encoded data in multiple groups is contained in, into the container layer and / or the audio stream layer (see Figure 14). [Explanation of symbols]

[0167] 10. Transmit / Receive System 100... Service Transmitter 110...Stream generation unit 112... Video Encoder 113... Audio Encoder 114...Multiplexer 200, 200A... Service receiver 201... Receiver 202... Demultiplexer 203...Video Decoder 204...Video processing circuit 205... Panel drive circuit 206...Display Panel 211-1~211-N···Multiplexed buffer 212... Combiner 213...3D Audio Decoder 214...Audio output processing circuit 215...Speaker System 221...CPU 222...Flash ROM 223···DRAM 224...Internal bus 225... Remote control receiver 226... Remote control transmitter 231...Stream Reconstruction Unit 232...Distribution Interface 300 devices

Claims

1. The receiving unit includes a container having a predetermined number of audio streams containing encoded data from multiple groups, The encoded data included in the predetermined number of audio streams are distinguished into groups by type, and the encoded data that can be selected from among the groups is registered in a switch group and encoded. The container layer and / or the audio stream layer include a group ID, which is an identifier for identifying the group; attribute information indicating the attributes of each of the encoded data of the multiple groups; a switch group ID, which is an identifier for identifying the switch group; and an identifier indicating the content type of each of the groups. The processing unit further comprises a processing unit that selects encoded data for a target group based on the information contained in the container layer and / or the audio stream layer for a predetermined number of audio streams, and performs a process of selectively decoding only the encoded data of the target group. Information processing device.

2. The container is MPEG-2 TS. The information processing apparatus according to claim 1.

3. The audio stream includes either or both channel-coded data and object-coded data. The information processing apparatus according to claim 1.

4. The object encoding data includes an audio object and metadata for mapping the audio object to the audio output unit and rendering it. The information processing apparatus according to claim 3.

5. The aforementioned metadata is included as an extension element (Ext_element), The information processing apparatus according to claim 4.

6. The predetermined number of audio streams consist of a main stream and substreams. The information processing apparatus according to claim 1.

7. The identifier indicating the type of content in each of the aforementioned groups takes a value between 0 and 15. The information processing apparatus according to claim 1.

8. The predetermined number of audio streams include identifiers for identifying the substreams. The information processing apparatus according to claim 6.

9. The object encoding data includes encoding data for speech language. The information processing apparatus according to claim 3.

10. Each of the predetermined number of audio streams includes a descriptor that describes information related to each of the audio streams. The information processing apparatus according to claim 1.

11. The descriptor has an 8-bit field "descriptor_tag" which is information indicating the descriptor type of the descriptor. The information processing apparatus according to claim 10.

12. The descriptor has an 8-bit field "descriptor_length" which is information indicating the length or size of the descriptor. The information processing apparatus according to claim 10.

13. Each of the predetermined number of audio streams corresponds to an audio elemental stream loop (audio ES loop), which includes a descriptor that describes information related to each of the audio streams. The information processing apparatus according to claim 1.

14. The receiving unit has a first step of receiving a container having a predetermined number of audio streams, which include encoded data from multiple groups. The encoded data included in the predetermined number of audio streams are distinguished into groups by type, and the encoded data that can be selected from among the groups is registered in a switch group and encoded. The container layer and / or the audio stream layer include a group ID, which is an identifier for identifying the group; attribute information indicating the attributes of each of the encoded data of the multiple groups; a switch group ID, which is an identifier for identifying the switch group; and an identifier indicating the content type of each of the groups. The processing unit further comprises a second step in which it selects encoded data for a target group based on the information contained in the container layer and / or the audio stream layer for a predetermined number of audio streams, and performs a process to selectively decode only the encoded data for the target group. Information processing methods.

15. The first step involves receiving a container having a predetermined number of audio streams, which include encoded data from multiple groups. The encoded data included in the predetermined number of audio streams are distinguished into groups by type, and the encoded data that can be selected from among the groups is registered in a switch group and encoded. The container layer and / or the audio stream layer include a group ID, which is an identifier for identifying the group; attribute information indicating the attributes of each of the encoded data of the multiple groups; a switch group ID, which is an identifier for identifying the switch group; and an identifier indicating the content type of each of the groups. The processing unit further comprises a second step in which it selects encoded data for a target group based on the information contained in the container layer and / or the audio stream layer for a predetermined number of audio streams, and performs a process to selectively decode only the encoded data for the target group. A program that causes a computer to execute an information processing method.