Transmitting device, receiving device, and program
The integration of multi-layer coding and object-based audio methods in video and audio transmission systems through layer and audio object associations in the multiplexing stage addresses the inefficiencies of existing technologies, enabling efficient decoding and user-customized program selection.
Patent Information
- Application Number
- JP2025003801
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-06
- Filing Date
- 2025-01-09
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies struggle to efficiently combine multi-layer coding and object-based audio methods in video and audio transmission systems, as they lack the ability to select layers and audio objects at the multiplexing level for efficient decoding.
A transmitting device and receiving device that utilize multi-layer compatible video encoding, object-based audio encoding, and multiplexing to include control information for layer and audio object associations, enabling efficient decoding by specifying layer and audio object combinations based on user preferences.
Enables efficient decoding processing and customization of broadcast programs by allowing users to select and combine layers and audio objects according to their preferences, enhancing the decoding efficiency and user experience.
Smart Images

Figure 2025146664000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a transmitting device, a receiving device, and a program used in a video and audio transmission system. [Background technology]
[0002] To improve the quality and functionality of terrestrial digital broadcasting, advanced terrestrial broadcasting, a transmission method for next-generation terrestrial broadcasting, is being studied. Advanced terrestrial broadcasting will introduce the latest technologies to realize IP and large capacity transmission, and will be able to respond to future changes in the media environment.
[0003] The aim is to realize a variety of use cases by adopting H.266 (VVC) for the video coding method to improve compression ratios and support multi-layer profiles (also known as "multi-layer coding"), and by adopting MPEG-H 3DA / AC-4 for the audio coding method to utilize object-based audio (see, for example, Non-Patent Documents 1 and 2). As for the multiplexing method, the adoption of the MMT-TLV method, which is used in the new 4K8K satellite broadcasting, is being considered (see, for example, Non-Patent Documents 3 and 4). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] ITU-T Recommendation H.266 Versatile video coding [Non-patent document 2] ISO / IEC 23008-3, “High efficiency coding and media delivery in heterogeneous environments Part 3: 3D audio” [Non-patent document 3] ARIB STD-B60, "Media Transport Method Using MMT in Digital Broadcasting" [Non-patent document 4] ISO / IEC 23008-1, “High efficiency coding and media delivery in heterogeneous environments Part 1: MPEG media transport (MMT)” Summary of the Invention [Problem to be solved by the invention]
[0005] In the advancement of terrestrial broadcasting, services that apply multi-layer coding are expected to include, for example, 8K / 4K / 2K services that combine broadcasting and communications, and services that transmit main content and sub-content in separate layers to customize broadcast programs to suit individual preferences. Furthermore, services that use object-based audio are expected to include customizing program audio, such as allowing viewers to watch sports by substituting commentary audio from the perspective of their favorite team, or by increasing the volume of just the commentary audio.
[0006] To realize these services, the receiving device needs information to recognize the content of each layer and determine which layers to combine for decoding according to the user's desired resolution and preferences, and which audio objects to combine for decoding in the selected layers.
[0007] The syntax of a bitstream (also called "video coded data") coded using a multi-layer coding method specifies information about multiple layers, but in order to speed up the video decoding process and reduce the load, it is desirable to be able to select layers at the multiplexing level, which is a stage before decoding the video coded data. Also, a bitstream (also called "audio coded data") coded using an object-based audio method includes metadata that can describe presentation patterns that combine various audio objects, but conventional technologies cannot indicate which pattern should be used to play audio for a selected video layer.
[0008] Therefore, an object of the present invention is to provide a transmitting device, a receiving device, and a program that enable efficient decoding processing while enabling a service that combines a multi-layer coding method and an object-based audio method. [Means for solving the problem]
[0009] The transmitting device of the first aspect is a transmitting device used in a video and audio transmission system, and comprises: a video encoding means for outputting video encoding data of multiple layers by performing video encoding processing corresponding to a multi-layer encoding method; an audio encoding means for outputting audio encoding data of multiple audio objects by performing audio encoding processing corresponding to an object-based audio method; and a multiplexing means for multiplexing the video encoding data, the audio encoding data, and control information including first information regarding the multiple layers and second information regarding the multiple audio objects, and outputting the multiplexed data.
[0010] A receiving device according to the second aspect is a receiving device used in a video and audio transmission system, and comprises: a receiving means for receiving multiplexed data and obtaining from the received data video encoding data of a plurality of layers, audio encoding data of a plurality of audio objects, and control information including first information regarding the plurality of layers and second information regarding the plurality of audio objects; a video decoding means for performing video decoding processing corresponding to a multi-layer encoding method on the video encoding data based on the control information; and an audio decoding means for performing audio decoding processing corresponding to an object-based audio method on the audio encoding data based on the control information.
[0011] A program according to the third aspect causes a computer to function as the transmitting device according to the first aspect or the receiving device according to the second aspect. [Effects of the Invention]
[0012] According to the present invention, it is possible to provide a transmitting device, a receiving device, and a program that enable efficient decoding processing and realize a service that combines a multi-layer coding method and an object-based audio method. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a transmission device according to an embodiment. [Figure 2] FIG. 10 is a diagram showing the configuration of an MMT package table (MPT) described in Non-Patent Document 3. [Figure 3] FIG. 2 is a diagram illustrating a first example of a configuration of a multilayer service descriptor (Multilayer_Service_descriptor) according to the embodiment. [Figure 4] FIG. 2 is a diagram illustrating a first example of the configuration of an audio configuration descriptor (Audio_config_descriptor) according to the embodiment. [Figure 5] FIG. 10 is a diagram illustrating a second example of a configuration of a multilayer service descriptor (Multilayer_Service_descriptor) according to the embodiment. [Figure 6] FIG. 10 is a diagram illustrating a second example of the configuration of an audio configuration descriptor (Audio_config_descriptor) according to the embodiment. [Figure 7] FIG. 10 is a diagram illustrating a third example of a configuration of a multilayer service descriptor (Multilayer_Service_descriptor) according to the embodiment. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of a receiving device according to an embodiment. [Figure 9] FIG. 2 is a diagram illustrating a service image of the first embodiment and a relationship between an OLS and an output layer in the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of description of audio metadata in the first embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of a description of a multi-layer service descriptor according to the first embodiment. [Figure 12] FIG. 10 is a diagram showing an example of a description of an audio configuration descriptor in the first embodiment. [Figure 13] FIG. 10 is a diagram illustrating a service image of the second embodiment and a relationship between the OLS and the output layer in the second embodiment. [Figure 14] FIG. 10 is a diagram showing an example of description of audio metadata in the second embodiment. [Figure 15] FIG. 10 is a diagram illustrating an example of a description of a multi-layer service descriptor in the second embodiment. [Figure 16] FIG. 10 is a diagram showing an example of a description of an audio configuration descriptor in the second embodiment. [Figure 17] FIG. 11 is a diagram showing an example of description of audio metadata in the third embodiment. [Figure 18] FIG. 11 is a diagram illustrating an example of a description of a multi-layer service descriptor in the third embodiment. [Figure 19] FIG. 11 is a diagram showing an example of a description of an audio configuration descriptor in the third embodiment. [Figure 20] FIG. 10 is a diagram showing the configuration of a video preset descriptor according to a modified example. [Figure 21] FIG. 10 is a diagram showing the configuration of an audio preset descriptor according to a modified example. [Figure 22] FIG. 10 is a diagram showing the configuration of an audio preset descriptor according to a modified example. [Figure 23] FIG. 10 is a diagram for explaining the meaning of an audio preset descriptor. [Figure 24] FIG. 10 is a diagram for explaining the meaning of an audio preset descriptor. [Figure 25] FIG. 10 is a diagram for explaining the meaning of an audio preset descriptor. DETAILED DESCRIPTION OF THE INVENTION
[0014] A video and audio transmission system according to an embodiment will be described with reference to the drawings. In the description of the drawings, the same or similar parts are denoted by the same or similar reference numerals.
[0015] (1) Example of transmitter configuration First, a configuration example of a transmission device according to this embodiment will be described. Fig. 1 is a diagram showing a configuration example of a transmission device 100 according to this embodiment.
[0016] The transmitting device 100 includes a multi-layer compatible video encoding unit 110 , an object-based audio compatible audio encoding unit 120 , a multiplexing unit 130 , and a transmitting unit 140 .
[0017] The multi-layer-compatible video encoding unit 110 outputs video encoded data of multiple layers by performing video encoding processing compatible with a multi-layer encoding method. In the example shown in the figure, the multi-layer-compatible video encoding unit 110 outputs a video bitstream of layer 0, a video bitstream of layer 1, and a video bitstream of layer 2 as video encoded data of multiple layers, but the number of layers may be two, or may be four or more.
[0018] Here, the multi-layer-compatible video encoding unit 110 encodes the input video to be encoded in a layer structure. The multi-layer-compatible video encoding unit 110 first encodes the lowest layer (layer 0) and uses the decoded video of layer 0 for prediction in the encoding process of higher layers (layer 1, layer 2, ...), thereby improving prediction efficiency of the higher layers and improving encoding efficiency compared to when each layer is encoded independently. The lowest layer is also referred to as a base layer, and the higher layers are also referred to as enhancement layers. Note that multi-layer encoding is also referred to as hierarchical encoding.
[0019] For example, spatial scalable coding (also referred to as "spatial hierarchical coding") is a typical use case of multi-layer coding, in which low-resolution video such as 2K video is coded as a base layer, and an enhancement layer, which is high-resolution video such as 4K or 8K video, is coded using a decoded image of the low-resolution video. Another use case is a technology that applies multi-layer coding to realize a service that adds sub-video (sub-content) (see, for example, JP 2022-53534 A). In this technology, a main video (main content) is coded as a base layer, and a video in which an add-on video (sub-video) such as a commentary video is superimposed on the main video is coded as an enhancement layer and transmitted. On the receiving side, the base layer and the enhancement layer are combined and decoded to display the video with the add-on information.
[0020] In this embodiment, it is assumed that the Multilayer Main 10 profile in VVC (Versatile Video Coding) is used as the multilayer encoding method. The Multilayer Main 10 profile is a profile that realizes the transmission of multiple layers compared to the Main 10 profile, and is capable of transmitting video of multiple layers. Alternatively, the Scalable Main 10 profile in HEVC (High Efficiency Video Coding) may be used as the multilayer encoding method. However, it is not limited to HEVC or VVC, and other video encoding methods may also be used.
[0021] The object-based audio-compatible audio encoder 120 outputs encoded audio data of a plurality of audio objects by performing audio encoding processing compatible with the object-based audio format. In the illustrated example, the object-based audio-compatible audio encoder 120 outputs an audio bitstream including audio metadata as encoded audio data of a plurality of audio objects.
[0022] The object-based audio-compatible audio encoder 120 encodes and outputs input audio to be encoded into audio materials (audio objects) such as commentary, background sounds, and sound effects. It transmits audio encoding data (audio bitstream) including each audio object and audio metadata, which is information on their configuration, playback position, etc., and on the receiving side, the audio objects are synthesized and played back according to the purpose.
[0023] In this embodiment, it is assumed that MPEG-H 3DA / AC-4 is used as the audio coding method compatible with object-based audio, but this is not limited to MPEG-H 3DA / AC-4 and other audio coding methods may also be used.
[0024] The multiplexing unit 130 multiplexes the video encoded data output by the multi-layer compatible video encoding unit 110, the audio encoded data output by the object-based audio compatible audio encoding unit 120, and control information, and outputs the multiplexed data.
[0025] In the illustrated example, the multiplexing unit 130 multiplexes the video bitstream of Layer 0, the video bitstream of Layer 1, the audio bitstream, and the control information, and outputs the result as a "video and audio stream + control information," and also outputs the video bitstream of Layer 2. However, the multiplexing unit 130 may also multiplex the video bitstream of Layer 2 as a "video and audio stream + control information," and output the result.
[0026] In this embodiment, it is assumed that the MPEG Media Transport (MMT) method (MMT-TLV method) is used as the multiplexing method. However, the Transport stream (TS) defined in MPEG-2 Systems may also be used as the multiplexing method. The multiplexing method is not limited to these, and any multiplexing method that can transmit multiple streams synchronously may be used.
[0027] The multiplexing unit 130 has a control information generating unit 131 that generates control information. The control information includes TLV-SI (Signaling Information) related to the TLV (Type Length Value) multiplexing method for multiplexing IP packets, and MMT-SI related to MMT, a media transport method. The control information consists of messages, tables, and descriptors.
[0028] The control information generator 131 generates control information including first information related to multiple layers in multi-layer coding and second information related to multiple audio objects in object-based audio. Specifically, the control information generator 131 adds a new descriptor that describes information necessary for realizing services that apply multi-layer coding and object-based audio to the descriptors that define the detailed parts of the control information (MMT-SI) in addition to those used in the new 4K8K satellite broadcasting. This makes it possible to describe and transmit necessary and sufficient information in a system layer (multiplexing layer) that is different from the coding layer.
[0029] In this embodiment, the control information generator 131 adds a multi-layer service descriptor as first information regarding multiple layers in multi-layer coding to an MPT descriptor area or an asset descriptor area of a video asset in an MMT package table (MPT) that describes information about broadcast programs.
[0030] Fig. 2 is a diagram showing the configuration of the MMT package table (MPT) described in Non-Patent Document 3. The meaning of the MMT package table (MPT) is defined in Non-Patent Document 3, but here we will mainly explain the contents related to the MPT descriptor area and asset descriptor area.
[0031] The MMT package table (MPT) includes table_id (table identification), MPT_mode (MPT mode), MMT_package_id_length (package ID length), MMT_package_id_byte (package ID byte), MPT_descriptors_length (MPT descriptor length), and MPT_descriptors_byte (MPT descriptor area). Note that the lowest 16 bits of the package ID are the same value as the service ID used to identify the service.
[0032] MPT_descriptors_byte (MPT descriptor area) is an area for storing MPT descriptors. In this embodiment, a multi-layer service descriptor can be added to MPT_descriptors_byte (MPT descriptor area).
[0033] The MMT Package Table (MPT) also includes number_of_assets (number of assets) indicating the number of assets for which this table provides information, and information for each asset. The information for each asset includes identifier_type (identifier type), asset_id_scheme (asset ID format), asset_id_length (asset ID length), asset_id_byte (asset ID byte) indicating the asset ID, asset_type (asset type) indicating the asset type (video, audio, timed text, application, synchronized general data, synchronized general data), asset_clock_relation_flag (clock information flag), location_count (number of locations) indicating the number of location information for the asset, MMT_general_location_info (location information) indicating the location information for the asset, asset_descriptors_length (asset descriptor length) indicating the total byte length of the subsequent descriptors, and asset_descriptors_byte (asset descriptor area) which is an area for storing asset descriptors.
[0034] In this embodiment, a multi-layer service descriptor can also be added to asset_descriptors_byte (asset descriptor area).
[0035] The multi-layer service descriptor is a new descriptor that indicates the structure of multiple layers in multi-layer coding. This enables the transmission of multi-layer service structure information at the system layer. On the receiving side, the multi-layer service structure information can be obtained before decoding the video coded data, enabling efficient video decoding processing such as selectively decoding only the necessary layers.
[0036] The multi-layer service descriptor describes, for example, a list of layer IDs related to the multi-layer service, each OLS (ols_index) that can be selected as a multi-layer service, the profile (PTL) required for the receiver, the layer and output layer required for decoding, the audio program ID (audio_programm_id) or audio tag (audio_tag) indicated by object-based audio, and text information such as an explanation of the OLS. Details of the multi-layer service descriptor will be described later.
[0037] Furthermore, in this embodiment, the control information generator 131 adds an audio configuration descriptor as second information related to multiple audio objects in object-based audio to the asset descriptor area of the audio asset of the MPT. The audio configuration descriptor is a new descriptor that indicates the configuration of multiple audio objects in object-based audio. This makes it possible to transmit the configuration information of object audio in the system layer.
[0038] The audio configuration descriptor describes, for example, the configuration of an audio program ID (audio_program_id), an audio object ID (audio_object_id), and an audio tag (audio_tag) specified in the audio metadata of object-based audio. The audio configuration descriptor will be described in detail later.
[0039] The MPT is a table included in the PA (Package Access) message. The PA message is the entry point for MMT-SI and transmits the MMT-SI table. The MPT is a table that provides information that configures a service (package) equivalent to a broadcast program. The MPT includes a list of assets, which are transmission units for video and audio multiplexed using the MMT method, as well as information on their locations.
[0040] Furthermore, the control information generator 131 includes, in the control information, association information indicating an association between a layer among the multiple layers in multi-layer coding and an audio object among the multiple audio objects in object-based audio, which enables the receiving side to specify, based on the association information, which audio object should be combined with the selected video layer for decoding.
[0041] In this embodiment, the control information generator 131 includes the association information in the multi-layer service descriptor and the audio configuration descriptor. Specifically, to associate an audio object with a video layer, an audio program ID or an audio tag is used as the association information. However, instead of including the association information in the multi-layer service descriptor and the audio configuration descriptor, the association information may be included in another descriptor (for example, an existing descriptor).
[0042] The transmitting unit 140 performs transmission channel coding and the like on the multiplexed data output by the multiplexing unit 130 and transmits the data via the transmission channel 10. In this embodiment, the transmitting unit 140 has a broadcast sending unit 141 that performs transmission via the broadcast transmission channel 10a and a communication unit 142 that performs transmission via the communication transmission channel 10b. However, the transmitting unit 140 is not limited to this configuration, and may have a configuration that includes only either the broadcast sending unit 141 or the communication unit 142.
[0043] In the illustrated example, the broadcast sending unit 141 transmits a "video and audio stream + control information" via the broadcast transmission path 10a. The communication unit 142 transmits a Layer 2 video bitstream via the communication transmission path 10b. Note that the broadcast transmission path 10a is a transmission path for one-way transmission. The communication transmission path 10b is a transmission path that allows two-way transmission, and may include, for example, the Internet.
[0044] (2) Example of control information configuration Next, configuration examples 1 to 3 of the control information according to this embodiment will be described.
[0045] (2.1) Control information configuration example 1 In this configuration example, the multi-layer service descriptor is placed in the MPT descriptor area of the MMT package table (MPT). Also, in this configuration example, the association between the video layer and the audio object is performed using a layer ID (layer_id) and an audio program ID (audio_program_id). That is, in this configuration example, the layer ID and the audio program ID correspond to association information that indicates the association between the video layer and the audio object.
[0046] 3 is a diagram showing a first configuration example of a multilayer service descriptor (Multilayer_Service_descriptor) according to this embodiment. The meaning of the multilayer service descriptor is as follows.
[0047] descriptor_tag: The descriptor tag is a 16-bit field that identifies the multi-layer service descriptor.
[0048] descriptor_length (descriptor length): This field is used to write the number of data bytes that follow. The bit length of the descriptor length field varies depending on the descriptor.
[0049] num_of_layer (number of layers): The number of layers that make up the service, with the upper limit assumed to be 8.
[0050] layer_id (layer ID): A unique ID assigned to each layer. List the layer_id (layer ID) for the number of layers indicated by num_of_layer.
[0051] ols_mode_idc (OLS mode): Indicates the mode of the Output Layer Set (OLS), with an upper limit of 4 assumed.
[0052] num_of_ols (number of OLS): Indicates the number of OLS, and the upper limit is assumed to be 16.
[0053] ols_index (OLS index): A unique ID assigned to each OLS.
[0054] profile_tier_level: Indicates the PTL (Profile Tier Level) applied to each OLS, with the upper limit assumed to be 8. Note that PTL is a parameter set that indicates the profile, tier, and level of the bitstream.
[0055] layer_used_for_decode_flag (layer flag for decoding): Indicates whether the layer is necessary for decoding. If it is a necessary layer, set to '1'.
[0056] output_layer_flag (output layer flag): Indicates whether it is an output layer. If it is an output layer, it is set to '1'. Note that output_layer_flag is applied when OLS mode = 2. When OLS mode = 2, the output layer is specified, and other layers are direct or indirect reference layers of the OLS output layer.
[0057] audio_program_id (audio program ID): An ID indicating an audio program assigned in object-based audio corresponding to each OLS.
[0058] ISO_639_language_code (language code): This 24-bit field indicates the language of the following text information field using a three-letter alphabetic code defined in ISO 639-2.
[0059] text_length (description length): This 8-bit field indicates the byte length of the following description.
[0060] text_char (description): This is an 8-bit field, a series of textual information that describes the layer.
[0061] The data structure (a) in FIG. 3 will be explained.
[0062] Enter a list of ols_index (OLS index) for the number of OLSs indicated by num_of_ols (number of OLSs).
[0063] For OLS with OLS mode = 0 or 1, the layer_used_for_decode_flag (layer flag to be decoded) is described for each layer. Here, for a layer required for decoding (layer_used_for_decode_flag = 1), the audio_program_id (audio program ID) is described to associate the layer with an audio object (audio program).
[0064] For OLS with OLS mode = 2, the layer_used_for_decode_flag (layer flag to be decoded) and output_layer_flag (output layer flag) are written for each layer. Here, for layers required for decoding (layer_used_for_decode_flag = 1), the audio_program_id (audio program ID) is written to associate the layer with an audio object (audio program).
[0065] Also, enter text information (text_char) for the number of OLSs indicated by num_of_ols (number of OLSs).
[0066] 4 is a diagram showing a first example of the configuration of the audio configuration descriptor (Audio_config_descriptor) according to this embodiment. The meaning of the audio configuration descriptor is as follows.
[0067] descriptor_tag: The descriptor tag is a 16-bit field that identifies the audio configuration descriptor.
[0068] descriptor_length (descriptor length): This field is used to write the number of data bytes that follow. The bit length of the descriptor length field varies depending on the descriptor.
[0069] num_of_audio_programme (number of audio programs): Indicates the number of audio programs that make up the service, with the upper limit assumed to be 16.
[0070] audio_program_id (audio program ID): An ID that indicates an audio program assigned in object-based audio.
[0071] avs_flag (AVS flag): Indicates whether AVS is enabled. If enabled, it is set to '1'.
[0072] num_of_audio_object (number of audio objects): Indicates the number of audio objects that make up the service, with the upper limit assumed to be 128.
[0073] audio_object_id (audio object ID): An ID indicating an audio object assigned in object-based audio.
[0074] The data structure (b) in FIG. 4 will be explained.
[0075] Enter the audio_programme_id (audio program ID), avs_flag (AVS flag), and num_of_audio_object (number of audio objects) for each audio program. Also, for each audio program, enter the audio_object_id (audio object ID) of each audio object that makes up the audio program for the number of num_of_audio_object (number of audio objects).
[0076] In this configuration example, the audio_program_id (audio program ID) in the audio configuration descriptor is associated with the audio_program_id (audio program ID) in the multi-layer service descriptor. Since the audio_program_id (audio program ID) in the multi-layer service descriptor is described for each video layer, each video layer is associated with an audio program consisting of one or more audio objects. As a result, the receiving side can identify which audio program (audio object) to combine with the selected video layer for decoding based on the audio_program_id (audio program ID).
[0077] Although the configuration in which video layers are associated with audio programs has been described, it is also possible to configure audio programs to be associated with OLSs consisting of one or more layers. Also, in this configuration example, only one audio_program_id (audio program ID) is described, but it is also possible to configure multiple audio_program_ids (audio program IDs) to be described. In any case, it is sufficient if the configuration allows video layers and audio objects to be directly or indirectly associated with each other.
[0078] (2.2) Control information configuration example 2 In this configuration example, the multi-layer service descriptor is placed in the MPT descriptor area of the MMT package table (MPT). Also, in this configuration example, the association between the video layer and the audio object is performed using a layer ID (layer_id) and an audio tag (audio_tag). That is, in this configuration example, the layer ID and audio tag correspond to association information that indicates the association between the video layer and the audio object.
[0079] 5 is a diagram showing a second configuration example of a multilayer service descriptor (Multilayer_Service_descriptor) according to this embodiment. Here, differences from the first configuration example shown in FIG. 3 will be described.
[0080] In this configuration example, audio_tag (audio tag) is described instead of audio_program_id (audio program ID) shown in Fig. 3. audio_tag (audio tag) indicates an audio tag assigned in object-based audio corresponding to each layer.
[0081] Fig. 6 is a diagram showing a second configuration example of the audio configuration descriptor (Audio_config_descriptor) according to this embodiment. Here, differences from the first configuration example shown in Fig. 4 will be mainly described. The meaning of the audio configuration descriptor is as follows:
[0082] num_of_audio_tag (number of audio tags): Indicates the number of audio tags that make up the service, with the upper limit assumed to be 256.
[0083] audio_tag (audio tag): Indicates the audio tag assigned in object-based audio.
[0084] num_of_audio_programme (number of audio programs): Indicates the number of audio programs that make up the service, with the upper limit assumed to be 16.
[0085] audio_program_id (audio program ID): An ID that indicates an audio program assigned in object-based audio.
[0086] avs_flag (AVS flag): Indicates whether AVS is enabled. If enabled, it is set to '1'.
[0087] num_of_audio_object (number of audio objects): Indicates the number of audio objects that make up the service, with the upper limit assumed to be 128.
[0088] audio_object_id (audio object ID): An ID indicating an audio object assigned in object-based audio.
[0089] The data structure (b) in FIG. 6 will be explained.
[0090] The number of audio_tags (audio tags) and num_of_audio_programs (number of audio programs) to be entered is the same as the number of num_of_audio_tags (number of audio tags). In other words, one audio tag is associated with one or more audio programs.
[0091] Enter the audio_programme_id (audio program ID), avs_flag (AVS flag), and num_of_audio_object (number of audio objects) for the number of num_of_audio_programs (number of audio programs). Also, for each audio program, enter the audio_object_id (audio object ID) of each audio object that makes up the audio program for the number of num_of_audio_object (number of audio objects).
[0092] In this configuration example, the audio_tag (audio tag) in the audio configuration descriptor is associated with the audio_tag (audio tag) in the multi-layer service descriptor. Since the audio_tag (audio tag) in the multi-layer service descriptor is described for each video layer, each video layer is associated with an audio tag indicating one or more audio programs. As a result, the receiving side can identify, based on the audio_tag (audio tag), which audio tag (audio program, audio object) to combine with the selected video layer for decoding.
[0093] Although the configuration in which audio tags are associated with each video layer has been described, it is also possible to associate audio tags with each OLS consisting of one or more layers. Also, in this configuration example, only one audio_tag (audio tag) is described, but it is also possible to configure in which multiple audio_tags (audio tags) can be described. In any case, it is sufficient if the configuration allows for direct or indirect association between video layers and audio objects.
[0094] (2.3) Control information configuration example 3 In this configuration example, the multi-layer service descriptor is placed in the asset descriptor area (i.e., for each asset) of the MMT package table (MPT). Also, the association between video layers and audio objects is performed using an ID (video_sets_id) indicating video options and an audio tag (audio_tag), which can be selected collectively in layer units or OLS units. That is, in this configuration example, the ID (video option ID) indicating video options and the audio tag correspond to association information indicating the association between video layers and audio objects.
[0095] The explanation of the audio tag (audio_tag) is the same as in the above-mentioned control information configuration example 2. However, instead of the audio tag, an audio program ID (audio_program_id) may be used as in the above-mentioned control information configuration example 1.
[0096] 7 is a diagram showing a third example of the configuration of a multilayer service descriptor (Multilayer_Service_descriptor) according to this embodiment. The meaning of the multilayer service descriptor is as follows.
[0097] descriptor_tag: The descriptor tag is a 16-bit field that identifies the multi-layer service descriptor.
[0098] descriptor_length (descriptor length): This field is used to write the number of data bytes that follow. The bit length of the descriptor length field varies depending on the descriptor.
[0099] num_of_video_sets (number of video options): Indicates the number of video options, with an upper limit of 16 assumed.
[0100] video_sets_id (video set ID): A unique ID assigned to each video set.
[0101] ols_index (OLS index): A unique ID assigned to each OLS.
[0102] profile_tier_level: Indicates the PTL (Profile Tier Level) applied to each OLS, with the upper limit assumed to be 8. Note that PTL is a parameter set that indicates the profile, tier, and level of the bitstream.
[0103] layer_used_for_decode_flag (layer flag for decoding): Indicates whether the layer is necessary for decoding. If it is a necessary layer, set to '1'.
[0104] output_layer_flag (output layer flag): Indicates whether it is an output layer. If it is an output layer, it is set to '1'. Note that output_layer_flag is applied when OLS mode = 2. When OLS mode = 2, the output layer is specified, and other layers are direct or indirect reference layers of the OLS output layer.
[0105] audio_tag (audio tag): An audio tag assigned by object-based audio corresponding to each video option (video option ID).
[0106] ISO_639_language_code (language code): This 24-bit field indicates the language of the following text information field using a three-letter alphabetic code defined in ISO 639-2.
[0107] text_length (description length): This 8-bit field indicates the byte length of the following description.
[0108] text_char (description): This is an 8-bit field, a series of textual information that describes the layer.
[0109] In this example, only one audio_tag (audio tag) is described, but multiple audio_tags (audio tags) may be described. In any case, it is sufficient that the video layer and the audio object, which is an audio option, can be directly or indirectly associated with each other.
[0110] Furthermore, in any of the above configuration examples 1 to 3, rather than indicating the correspondence between video layers and audio objects in the multi-layer service descriptor, it is also possible to construct an independent descriptor that indicates only the correspondence between video layers and audio objects.
[0111] (3) Example of receiving device configuration Next, an example of the configuration of a receiving device according to this embodiment will be described. Fig. 8 is a diagram showing an example of the configuration of a receiving device 200 according to this embodiment.
[0112] The receiving device 200 includes a receiving unit 210 , a multi-layer compatible video decoding unit 220 , an object-based audio compatible audio decoding unit 230 , a UI (User Interface) unit 240 , and a control unit 250 .
[0113] The receiving unit 210 receives the multiplexed data via the transmission path 10. The receiving unit 210 obtains and outputs from the received data video coded data of multiple layers in multi-layer coding, audio coded data of multiple audio objects in object-based audio, and control information. As described above, the control information includes first information (multi-layer service descriptor) related to the multiple layers and second information (audio configuration descriptor) related to the multiple audio objects.
[0114] In the illustrated example, the receiving unit 210 outputs a video bitstream of layers 0 to 2 corresponding to the video encoding data to a multi-layer compatible video decoding unit 220, outputs an audio bitstream (including audio metadata) corresponding to the audio encoding data to an object-based audio compatible audio decoding unit 230, and outputs control information to a control unit 250.
[0115] In this embodiment, the receiving unit 210 has a broadcast receiving unit 211, a communication unit 212, and a demultiplexing unit 213. The broadcast receiving unit 211 receives broadcast data (broadcast signals) via the broadcast transmission path 10a and outputs the received data to the demultiplexing unit 213. This received data corresponds to the above-mentioned "video and audio stream + control information." The demultiplexing unit 213 performs demultiplexing (MMT decoding) processing compatible with the MMT method, and separates the "video and audio stream + control information" into a video bitstream of layer 0, a video bitstream of layer 1, an audio bitstream, and the control information. The communication unit 212 receives a video bitstream of layer 2 via the communication transmission path 10b and outputs the received video bitstream.
[0116] The multi-layer-compatible video decoding unit 220 performs video decoding processing compatible with the multi-layer encoding method on the video encoded data based on the control information. In the example shown in the figure, the multi-layer-compatible video decoding unit 220 performs video decoding processing on the video bitstreams of layers 0 to 2, and outputs the video signals obtained by the video decoding processing.
[0117] In this embodiment, the multi-layer-compatible video decoding unit 220 has a bitstream synthesizing unit 221 and a video decoding unit 222. The bitstream synthesizing unit 221 synthesizes video bitstreams of each layer (i.e., each layer constituting the OLS) selected in accordance with an OLS index specified by the control unit 250, and outputs the synthesized video bitstream to the video decoding unit 222. The video decoding unit 222 selects a decoding layer in accordance with an output layer ID specified by the control unit 250, decodes the video bitstream of the selected layer, and outputs a video signal. Alternatively, the video decoding unit 222 may decode each layer constituting the OLS to generate a video signal of each layer, and then select and output the video signal of the layer specified by the output layer ID.
[0118] The object-based audio-compatible audio decoding unit 230 performs audio decoding processing compatible with the object-based audio format on the audio encoded data based on the control information. In the example shown in the figure, the object-based audio-compatible audio decoding unit 230 performs audio decoding processing on an audio bitstream including audio metadata, and outputs the audio signal obtained by the audio decoding processing.
[0119] In this embodiment, the object-based audio-compatible audio decoding unit 230 selects decoded audio objects in accordance with an audio program ID or audio tag specified by the control unit 250, taking audio metadata into consideration, and decodes (and synthesizes) the selected audio objects to output an audio signal. Alternatively, the control unit 250 may specify an audio object ID, and the object-based audio-compatible audio decoding unit 230 may decode (and synthesize) the specified audio object.
[0120] The UI unit 240 accepts operations from the user (i.e., the viewer) and presents (outputs) video and audio to the user. The UI unit 240 has an operation unit 241, a display unit 242, and an audio output unit 243. The operation unit 241 accepts an operation from the user to select a service (program, sub-content, resolution, etc.) that the user wishes to view, and outputs an operation signal indicating the operation content to the control unit 250. The operation unit 241 may be configured to include a remote control, or may be configured to include input means such as a touchpad. The display unit 242 displays video based on the video signal output by the multi-layer-compatible video decoding unit 220. The audio output unit 243 outputs audio based on the audio signal output by the object-based audio-compatible audio decoding unit 230.
[0121] The control unit 250 controls the overall operation of the receiving device 200. In this embodiment, the control unit 250 identifies a decoding layer (OLS, output layer) corresponding to a service selected by a user, based on a multi-layer service descriptor included in control information output by the demultiplexing unit 213 of the receiving unit 210 and an operation signal output by the operation unit 241, and outputs an OLS index and an output layer ID indicating the identified layer to the multi-layer-compatible video decoding unit 220. Furthermore, the control unit 250 identifies an audio object (audio program, audio tag) to be decoded (played) in combination with the identified layer (OLS, output layer) based on association information (multi-layer service descriptor and audio configuration descriptor) included in the control information, and outputs an audio program ID or an audio tag indicating the identified audio object to the object-based audio-compatible audio decoding unit 230.
[0122] In addition, the control unit 250 may adaptively select a service (program, sub-content, resolution, etc.) based on at least one of preset information, viewing environment information, and viewing device information from the user, in addition to or instead of the operation signal from the operation unit 241.
[0123] (4) Working Example Next, Examples 1 to 3 will be described.
[0124] (4.1) Example 1 The first embodiment is an embodiment in the case of content hierarchical coding.
[0125] 9(a) is a diagram showing a service image of Example 1. In this example, the base layer is used to provide the main content, the enhancement layer 1 is used to provide the main content and sub-content 1, and the enhancement layer 2 is used to provide the main content and sub-content 2. Each sub-content may be content related to the main content, such as content from another perspective or commentary content, or may be content unrelated to the main content.
[0126] In this embodiment, the main content is a service of regular 4K program video and main audio. The sub-content 1 is a service of sub-video 1 and sub-audio 1 that are added to regular 4K program video. The sub-content 2 is a service of sub-video 2 and sub-audio 2 that are added to regular 4K program video. The user can select a viewing service from these three services.
[0127] 9(b) is a diagram showing the relationship between the OLS and the output layer in the first embodiment. OLS 0 corresponds to a service without sub-content and is composed of a base layer. OLS 1 corresponds to a service with sub-content and is composed of a base layer, enhancement layer 1, and enhancement layer 2. The base layer is output layer 0, enhancement layer 1 is output layer 1, and enhancement layer 2 is output layer 2. On the receiving side, the presence or absence of sub-content is distinguished by the OLS, and switching between sub-contents 1 and 2 is performed by switching the output layer.
[0128] 10 is a diagram showing an example of how audio metadata is described in Example 1. audioProgrammeID (audio program ID) = '1001' corresponds to "main audio" and audioObjectID (audio object ID) = '1010'. audioProgrammeID (audio program ID) = '1002' corresponds to "main audio + sub audio 1" and audioObjectID (audio object ID) = '1010' for the main audio and audioObjectID (audio object ID) = '1020' for the sub audio 1. audioProgrammeID (audio program ID) = '1003' corresponds to "main audio + sub audio 2" and audioObjectID (audio object ID) = '1010' for the main audio and audioObjectID (audio object ID) = '1030' for the sub audio 2.
[0129] FIG. 11 is a diagram illustrating an example of a description of a multi-layer service descriptor according to the first embodiment.
[0130] In this embodiment, the number of layers in multi-layer coding is 3, so num_of_layer (number of layers) = '3'. The layer_id (layer ID) of the base layer is '101', the layer_id (layer ID) of enhancement layer 1 is '102', and the layer_id (layer ID) of enhancement layer 2 is '103'. Also, ols_mode_idc (OLS mode) is '1', and num_of_ols (number of OLSs) is '2' (i.e., a total of two, OLS 0 and OLS 1). In embodiment 1, OLS mode = '1', so output_layer_flag (output layer flag) is not provided.
[0131] The ols_index (OLS index) of OLS 0 is '0', and the profile_tier_level (profile tier level) of OLS 0 is '1'. The layer_uesd_for_decode_flag (layer flag to be decoded) in OLS 0 is provided for each layer, with the base layer (layer_id='101') being '1' (on), and enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') both being '0' (off). The base layer (layer_id='101') with layer_uesd_for_decode_flag='1' (on) has an audio_programm_id (audio program ID), which is '1001' (i.e., main audio). In this way, audio program ID='1001' (main audio) is associated with the base layer (layer_id='101') in OLS 0. Note that text_char in OLS 0 states 'main program'.
[0132] The ols_index (OLS index) of OLS 1 is '1', and the profile_tier_level (profile tier level) of OLS 1 is '2'. The layer_uesd_for_decode_flag (layer flag to be decoded) in OLS 1 is provided for each layer, and the base layer (layer_id='101'), enhancement layer 1 (layer_id='102'), and enhancement layer 2 (layer_id='103') are all set to '1' (on). For the base layer (layer_id='101'), audio_programm_id (audio program ID)='1001' (i.e., main audio) is written. For the enhancement layer 1 (layer_id='102'), audio_programm_id (audio program ID)='1002' (i.e., main audio + sub audio 1) is written. For enhancement layer 2 (layer_id='103'), audio_programm_id (audio program ID) = '1003' (i.e., main audio + sub audio 2) is recorded. Note that the text_char of OLS 1 is recorded as 'main program + sub program'.
[0133] FIG. 12 is a diagram illustrating an example of a description of the audio configuration descriptor in the first embodiment.
[0134] In this embodiment, the number of audio programs in the object-based audio is three (specifically, a total of three: "main audio," "main audio + sub audio 1," and "main audio + sub audio 2"), so num_of_audio_program (number of audio programs) = '3'. For "main audio," audio_programme_id (audio program ID) = '1001', num_of_audio_object (number of audio objects) = '1', and audio object ID (audio_object_id) = '1010' are entered. For "main audio + sub audio 1," audio_programme_id (audio program ID) = '1002', num_of_audio_object (number of audio objects) = '2', and audio object ID (audio_object_id) = '1010' and '1020' are entered. For "main audio + sub audio 2", audio_programme_id (audio program ID) = '1003', num_of_audio_object (number of audio objects) = '2', and audio object IDs (audio_object_id) = '1010' and '1030' are written.
[0135] (4.2) Example 2 The second embodiment is an example of spatial hierarchical coding, in which an audio program ID (audio_program_id) is used as association information.
[0136] 13(a) is a diagram showing a service image of Example 2. In this example, low-resolution main content is provided using the base layer, high-resolution main content is provided using the enhancement layer 1, and high-resolution main content and sub-content are provided using the enhancement layer 2.
[0137] In this embodiment, the low-resolution main content corresponding to the base layer is a service of 2K program video and 2ch audio. The high-resolution main content corresponding to enhancement layer 1 is a service of 4K program video and 5.1ch audio. The high-resolution main content and sub-content corresponding to enhancement layer 2 are a service of 4K program video with sub-video and 5.1ch audio. The user can select a viewing service from among these three services. The sub-content in the enhancement layer is switched on and off by switching the output layer.
[0138] 13(b) is a diagram showing the relationship between OLSs and output layers in Example 2. OLS 0 corresponds to a 2K program video service and is composed of a base layer. OLS 1 corresponds to a 4K program video service and is composed of a base layer, enhancement layer 1, and enhancement layer 2. The base layer is output layer 0, enhancement layer 1 is output layer 1, and enhancement layer 2 is output layer 2.
[0139] 14 is a diagram showing an example of how audio metadata is described in Example 2. audioProgrammeID (audio program ID) = '1001' corresponds to "main audio (2ch)", and the audioObjectID (audio object ID) of the main audio (2ch) = '1010'. audioProgrammeID (audio program ID) = '1002' corresponds to "main audio (5.1ch)", and the audioObjectID (audio object ID) of the main audio (5.1ch) = '1020'. audioProgrammeID (audio program ID) = '1003' corresponds to "main audio (5.1ch) + sub audio", and the audioObjectID (audio object ID) of the main audio (5.1ch) = '1020' and the audioObjectID (audio object ID) of the sub audio = '1030'.
[0140] FIG. 15 is a diagram illustrating an example of a description of a multi-layer service descriptor according to the second embodiment.
[0141] In this embodiment, the number of layers in multi-layer coding is 3, so num_of_layer (number of layers) = '3'. The layer_id (layer ID) of the base layer is '101', the layer_id (layer ID) of enhancement layer 1 is '102', and the layer_id (layer ID) of enhancement layer 2 is '103'. Also, ols_mode_idc (OLS mode) is '2', and num_of_ols (number of OLSs) is '2' (i.e., a total of two, OLS 0 and OLS 1). In this embodiment, since the OLS mode is '2', output_layer_flag (output layer flag) is provided.
[0142] The ols_index (OLS index) of OLS 0 is '0', and the profile_tier_level (profile tier level) of OLS 0 is '0'. The layer_uesd_for_decode_flag (layer flag to be decoded) in OLS 0 is provided for each layer, with the base layer (layer_id='101') being '1' (on) and enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') both being '0' (off). Also, the output_layer_flag (output layer flag) in OLS 0 is provided for each layer, with the base layer (layer_id='101') being '1' (on) and enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') both being '0' (off). The base layer (layer_id='101') with output_layer_flag='1' (on) has audio_programm_id (audio program ID), and this audio program ID is '1001' (i.e., main audio (2ch)). In this way, in OLS 0, the base layer (layer_id='101') is associated with audio program ID='1001' (main audio (2ch)). Note that the text_char of OLS 0 states '2K program'.
[0143] The ols_index (OLS index) of OLS 1 is '1', and the profile_tier_level (profile tier level) of OLS 1 is '1'. The layer_uesd_for_decode_flag (layer flag to be decoded) in OLS 1 is provided for each layer, and the base layer (layer_id='101'), enhancement layer 1 (layer_id='102'), and enhancement layer 2 (layer_id='103') are all '1' (on). In addition, the output_layer_flag (output layer flag) in OLS 1 is provided for each layer, and the base layer (layer_id='101') is '0' (off), and the enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') are all '1' (on). Enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') with output_layer_flag='1' (on) have audio_programm_id (audio program ID). For enhancement layer 1 (layer_id='102'), audio program ID='1002' (i.e., main audio (5.1ch)), and for enhancement layer 2 (layer_id='103'), audio program ID='1003' (i.e., main audio (5.1ch) + sub audio). Thus, in OLS 1, enhancement layer 1 (layer_id='102') is associated with audio program ID='1002' (main audio (5.1ch)), and enhancement layer 2 (layer_id='103') is associated with audio program ID='1003' (main audio (5.1ch) + sub audio). Note that the text_char of OLS 1 states '4K program'.
[0144] FIG. 16 is a diagram illustrating an example of a description of the audio configuration descriptor in the second embodiment.
[0145] In this embodiment, the number of audio programs in object-based audio is three (specifically, a total of three: "main audio (2ch)", "main audio (5.1ch)", and "main audio (5.1ch) + sub audio"), so num_of_audio_program (number of audio programs) = '3'. For "main audio (2ch)", audio_programme_id (audio program ID) = '1001', num_of_audio_object (number of audio objects) = '1', and audio_object_id (audio object ID) = '1010' are entered. For "main audio (5.1ch)", audio_programme_id (audio program ID) = '1002', num_of_audio_object (number of audio objects) = '1', and audio_object_id (audio object ID) = '1020' are entered. For "main audio (5.1ch) + sub audio", audio_programme_id (audio program ID) = '1003', num_of_audio_object (number of audio objects) = '2', and audio_object_id (audio object ID) = '1020' and '1030' are written.
[0146] (4.3) Example 3 The third embodiment is an example of spatial hierarchical coding similar to the second embodiment, in which an audio tag (audio_tag) is used as association information. The relationship between the service image, the OLS, and the output layer is similar to that of the second embodiment (see FIG. 13).
[0147] 17 is a diagram showing an example of how audio metadata is described in Example 3. audio_tag (audio tag) = '1' is associated with audioProgrammeID (audio program ID) = '1001', which corresponds to "main audio (2ch)" and the audioObjectID (audio object ID) of the main audio (2ch) = '1010'. audio_tag (audio tag) = '2' is associated with audioProgrammeID (audio program ID) = '1002', which corresponds to "main audio (5.1ch)" and the audioObjectID (audio object ID) of the main audio (5.1ch) = '1020'. audio_tag (audio tag) = '1' corresponds to audioProgrammeID (audio program ID) = '1003', which corresponds to "main audio (5.1ch) + sub audio", where the main audio (5.1ch) has audioObjectID (audio object ID) = '1020' and the sub audio has audioObjectID (audio object ID) = '1030'.
[0148] 18 is a diagram showing an example of a description of a multi-layer service descriptor in the third embodiment. Here, differences from the second embodiment will be described.
[0149] In OLS 0, the base layer (layer_id='101') with output_layer_flag='1' (on) has an audio_tag (audio tag), and this audio_tag (audio tag) is '1' (i.e., main audio (2ch)). In this way, in OLS 0, the base layer (layer_id='101') is associated with audio_tag (audio tag)='1' (i.e., main audio (2ch)).
[0150] In OLS 1, enhancement layer 1 (layer_id='102') and enhancement layer 2 (layer_id='103') with output_layer_flag='1' (on) have audio_tags. For enhancement layer 1 (layer_id='102'), audio_tag='2' (i.e., main audio (5.1ch)), and for enhancement layer 2 (layer_id='103'), audio_tag='3' (i.e., main audio (5.1ch) + sub audio). Thus, in OLS 1, enhancement layer 1 (layer_id='102') is associated with audio_tag='2' (main audio (5.1ch)), and enhancement layer 2 (layer_id='103') is associated with audio_tag='3' (main audio (5.1ch) + sub audio).
[0151] FIG. 19 is a diagram illustrating an example of a description of the audio configuration descriptor in the third embodiment.
[0152] In this embodiment, the number of audio tags in the object-based audio is three (specifically, a total of three: "main audio (2ch)", "main audio (5.1ch)", and "main audio (5.1ch) + sub audio"), so num_of_audio_tag (number of audio tags) = '3'. For "main audio (2ch)", audio_tag (audio tag) = '1', audio_programme_id (audio program ID) = '1001', num_of_audio_object (number of audio objects) = '1', and audio_object_id (audio object ID) = '1010' are entered. For "main audio (5.1ch)", audio_tag (audio tag) = '2', audio_programme_id (audio program ID) = '1001', num_of_audio_object (number of audio objects) = '1', and audio_object_id (audio object ID) = '1020' are entered. For "main audio (5.1ch) + sub audio", audio_tag (audio tag) = '3', audio_programme_id (audio program ID) = '1001', num_of_audio_object (number of audio objects) = '2', and audio_object_id (audio object ID) = '1020' and '1030' are written.
[0153] (5) Example of control information changes A modification example of the control information according to the above embodiment will be described.
[0154] (5.1) First information In the above-described embodiment, the first information regarding the multiple layers in multi-layer coding (hierarchical coding) is a multi-layer service descriptor (Multilayer_Service_descriptor) arranged in the MPT descriptor area of the MMT package table (MPT).
[0155] In contrast to this, in this modified example, a descriptor similar to the multi-layer service descriptor is a "video preset descriptor (video_preset_descriptor)." Like the multi-layer service descriptor, the video preset descriptor is placed in the MPT descriptor area of the MMT package table (MPT), but the contents of the descriptor are slightly different from those of the multi-layer service descriptor.
[0156] Fig. 20 is a diagram showing the configuration of a video preset descriptor according to this modified example. The video preset descriptor is used to describe videos (referred to as "video presets") that a receiver can select and present during multi-layer coding (hierarchical coding). Note that a video preset represents a selection of videos in multiple layers, and an OLS represents a set of output layers in a video coding layer. If the OLS is different, the same output layer will be treated as a different video preset.
[0157] The video preset descriptor is placed in the MPT descriptor area of the MPT. The meaning of the video preset descriptor is as follows:
[0158] descriptor_tag: The descriptor tag is a 16-bit field that identifies the video preset descriptor.
[0159] descriptor_length (descriptor length): This field is used to write the number of data bytes that follow. The bit length of the descriptor length field varies depending on the descriptor.
[0160] ISO_639_language_code (language code): This 24-bit field represents the language of the following character information field using a 3-letter alphabetic code specified in ISO 639-2. Each character is coded in 8 bits according to ISO 8859-1 and inserted into the 24-bit field in that order.
[0161] number_of_video_assets (number of assets): This 7-bit field indicates the number of video assets included in the service. Note that in the transmission of a hierarchically coded bitstream with a layer structure using the VVC method, the hierarchically coded bitstream is composed of multiple assets (multiple video assets), and one layer corresponds to one asset (video asset).
[0162] component_tag (component tag): A 16-bit field. The component tag is a label for identifying a component stream. In the case of a multi-layer service, a bit stream corresponding to one layer corresponds to one component stream.
[0163] video_asset_id (asset identification number): This 7-bit field is the identification number of the video asset, and has a one-to-one correspondence with the component tag.
[0164] video_preset_id (video preset number): This 6-bit field is the video preset number. It is used to identify the video that can be presented in the service.
[0165] number_of_dependent_video_assets (number of dependent assets): This 3-bit field indicates the number of video assets required to decode the corresponding video asset.
[0166] dependent_video_asset_id (dependent asset identification number): This 7-bit field indicates the identification number of the video asset required to decode the corresponding video asset.
[0167] video_preset_label_length (video preset label length): This 8-bit field indicates the byte length of the string that describes the video preset.
[0168] video_preset_label_byte (video preset label byte): An 8-bit field. A series of label byte areas describes the label that describes the video preset.
[0169] number_of_olss (number of OLSs): This 8-bit field indicates the number of Output Layer Sets (OLSs) included in the service.
[0170] OLSidx (OLS number): This 8-bit field indicates the OLS number.
[0171] number_of_output_video_presets (number of output video presets): This 3-bit field indicates the number of video presets that can be selected and presented and are included in the corresponding OLS.
[0172] output_video_preset_id (output video preset number): This 6-bit field indicates the number of a video preset that is included in the corresponding OLS and can be selected and presented.
[0173] OLS_label_length (label length): This 8-bit field indicates the byte length of the string that describes the OLS.
[0174] OLS_label_byte (label byte): An 8-bit field. A series of label byte areas describes the label that describes the OLS.
[0175] reserved_future_use (reserved for future use): Reserved area for future expansion.
[0176] (5.2)Second information In the above embodiment, the second information about the multiple audio objects in the object-based audio is an audio configuration descriptor (Audio_config_descriptor) located in the asset descriptor area of the audio asset in the MMT package table (MPT).
[0177] In contrast to this, in this modified example, a descriptor similar to the audio configuration descriptor is an “audio preset descriptor (audio_preset_descriptor).” The audio preset descriptor is placed in the MPT descriptor area of the MMT package table (MPT), not in the asset descriptor area of the MMT package table (MPT).
[0178] 21 and 22 are diagrams showing the configuration of an audio preset descriptor according to this modified example. The audio preset descriptor is used in object-based audio to describe audio (referred to as "audio presets") that can be selected and presented by a receiver. The audio preset descriptor is placed in the MPT descriptor area of the MPT. The meaning of the audio preset descriptor is as follows:
[0179] As shown in FIG. 21, descriptor_tag: The descriptor tag is a 16-bit field that identifies the audio preset descriptor.
[0180] descriptor_length (descriptor length): This field is used to write the number of data bytes that follow. The bit length of the descriptor length field varies depending on the descriptor.
[0181] ISO_639_language_code (language code): This 24-bit field indicates the language code of the string used in the audio preset descriptor as a three-letter alphabetic code conforming to ISO 639-2.
[0182] number_of_audio_preset (number of audio presets): This 6-bit field indicates the number of all audio presets that can be presented and are included in the audio components that make up the package.
[0183] audio_preset_id (audio preset number): This 6-bit field indicates the number of the audio preset included in the audio component that makes up the package. The audio preset number must be unique within the package.
[0184] audio_preset_type (audio preset type): This 2-bit field indicates the relationship between the audio preset and the reference audio preset. The audio preset type is coded according to Figure 23(a).
[0185] reference_audio_preset_id (reference audio preset number): When the audio preset type is 2 or 3, this 6-bit field indicates the number of the audio preset (reference audio preset) in the higher hierarchy.
[0186] main_component_tag (main audio component tag): This 16-bit field is a label for identifying the component stream of the audio component (main audio component) that includes the main audio object that makes up the audio preset.
[0187] stream_content (component content): This 4-bit field indicates the type of audio stream that includes the audio preset.
[0188] number_of_sub_component (number of sub audio components): This 4-bit field indicates the number of sub audio components other than the main audio component when the audio object that makes up the audio preset is made up of multiple audio components.
[0189] sub_component_tag (sub audio component tag): This 16-bit field indicates the component tag of the sub audio component.
[0190] 3da_group_preset_id (3DA group preset number): This 4-bit field indicates the number of the MPEG-H 3DA group preset when the component content of the audio component is MPEG-H 3D audio (3DA).
[0191] number_of_3da_switch_group (number of 3DA switch groups): This 4-bit field indicates the number of MPEG-H 3DA switch groups when the component content of the audio component is MPEG-H 3DA.
[0192] 3da_swg_id (3DA switching group number): This 4-bit field indicates the number of the MPEG-H 3DA switching group when the component content of the audio component is MPEG-H 3DA.
[0193] 3da_swg_object_id (3DA switching object number): This 6-bit field indicates the number of the audio object in the switching group specified by the 3DA switching group number when the component content of the audio component is MPEG-H 3DA.
[0194] ac4_presentation_id (AC-4 presentation number): This 6-bit field indicates the AC-4 presentation number if the component content of the audio component is AC-4.
[0195] As shown in FIG. 22, number_of_audio_preset_switch_group (number of audio preset switching groups): This 6-bit field indicates the number of audio preset switching groups.
[0196] audio_preset_switch_group_id (audio preset switch group number): This 6-bit field indicates the number of the audio preset switch group.
[0197] number_of_flagged_video_preset (number of flagged video presets): This 6-bit field indicates the number of video presets that have a flag (association_type) indicating an association with the audio preset.
[0198] video_preset_id (video preset number): This 6-bit field indicates the number of the video preset that indicates the association with the audio preset.
[0199] Here, video_preset_id (video preset number) is associated with audio_preset_id (audio preset number), and corresponds to association information indicating the association between a layer in multi-layer coding (hierarchical coding) and an audio object in object-based audio. Note that association_type may also correspond to association information.
[0200] association_type (relationship type): This 3-bit field indicates the relationship between the audio preset and the video preset. The relationship type is coded according to Figure 23(b).
[0201] audio_ISO_639_language_code (audio language code): This 24-bit field indicates the language of the spoken voice of the audio preset in a three-letter alphabetic code specified in ISO 639-2.
[0202] dialogue_control_type (dialog control type): This 3-bit field indicates the type of control method for the speech of the corresponding voice preset. The dialogue control type is coded according to Figure 24(a).
[0203] added_object_type (added object type): This 3-bit field indicates the type of added audio object when the audio preset is composed of a combination of an audio object of the reference audio preset and another audio object. The added object type is coded according to Figure 24(b).
[0204] audio_accessibility_type (audio accessibility type): This 2-bit field indicates the type of audio preset if it is content intended to improve accessibility. The audio accessibility type is coded according to Figure 24(c).
[0205] audio_system_type (audio system type): This 2-bit field indicates the audio format of the audio preset. The audio system type is coded according to Figure 25(a).
[0206] presentation_control_type (listening restriction type): This 3-bit field indicates the type of listening restriction for the audio preset, and is coded according to FIG. 24(b).
[0207] preset_priority (preset priority): This 2-bit field indicates the priority of the audio preset to be presented to the user. If an audio preset with a higher priority than the audio preset can be presented to the user as a choice, the audio preset will not be presented as a choice. 3 is the highest priority, and 0 is the lowest priority, meaning that the audio preset will not be presented to the user.
[0208] dialogue_user_control_flag (dialogue user control flag): This 1-bit field indicates that user dialogue control is prohibited for all voice objects in the voice preset.
[0209] label_flag (label flag): If this 1-bit field is 1, the label indicated by the label_byte of the audio preset is used; if it is 0, the label of the audio preset is not used.
[0210] keyword_flag (keyword flag): If this 1-bit field is 1, the keyword of the voice preset is used; if it is 0, the keyword of the voice preset is not used.
[0211] label_length (label length): This 8-bit field indicates the byte length of label_byte.
[0212] label_byte: A label that describes the content of the audio preset to be presented to the user. (Example: Main Audio (Japanese), Main Audio (with audio description), etc.)
[0213] keyword_length (keyword length): This 8-bit field indicates the byte length of keyword_byte.
[0214] keyword_byte (keyword byte): Enter the keywords that describe the content of the voice preset. Add ## to the beginning of each keyword. (Example: ##Team name##Player name)
[0215] number_of_reserved_byte (number of reserved bytes): Indicates the number of bytes reserved for future expansion.
[0216] reserved_future_use (reserved for future use): Reserved area for future expansion.
[0217] In this way, in this modification, the audio preset descriptor includes a video preset number (video_preset_id) for each audio preset (audio_preset_id). As a result, it is possible to associate layers in multi-layer coding (hierarchical coding) with audio objects in object-based audio.
[0218] (6) Other embodiments In the above embodiment, an example has been described in which MMT is used as the multiplexing method and the first information, second information, and association information are included in MMT-SI (MPT). However, when CMAF (Common Media Application Format) is used as the media format, the first information, second information, and association information may be included in metadata of ISOBMFF (ISO Base Media File Format), for example.
[0219] A program may be provided that causes a computer to execute each process performed by the transmitting device 100 or the receiving device 200. The program may be recorded on a computer-readable medium. The program can be installed on a computer using the computer-readable medium. Here, the computer-readable medium on which the program is recorded may be a non-transitory storage medium. The non-transitory storage medium is not particularly limited, and may be, for example, a storage medium such as a CD-ROM or a DVD-ROM. Furthermore, circuits that execute each process performed by the transmitting device 100 or the receiving device 200 may be integrated, and at least a part of the transmitting device 100 or the receiving device 200 may be configured as a semiconductor integrated circuit (chip set, SoC).
[0220] As used in this disclosure, the terms "based on" and "depending on / in response to" do not mean "based only on" or "depending only on," unless expressly stated otherwise. The term "based on" means both "based only on" and "based at least in part on." Similarly, the term "depending on" means both "depending only on" and "depending at least in part on." The terms "include," "comprise," and variations thereof do not mean including only the listed items, but may mean including only the listed items or may include additional items in addition to the listed items. Additionally, the term "or," as used in this disclosure, is not intended to mean an exclusive or. Furthermore, any reference to elements using designations such as "first," "second," etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way. In this disclosure, where articles are added by translation, such as a, an, and the in English, these articles shall include the plural unless the context clearly indicates otherwise.
[0221] Although the embodiments have been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design modifications are possible within the scope of the gist.
[0222] (7) Supplementary Notes Additional notes will be given regarding the features of the above-described embodiment.
[0223] Appendix 1 A transmitting device (100) for use in a video and audio transmission system, a video encoding means (110) for performing video encoding processing corresponding to a multi-layer encoding method to output video encoded data of a plurality of layers; an audio encoding means (120) for outputting audio encoded data of a plurality of audio objects by performing audio encoding processing corresponding to the object-based audio system; a multiplexing means (130) for multiplexing the video encoded data, the audio encoded data, and control information including first information on the plurality of layers and second information on the plurality of audio objects, and outputting the multiplexed data. Transmitting device.
[0224] Appendix 2 The control information is provided in the MMT package table in the MMT-SI. 2. The transmitting device according to claim 1.
[0225] Appendix 3 the first information is a multi-layer service descriptor indicating a configuration of the plurality of layers, The second information is an audio configuration descriptor that indicates the configuration of the plurality of audio objects. 3. A transmitting device according to claim 1 or 2.
[0226] Appendix 4 The control information includes association information indicating association between a layer among the plurality of layers and an audio object among the plurality of audio objects. 4. A transmitting device according to any one of Supplementary notes 1 to 3.
[0227] Appendix 5 A receiving device (200) used in a video and audio transmission system, a receiving means (210) for receiving multiplexed data and obtaining from the received data video coded data of a plurality of layers, audio coded data of a plurality of audio objects, and control information including first information relating to the plurality of layers and second information relating to the plurality of audio objects; a video decoding means (220) for performing a video decoding process corresponding to a multi-layer coding method on the video coded data based on the control information; and an audio decoding means (230) for performing an audio decoding process corresponding to the object-based audio system on the audio encoded data based on the control information. Receiving device.
[0228] Appendix 6 The control information is provided in the MMT package table in the MMT-SI. 6. A receiving device according to claim 5.
[0229] Appendix 7 the first information is a multi-layer service descriptor indicating a configuration of the plurality of layers, The second information is an audio configuration descriptor that indicates the configuration of the plurality of audio objects. 7. A receiving device according to claim 5 or 6.
[0230] Appendix 8 The control information includes association information indicating association between a layer among the plurality of layers and an audio object among the plurality of audio objects. 8. A receiving device according to any one of Supplementary Notes 5 to 7.
[0231] Appendix 9 A computer is caused to function as a transmitting device according to any one of Supplementary Notes 1 to 4 or a receiving device according to any one of Supplementary Notes 5 to 8. program. [Explanation of symbols]
[0232] 10: Transmission line 10a: Broadcast transmission path 10b: Communication transmission line 100: Transmitting device 110: Multi-layer compatible video encoding unit 120: Object-based audio compatible audio coding unit 130: Multiplexing section 131: Control information generation unit 140: Transmitter 141: Broadcast transmission unit 142: Communications Department 200: Receiving device 210: Receiving unit 211: Broadcast receiving unit 212: Communications Department 213: Demultiplexer 220: Multi-layer compatible video decoding unit 221: Bitstream synthesis unit 222: Video decoding unit 230: Object-based audio compatible audio decoding unit 240:UI section 241:Operation unit 242:Display section 243: Audio output section 250: Control unit
Claims
1. A transmitting device used in a video and audio transmission system, a video encoding means for performing a video encoding process corresponding to a multi-layer encoding method to output video encoded data of a plurality of layers; an audio encoding means for performing an audio encoding process corresponding to the object-based audio system to output audio encoded data of a plurality of audio objects; a multiplexing means for multiplexing the video encoded data, the audio encoded data, and control information including first information on the plurality of layers and second information on the plurality of audio objects, and outputting the multiplexed data. Transmitting device.
2. The control information is provided in the MMT package table in the MMT-SI. The transmitting device according to claim 1 .
3. the first information is a multi-layer service descriptor indicating a configuration of the plurality of layers, The second information is an audio configuration descriptor that indicates the configuration of the plurality of audio objects.
3. The transmitting device according to claim 1 or 2.
4. The control information includes association information indicating association between a layer among the plurality of layers and an audio object among the plurality of audio objects.
3. The transmitting device according to claim 1 or 2.
5. A receiving device used in a video and audio transmission system, a receiving means for receiving multiplexed data and obtaining from the received data video encoded data of a plurality of layers, audio encoded data of a plurality of audio objects, and control information including first information on the plurality of layers and second information on the plurality of audio objects; a video decoding means for performing a video decoding process corresponding to a multi-layer coding method on the video coded data based on the control information; and an audio decoding means for performing an audio decoding process corresponding to the object-based audio system on the audio encoded data based on the control information. Receiving device.
6. The control information is provided in the MMT package table in the MMT-SI.
6. The receiving device according to claim 5.
7. the first information is a multi-layer service descriptor indicating a configuration of the plurality of layers, The second information is an audio configuration descriptor that indicates the configuration of the plurality of audio objects.
7. The receiving device according to claim 5 or 6.
8. The control information includes association information indicating association between a layer among the plurality of layers and an audio object among the plurality of audio objects.
7. The receiving device according to claim 5 or 6.
9. A computer is caused to function as the transmitting device according to claim 1 or the receiving device according to claim 5. program.
Citation Information
Patent Citations
IEC23008-3、
IEC23008-1、