Receiver
The broadcasting system addresses the issue of receivers not processing object-based audio by multiplexing identification information, enabling appropriate audio reproduction based on receiver capabilities through a separation and decoding process.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SHARP KK
- Filing Date
- 2022-07-27
- Publication Date
- 2026-07-24
Smart Images

Figure 0007894758000001 
Figure 0007894758000002 
Figure 0007894758000003
Abstract
Description
Technical Field
[0001] The present invention relates to Receiver .
Background Art
[0002] In broadcasting, the use of object-based audio signals such as AC-4 (ETSI TS 103 190) has been considered. Patent Document 1 describes that an audio signal of an audio object (object-based audio signal) is used as a priority signal and is preferentially reproduced.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In Patent Document 1, an audio encoder encodes an object-based audio signal, and an audio decoder performs decoding processing on the bitstream. However, there are also receivers for receiving broadcasts that do not have the ability to process object-based audio signals. It is desired that appropriate audio reproduction can be performed according to the capabilities of the receiver.
[0005] The present invention has been made in view of such circumstances, and provides a broadcast system, a receiver, a reception method, and a program capable of performing appropriate audio reproduction according to the capabilities of the receiver.
Means for Solving the Problems
[0006] This invention was made to solve the above-mentioned problems, and one aspect of the present invention is a broadcasting system that broadcasts an audio component, comprising: a broadcasting device that broadcasts a broadcast wave in which identification information indicating whether or not the audio component is an AC-4 audio component is multiplexed on a multiplexing layer; and a receiver that receives the broadcast wave, wherein the receiver comprises: a separation unit that acquires the identification information from the broadcast wave on the multiplexing layer; a control unit that selects an audio component according to the capabilities of the receiver based on the identification information; and a decoding unit that decodes the audio data of the selected audio component.
[0007] Another aspect of the present invention is the broadcasting system described above, wherein the identification information is placed in an audio component descriptor, which is an MMT (MPEG Media Transport) descriptor that describes parameters relating to an audio signal among program elements.
[0008] Another aspect of the present invention is the broadcasting system described above, wherein the broadcast includes a plurality of audio components, each including an AC-4 audio audio component and an MPEG-4 audio audio component, the control unit selects one audio component from the plurality of audio components based on the identification information placed in the audio component descriptor, the separation unit separates the selected audio component, and the decoding unit decodes the audio data of the selected audio component.
[0009] Another aspect of the present invention is the broadcasting system described above, wherein the broadcasting device broadcasts a broadcast wave in which information indicating sound materials included in the AC-4 audio component is multiplexed on a multiplexing layer when the audio component is an AC-4 audio component, the separation unit acquires information indicating the sound materials from the broadcast wave using the multiplexing layer, the control unit selects a combination of sound materials based on the information indicating the sound materials, and the decoding unit decodes the audio data of the selected sound materials.
[0010] Another aspect of the present invention is a receiver comprising: a separation unit that acquires identification information from a broadcast wave, in a multiplexing layer, indicating whether or not an audio component is an AC-4 audio component; a control unit that selects an audio component according to the capabilities of the device based on the identification information; and a decoding unit that decodes the audio data of the selected audio component.
[0011] Another aspect of the present invention is a receiving method comprising the steps of: acquiring identification information from a broadcast wave at a multiplexing layer indicating whether or not an audio component is an AC-4 audio component; selecting an audio component according to the capabilities of the device based on the identification information; and decoding the audio data of the selected audio component.
[0012] Another aspect of the present invention is a program for causing a receiver computer, which includes a separation unit that acquires identification information from a broadcast wave at a multiplexing layer indicating whether or not an audio component is an AC-4 audio component, and a decoding unit that decodes the audio data of the selected audio component, to function as a control unit that selects an audio component according to the capabilities of the device based on the identification information. [Effects of the Invention]
[0013] According to this invention, appropriate audio reproduction can be performed according to the capabilities of the receiver. [Brief explanation of the drawing]
[0014] [Figure 1] This figure shows an example of the configuration of a broadcasting system Sys according to an embodiment of the present invention. [Figure 2] This figure shows an example of a broadcasting system Sys according to the same embodiment. [Figure 3] This figure shows a comparative example of the broadcasting system Sys according to the same embodiment. [Figure 4]It is a diagram showing another example of the broadcast system Sys according to the same embodiment. [Figure 5] It is an explanatory diagram explaining the outline of the broadcast system Sys according to the same embodiment. [Figure 6] It is a diagram showing an example of the structure of the protocol stack according to the same embodiment. [Figure 7] It is a diagram showing the data structure of MPT according to the same embodiment. [Figure 8] It is a schematic diagram showing the hardware configuration of the receiver 2 according to the same embodiment. [Figure 9] It is a schematic diagram showing an example of the signal processing flow in the receiver according to the same embodiment. [Figure 10] It is a diagram showing an example of the audio switching menu according to the same embodiment. [Figure 11] It is a schematic diagram showing an example of the structure of the MH-audio component descriptor according to the same embodiment. [Figure 12] It is a table showing an example of the value of nga_type according to the same embodiment. [Figure 13] It is a table showing an example of the value of nga_level according to the same embodiment. [Figure 14] It is a schematic diagram showing another example of the structure of the MH-audio component descriptor according to the same embodiment. [Figure 15] It is a table showing an example of the value of stream_content according to the same embodiment. [Figure 16] It is a table showing another example of the value of nga_level according to the same embodiment. [Figure 17] It is a schematic diagram showing an example of the structure of the MH-AC-4 audio descriptor according to the same embodiment. [Figure 18] It is a schematic diagram showing an example of the structure of presentation() according to the same embodiment. [Figure 19] It is a schematic diagram showing another example of the structure of presentation() according to the same embodiment. [Figure 20] It is a flowchart showing a detailed example of switching according to the same embodiment. [Modes for carrying out the invention]
[0015] Embodiments of the present invention will be described in detail below with reference to the drawings.
[0016] [System Configuration] Figure 1 shows an example of the configuration of a broadcasting system Sys according to an embodiment of the present invention. The broadcasting system Sys comprises a broadcasting station's broadcasting equipment 1 ("broadcasting station 1"), a relay station Sa, a receiver 2, a broadcasting station server 3, and a carrier server 4. The broadcast is, for example, terrestrial digital broadcasting, but may also be, for example, advanced BS (Broadcasting Satellites) digital broadcasting or advanced broadband CS (Communication Satellites) digital broadcasting. Furthermore, the present invention is not limited to these broadcasts, and the broadcast may be a broadcast that does not use a relay station Sa. The broadcast may also be a wired broadcast such as cable television. The relay station Sa may be, for example, a digital relay station, but may also be a broadcasting satellite.
[0017] In the broadcasting system Sys, broadcasting station 1 transmits digital broadcast signals, application control information, and display-related control information via broadcast waves. Service providers provide program-related metadata and video content from their service provider server 4. Application control information informs compatible receivers of applications linked to the program, and sends commands and control information for starting and stopping them. The control information related to presentation includes information regarding the overlaying of the application and the broadcast program on the same TV screen, as well as control information regarding whether or not the application can be presented. The broadcasting station operates Broadcasting Station Server 3 within the Broadcasting System Sys. Broadcasting Station Server 3 provides metadata such as program title, program ID, program summary, cast, and broadcast date and time. The information that the broadcasting station provides to service providers is provided through the API (Application Programming Interface) provided by Broadcasting Station Server 3.
[0018] The service provider is the provider of services through the broadcasting system Sys, and is responsible for the production and distribution of content and applications for providing services, as well as the operation of the broadcasting station server 3 to realize individual services. Here, the services include broadcast-communication integration services that link broadcasting and communication. Broadcast server 3 sends applications to receiver 2 for "application management and distribution." Broadcast server 3, acting as a "service-specific server," provides server functions to implement individual services (MPEG-H service, VOD program recommendation service, multilingual subtitle service, etc.).
[0019] MPEG-H is a digital container standard, video compression standard, audio compression standard, and two conformance testing standards, as defined by ISO / IEC Moving Picture Experts. MPEG-H is a set of standards developed under the MPEG Group. MPEG-H services can include AC-4 audio components. AC-4 audio, for example, enables object-based audio. In object-based audio, an "object" refers to each individual sound element that makes up a program, such as music or human voices. In object-based audio, an audio signal is recorded for each sound element, allowing for individual sound control of each element. Furthermore, when playing back on receiver 2, it is possible to play the program based on the playback position information of the elements, adjusting it to the actual speaker positions.
[0020] Broadcasting station server 3 not only implements the functional aspects of these services but also transmits the content that makes up the services (AC-4 audio data, VOD content, subtitle data, etc.). Broadcasting station server 3 acts as a "repository" to register applications for distribution of the broadcasting system Sys, and provides and searches for a list of available applications in response to inquiries from receiver 2.
[0021] Receiver 2 includes functions for realizing broadcast-communication collaboration services, in addition to the function for receiving existing digital broadcasts. In addition to broadband network connectivity, receiver 2 has the following functions: • A function to execute applications in response to application control signals from broadcasts. • A function to provide information through coordination between broadcasts and communications. • Device linking function Here, the term "terminal" includes, for example, user devices such as smartphones and smart speakers. The terminal connectivity function of receiver 2 accesses broadcast resources such as program information and invokes receiver functions such as playback control in response to requests from other terminals. Another example of an application is the AC-4 audio digital mixer. The user (also called the "receiver") can use the digital mixer received from the service provider server 4 to adjust the volume or effects of each audio signal, as well as the balance between audio signals. These adjustments can also be made for each speaker.
[0022] More specifically, receiver 2 has the following functions: Receiver 2 has a "broadcast reception and playback" function, which allows it to receive broadcast signals, select a specific broadcast service, and synchronously play back the video, audio, subtitles, and data broadcasts that make up the service. Receiver 2 has a "communication content reception and playback" function, which allows it to access video content stored on a server (e.g., carrier server 4) on the communication network, receive it as VOD streaming, and synchronize playback of the video, audio, and subtitles that make up the content. Receiver 2, as an "application control" function, interacts with the application engine, primarily with respect to managed applications, based on application control information acquired from servers or broadcast signals on the communication network, and has the function of controlling and managing the lifecycle and events of each application. Receiver 2 has the function of acquiring and executing applications as an "application engine." This function is implemented, for example, in an HTML5 browser. Receiver 2 has a "presentation synchronization control" function that controls the presentation synchronization of video and audio streams received from broadcast and video and audio streams received via streaming. Receiver 2 has an "application launcher" function, which is a navigation function that allows the user to select and launch mainly non-broadcast managed applications.
[0023] Figure 2 shows an example of the broadcasting system Sys according to this embodiment. Receiver 2 in Figure 2 is a receiver that supports AC-4 audio.
[0024] Broadcasting station 1 multiplexes video and audio signals and transmits the multiplexed signals. The multiplexing method used is MMT (MPEG Media Transport) TLV (Type Length Value). Broadcasting station 1 generates and transmits an audio signal A1 (also referred to as the "advanced audio signal") which includes both MPEG-4 audio (channel-based audio: e.g., audio A12 to audio A14) and AC-4 audio (object-based audio: audio A11) signals. In this way, broadcasting station 1 transmits the high-resolution audio signal and the MPEG-4 audio signal in parallel.
[0025] More specifically, broadcast station 1 uses AC-4 audio as an audio component (asset) and generates an advanced audio signal A1 by multiplexing this audio component with each audio component of MPEG-4 audio. This multiplexing is performed after the audio data sequence of each audio component has been encoded. Broadcast station 1 transmits a broadcast wave to which this advanced audio signal A1 has been multiplexed. Broadcasting station 1 describes in a table describing asset information whether or not an asset (audio component) is a component of AC-4 audio (object-based audio: e.g., audio A11) using a descriptor. The descriptor indicating whether or not it is an AC-4 audio component may also be a descriptor indicating whether or not an advanced audio signal A1 exists, or it may indicate whether or not an AC-4 audio signal is included. In MMT, components such as video and audio are defined as assets. An example of AC-4 audio A11 is that it has up to 11.1 channels of audio, with dialogue in Japanese or English, and commentary in Japanese or English.
[0026] Receiver 2 comprises a tuner 211, a demux (demultiplexer) 22, a selector 231, an audio decoder (decoding unit) 232, a mixer 233, and a video decoder 241. The detailed configuration of receiver 2 will be described later.
[0027] The tuner 211 receives broadcast waves via the antenna and tunes (selects) to the channel selected based on user input. The tuned signal is demodulated and input as data to the Demux22. Demux22 separates the input data into video data streams, audio data streams, text superimposition data streams, subtitle data streams, etc. The separated audio data streams are output to selector 231. The separated video data streams are output to video decoder 241.
[0028] Here, Demux22 separates the audio data sequence into AC-4 audio A11 and the audio data sequences of each MPEG-4 audio component, A12, A13, and A14. More specifically, Demux22 determines whether an AC-4 audio signal exists using a descriptor in the table describing the asset information. If Demux22 determines that an AC-4 audio signal exists, and receiver 2 has the capability to decode AC-4 audio, it separates audio A11, A12, A13, and A14 from the advanced audio signal data. If Demux22 determines that an AC-4 audio signal does not exist, it separates only audio A12, A13, and A14. Alternatively, Demux22 may separate only the audio A11, A12, A13, and A14 that are to be decoded.
[0029] The audio data streams of each audio component output from Demux22 are input to selector 231. Selector 231 selects the audio data streams of the audio components according to user operation or the capabilities of receiver 2. The capabilities of receiver 2 include, for example, the number of channels that can be decoded simultaneously, or the types and capabilities of speakers that can be played back. Selector 231 outputs the selected audio data streams to audio decoder 232. The audio decoder 232 decodes the audio data sequence of the audio component input from the selector 231. If the audio data sequence decoded by the audio decoder 232 is an AC-4 audio data sequence, the mixer 233 synthesizes the audio from each sound source and performs downmixing. The downmixed audio data sequence is converted back into sound and output from the speaker. If the audio data sequence decoded by the audio decoder 232 is an MPEG-4 audio data sequence, the audio data sequence is converted back into sound and output from the speaker. In other words, for MPEG-4 audio data sequences, neither audio synthesis of each sound source nor downmixing is performed.
[0030] The video data stream output from Demux22 is input to the video decoder 241. The video decoder 241 decodes the input video data sequence. The decoded video data sequence undergoes color space conversion processing as needed and is used to display the video on the display. The character superdata sequence and subtitle data sequence separated into Demux 22 are decoded by the character superdecoder and subtitle decoder (not shown), respectively, and the decoded strings are superimposed on the video.
[0031] As described above, in this embodiment, the receiver 2 acquires information at the multiplexing layer indicating that a broadcast containing AC-4 audio (speech) is an AC-4 audio component. The receiver 2 selects the audio component according to its own capabilities, and therefore can perform appropriate audio playback according to the capabilities of the receiver 2.
[0032] Figure 3 shows a comparative example of the broadcasting system Sys according to this embodiment. This diagram shows an example where broadcaster C1 transmits only AC-4 audio. In this example, the AC-4 audio is not multiplexed with the MPEG-4 audio. In this example, AC-4 audio is used as the sole audio component. Therefore, the asset information table does not include a descriptor indicating whether or not it is an AC-4 audio component. In this case, DemuxC22 can acquire AC-4 audio as an audio component (audio configuration), but it cannot determine whether or not it is processable (level, etc.). The audio data stream output from DemuxC22 is decoded by the audio decoder C232 and output to the mixer C233.
[0033] In contrast to the comparative example in Figure 3, the broadcasting station 1 according to this embodiment transmits AC-4 audio signals and MPEG-4 audio signals in parallel. The receiver 2 first determines from the descriptor whether or not it is an AC-4 audio signal, and selects either AC-4 or MPEG-4 audio according to its audio decoding capability. As a result, receiver 2 can perform appropriate audio playback, either AC-4 or MPEG-4, depending on its own capabilities.
[0034] Figure 4 shows another example of the broadcasting system Sys according to this embodiment. In this figure, receiver 2a is a receiver that does not support AC-4. The receiver 2a in this figure comprises a tuner 211, a demux 22a, a selector 231a, an audio decoder 232a, and a video decoder 241. In this figure, the same reference numerals are used for the same functional parts as in the receiver 2 of Figure 3, and their descriptions are omitted.
[0035] Demux22a separates the input data into video data streams, audio data streams, text superimposition data streams, subtitle data streams, etc. The separated audio data stream is output to selector 231a. The separated video data stream is output to video decoder 241. Here, Demux22a separates the audio data sequence into the individual MPEG-4 audio tracks A12, A13, and A14, and the audio data sequences of each audio component. More specifically, in the table describing the asset information, Demux22a uses a descriptor to determine whether each audio component is an AC-4 audio component. Demux22a determines that audio track A11, which is an AC-4 audio component, is not playable. Demux22 separates audio A12, A13, and A14 from the data of the advanced audio signal A1. Note that such signal selection (selection of MPEG-4 audio signal, or selection of simulcast of only MPEG-4 audio signal) may be performed by selector 231.
[0036] Figure 5 is an explanatory diagram illustrating the outline of the broadcasting system Sys according to this embodiment. In the broadcasting system Sys, broadcasting station 1 is composed of an AC-4 encoder 11, MPEG-4 encoders 111-113, and a Mux (multiplexer) 12. Broadcasting station 1 also has other functional units necessary for broadcasting. While this diagram shows an example with three MPEG-4 encoders, broadcasting station 1 may have two or fewer MPEG-4 encoders, or four or more.
[0037] Receiver 2 comprises Demux 22, selector 231, AC-4 decoder 232-1, MPEG-4 decoder 232-2, AC-4 renderer 233-1, and mixer 233-2. In this figure, the same functional parts as those in Receiver 2 in Figure 2 are denoted by the same reference numerals. Note that AC-4 decoder 232-1 and MPEG-4 decoder 232-2 correspond to audio decoder 232 in Figure 2. AC-4 renderer 233-1 and mixer 233-2 correspond to mixer 233 in Figure 2.
[0038] At broadcasting station 1, the following audio sources for AC-4 audio are input to AC-4 encoder 11: background sound (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), commentary (Japanese), and commentary (English). Additionally, the following audio sources for MPEG-4 audio are input to MPEG-4 encoders 111, 112, and 113, respectively: 7.1ch audio including Japanese dialogue, stereo audio including Japanese dialogue, and stereo audio including English dialogue.
[0039] The AC-4 encoder 11 outputs an AC-4 audio stream St1 by encoding the input audio. This stream is also called an AC-4 stream and is a single elementary stream containing multiple audio objects (background sound, dialogue (Japanese), dialogue (English), commentary (Japanese), and commentary (English)). The MPEG-4 encoders 111, 112, and 113 encode the input audio and output MPEG-4 audio streams St2, St3, and St4, respectively.
[0040] Mux12 receives the video stream, SI (Signaling Information), MPEG-H audio stream St1, and MPEG-4 audio streams St2, St3, and St4 as input. Mux12 multiplexes this data. The multiplexed data is then modulated, and the modulated signal is broadcast as a broadcast wave.
[0041] The broadcast wave received by receiver 2 is demodulated, and the demodulated data is input to Demux22. Demux22 separates the input data into a video stream, SI, AC-4 audio stream St1, and MPEG-4 audio streams St2, St3, and St4. The AC-4 audio stream St1 and MPEG-4 audio streams St2, St3, and St4 are input to selector 231. From the SI's MPT (MMT Package Table), descriptors are extracted indicating whether or not the audio components are AC-4 audio components.
[0042] The selector 231 determines whether an AC-4 audio signal exists based on the extracted descriptor. If an AC-4 audio signal exists, the selector 231 outputs the AC-4 audio stream St1 to the AC-4 decoder 232-1. The selector 231 outputs the MPEG-4 audio stream St2, St3, or St4 to the MPEG-4 decoder 232-2 based on MPT.
[0043] The AC-4 decoder 232-1 decodes the AC-4 audio stream St1 to extract data for background sound (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), commentary (Japanese), and commentary (English). The AC-4 renderer 233-1 is an audio renderer for AC-4 audio. It renders the audio data extracted by the AC-4 decoder 232-1 (including down-converting or up-converting) and outputs it to the mixer 233. The MPEG-4 decoder 232-2 decodes the MPEG-4 audio streams St2, St3 or St4 to extract 7.1ch audio including Japanese subtitles, stereo audio including Japanese subtitles, and stereo data including English subtitles, and outputs them to the mixer 233. The mixer 233-2 synthesizes the audio of the input data, and the synthesized audio is output from each speaker or headphones, etc.
[0044] [Regarding the control information of the broadcast wave] The broadcast wave according to this embodiment will be described. In the broadcast wave, the control information is superimposed and transmitted by each broadcaster on its broadcast signal, which is the TLV stream. The control information includes TLV-SI (TLV-Signaling Information) related to the TLV multiplexing method and MMT-SI (MMT-Signaling Information) related to MMT, which is the media transport method. Hereinafter, a "component" (of video or audio) will also be referred to as an "asset".
[0045] [The protocol stack structure of the system using MMT] In the system using MMT, an example of the protocol stack structure in which the control information is arranged will be described. FIG. 6 is a diagram showing an example of the protocol stack structure according to this embodiment. As shown in this figure, the protocol stack used in the broadcast system is TMCC (Transmission and Multiplexing Configuration It consists of the following: control, time information, encoded video data, encoded audio data, encoded subtitle data, MMT-SI, an application written in HTML5 standard (also simply called an app), EPG (Electronic Program Guide), content download data, etc. The encoding of the video and audio signals of a broadcast program is MFU (Media Fragment Unit) / MPU. The MFU / MPU is then loaded onto an MMTP payload, packetized into MMTP packets by broadcaster 1, and transmitted by broadcaster 1 as IP packets. For data content transmission, the data is packetized into MMTP packets by broadcaster 1 and transmitted by broadcaster 1 as IP packets. When these IP packets are broadcast using a broadcast transmission path, they are transmitted by broadcaster 1 in the form of TLV packets. One IP packet or one header-compressed IP packet is transmitted by broadcaster 1 as one TLV packet.
[0046] Furthermore, the protocol stack used in the broadcasting system includes two types of control information: MMT-SI and TLV-SI. MMT-SI is control information that indicates the structure of broadcast programs, etc. MMT-SI is in the format of an MMT control message, is placed in an MMTP payload by broadcasting station 1, is converted into an MMTP packet, and is transmitted by broadcasting station 1 as an IP packet. TLV-SI is control information related to IP packet multiplexing, and provides information for channel selection and information on the correspondence between IP addresses and services.
[0047] Furthermore, TMCC refers to the control information that is inserted into the transmission frame and transmitted in a hierarchical modulation scheme that specifies the modulation scheme and error correction scheme for each unit (slot) of signal on the transmission path. HEVC (High Efficiency Video Coding) is a method for encoding video signals. VVC (Versatile Video Codec) may also be used as a method for encoding video signals. AAC (Advanced Audio Coding), ALS (Audio Lossless Coding), and AC-4 are methods for encoding audio signals. UDP / IP (User Datagram Protocol / Internet Protocol) is one of the protocols used for communication. TLV (Type Length Value) is one of the data multiplexing methods. TLV (Data Level Variable) encoding consists of three elements: data type, length, and value.
[0048] <Message Types and Identification> MMT-SI includes messages, tables, and descriptors. Messages include Package Access (PA) messages, M2 section messages, CA messages, M2 short section messages, data transmission messages, and messages set by the carrier. The MMT-SI message used for broadcasting is as follows:
[0049] The "PA message" transmits the PLT and MPT to indicate the service entry point. "M2 section messages" transmit the section extension format of MPEG-2 Systems. The "CA message" transmits information regarding restricted access. "M2 Short Section Messages" transmit the section short format of MPEG-2 Systems. The "data transmission message" transmits a table related to data transmission.
[0050] The TLV-SI table used for broadcasting is as follows:
[0051] “TLV-NIT(Network Information Table for TLV) transmits information that associates the broadcast program with transmission path information such as modulation frequency in TLV packet transmission. The "AMT (Address Map Table)" transmits information that associates a service identifier, which identifies the broadcast program number, with an IP packet. The MMT-SI table used for broadcasting is as follows: The "MPT (MMT Package Table)" provides information that makes up a package, such as a list of assets and their locations. The "PLT (Package List Table)" is a list of packet IDs that transmit PA messages, including MPTs, for services provided as broadcast services. The "ECM (EntitLement Control Message)" transmits common information consisting of program information (information about the program and keys for descrambling, etc.) and control information (commands to force the decoder's scrambling function on / off).
[0052] "EMM (Entrance Management Message) transmits individual information, including contract information for each subscriber and a work key for decrypting common information." "CAT(MH)(Conditional Access Table)" specifies the packet identifier of the MMTP packet that transmits individual information among the related information that constitutes restricted access broadcasting. "MH-EIT (MH-Event Information Table)" transmits information about a program, such as its name, broadcast date and time, and a description of its content. "MH-AIT (MH-Application Information Table)" transmits dynamic control information and additional information necessary for the execution of an application. "MH-BIT (MH-Broadcaster Information Table)" is used to display information about broadcasters present on the network.
[0053] "MH-SDTT (MH-Software Download Trigger Table)" transmits notification information such as the download service ID, schedule information, and the type of receiver to be updated. "MH-SDT (MH-Service Description Table)" transmits information about programming channels, such as the name of the programming channel and the name of the broadcasting company. "MH-TOT (MH-Time Offset Table)" transmits the current date and time, as well as the difference between the actual time and the time displayed to the human system. "MH-CDT (MH-Common Data Table)" transmits data that is commonly required by receivers, such as carrier logos, and is intended to be stored in non-volatile memory. "DDMT (Data Directory Management Table)" provides the directory structure for the files that make up an application.
[0054] "DAMT (Data Asset Management Table)" provides the configuration of MPUs within an asset and version information for each MPU. The "DCCT (Data Content Configuration Table)" provides configuration information for files as data content. The Event Message Table (EMT) is used to transmit information related to event messages.
[0055] <MMTパッケージテーブル> The MPT (MMT Package Table) provides information that makes up a package, such as a list of assets and their location on the network. Figure 7 shows the data structure of the MPT according to this embodiment. The "table_id" (table identifier) is an 8-bit field that identifies each table. The "version" field is where the table's version number is written. The "length" field (table length) is the area where the number of data bytes following this field will be written.
[0056] "MMT_package_id_length" indicates the length of the package ID byte in bytes. "MMT_package_id_byte" represents the package ID. This should be the same value as the service identifier used to identify the service. "MPT_descriptors_length" indicates the length of the MPT descriptor area in bytes. The "MPT_descriptors_byte" (MPT descriptor area) is the area where MPT descriptors are stored. If the program is a multiview program, the MPT descriptor area will include an MH-Component_Group_Descriptor(). Conversely, if the program is not a multiview program, the MPT descriptor area will not include an MH-Component_Group_Descriptor(). "number_of_assets" indicates the number of assets for which this table provides information.
[0057] The MPT has an area that describes each of one or more assets. This area contains the following fields for each asset: The "identifier_type" (identifier type) indicates the ID scheme of the MMTP packet flow. If the ID scheme indicates an asset ID, it should be set to a specific value (0x00). "asset_id_scheme" (asset ID format) indicates the format of the asset ID. Receiver 2 uses the component_tag value for receiving "asset_id". Receiver 2 uses the component_tag value to identify the asset. "asset_id_length" indicates the length of the asset ID byte in bytes. "asset_id_byte" (asset ID byte) indicates the asset ID.
[0058] "asset_type" indicates the type of asset. The asset type may include, for example, hvc1 indicating video data encoded in HEVC, mp4a indicating audio data encoded in MPEG-4 audio, or mha1, mha2, mhm1, mhm2 indicating audio data encoded in MPEG-H audio, or ac-4 indicating audio data encoded in AC-4.
[0059] The "asset_clock_relation_flag" (clock information flag) indicates whether or not the asset has a clock information field. "location_count" indicates the number of location entries for the asset. "MMT_general_location_info" (location information) indicates the location information of the asset. "asset_descriptors_length" indicates the total byte length of the subsequent descriptor. The "asset_descriptors_byte" (asset descriptor area) is the area where asset descriptors are stored.
[0060] <Types and identification of descriptors> The TLV-SI descriptors used in broadcasting are as follows: A "Service List Descriptor" is a description of a list of organizing channels and their types. Satellite Delivery System Descriptor A "Descriptor" is a description of the physical conditions of a satellite transmission line. A "System Management Descriptor" is used to identify whether something is broadcast or not. A "Network Name Descriptor" is a description of the network name.
[0061] The MMT-SI descriptors used in broadcasting are as follows: The "Remote Control Key Descriptor" uniquely provides a service that assigns it to one-touch keys on the receiver's remote control (remote controller). The "asset group descriptor" provides the group relationships of assets and their priority within those groups. The "MPU timestamp descriptor" provides the presentation time of the MPU. The "access control descriptor" identifies the restricted access method. The "scrambling scheme descriptor" identifies the scrambling subsystem. The "Emergency Information Descriptor (MH)" provides a description of the necessary information and functions as an emergency warning signal. The "MH-Event Group Descriptor" describes the grouping information for multiple events. The "MH-Service List Descriptor" describes a list of organized channels and their types. The "MH-Short Event Descriptor" describes the program name and a brief description of the program. The "MH-Extended Event Descriptor" describes detailed information about the program.
[0062] A "video component descriptor" describes parameters, descriptions, and other information related to the video signal among the program's elemental signals. The "MH-stream identification descriptor" is used to identify individual program element signals. The "MH-Content Descriptor" describes the program genre. The "MH-Parental Rate Descriptor" describes the age restriction for viewing permission. The "MH-Audio Component Descriptor" describes the parameters related to the audio signal within the program elements. The "MH-Target Area Descriptor" describes the target area. The "MH-series descriptor" describes series information that spans multiple events. The "MH-SI transmission parameter descriptor" describes the parameters of SI transmission (such as period group and retransmission period). The "MH-Broadcaster Name Descriptor" is used to describe the broadcaster name. The "MH-Service Descriptor" describes the programming channel name and the name of the service provider.
[0063] The "MH-Data Encoding Scheme Descriptor" is used to identify the data encoding scheme. The "UTC-NPT reference descriptor" communicates the relationship between NPT and UTC. An "event message descriptor" conveys general information about event messages. The "MH-Local Time Offset Descriptor" describes the difference between the actual time and the time displayed to human systems when daylight saving time is in effect. The "MH-Logo Transmission Descriptor" describes a string for a simplified logo, pointing to a logo in CDT format, and other similar information. The "MPU Extended Timestamp Descriptor" provides information such as the decryption time of the access unit within the MPU. The "MPU Download Content Descriptor" describes the attribute information of content downloaded using the MPU. The "MH-Application Descriptor" describes information about the application. The "MH-Transmission Protocol Descriptor" specifies the transmission protocol and describes the location information of applications that depend on that transmission protocol. The "MH-Simplified Application Location Descriptor" describes the details of where to obtain the application.
[0064] The "MH-Application Boundary Permission Descriptor" describes the settings for application boundaries and broadcast resource access permissions for each region (URL). The "Linked PU Descriptor" describes the information of the linked presentation unit. An "application service descriptor" describes entry information and other details of an application related to a service. The "MPU node descriptor" indicates that the MPU corresponds to a directory node defined in the data directory management table. The "PU configuration descriptor" shows a list of MPUs that make up the presentation unit. The "MH-Hierarchical Encoding Descriptor" describes information for identifying hierarchically encoded video stream components.
[0065] The "Content Copy Control Descriptor" is placed to indicate control information regarding digital copies for the entire service, or to describe the maximum transmission rate. The "Content Usage Control Descriptor" is used to describe control information regarding storage and output for the program in question. It is also used to specify whether or not to implement "Quantity-Limited Copying Allowed" for the program or asset in question. The "Related Broadcaster Descriptor" indicates the identification values of the BS / broadband CS digital broadcaster and terrestrial digital broadcaster series required to access NVRAM. The "Multimedia Service Information Descriptor" describes detailed information about individual content within a multimedia service, such as whether or not data content is available and whether or not subtitles are available. The "Emergency News Descriptor" indicates that an emergency news flash related to safety and security (earthquake early warning, breaking news, news overlay) is currently being broadcast. The "MH-CA contract information descriptor" describes information for confirming that a service or event can be reserved. The "MH-CA service descriptor" indicates the composition channel of the entity that operates the automatic display message and describes the display control information of the message. The "MH-AC-4 audio descriptor" describes parameters related to the audio components of AC-4.
[0066] <Arrangement of MH-Audio Component Descriptor and MH-AC-4 Audio Descriptor> The "MH-Audio Component Descriptor" and the "MH-AC-4 Audio Descriptor" are arranged in the following table. ·MPT (Asset Descriptor Area) ·MH-EIT[p / f actual] (MH-EIT[p / f]) ·MH-EIT[schedule actual basic] (MH-EIT[schedule basic])
[0067] "MPT" is stored in the "PA message". "MH-EIT[p / f]" is time-series information regarding the current and next events, where the former is called present and the latter is called following. "MH-EIT[p / f actual]" and "MH-EIT[schedule actual basic]" are tables that describe events included in the service operating in the self-TLV stream and are stored in the "M2 section message".
[0068] Furthermore, "MH-AIT" is also a table that shows control information indicating the application's lifecycle, constraints, etc. "MMT" is also a multiplexing method that enables integrated transmission over multiple transmission paths. "MP4 ACC" is an audio coding method defined by ISO / IEC 14496-3. "MP4 ALS" (ALS: Audio Lossless Coding) is an audio lossless coding method defined by ISO / IEC 14496-3. "MPT" is an abbreviation for MMT package table. "MPT" is a table that provides information that constitutes a service (package), such as a list of assets and their locations. It has elements and attributes that indicate specific information. The "table" is stored in a message and transmitted in an MMTP packet. The message that stores the "table" is determined according to the table. In the MMT standard, a "package" refers to a unit of content. A "message" stores tables and descriptors. Messages are stored in the MMTP payload and transmitted using MMTP packets.
[0069] "SI information" is also information that describes the content of the multiplexed information, identification information, etc. Receiver 2 is, for example, a "terrestrial digital broadcast receiver" and has the function of selecting and demodulating the receiving channel from the IF signal, selecting and decoding the desired program, and outputting a baseband signal. However, receiver 2 may also be an "advanced BS digital broadcast receiver," in which case, in addition to having these functions, it is a device capable of receiving advanced BS digital broadcasts in the frequency band of 11.7GHz to 12.75GHz. Receiver 2 is also sometimes referred to as STB or IRD. An "item" is the smallest unit of transmission that constitutes an MPU in application data transmission based on the MMT transmission method. An "item" is equivalent to a file. An "MPU" is a transmission unit composed of a collection of items contained within a single component. It is envisioned that an "MPU" will correspond to a presentation unit (PU), update unit, or storage control unit.
[0070] A "component" (asset) is a unit that shares the same packet ID within a single IP data flow. In MPT, it is referred to as an asset. Components are identified by the component_tag, which will be described later. The set of applications being transmitted switches depending on the data event. An "asset" is a transmission unit of video, audio, etc., multiplexed using the MMT method. An "asset type" is a type that indicates the content being transmitted in each asset. "Simultaneous audio" is the simultaneous transmission of multiple different audio modes within the same event. An "event" is a collection of streams with fixed start and end times within the same service (programming channel), such as news or dramas.
[0071] [Hardware configuration of receiver 2] Figure 8 is a schematic diagram showing the hardware configuration of the receiver 2 according to this embodiment. Receiver 2 includes a tuner 211, demodulator 212, separator 22, selector 231, audio decoder 232, speaker 234, video decoder 241, presentation processor 242, display 243, input / output device 251, auxiliary storage device 252, and ROM (Read Only). It consists of a Memory 253, a Random Access Memory (RAM) 254, a CPU (Central Processing Unit) 255, and a communication chip 256. The demodulator 212, separator 22, selector 231, audio decoder 232, and speaker 234 are also referred to as the audio processing unit M. Note that the data processing configuration (for example, separator 22, selector 231, audio decoder 232, video decoder 241, presentation processor 242) may be implemented in software (calculation processing by CPU 255). For the hardware configurations corresponding to each configuration of receiver 2 in Figure 2, Figure 4, or Figure 5, the same numbers as those assigned to the configurations in Figure 2, Figure 4, or Figure 5 are used in Figure 8.
[0072] The digital broadcast signal received by the antenna is input to the receiver 2 via the input terminal, converted into a TLV stream by the tuner 211 and demodulator 212, and separated into video, audio, other assets, and various message tables of MMT after TLV / MMT separation processing by the separator 22. The scrambled assets are processed by a CAS module (not shown) using the EMM / ECM extracted in the TLV / MMT separation process, and decoded by a descrambler using the obtained key. The video assets are decoded by the video decoder 241, and output after displaying text and graphics images. The audio assets are output after audio decoding processing by the audio decoder 232. For video and audio output, the receiver body may be equipped with video and audio output means (display 243 and speaker 234), or it may be equipped with a digital video and audio output that outputs the decoded video and audio signals to an external device, or a digital audio output that outputs only the audio to an external device. Furthermore, a high-speed digital interface may be provided. Furthermore, the receiver may be equipped with an auxiliary storage device 252 (HDD, etc.) or other storage means inside, and may have a broadcast storage function. The receiver 2 has memory including RAM 254 used for receiver applications such as EPG and multimedia services, auxiliary storage device 252 (non-volatile memory: NVRAM, etc.) for storing service logo data and EPG data, and ROM 253 (which can also be replaced with NVRAM) for storing fonts, etc.
[0073] The separator 22 determines whether an AC-4 audio signal exists in the table describing asset information using a descriptor. An AC-4 audio signal includes an AC-4 voice asset (also referred to as an "AC-4 voice asset"). If the separator 22 determines that an AC-4 audio signal is present, and the receiver 2 has the capability to decode AC-4 audio, it separates the AC-4 audio asset and one or more MPEG-4 audio assets (MPEG-4 assets). If the separator 22 determines that an AC-4 audio signal is not present, it separates one or more MPEG-4 audio assets from the MPEG-4 audio signal data.
[0074] <Input terminals, tuner, demodulator> Receiver 2 has two terminals for inputting digital broadcast signals: an IF input and an optical input. It has various types. However, receiver 2 may have an IF input but not an optical input. Tuner 211 supports either a right-hand polarized band IF frequency, a left-hand polarized band IF frequency, or both. The demodulator 212 performs front-end signal processing.
[0075] <Separator / Video Decoder> The TLV / MMT separation process by the separator 22 consists of two processes: TLV separation and MMT separation. In broadcast transmission, the receiver 2 has the capability to simultaneously process at least 12 assets per service. The receiver 2 may have a maximum of 22 assets per service. Video assets may be screen-split encoding. In addition, the receiver 2 may not have built-in video decoding processing and may be equipped with functions such as streaming distribution from a high-speed digital interface. The receiver 2 may output HDR (High Dynamic Range) video to an SDR (Standard Dynamic Range) compatible display.
[0076] <Audio Decoder> In receiver 2, which has optional downmix processing for an external pseudo-surround processor and downmix processing for stereo sound field expansion, the display processor 242 displays the downmix setting status on the display 243. This allows the receiver to understand the setting status. If receiver 2 is equipped with a digital audio output for MPEG-4 AAC (Advanced Audio Coding) audio streams, it will output in a format that conforms to the AAC extension and is multiplexed using the broadcast format LATM / LOAS (Low-overhead MPEG-4 Audio Transport Multiplex / Low Overhead Audio Stream). If receiver 2 is equipped with a digital audio output for MPEG-4 ALS audio streams, it will output in a format that conforms to the ALS extension and is multiplexed using the broadcast format LATM / LOAS.
[0077] <Output terminals> The output terminals provided by receiver 2, namely the digital video and audio output terminal and the digital audio output terminal, are described below. However, receiver 2 may be equipped with a high-speed digital interface instead of these output terminals. In the case of receiver 2, which has a display device built into the main unit, it is not necessary to equip it with a digital video and audio output terminal. If receiver 2 does not have a display device (display 243) such as an STB, it shall be equipped with one of the following as a digital video and audio output terminal: an HDMI (registered trademark, hereinafter the same) terminal, a terminal for MHL / superMHL output, or a terminal for wireless digital video and audio output function.
[0078] The receiver may be equipped with an optical digital audio output terminal or a coaxial digital audio output terminal as a digital audio output terminal. It may also be equipped with an HDMI terminal and have a digital audio output function using the HDMI Audio Return Channel (HDMI-ARC) as defined in HDMI 1.4. When outputting an MPEG-4 AAC audio stream to the digital audio output terminal, compliance with the AAC extension is required, however, for 22.2ch multi-channel audio output, TBD (Total Byte) is also acceptable. When outputting an MPEG-4 ALS audio stream to the digital audio output terminal, compliance with the ALS extension is required, however, for MPEG-4 ALS stream output, TBD (Total Byte) is also acceptable.
[0079] NVRAM is used as memory for downloading receiver software and common data for all receivers, as well as for downloading data transmitted using the MH-CDT method, such as logo data. NVRAM stores data types, common data for all receivers (genre code table, program characteristic code table, reserved word table), logo data, multimedia services, and email reception, and for example, the digital mixer for AC-4 audio is stored there.
[0080] <Signal processing flow in the audio processing unit M> Figure 9 is a schematic diagram showing an example of the signal processing flow within the receiver according to this embodiment. This diagram shows an example of an audio processing unit M. The audio processing unit M consists of a demodulator 212, a TLV / MMT separation unit 22, a selector 231, a decoder unit 232, a mixer unit 2331, a downmixer (DMIX) unit 2332, a switch (SW) unit 2333, a DAC (Digital-Analog Converter) unit 2334, and an external output I / F (interface) unit 251. For each configuration of receiver 2 shown in Figure 2, Figure 4, or Figure 5, the same numerical part of the numbers assigned to the configurations in Figure 2, Figure 4, or Figure 5 is used in Figure 9. Note that the mixer unit 2331, the downmixer unit 2332, the switch unit 2333, and the DAC unit 2334 correspond to the mixer 233 in Figure 2.
[0081] This diagram shows the flow of audio signal processing within receiver 2. Receiver 2 extracts multiple audio assets via the TLV / MMT separation processing unit 22. Selector 231 selects the audio asset to be output from these assets, decodes it in decoder unit 232, and outputs the sound. Here, the audio asset selected by selector 231 is input to decoder unit 232, and decoding is performed according to its audio mode (audio codec).
[0082] The switch unit 2333 selects an audio asset from among several audio assets that is suitable for the external AV amplifier and outputs it to the decoder unit 232 and the external output I / F unit 251. The decoder unit 232 decodes the audio asset according to the input audio asset. If the decoded data sequence is an AC-4 data sequence, that is, if the audio asset is an AC-4 audio asset, the decoder unit 232 outputs the data sequence to the mixer unit 2331. If the decoded data sequence is a 5.1ch PCM data sequence, the decoder unit 232 outputs the data sequence to the downmixer unit 2332. If the decoded data sequence is a 2ch PCM data sequence, the decoder unit 232 outputs the data sequence to the switch unit 2333. The mixer unit 2331 synthesizes the audio from each sound material within the AC-4 audio asset into the input data stream and performs downmixing. The downmixer unit 2331 performs downmixing to convert the input data stream into 2-channel PCM data. The data sequence that has undergone downmixing is output to the switch unit 2333. The switch unit 2333 outputs a data stream to the DAC unit 2334 or the external output I / F (interface) unit 251 in response to instructions based on control information from the selector 231. The DAC unit 2334 converts the input data stream into an analog audio signal and outputs it to the speaker 234.
[0083] <Voice Switching Menu> Figure 10 shows an example of the audio switching menu according to this embodiment. The audio switching menu is a menu for selecting one of the simultaneous audio types. Audio switching menu F81 is an example of an audio switching menu displayed on receiver 2, which supports AC-4. Audio switching menu F82 is an example of an audio switching menu displayed on receiver 2a, which does not support AC-4. These are examples where one AC-4 audio asset and three MPEG-4 audio assets are being transmitted, and the control unit of receiver 2 generates these audio switching menus F81 and F82 by referring to the MH-audio component descriptor located in the MPT or MH-EIT. When creating audio switching menu F81, the MH-AC-4 audio descriptor, described later, may also be referred to. This one AC-4 audio asset includes 11.1ch background sound, dialogue (Japanese), dialogue (English), commentary (Japanese), and commentary (English). The three MPEG-4 audio assets are 5.1ch Japanese, 2ch (stereo) Japanese, and 2ch English, respectively.
[0084] Audio type F811 is an audio type for selecting background sound (11.1ch) and dialogue (Japanese) included in the AC-4 audio asset. Audio type F812 is an audio type for selecting background sound (11.1ch) and dialogue (English) included in the AC-4 audio asset. Audio type F813 is an audio type for selecting background sound (11.1ch) and commentary audio (Japanese) included in the AC-4 audio asset. Audio type F814 is an audio type for selecting background sound (11.1ch) and commentary audio (English) included in the AC-4 audio asset.
[0085] Audio type F815 is the audio type for selecting 5.1ch Japanese, one of the three MPEG-4 audio assets. Audio type F816 is the audio type for selecting 2ch (stereo) Japanese, one of the three MPEG-4 audio assets. Audio type F817 is the audio type for selecting 2ch (stereo) English, one of the three MPEG-4 audio assets. In the audio switching menu F81, audio type F811 is currently selected. When any of audio types F811 through F814 are selected, the audio objects included in the selected audio type are mixed in mixer 233.
[0086] Audio type F821 is the audio type for selecting 5.1ch Japanese, one of the three MPEG-4 audio assets. Audio type F822 is the audio type for selecting 2ch (stereo) Japanese, one of the three MPEG-4 audio assets. Audio type F823 is the audio type for selecting 2ch (stereo) English, one of the three MPEG-4 audio assets. In the audio switching menu F82, it indicates that audio type F821 is selected.
[0087] Furthermore, for receiver 2, which supports 5.1ch, the audio type for selecting 2ch may or may not be displayed. Also, the audio type used is the audio notation described in the text_char area of the MH-audio component descriptor. In the case of AC-4 audio assets where multiple languages exist, multiple audio types (audio notations) may be described in the text_char area.
[0088] [MH-Audio Component Descriptor] Figure 11 is a schematic diagram showing an example of the structure of an MH-voice component descriptor according to this embodiment. MH-Audio Component Descriptors are used to describe each parameter of an audio elementary stream in an asset and to represent the elementary stream in character form. MPEG-4 audio is multiplexed as an audio elementary stream for each audio configuration (e.g., language, number of channels). AC-4 audio contains various audio configurations in a single audio elementary stream. When an MH-Audio Component Descriptor is placed in an MPT, it is placed in the asset_descriptors_byte of the corresponding audio asset within the MPT.
[0089] In the MH-Audio Component descriptor, the meaning of each field is as follows. Note that in Figure 11, etc., "uimsbf" represents an unsigned integer most significant bit first, and "bslbf" represents a bit string left bit first. The "descriptor_tag" field contains a fixed value that indicates it is an MH-audio component descriptor. "descriptor_length" specifies the descriptor length of the MH-audio component descriptor.
[0090] The "nga_type" field (field F91) indicates the type of next-generation audio for the audio asset corresponding to this MH-Audio component descriptor. Figure 12 is a table showing examples of nga_type values according to this embodiment. In the example in Figure 12, a value of 0b0 for "nga_type" indicates that the next-generation audio type is MPEG-H 3D Audio (Baseline Profile), and 0b1 indicates that it is AC-4.
[0091] The "nga_level" (field F92) indicates the level of the audio asset corresponding to this MH-audio component descriptor. The level is information indicating the processing power required to decode the audio asset. The level may correspond to, for example, the processing load or the amount of memory used, but it may also correspond to the number of channels. The type and level pair may specify, for example, the performance of the equipment or the performance required to decode the bitstream. Figure 13 is a table showing examples of nga_level values according to this embodiment. In the example in Figure 13, when the value of "nga_level" is 0b000, it indicates that it is not NGA (Next Generation Audio). In this case, the value of "nga_type" is meaningless. Also, when the value of "nga_level" is 0b001, it indicates that the level is 1 (Level 1). Similarly, when it is 0b010, it indicates that the level is 2 (Level 2). When it is 0b011, it indicates that the level is 3 (Level 3). When it is 0b100, it indicates that the level is 4 (Level 4). Also, 0b101 through 0b111 are unused. Note that "0b" indicates that the following sequence of numbers is in binary.
[0092] For MPEG-4 AAC audio streams, a specific value (0x03) is set for "stream_content," while a different value (0x04) is set for MPEG-4 ALS audio streams. Furthermore, for AC-4 audio streams, yet another value (for example, 0x07) may be set.
[0093] "component_type" defines the type of audio component, defining 8 bits (b7-b0) as follows: b7: dialogue control, b6-b5: voice for disabled persons, b4-b0: voice mode. Note that "component_type" may have more bits and additional values (e.g., b8), and the added value may be defined as AC-4. The "component_tag" is a label used to identify a component stream, and it has the same value as the component tag in the MH-stream identification descriptor. The "stream_type" field should contain a fixed value indicating that the stream is in LATM / LOAS format.
[0094] The "simulcast_group_tag" assigns the same number to components that are performing simulcast (transmitting the same content using different encoding methods or audio modes). Components that are not performing simulcast are set to a specific value (0xFF). The "main_component_flag" is set to a specific value when the audio component is the main audio component. "quality_indicator" represents the sound quality mode. "sampling_rate" indicates the sampling frequency. "ISO_639_language_code" indicates the language of the speech component. In ES multilingual mode, it indicates the language of the first speech component. The language code is represented by a 3-letter alphabet code. Each character is written in 8 bits and inserted into a 24-bit field in that order. The "text_char" field specifies the voice type name. This field can be omitted if it is the default string.
[0095] To distinguish between stereo audio transmitted simultaneously with AC-4, 22.2ch surround, or 5.1ch surround, and MPEG-4 AAC stereo audio transmitted simultaneously with ALS encoding, the simulcast_group_tag (simulcast group identification) is used on the receiver side. The same simulcast_group_tag value is used for all simultaneously transmitted audio.
[0096] Note that "ISO_639_language_code2" may indicate one or more language names for the AC-4 audio asset. Specifically, the first language (e.g., Japanese) may be specified in "ISO_639_language_code" and the second language (e.g., English) in "ISO_639_language_code2". Furthermore, for AC-4 audio assets, multiple languages (e.g., Japanese and English) may be specified in "ISO_639_language_code2".
[0097] In the transmission operation of the MH-Audio Component Descriptor, when the parameters of the audio stream are updated within the same event, broadcaster 1 will, in principle, change the contents of the MPT's MH-Audio Component Descriptor and update the MPT version. However, as an exception, there may be cases where the transmission operation does not update this descriptor. In this case, the contents of the audio stream and the MH-Audio Component Descriptor will temporarily be inconsistent. For example, this may occur when transitioning from the main program to commercials, or during flexible programming. In this case, since broadcaster 1 does not update the MPT version, the receiver will continue to play the audio stream with the same component tag value. This practice of not updating this descriptor when updating audio stream parameters is only permissible when the audio encoding scheme is AAC and when switching between audio modes of 5.1ch or less.
[0098] During the reception processing of MH-audio component descriptors, if the MPT version is updated and the number of audio streams or the contents of this descriptor are updated, receiver 2 will appropriately play the audio according to the contents of this descriptor. If the MPT version has not been updated, receiver 2 will, in principle, continue to play the audio stream with the same component tag value. When switching between audio modes of 5.1ch or less, the contents of the audio stream and this descriptor may differ. In that case, receiver 2 will prioritize decoding the contents of the audio stream.
[0099] [Selecting an audio asset] Broadcasting utilizes multiple audio modes (AC-4, MPEG-H, MPEG-4 AAC2ch, AAC5.1ch, AAC7.1ch, AAC22.2ch, ALS2ch, ALS5.1ch).
[0100] Receiver 2a has the following functions as an audio decoding function: • MPEG-4 AAC 2ch playback • Downmix playback function from MPEG-4 AAC 5.1ch to 2ch To meet these conditions, AAC2ch will be used in simulcast mode for AC-4, MPEG-H, MPEG-4 AAC7.1ch, or AAC22.2ch audio modes (AAC5.1ch may also be used in simulcast mode). In addition, AAC2ch or AAC5.1ch will be used in simulcast mode for AC-4, MPEG-H audio modes, or ALS audio modes.
[0101] Receiver 2 has the ability to switch and select between multiple audio assets as follows when operating them. When playback is initiated on the receiver 2 itself, it determines the audio modes that can be played on the receiver 2 and switches to play assets prioritizing those with the smallest component tag values. When selecting a channel, the asset with the smallest playable component tag value is played as the default audio. Receiver 2, in a playback environment with up to 2 channels, will prioritize playback of AAC2ch audio if AC-4, MPEG-H, or AAC5.1ch are being used in conjunction with AAC2ch audio. Receiver 2, in a playback environment up to a specific level, will prioritize playback of the highest or lowest level among those below that specific level if AC-4 is being used. However, if the audio mode switches without updating the MPT version, the currently playing asset will continue to play as is.
[0102] Receiver 2 determines multi-language operation by referring to the simulcast group identification and plays the asset (language) with the smallest component tag value as the default language. Receiver 2 may also play the audio for AC-4 in a predetermined default language. Receiver 2 will revert to the default language when re-tuning, even if the language has been switched. Receiver 2 may also have a language lock mode for AC-4. In receiver 2, the selection of active audio assets can be cyclically switched using the audio button on the remote control, etc. For example, in receiver 2, the selection between an AC-4 audio asset and one or more MPEG-4 audio assets can be cyclically switched. In the user interface where the receiver selects a voice from a menu, the voice information should be displayed according to the information in the MH-Voice Component Descriptor. The voice type notation should prioritize the voice notation specified in the text_char area of the MH-Voice Component Descriptor. However, for AC-4, receiver 2 may prioritize a predetermined voice notation.
[0103] Receiver 2 switches audio modes within the same audio asset, and also switches to audio from a different asset automatically, in a way that does not cause any unnaturalness to the recipient. The switching operation of Receiver 2 is as follows: (1) Receiver 2, having grasped the audio mode and asset switch based on the MPT update approximately 0.5 seconds prior, fades out the output of the preceding audio and then mutes it. (2) After the receiver 2 has performed the necessary processing for switching, it unmutes and resumes outputting the subsequent audio. The time required for the switching process varies depending on whether or not the audio asset is being switched and the type of audio mode being updated. Generally, the encoding method switching time is the longest. During the switching process, a silent period is provided on the transmitting side. (3) When the receiver 2 switches from an MPEG-4 audio asset to an AC-4 audio asset, it displays the digital mixer for the AC-4 audio.
[0104] Figure 14 is a schematic diagram showing another example of the structure of an MH-audio component descriptor according to this embodiment. The only difference between the example in Figure 14 and the example in Figure 11 is that it does not have "nga_type" (field F91) and "nga_level" (field F92), but does have a 4-bit "nga_level" (field F93). "nga_level" (field F93) indicates the level of the audio asset corresponding to this MH-audio component descriptor. In the example in Figure 14, "stream_content" identifies whether or not it is an AC-4 audio asset, and if it is an AC-4 audio asset, its level is identified by "nga_level".
[0105] Figure 15 is a table showing examples of stream_content values according to this embodiment. In the example in Figure 15, a value of 0x3 for "stream_content" indicates that the audio asset is MPEG-4 AAC. Similarly, a value of 0x4 indicates that the audio asset is MPEG-4 ALS. A value of 0x6 indicates that the audio asset is MPEG-H 3D Audio (Baseline Profile). A value of 0x7 indicates that the audio asset is AC-4. 0x5 is unused.
[0106] Figure 16 is a table showing another example of the nga_level value according to this embodiment. In the example in Figure 16, when the value of "nga_level" is 0b0001, it indicates that the level is 1 (Level 1). Similarly, when it is 0b0010, it indicates that the level is 2 (Level 2). When it is 0b0011, it indicates that the level is 3 (Level 3). When it is 0b0100, it indicates that the level is 4 (Level 4). 0b0000 and 0b0101 through 0b1111 are unused.
[0107] Figure 17 is a schematic diagram showing an example of the structure of an MH-AC-4 speech descriptor according to this embodiment. Like MH-speech component descriptors, MH-AC-4 speech descriptors are placed in the asset_descriptors_byte of the corresponding asset in the MPT.
[0108] The "descriptor_tag" field contains a fixed value that indicates it is an MH-AC-4 audio descriptor. "descriptor_length" specifies the descriptor length of the MH-AC-4 audio descriptor. "nga_type" and "nga_level" are the same as "nga_type" and "nga_level" in Figure 11. Alternatively, receiver 2 may refer to the "nga_type" and "nga_level" in the MH-AC-4 audio descriptor, and not include "nga_type" and "nga_level" in the MH-audio component descriptor.
[0109] "presentation()" indicates the sound materials included in the AC-4 audio asset. Furthermore, "presentation()" may also specify the combination of sound materials to be mixed. For example, the combination of background sound (11.1ch) and dialogue (Japanese) for audio type F811 in Figure 10, the combination of background sound (11.1ch) and dialogue (English) for audio type F812, the combination of background sound (11.1ch) and commentary (Japanese) for audio type F813, and the combination of background sound (11.1ch) and commentary (English) for audio type F814 may each be defined in "presentation()". Alternatively, it may be ac4_presentation_info as defined in ETSI TS 103 190-1 and ETSI TS 103 190-2.
[0110] "dialogue_enhancement()" indicates whether or not the AC-4 voice asset has a dialogue enhancement function, and may also indicate information available for dialogue enhancement. Note that "presentation()" and "dialogue_enhancement()" may be included multiple times (N times) in the MH-AC-4 voice descriptor, as shown in Figure 17. When multiple times are included, each may correspond to one of the options included in the voice switching menu.
[0111] Figure 18 is a schematic diagram showing an example of the structure of presentation() according to this embodiment. The example shown in Figure 18 is an excerpt from section 4.2.3.2, ac4_presentation_info - AC-4 presentation information, of ETSI TS 103 190-1 V1.3.1 (2018-02) Digital Audio Compression (AC-4) Standard; Part 1: Channel based coding.
[0112] Figure 19 is a schematic diagram showing another example of the structure of presentation() according to this embodiment. The example shown in Figure 19 is an excerpt from section 6.2.1.2, ac4_presentation_info, of ETSI TS 103 190-2 V1.2.1 (2018-02) Digital Audio Compression (AC-4) Standard; Part 2: Immersive and personalized.
[0113] [Audio switching operation] Figure 20 is a flowchart illustrating a detailed example of the switching process according to this embodiment. This diagram illustrates the audio switching operation of receiver 2. The processing in the next steps S101-S104, S112-S114, and S121, and the control in steps S122 and S123 are performed by the computer (CPU255: control unit) of receiver 2. Note that the example in Figure 20 is a flowchart for the case where the structure of the MH-audio component descriptor is the example shown in Figure 11.
[0114] (Step S101) The receiver selects a channel instructed by the receiver via remote control, or a channel automatically designated by receiver 2, i.e., the tuner 211 and Demux 22 are configured to receive the channel. After that, the process in step S102 is performed. (Step S102) Receiver 2 updates the MPT. Then, the process in step S103 is performed. (Step S103) Receiver 2 checks the default asset. Specifically, the default asset is an asset whose "component_tag" is a specific value. The default asset is predetermined for each asset type. If the asset type is "broadcast transmission audio", the specific value "0x0010" is assigned to the default asset. This specific value is set as the initial value of variable i (i=0x0010). After that, the process in step S104 is performed.
[0115] (Step S104) Receiver 2 determines whether "component_tag" is a broadcast transmission audio asset. Assets of the asset type "broadcast transmission audio" are assigned values from "0x0010" to "0x002F". Receiver 2 determines whether it is a broadcast transmission audio asset by checking whether the value of variable i is less than or equal to "0x002F". Note that AC-4 audio assets may be assigned a smaller value in "component_tag" than MPEG-4 audio assets. In this case, the playability of the AC-4 audio asset will be determined first. However, AC-4 audio assets may be assigned a larger value in "component_tag" than MPEG-4 audio assets. If it is determined that the asset is a broadcast transmission audio asset (Yes), the process in step S1111 is performed. On the other hand, if it is determined that the asset is not a broadcast transmission audio asset (No), the process in step S121 is performed.
[0116] (Step S1111) Receiver 2 determines whether the "nga_level" of the MH-audio component descriptor is "0". This determines whether the corresponding asset is a next-generation audio voice asset. If "nga_level" is "0" (Yes), i.e., it is not a next-generation audio voice asset, the process in step S1114 is performed. On the other hand, if "nga_level" is not "0" (No), i.e., it is a next-generation audio voice asset, the process in step S1112 is performed.
[0117] (Step S1112) Receiver 2 determines whether the type of next-generation audio indicated by "nga_type" is a type that its device can play (support). If it is not a playable type (No), the variable i is incremented and then the process in step S104 is performed. On the other hand, if it is a playable type (Yes), the process in step S1113 is performed.
[0118] (Step S1113) Receiver 2 determines whether the level indicated by "nga_level" is a level that its device can regenerate (handle) (whether it is below a regenerative level). If it is not a regenerative level (No), the variable i is incremented and then the process in step S104 is performed. If it is a regenerative level (Yes), the process in step S113 is performed. In addition, if the level is playable (Yes) during the processing in step S1113, the processing in step S114 or S112 may be performed.
[0119] (Step S114) Receiver 2 determines whether the stream is playable. Specifically, Receiver 2 uses "stream_content", "component_type", and "stream_type" to determine whether the stream is playable. If the stream is not playable (No), the variable i is incremented and then the process in step S104 is performed. If the stream is playable (Yes), the process in step S112 is performed.
[0120] (Step S112) Receiver 2 checks for the presence or absence of simulcast and the language, etc. Specifically, receiver 2 performs this process using "simulcast_group_tag", "ES_multi_lingual_flag", "main_component_flag", "ISO_639_language_code", "ISO_639_language_code2", and "text_char".
[0121] (Step S113) The receiver 2 adds the asset information obtained in step S112 or the asset information determined to be at a playable level in step S1113 to a list in memory (RAM 254 or auxiliary storage device 252). This lists playable MPEG-4 audio assets and AC-4 sound material combinations. The information added to the list is, for example, the information indicated by the presentation() of the MH-AC-4 audio descriptor for AC-4 audio assets. After that, the receiver 2 increments the variable i and performs the process in step S104 for the asset with the next "component_tag" value.
[0122] (Step S121) Receiver 2 selects a combination of MPEG-4 assets or AC-4 sound material from the list compiled in step S113. The selection may be made automatically by Receiver 2 or manually by the receiver. If the receiver makes a manual selection, Receiver 2 displays an audio switching menu, for example, as shown in Figure 10, based on the list compiled in step S113, and allows the receiver to make a selection. Then, the process in step S122 is performed. (Step S122) Receiver 2 displays the selected language, etc. The selection may be made automatically by receiver 2 or manually by the recipient. Then proceed to step S123. In AC-4 audio, a single asset may store data sequences for multiple languages. In this case, if an AC-4 audio asset is selected in step S121, the receiver 2 can select a language without changing (selecting) the asset. (Step S123) Receiver 2 plays the combination of audio assets or sound materials selected in steps S121 and S122. Then, after the process in step S102 is performed, receiver 2 repeats the operation of this flowchart.
[0123] Furthermore, in step S1111, the receiver 2 may determine whether the stream_content of the MH-audio component descriptor is 0x3 or 0x4. That is, if stream_content is 0x3 or 0x4, it may be determined that it is not next-generation audio and the process in step S114 may be performed. Alternatively, if stream_content is not 0x3 or 0x4, or if it is 0x6 or 0x7, it may be determined that it is next-generation audio and the process in step S1111 may be performed.
[0124] Furthermore, in step S1112, the receiver 2 may determine whether it is compatible by referring to the stream_content of the MH-audio component descriptor. That is, when stream_content is 0x6, the receiver 2 determines whether it is compatible with MPEG-H 3D Audio (Baseline Profile). Also, when stream_content is 0x7, the receiver 2 determines whether it is compatible with AC-4.
[0125] As described above, in this embodiment, the broadcasting system Sys broadcasts, which includes AC-4 audio as an audio asset (an example of a component). Broadcasting station 1 broadcasts a broadcast wave in which identification information indicating whether or not an audio component is an AC-4 audio component is multiplexed into the multiplexing layer. The separator 22 of receiver 2 acquires this identification information from the broadcast wave in the multiplexing layer. The CPU 255 (an example of a control unit) selects an audio asset according to the capabilities of receiver 2 based on this identification information. The audio decoder 232 decodes the audio data of the selected audio asset.
[0126] In this embodiment, the broadcasting system Sys broadcasts AC-4 audio as an audio asset (an example of a component). Broadcasting station 1 broadcasts a broadcast wave containing MMT control information that includes identification information indicating whether or not an audio component is an AC-4 audio component. The separator 22 of receiver 2 separates the MMT control information containing this identification information from the broadcast wave. The CPU 255 (an example of a control unit) selects an audio asset according to the capabilities of receiver 2 based on this identification information. The audio decoder 232 decodes the audio data of the selected audio asset.
[0127] In the above embodiment, the multiplexing layer on which identification information indicating whether or not an audio component is an AC-4 audio component is multiplexed is an MMT layer, but it may also be a UDP / IP or TLV layer. Furthermore, the identification information indicating whether or not an audio component is an AC-4 audio component may be the data nga_type or stream_content, i.e., identification information included in the control information, or it may be a descriptor (MH-audio component descriptor) indicating whether or not an audio component is an AC-4 audio component, or it may be control information (MMT-SI, MPT, MH-EIT) indicating whether or not an audio component is an AC-4 audio component. As a result, in the broadcasting system Sys, receiver 2 can perform appropriate audio playback according to the capabilities of its own device.
[0128] In the above embodiment, "sound" may be replaced with "audio." "Asset" may be replaced with "component," and "component" may be replaced with "asset."
[0129] Furthermore, some parts of the broadcasting station 1 (broadcasting equipment), receiver 2, broadcasting station server 3, and carrier server 4 in the above-described embodiment, such as the receiver 2's separators (Demux, TLV / MMT separation units) 22, 22a, selector 231, audio decoder (decoder unit) 232, mixer unit 2331, downmixer unit 2332, switch unit 2333, DAC unit 2334, mixer 233, video decoder 241, presentation processor 242, input / output device 251, CPU 255, and communication chip 256, may be implemented using a computer. In that case, the program for implementing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Hereinafter, "computer system" refers to a computer system built into the broadcasting station 1, receiver 2, broadcasting station server 3, or carrier server 4, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. In addition, "computer-readable recording media" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs over networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside computer systems that act as servers or clients in such cases. Moreover, the above-mentioned programs may be for the purpose of realizing some of the functions described above, and may also be able to realize the above-mentioned functions in combination with programs already recorded in the computer system. Furthermore, some or all of the broadcasting station 1, receiver 2, broadcasting station server 3, and carrier server 4 in the above-described embodiment may be implemented as integrated circuits such as LSIs (Large Scale Integration). Each functional block of the broadcasting station 1, receiver 2, broadcasting station server 3, and carrier server 4 may be individually processorized, or some or all of them may be integrated into a single processor. In addition, the method of implementing integrated circuits is not limited to LSIs; dedicated circuits or general-purpose processors may also be used. Furthermore, if advances in semiconductor technology lead to the emergence of integrated circuit technologies that can replace LSIs, integrated circuits using such technologies may be used.
[0130] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention. [Explanation of symbols]
[0131] Broadcasting System Sys Relay station Sa Broadcasting station (broadcasting equipment) 1 AC-4 encoder 11 MPEG-4 encoders 111, 112, 113 Mux (Multiplexer) 12 Receiver 2, 2a Broadcasting station server 3 Operator Server 4 Tuner 211 Demodulator 212 Separator (Demux, TLV / MMT separation section) 22, 22a Selector 231 Audio decoder (decoder unit) 232 AC-4 Decoder 232-1 MPEG-4 Decoder 232-2 Mixer section 2331 Down mixer section 2332 AC-4 Renderer 233-1 Mixer 233-2 Switch section 2333 DAC section 2334 Mixer 233 Speaker 234 Video Decoder 241 Presentation processor 242 Display 243 Input / Output Device (External Output I / F) 251 Auxiliary storage 252 ROM 253 RAM 254 CPU 255 Communication chip 256
Claims
1. An acquisition unit that acquires, in a multiplexing layer, identification information indicating whether or not an audio component is an AC-4 audio audio component from a broadcast wave, information indicating multiple combinations of sound material for one of the AC-4 audio audio components, and information regarding dialogue emphasis for each of the multiple combinations. A selection unit that selects an audio component and a combination of sound materials according to the capabilities of the device, based on the aforementioned identification information and information indicating multiple combinations of sound materials. A decoding unit that decodes the audio data of the combination of sound materials of the selected audio components. Equipped with, Receiving device.
2. The aforementioned identification information is included in an audio component descriptor, which is an MMT (MPEG Media Transport) descriptor that describes parameters related to an audio signal. The receiving device according to claim 1.
3. The aforementioned broadcast wave includes multiple audio components, including an AC-4 audio audio component and an MPEG-4 audio audio component. The selection unit selects one audio component from the plurality of audio components based on the identification information contained in the audio component descriptor. The acquisition unit acquires the selected audio component, The decoding unit decodes the audio data of the selected audio component. The receiving device according to claim 2.
4. The aforementioned broadcast wave includes a descriptor that describes information about object-based acoustics, The acquisition unit acquires descriptors that describe information about the object-based acoustics. The receiving device according to claim 2.
5. The descriptor describing the object-based audio information includes a profile representing a level or set of functions that indicates the processing capability of AC-4 audio, The selection unit selects an audio component according to the capabilities of the receiving device based on the level or profile. The receiving device according to claim 4.