Broadcast system, receiver, transmission device, reception method, and transmission method

The broadcasting system and receiver adapt to receiver capabilities by acquiring and decoding MPEG-H audio information, allowing appropriate audio reproduction.

JP2026035674APending Publication Date: 2026-03-04SHARP KK
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing receivers lack the capability to process object-based audio signals such as MPEG-H audio, necessitating a solution for appropriate audio reproduction based on the receiver's capabilities.

Method used

A broadcasting system and receiver that include a receiving unit to acquire information about MPEG-H audio presence, a selecting unit to choose an audio component based on receiver capabilities, and a decoding unit to decode the appropriate audio data, with a transmitting device providing information about MPEG-H audio and object-based audio descriptors.

Benefits of technology

Enables appropriate audio reproduction according to the receiver's capabilities, ensuring seamless playback of MPEG-H audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035674000001_ABST
    Figure 2026035674000001_ABST
Patent Text Reader

Abstract

To perform appropriate sound reproduction according to the capability of a receiver.SOLUTION: Receiving a broadcast including an MPEG-H audio in an audio component, obtaining information indicating that the MPEG-H audio is present from the broadcast, selecting an audio component based on the information, decoding audio data of the selected audio component, wherein the broadcast includes a descriptor describing information related to an object-based audio, and obtaining the descriptor describing the information related to the object-based audio.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a broadcasting system, a receiver, a transmitting device, a receiving method, and a transmitting method. [Background technology]

[0002] In broadcasting, the use of object-based audio signals such as MPEG-H audio is being considered. Patent Document 1 describes that an audio signal of an audio object (an object-based audio signal) is treated as a priority signal and is played back with priority. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-124719 Summary of the Invention [Problem to be solved by the invention]

[0004] In Patent Document 1, an audio encoding device encodes an object-based audio signal, and an audio decoding device performs decoding processing on the resulting bitstream. However, some receivers that receive broadcasts do not have the capability to process object-based audio signals, so it is desirable to be able to perform appropriate audio reproduction depending on the capabilities of the receiver.

[0005] The present invention has been made in view of the above points, and provides a broadcasting system, a receiver, a transmitting device, a receiving method, and a transmitting method that are capable of performing appropriate audio reproduction according to the capabilities of the receiver. [Means for solving the problem]

[0006] (1) The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention comprises a receiving unit that receives a broadcast that includes MPEG-H audio in an audio component, an acquiring unit that acquires information indicating the presence of MPEG-H audio from the broadcast, a selecting unit that selects an audio component based on the information, and a decoding unit that decodes audio data of the selected audio component, wherein the broadcast includes a descriptor that describes information related to object-based audio, and the acquiring unit acquires the descriptor that describes the information related to the object-based audio. It is a receiver.

[0007] (2) Another aspect of the present invention is a transmitting device that broadcasts including at least an audio component, and when MPEG-H audio is included as the audio component, the transmitting device includes in the broadcast information indicating the presence of MPEG-H audio and a descriptor describing information regarding object-based audio.

[0008] (3) Another aspect of the present invention is a receiving method in a receiver, in which the receiver receives a broadcast that includes MPEG-H audio in an audio component, obtains information from the broadcast indicating the presence of MPEG-H audio, selects an audio component based on the information, decodes the audio data of the selected audio component, and the broadcast includes a descriptor that describes information related to object-based audio, and the receiving method obtains the descriptor that describes the information related to the object-based audio.

[0009] (4) Another aspect of the present invention is a transmission method in a transmission device that broadcasts including at least an audio component, wherein when the transmission device includes MPEG-H audio as the audio component, the transmission device includes in the broadcast information indicating the presence of MPEG-H audio and a descriptor describing information regarding object-based audio. [Effects of the Invention]

[0010] According to the present invention, appropriate audio reproduction can be performed according to the capabilities of the receiver. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating an example of the configuration of a broadcasting system Sys according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of a broadcasting system according to an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram illustrating a comparative example of a broadcasting system according to the present embodiment. [Figure 4] FIG. 2 is a diagram illustrating another example of the broadcasting system Sys according to the present embodiment. [Figure 5] 1 is an explanatory diagram illustrating an outline of a broadcasting system Sys according to an embodiment of the present invention. [Figure 6] FIG. 2 is a diagram illustrating an example of a protocol stack structure according to the present embodiment. [Figure 7] FIG. 2 is a diagram showing the data structure of an MPT according to the present embodiment. [Figure 8] FIG. 2 is a schematic diagram illustrating a hardware configuration of a receiver according to the present embodiment. [Figure 9] FIG. 4 is a diagram illustrating an example of a list of audio modes according to the embodiment. [Figure 10] FIG. 2 is a schematic diagram illustrating an example of a flow of signal processing in a receiver according to the present embodiment. [Figure 11] FIG. 4 is a diagram showing an example of an audio switching menu according to the embodiment. [Figure 12] 10] FIG. 10 is a schematic diagram showing an example of the structure of an MH-audio component descriptor according to the present embodiment. [Figure 13] FIG. 2 is a schematic diagram illustrating an example of a transmission operation rule according to the present embodiment. [Figure 14] FIG. 2 is a schematic diagram illustrating an example of a receiving processing standard according to the present embodiment. [Figure 15] 10 is a flowchart illustrating a detailed example of switching according to the present embodiment. [Figure 16] 10 is a flowchart illustrating another detailed example of switching according to the present embodiment. [Figure 17] FIG. 10 is a schematic diagram showing an example of the structure of an MH-MPEG-H audio descriptor according to this modification. [Figure 18] FIG. 10 is a schematic diagram illustrating an example of a transmission operation rule according to the present modification. [Figure 19] FIG. 10 is a schematic diagram illustrating an example of a reception processing standard according to the present modification. [Figure 20] FIG. 1 is a schematic diagram showing an example of simultaneous audio operation according to the present embodiment. [Figure 21] FIG. 10 is a schematic diagram showing another example of simultaneous audio operation according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0013] [System Configuration] FIG. 1 is a diagram showing an example of the configuration of a broadcasting system Sys according to an embodiment of the present invention. The broadcasting system Sys includes a broadcasting device 1 of a broadcasting station (referred to as "broadcasting station 1"), a relay station Sa, a receiver 2, a broadcasting station server 3, and a carrier server 4. The broadcasting is, for example, terrestrial digital broadcasting, but may also be, for example, advanced BS (Broadcasting Satellites) digital broadcasting or advanced wideband CS (Communication Satellites) digital broadcasting. The present invention is not limited to these broadcasting methods, and the broadcasting may also be broadcasting that does not use a relay station Sa. The broadcasting may also be wired broadcasting such as cable television. The relay station Sa is, for example, a digital relay station, but may also be a broadcasting satellite.

[0014] In the broadcasting system Sys, a broadcasting station 1 transmits digital broadcasting signals, application control information, presentation control information, etc. via broadcast waves. A service provider provides metadata and video content related to programs from a provider server 4. The application control information notifies receivers compatible with this system of applications linked to programs, and also sends commands and control information for starting and stopping them. The control information regarding presentation is transmitted as control information regarding whether or not an application can be presented, and whether or not an application and a broadcast program can be superimposed on the same TV screen. The broadcasting station operates a broadcasting station server 3 in the broadcasting system Sys. The broadcasting station server 3 provides metadata such as program title, program ID, program summary, cast, and broadcast date and time. The information provided by the broadcasting station to the service provider is provided via an API (Application Programming Interface) provided by the broadcasting station server 3.

[0015] A service provider is a party that provides services through the broadcasting system Sys, and is responsible for producing and distributing content and applications for the services, and for operating the broadcasting station server 3 to realize each individual service. Here, the services include broadcasting and communication integration services that integrate broadcasting and communication. The broadcasting station server 3 "manages and distributes applications" by sending them to the receiver 2. As a "server for each service," the broadcasting station server 3 provides server functions for realizing individual services (MPEG-H service, VOD program recommendation service, multilingual subtitle service, etc.).

[0016] MPEG-H is a set of standards under development by the ISO / IEC Moving Picture Experts Group (MPEG) for digital container standards, video compression standards, audio compression standards, and two conformance testing standards. MPEG-H audio, for example, enables object-based audio. In object-based audio, an "object" is each sound element that makes up a program, such as music or human voice. With object-based audio, an audio signal is recorded for each sound element, allowing for audio control for each element. Furthermore, when playing back on receiver 2, it is possible to play back a program based on the actual speaker position, using playback position information for the element.

[0017] The broadcasting station server 3 not only realizes the functional aspects of these services, but also transmits the content that makes up the services (MPEG-H audio data, VOD content, subtitle data, etc.) The broadcasting station server 3 acts as a "repository" to register applications for the broadcasting system Sys for distribution, and provides and searches for a list of available applications in response to inquiries from the receiver 2.

[0018] In addition to the function of receiving existing digital broadcasts, receiver 2 also includes a function for realizing broadcast and communication integrated services. In addition to the broadband network connection function, receiver 2 has the following functions. - Function to execute applications in response to application control signals from broadcasting - Function to present information through collaboration between broadcasting and communications - Device linking function Here, the terminals include, for example, user terminals such as smartphones, smart speakers, etc. The terminal linking function of the receiver 2 accesses broadcast resources such as program information in response to requests from other terminals, and invokes receiver functions such as playback control. Another example of the application is a digital mixer for MPEG-H audio. A user (also called a "receiver") can use the digital mixer received from the provider's server 4 to adjust the volume or effects of the audio signal for each sound material, or to adjust the balance between sound materials. These adjustments can also be made for each speaker.

[0019] More specifically, the receiver 2 has the following functions. The receiver 2 has a "broadcast reception and playback" function that receives broadcast radio waves, selects a specific broadcast service, and synchronously plays back the video, audio, subtitles, and data broadcast that make up the service. The receiver 2 has a "communication content reception and playback" function that accesses video content stored on a server on the communication network (for example, the provider's server 4), receives it as VOD streaming, and synchronizes and plays back the video, audio, and subtitles that make up the content.The receiver 2 has an "application control" function that acts on the application engine, mainly for managed applications, based on application control information obtained from a server on the communication network or from a broadcast signal, and controls and manages the life cycle and events of each application. The receiver 2 has an "application engine" function that acquires and executes applications. This function is realized by, for example, an HTML5 browser. The receiver 2 has a "presentation synchronization control" function that controls the stream presentation synchronization of video, audio, etc. received by broadcasting and video, audio, etc. received by streaming. The receiver 2 has an "application launcher" function, which is a navigation function that allows the user to select and launch managed applications outside of broadcasting.

[0020] FIG. 2 is a diagram showing an example of a broadcasting system Sys according to this embodiment. The receiver 2 in FIG. 2 is a receiver compatible with MPEG-H audio.

[0021] Broadcast station 1 multiplexes the video and audio signals and transmits the multiplexed signal. The multiplexing method used is MMT (MPEG Media Transport)·TLV (Type Length Value). Broadcasting station 1 generates and transmits an audio signal A11 (also referred to as an "advanced audio signal") that includes both MPEG-4 audio (channel-based audio: e.g., audio 12 to audio 14) signals and MPEG-H audio (object-based audio: audio 11) signals. In this way, the broadcasting station 1 transmits the MPEG-H audio signal and the MPEG-4 audio signal in parallel.

[0022] More specifically, the broadcasting station 1 uses MPEG-H audio as an audio component (asset) and generates an advanced audio signal A11 by multiplexing this audio component with each audio component of MPEG-4 audio. This multiplexing is performed after the audio data string of each audio component is encoded. The broadcasting station 1 transmits a broadcast wave on which this advanced audio signal A11 is multiplexed. In a table describing asset information, the broadcasting station 1 uses a descriptor to describe whether an MPEG-H audio signal (object-based audio: audio 11) is present (see FIG. 12). Note that the descriptor indicating whether an MPEG-H audio signal is present may be a descriptor indicating whether an advanced audio signal A11 is present, or may indicate whether an MPEG-H audio signal is included. In MMT, components such as video and audio are defined as assets.

[0023] An example of MPEG-H audio 11 is level 3, with a maximum of 11.1 channels, dialogue in Japanese or English, and commentary in Japanese. At level 3, the receiver 2 must be capable of simultaneously decoding 16 channels in addition to processing MPEG-H. Another example of MPEG-H audio 11 is level 4, with a maximum of 22.2 channels, dialogue in Japanese or English, and commentary in Japanese. At level 4, the receiver 2 must be capable of simultaneously decoding 28 channels in addition to processing MPEG-H. In an example of MPEG-4 audio, audio 12 is 7.1ch Japanese, audio 13 is stereo Japanese, and audio 14 is 7.1ch English. Audio 13 is a simulcast (simultaneous parallel broadcast) of audio 12.

[0024] The receiver 2 includes a tuner 211, a Demux (demultiplexer) 22, a selector 231, an audio decoder 232, a mixer 233, and a video decoder 241. The detailed configuration of the receiver 2 will be described later.

[0025] The tuner 211 receives broadcast waves via an antenna and tunes (selects) to a channel selected by a user operation. The tuned signal is demodulated and input to the Demux 22 as data. The Demux 22 separates the input data into a video data string, an audio data string, a superimposed text data string, a subtitle data string, etc. The separated audio data string is output to the selector 231. The separated video data string is output to the video decoder 241.

[0026] Here, the Demux 22 separates the audio data stream into audio data streams of each audio component: MPEG-H audio 11 and MPEG-4 audio 12, 13, and 14. More specifically, the Demux 22 determines whether an MPEG-H audio signal is present by using a descriptor in a table describing asset information. If the Demux 22 determines that an MPEG-H audio signal is present, and if the receiver 2 has MPEG-H audio decoding capability, it separates audio 11, audio 12, audio 13, and audio 14 from the advanced audio signal data. If the Demux 22 determines that an MPEG-H audio signal is not present, it separates only audio 12, audio 13, and audio 14.

[0027] The audio data string of each audio component output from the Demux 22 is input to the selector 231. The selector 231 selects the audio data string of the audio component in accordance with a user operation or the capabilities of the receiver 2. The capabilities of the receiver 2 include, for example, the number of channels that can be decoded simultaneously, or the type and capabilities of speakers that can be played back. The selector 231 outputs the selected audio data string to the audio decoder 232. The audio decoder 232 decodes the audio data string of the audio component input from the selector 231 . If the audio data sequence decoded by the audio decoder 232 is an MPEG-H audio data sequence, the mixer 233 synthesizes the audio for each sound material and performs downmixing. The downmixed audio data sequence is converted into audio and output from a speaker. If the audio data sequence decoded by the audio decoder 232 is an MPEG-4 audio data sequence, the audio data sequence is converted into audio and output from a speaker. In other words, audio synthesis for each sound material and downmixing are not performed on MPEG-4 audio data sequences.

[0028] The video data stream output from Demux 22 is input to the video decoder 241. The video decoder 241 decodes the input video data stream. The decoded video data stream undergoes color space conversion processing as necessary and is used to display the video on a display. The superimposed text data stream and subtitle data stream separated by Demux 22 are decoded by a superimposed text decoder and subtitle decoder (not shown), respectively, and the decoded character strings are superimposed on the video.

[0029] As described above, the receiver 2 according to this embodiment acquires information indicating the presence of an MPEG-H audio signal in a broadcast that includes MPEG-H audio (sound) in the audio component at the multiplex layer. The receiver 2 selects an audio component according to its own capabilities, allowing appropriate audio playback according to the capabilities of the receiver 2.

[0030] FIG. 3 is a diagram showing a comparative example of the broadcasting system Sys according to the present embodiment. This diagram shows an example in which broadcasting station C1 transmits only MPEG-H audio. In this example, the MPEG-H audio is not multiplexed with the MPEG-4 audio. In this example, MPEG-H audio is used as the only audio component. Therefore, the table describing asset information does not include a descriptor indicating whether an MPEG-H audio signal is present. In this case, Demux C22 can acquire MPEG-H audio as an audio component (audio configuration), but cannot determine whether it can be processed (level, etc.). The audio data stream output from Demux C22 is decoded by the audio decoder C232 and output to the mixer C233.

[0031] In contrast to the comparative example in Figure 3, broadcasting station 1 in this embodiment transmits MPEG-H audio signals and MPEG-4 audio signals in parallel. Receiver 2 first checks the descriptor to see if an MPEG-H audio signal is present, and then selects either MPEG-H or MPEG-4 audio depending on its audio decoding capabilities. This allows the receiver 2 to play back audio in the appropriate format, either MPEG-H or MPEG-4, depending on the capabilities of the receiver 2 itself.

[0032] 4 is a diagram showing another example of the broadcasting system Sys according to this embodiment, in which the receiver 2 is not compatible with MPEG-H. The receiver 2 in this figure includes a tuner 211, a Demux 22a, a selector 231a, an audio decoder 232a, and an audio decoder 241. In this figure, the same functional units as those in the receiver 2 in Figure 3 are denoted by the same reference numerals, and their description will be omitted.

[0033] The Demux 22a separates the input data into a video data string, an audio data string, a superimposed text data string, a subtitle data string, etc. The separated audio data string is output to the selector 231a. The separated video data string is output to the video decoder 241. Here, the Demux 22a separates the audio data stream into MPEG-4 audio 12, 13, and 14 and audio data streams of each audio component. More specifically, the Demux 22a determines whether an MPEG-H audio signal 11 exists based on a descriptor in a table describing asset information. The Demux 22a determines that the MPEG-H audio signal 11 cannot be played. Demux 22 separates voice 12, voice 13, and voice 14 from the data of audio signal A11. Note that the selection of such a signal (selection of an MPEG-4 audio signal or selection of a simulcast of only an MPEG-4 audio signal) may be performed by the selector 231.

[0034] FIG. 5 is an explanatory diagram illustrating an outline of the broadcasting system Sys according to this embodiment. In the broadcasting system Sys, the broadcasting station 1 is configured to include an MPEG-H encoder 11, MPEG-4 encoders 111 to 113, and a Mux (multiplexer) 12. The broadcasting station 1 also has other functional units necessary for broadcasting. Although the figure shows an example in which there are three MPEG-4 encoders, the broadcasting station 1 may be equipped with two or less MPEG-4 encoders, or may be equipped with four or more MPEG-4 encoders.

[0035] The receiver 2 is configured to include a Demux 22, a selector 231, an MPEG-H audio core decoder 232-1, an MPEG-4 decoder 232-2, an MPEG-H audio renderer 233-1, and a mixer 233-2. In this figure, the same functional units as those in the receiver 2 in Figure 2 are assigned the same reference numerals. The MPEG-H audio core decoder 232-1 and MPEG-4 decoder 232-2 correspond to the audio decoder 232 in Figure 2. The MPEG-H audio renderer 233-1 and mixer 233-2 correspond to the mixer 233 in Figure 2.

[0036] At broadcasting station 1, as MPEG-H audio sound materials, data of background sounds (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), and commentary audio (Japanese) are input to MPEG-H encoder 11. Also, as MPEG-4 audio sound materials, data of 7.1ch audio including Japanese dialogue, stereo audio including Japanese dialogue, and stereo data including English dialogue are input to MPEG-4 encoders 111, 112, and 113, respectively.

[0037] The MPEG-H encoder 11 encodes the input audio and outputs an MPEG-H audio stream St1. This stream is also called an MPEG-H 3D audio stream. The MPEG-4 encoders 111, 112, and 113 encode the input audio and output MPEG-4 audio streams St2, St3, and St4, respectively.

[0038] The video stream, SI (Signaling Information), MPEG-H audio stream St1, and MPEG-4 audio streams St2, St3, and St4 are input to Mux 12. Mux 12 multiplexes these data. The multiplexed data is modulated, and the modulated signal is broadcast as a broadcast wave.

[0039] The broadcast waves received by the receiver 2 are demodulated, and the demodulated data is input to the Demux 22. The Demux 22 separates the input data into a video stream, SI, an MPEG-H audio stream St1, and MPEG-4 audio streams St2, St3, and St4. The MPEG-H audio stream St1 and the MPEG-4 audio streams St2, St3, and St4 are input to a selector 231. A descriptor indicating whether an MPEG-H audio signal is present is extracted from the MPT (MMT Package Table) of the SI.

[0040] The selector 231 determines whether an MPEG-H audio signal is present based on the extracted descriptor, and if an MPEG-H audio signal is present, the selector 231 outputs the MPEG-H audio stream St1 to the MPEG-H audio core decoder 232-1. The selector 231 outputs the MPEG-4 audio stream St2, St3 or St4 to the MPEG-4 decoder 232-2 based on the MPT.

[0041] The MPEG-H audio core decoder 232-1 extracts data of background sounds (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), and commentary audio (Japanese) by decoding the MPEG-H audio core decoder 232-1. The MPEG-H audio renderer 233-1 is an audio renderer for MPEG-H audio, and renders (including down-converts or up-converts) the audio data extracted by the MPEG-H audio core decoder 232-1, and outputs it to the mixer 233. The MPEG-4 decoder 232-2 decodes the MPEG-4 audio streams St2, St3 or St4 to extract 7.1ch audio including Japanese subtitles, stereo audio including Japanese subtitles, and stereo data including English subtitles, and outputs them to the mixer 233. The mixer 233-2 synthesizes the audio of the input data, and the synthesized audio is output from each speaker or headphones, etc.

[0042] [Regarding the control information of the broadcast wave] The broadcast wave according to this embodiment will be described. In the broadcast wave, the control information is superimposed and transmitted by each broadcaster on its broadcast signal, which is a TLV stream. The control information includes TLV-SI (TLV-Signaling Information) related to the TLV multiplexing method and MMT-SI (MMT-Signaling Information) related to MMT, which is a media transport method. Hereinafter, a “component” (of video or audio) is also referred to as an “asset”.

[0043] [The protocol stack structure of the system using MMT] In the system using MMT, an example of the structure of the protocol stack in which the control information is arranged will be described. FIG. 6 is a diagram showing an example of the structure of the protocol stack according to this embodiment. As shown in this diagram, the protocol stack used in the broadcasting system includes TMCC (Transmission and Multiplexing Configuration Control), time information, encoded video data, encoded audio data, encoded subtitle data, MMT-SI, applications written in HTML5 standard (also simply called apps), EPG (Electronic Program Guide), and content download data. The video and audio signals of broadcast programs are encoded as MFUs (Media Fragment Units) / MPUs. The MFUs / MPUs are then loaded onto MMTP payloads and packetized by broadcast station 1 as MMTP packets, which are then transmitted by broadcast station 1 as IP packets. For data content transmission, data is packetized by broadcast station 1 as MMTP packets and transmitted by broadcast station 1 as IP packets. When IP packets configured in this way are broadcast over a broadcast transmission path, they are transmitted by broadcast station 1 in the form of TLV packets. One IP packet or one header-compressed IP packet is transmitted by broadcast station 1 as one TLV packet.

[0044] Furthermore, the protocol stack used in the broadcasting system provides two types of control information: MMT-SI and TLV-SI. MMT-SI is control information that indicates the configuration of a broadcast program, etc. MMT-SI is in the form of an MMT control message, placed in the MMTP payload by broadcast station 1, converted into MMTP packets, and transmitted by broadcast station 1 in IP packets. TLV-SI is control information related to multiplexing of IP packets, and provides information for channel selection and information corresponding to IP addresses and services.

[0045] TMCC is a hierarchical modulation method that specifies the modulation method and error correction method for each signal unit (slot) on the transmission path, and this control information is inserted into the transmission frame for transmission. HEVC (High Efficiency VIDEO Coding) is a method of encoding video signals. AAC (Advanced Audio Coding) and ALS (Audio Lossless Coding) are methods of encoding audio signals. UDP / IP (User Datagram Protocol / Internet Protocol) is one of the protocols used for communication. TLV (TYPE LENGTH VALUE) is a data multiplexing method. TLV consists of three parts for encoding data: data type (Type), length (Length), and value (Value).

[0046] <Message type and identification> MMT-SI includes messages, tables, and descriptors. The messages include Package Access (PA) messages, M2 section messages, CA messages, M2 short section messages, data transmission messages, and messages set by the operator. The MMT-SI messages used in broadcasting are as follows:

[0047] The "PA message" carries the PLT and MPT to indicate the entry point of the service. The "M2 Section Message" transmits the section extension format of MPEG-2 Systems. The "CA message" transmits information about the conditional access method. The "M2 Short Section Message" transmits the MPEG-2 Systems section short format. The "data transmission message" transmits a table relating to data transmission.

[0048] The TLV-SI table used in broadcasting is as follows:

[0049] "TLV-NIT (Network Information Table for TLV)" transmits information that associates transmission path information, such as modulation frequency, with broadcast programs when transmitted via TLV packets. The "AMT (Address Map Table)" transmits information that associates a service identifier that identifies a broadcast program number with an IP packet. The "MPT (MMT Package Table)" provides information that makes up a package, such as a list of assets and their locations. "PLT (Package List Table)" indicates a list of packet IDs that transmit PA messages including MPTs of services provided as broadcasting services. "ECM (Entity Control Message)" transmits common information consisting of program information (information about the program and keys for descrambling, etc.) and control information (forced on / off command for the decoder's scrambling function).

[0050] The "EMM (Entity Management Message)" transmits individual information including the contract information for each subscriber and a work key for decrypting common information. The "CAT (MH) (Conditional Access Table)" specifies the packet identifier of the MMTP packet that transmits individual information among the related information that constitutes the conditional reception broadcasting. The "MH-EIT (MH-Event Information Table)" transmits information about the program, such as the program name, broadcast date and time, and a description of the content. The "MH-AIT (MH-Application Information Table)" transmits dynamic control information related to the application and additional information required for execution. The "MH-BIT (MH-Broadcaster Information Table)" is used to present information about broadcasters present on the network.

[0051] The "MH-SDTT (MH-Software DownLoad Trigger Table)" transmits notification information such as the download service ID, schedule information, and the type of receiver to be updated. The "MH-SDT (MH-Service Description Table)" transmits information about the channel, such as the name of the channel and the name of the broadcasting company. The "MH-TOT (MH-Time Offset Table)" transmits the current date and time, as well as the difference between the actual time and the time displayed to the human system. The "MH-CDT (MH-Common Data Table)" transmits data that is commonly required by receivers, such as operator logos, and is intended to be stored in non-volatile memory. The "DDMT (Data Directory Management Table)" provides the directory structure of the files that make up the application.

[0052] The "DAMT (Data Asset Management Table)" provides the configuration of the MPUs in the asset and version information for each MPU. "DCCT (Data Content Configuration Table)" provides configuration information of files as data contents. The "EMT (Event Message Table)" is used to transmit information about event messages.

[0053] <MMTパッケージテーブル> The MPT (MMT Package Table) provides information that configures a package, such as a list of assets and their locations on the network. FIG. 7 is a diagram showing the data structure of the MPT according to this embodiment. "tabLe_ID" (table identifier) ​​is an 8-bit field that identifies each table. "Version" is an area where the version number of the table is written. "Length" (table length) is an area where the number of data bytes following this field is written.

[0054] "MMT_package_ID_Length" indicates the length of the package ID byte in bytes. "MMT_package_ID_byte" indicates the package ID. It should have the same value as the service ID used to identify the service. "MPT_descriptors_Length" indicates the length of the MPT descriptor area in bytes. "MPT_descriptors_byte" (MPT descriptor area) is an area that stores MPT descriptors. If the program is a multiview program, the descriptor area of ​​the MPT includes the MH-Component_Group_Descriptor(). On the other hand, if the program is not a multiview program, the descriptor area of ​​the MPT does not include the MH-Component_Group_Descriptor(). "number_of_assets" indicates the number of assets for which this table provides information. "IDentifier_type" indicates the ID system of the MMTP packet flow. If the ID system indicates an asset ID, a specific value (0x00) is set.

[0055] The MPT has an area that describes one or more assets. This area contains the following fields for each asset: "asset_ID_scheme" (asset ID format) indicates the format of the asset ID. For "asset_ID", the receiver 2 uses the component_tag value for reception operations. The receiver 2 uses the component_tag value to identify the asset. "asset_ID_length" (asset ID length) indicates the length of the asset ID bytes in bytes. "asset_ID_byte" (asset ID byte) indicates the asset ID.

[0056] "asset_type" indicates the type of asset. The asset type may be, for example, hcv1, which indicates video data encoded in HEVC, mp4a, which indicates audio data encoded in MPEG-4 audio, or mha1, mha2, mhm1, or mhm2, which indicates audio data encoded in MPEG-H audio.

[0057] "asset_clock_relation_flag" (clock information flag) indicates whether or not the asset has a clock information field. "location_count" (number of locations) indicates the number of location information items for the asset. "MMT_general_location_info" (location information) indicates the location information of the asset. "asset_descriptors_length" indicates the total byte length of the subsequent descriptors. "asset_descriptors_byte" (asset descriptor area) is an area for storing asset descriptors.

[0058] <Descriptor types and identification> The TLV-SI descriptors used in broadcasting are as follows: The "Service List Descriptor" is a description of a list of programming channels and their types. "Satellite Delivery System Descriptor" is a description of the physical conditions of a satellite transmission path. The "System Management Descriptor" is an identification such as broadcast / non-broadcast. A "Network Name Descriptor" is a description of a network name.

[0059] The MMT-SI descriptors used in broadcasting are as follows: The "remote control key descriptor" uniquely provides a service for assigning to one-touch keys on the receiver's remote control (remote controller). The "asset group descriptor" provides the group relationship of assets and their priority within the group. The "MPU timestamp descriptor" provides the presentation time of the MPU. The "access control descriptor" identifies the conditional access method. The "Scrambling Scheme Descriptor" identifies the scrambling subsystem. The "Emergency Information Descriptor (MH)" provides a description of the information and functions required for an emergency warning signal. The "MH-Event Group Descriptor" describes grouping information for multiple events. The "MH-Service List Descriptor" describes a list of programming channels and their types. The "MH-Short Event Descriptor" describes the program name and a brief description of the program. The "MH-extended event descriptor" describes detailed information about the program.

[0060] The "video component descriptor" describes parameters, explanations, etc. relating to the video signal among the program element signals. The "MH-stream identification descriptor" is used to identify individual program element signals. The "MH-Content Descriptor" describes the program genre. The "MH-Parental Rate Descriptor" describes the age restriction for viewing. The "MH-audio component descriptor" describes parameters related to the audio signal among the program elements. The "MH-Target Area Descriptor" describes the target area. The "MH-series descriptor" describes series information spanning multiple events. The "MH-SI transmission parameter descriptor" describes the parameters of SI transmission (cycle group, retransmission cycle, etc.). The "MH-broadcaster name descriptor" describes the broadcaster name. The "MH-service descriptor" describes the name of the channel and the name of its operator.

[0061] The "MH-Data Encoding Descriptor" is used to identify the data encoding method. The "UTC-NPT Reference Descriptor" conveys the relationship between NPT and UTC. The "Event Message Descriptor" conveys information about the event message in general. The "MH-Local Time Offset Descriptor" describes the difference in time between the actual time and the time displayed in the human system when daylight saving time is in effect. The "MH-Logo Transmission Descriptor" describes the simple logo character string, pointing to the CDT format logo, etc. The "MPU extended timestamp descriptor" provides the decoding time of an access unit within the MPU, etc. The "MPU download content descriptor" describes attribute information of content downloaded using the MPU. The "MH-Application Descriptor" describes information about an application. The "MH-transmission protocol descriptor" specifies the transmission protocol and describes the location information of the application depending on the transmission protocol. The "MH-Simple Application Location Descriptor" describes in detail where to obtain the application.

[0062] The "MH-application boundary authority setting descriptor" describes the application boundary setting and the broadcast resource access authority setting for each area (URL). The "link destination PU descriptor" describes information about the link destination presentation unit. The "application service descriptor" describes entry information about the application related to the service. The "MPU node descriptor" indicates that the corresponding MPU corresponds to the directory node defined in the data directory management table. The "PU configuration descriptor" indicates a list of MPUs that make up the presentation unit. The "MH-hierarchical coding descriptor" describes information for identifying a hierarchically coded video stream component.

[0063] The "content copy control descriptor" is placed when indicating control information regarding digital copying for the entire service or when describing the maximum transmission rate. The "content usage control descriptor" is placed when describing control information regarding storage and output for the program. It is also placed when specifying whether "copyable with quantity limit" is applied to the program or asset. The "related broadcaster descriptor" indicates the identification values of the broadcaster of BS / broadband CS digital broadcasting and the series of terrestrial digital broadcasting necessary for access to NVRAM. The "multimedia service information descriptor" describes detailed information regarding individual contents of the multimedia service, such as the presence or absence of data content and subtitles. The "emergency news descriptor" indicates that an emergency news bulletin (earthquake early warning, special news, news flash) related to safety is being broadcast. The "MH-CA contract information descriptor" describes information for confirming that a service or event is reservable. The "MH-CA service descriptor" indicates the organized channel of the entity that uses the automatic display message and describes the display control information of the message.

[0064] <Placement of MH-audio component descriptor> The "MH-audio component descriptor" is placed in the following table. ·MPT (asset descriptor area) ·MH-EIT[p / f actual] (MH-EIT[p / f]) ·MH-EIT[schedule actual basic](MH-EIT[schedule basic])

[0065] "MPT" is stored in "PA message". "MH-EIT[p / f]" is time-series information about the current and next events, the former being called "present" and the latter being called "following." "MH-EIT[p / f actual]" and "MH-EIT[schedule actual basic]" are tables that describe events included in the service operated in the TLV stream, and are stored in the "M2 section message."

[0066] The "MH-AIT" is also a table that indicates control information such as application lifecycle and constraints. "MMT" is a multiplexing method that enables integrated transmission over multiple transmission paths. "MP4 AAC" is an audio coding method defined by ISO / IEC 14496-3. "MP4 ALS" (Audio Lossless Coding) is a lossless audio coding method defined by ISO / IEC 14496-3. "MPT" stands for MMT Package Table. An MPT is a table that provides information that constitutes a service (package), such as a list of assets and their locations. It contains elements and attributes that indicate specific information. "Tables" are stored in messages and transmitted in MMTP packets. The messages that store "tables" are determined by the table. In the MMT standard, a "package" refers to a unit of content. "Messages" store tables and descriptors. Messages are stored in MMTP payloads and transmitted using MMTP packets.

[0067] "SI information" is also information that describes the content of the multiplexed information, identification information, etc. Receiver 2 is, for example, a "terrestrial digital broadcasting receiver" that has the function of selecting and demodulating a receiving channel from the IF signal, selecting and decoding a desired program, and outputting a baseband signal. However, receiver 2 may also be an "advanced BS digital broadcasting receiver," in which case, in addition to having these functions, it is a device that can receive advanced BS digital broadcasting in the 11.7 GHz to 12.75 GHz frequency band. Receiver 2 is also called an STB or IRD. An "item" is the smallest unit of transmission that constitutes an MPU in application data transmission based on the MMT transmission method. An "item" corresponds to a file. An "MPU" is a transmission unit that is comprised of a collection of items contained within one component. It is expected that an "MPU" will be operated in correspondence with a presentation unit (PU), update unit, or storage control unit.

[0068] A "component" (asset) is a unit with the same packet ID in one IP data flow. In MPT, it is referred to as an asset. A "component" is identified by the component_tag, which will be described later. The application set being transmitted changes depending on the data event. An "asset" is a transmission unit for video, audio, etc. multiplexed using the MMT method. An "asset type" indicates the type of content being transmitted in each asset. "Simultaneous audio" is the simultaneous transmission of multiple different audio modes within the same event. An "event" is a collection of streams with fixed start and end times within the same service (programming channel), such as news or drama.

[0069] [Receiver 2 hardware configuration] FIG. 8 is a schematic diagram showing the hardware configuration of the receiver 2 according to this embodiment. The receiver 2 is composed of a tuner 211, a demodulator 212, a separator 22, a selector 231, an audio decoder 232, a speaker 234, a video decoder 241, a presentation processor 242, a display 243, an input / output device 251, an auxiliary memory device 252, a ROM (Read Only Memory) 253, a RAM (Random Access Memory) 254, a CPU (Central Processing Unit) 255, and a communication chip 256. The demodulator 212, the separator 22, the selector 231, the audio decoder 232, and the speaker 234 are also referred to as an audio processing unit M. Note that the configuration for processing data (for example, the separator 22, the selector 231, the audio decoder 232, the video decoder 241, and the presentation processor 242) may be realized by software (arithmetic processing by the CPU 255). 8, the hardware configurations corresponding to the respective components of the receiver 2 in FIG. 2, FIG. 4, or FIG. 5 are assigned the same numbers as the numerical portions of the numbers assigned to the components in FIG. 2, FIG. 4, or FIG.

[0070] A digital broadcast signal received via an antenna is input to the receiver 2 via an input terminal, converted into a TLV stream by the tuner 211 and demodulator 212, and separated into video, audio, other assets, and various MMT message tables through a TLV / MMT separation process by the separator 22. The scrambled assets are decoded by a descrambler using the key obtained by processing the EMM / ECM extracted during the TLV / MMT separation process in a CAS module (not shown). The video assets are decoded by a video decoder 241 and output after processing to present text and graphics images. The audio assets are decoded by an audio decoder 232 and then output. Regarding video and audio output, the receiver itself may be equipped with video and audio output means (display 243 and speaker 234), or may be equipped with a digital video and audio output that outputs decoded video and audio signals to an external device, or a digital audio output that outputs only audio to an external device. A high-speed digital interface may also be provided. The receiver may also be provided with a storage means such as an auxiliary storage device 252 (HDD, etc.) inside, providing a broadcast storage function. The receiver 2 has the following memories: RAM 254 used by receiver applications such as EPG and multimedia services, auxiliary storage device 252 (non-volatile memory: NVRAM, etc.) for storing service logo data and EPG data, and ROM 253 (NVRAM can be used instead) for storing fonts, etc.

[0071] The separator 22 determines whether an MPEG-H audio signal is present based on a descriptor in a table describing asset information. The MPEG-H audio signal includes an MPEG-H audio asset (also referred to as an "MPEG-H audio asset"). If separator 22 determines that an MPEG-H audio signal is present, and if receiver 2 has MPEG-H audio decoding capability, separator 22 separates an MPEG-H audio asset and one or more MPEG-4 audio assets (MPEG-4 assets) from the data of the MPEG-H audio signal. If separator 22 determines that an MPEG-H audio signal is not present, separator 22 separates one or more MPEG-4 audio assets from the data of the MPEG-4 audio signal.

[0072] <Input terminal, tuner, demodulator> The receiver 2 has two types of terminals for inputting digital broadcast signals: an IF input and an optical input. However, the receiver does not necessarily have to have the IF input but not the optical input. The tuner 211 supports IF frequencies for the right-hand circular band, IF frequencies for the left-hand circular band, or both. Demodulator 212 performs front-end signal processing.

[0073] <Separator / Video Decoder> The TLV / MMT separation process by the separator 22 consists of two processes: TLV separation and MMT separation. The receiver 2 in broadcast transmission has the ability to simultaneously process a minimum of 12 assets per service. The receiver 2 may handle a maximum of 22 assets per service. Video assets may be split-screen encoded. Furthermore, the receiver 2 may not have built-in video decoding processing, but may instead be equipped with a function for streaming via a high-speed digital interface. The receiver 2 may output HDR (High Dynamic Range) video to an SDR (Standard Dynamic Range) compatible display. The receiver 2 monitors the vIDeo_transfer_characteristics value in the video component descriptor to identify the transfer characteristics of the received video signal.

[0074] <Audio decoder> In a receiver 2 that has downmix processing for an external pseudo surround processor and downmix processing for stereo sound field expansion added as options, the presentation processing unit 242 displays the downmix setting status on the display 243. This allows the receiver to understand the setting status. If receiver 2 is equipped with a digital audio output for MPEG-4 AAC (Advanced Audio Coding) audio streams, it will output the audio in a format that complies with the AAC extension and is multiplexed in the LATM / LOAS (Low-overhead MPEG-4 Audio Transport Multiplex / Low Overhead Audio Stream) broadcast format.If receiver 2 is equipped with a digital audio output for MPEG-4 ALS audio streams, it will output the audio in a format that complies with the ALS extension and is multiplexed in the LATM / LOAS broadcast format.

[0075] <Output terminal> The following describes a digital video and audio output terminal and a digital audio output terminal as output terminals provided on the receiver 2. However, the receiver 2 may be equipped with a high-speed digital interface instead of these output terminals. Note that a receiver 2 with a built-in display device does not need to be equipped with a digital video and audio output terminal. In the case of a receiver 2 that does not have a display device (display 243), such as an STB, the receiver 2 is equipped with an HDMI (registered trademark, the same applies below) terminal, a terminal for MHL / superMHL output, or a terminal for wireless digital video and audio output function as the digital video and audio output terminal.

[0076] The receiver may be equipped with an optical digital audio output terminal or a coaxial digital audio output terminal as a digital audio output terminal. It may also be equipped with an HDMI terminal and have a digital audio output function using the HDMI Audio Return Channel (HDMI-ARC) defined in HDMI 1.4. When outputting an MPEG-4 AAC audio stream to the digital audio output terminal, it must comply with the AAC extension, but 22.2ch multi-channel audio output may use TBD.When outputting an MPEG-4 ALS audio stream to the digital audio output terminal, it must comply with the ALS extension, but MPEG-4 ALS stream output may use TBD.

[0077] NVRAM is used as memory for downloading receiver software and data common to all receivers, as well as downloading data transmitted using the MH-CDT method, such as logo data. NVRAM stores data types, data common to all receivers (genre code table, program characteristic code table, reserved word table), logo data, multimedia services, received emails, etc., and also stores, for example, the digital mixer for MPEG-H audio.

[0078] [Audio asset selection and switching] FIG. 9 is a diagram showing an example of a list of audio modes according to this embodiment. In broadcasting, the audio modes shown in Figure 9 are used. Of these, the three that need to be decoded by the receiver 2 alone are MPEG-4 AAC1ch (only used in advanced wideband CS digital broadcasting), AAC2ch, and AAC5.1ch (2ch downmix processing possible). Audio assets in other modes are not decoded within the receiver 2, but are output in stream format to an external amplifier and may be decoded by the external amplifier. On the other hand, assuming a program that uses an audio mode other than AAC1ch, AAC2ch, or AAC5.1ch as the main audio is being received by a receiver 2 that is not connected to an external amplifier, when audio in a mode other than these three is broadcast, simulcast audio that can be decoded by all receivers 2 is broadcast simultaneously using different audio assets.

[0079] For example, if the audio codec is "MPEG-H," any of MPEG-4 AAC 1ch, 2ch, and 5.1ch will be broadcast simultaneously using different audio assets as a simulcast audio combination. Note that an audio mode may include multiple MPEG-H audio assets, and in this case, each different audio asset may include MPEG-H audio assets of different levels. For example, when an audio mode with a high level (Level 3) is broadcast, simulcast audio of an audio mode with a lower specific level (e.g., Level 1, 2: levels with fewer channels) may be broadcast simultaneously using different audio assets. This means that when MPEG-H of a particular level (e.g., level 3) is broadcast, even a receiver 2 that only supports a lower level (e.g., level 2) can play MPEG-H audio by selecting an audio asset of that lower level.

[0080] A receiver 2 that is not connected to an external amplifier selects one asset to output sound from among these multiple audio assets and outputs it to the audio output (such as speaker 234). The asset to be output can be selected arbitrarily by the receiver, or the receiver 2 automatically selects an asset with a compatible audio mode. In the case of automatic selection, priority is basically given to assets with a smaller component tag value among assets with compatible audio modes, but the decision may be changed depending on the selection status of language and audio type (such as commentary audio). The receiver 2 may automatically select an MPEG-H asset by minimizing the component tag value of the MPEG-H asset. Conversely, the receiver 2 may automatically select an MPEG-4 asset by minimizing the component tag value of the MPEG-4 asset. If an MPEG-H asset is present and an MPEG-4 asset is selected, the receiver 2 may display a message indicating that an MPEG-H asset is present.

[0081] <Signal processing flow in the audio processing unit M> 10 is a schematic diagram showing an example of the flow of signal processing in the receiver according to this embodiment. This diagram shows an example of an audio processing unit M. The audio processing unit M is configured to include a demodulation unit 212, a TLV / MMT separation unit 22, an audio asset selection unit 231, a decoder unit 232, a mixer unit 2331, a downmixer (DMIX) unit 2332, a switch (SW) unit 2333, a DAC (Digital-Analog Converter) unit 2334, and an external output I / F (interface) unit 251. 10, the components corresponding to the components of the receiver 2 in Fig. 2, 4, or 5 are given the same numbers as the numerical portions of the numbers given to the components in Fig. 2, 4, or 5. Note that the mixer unit 2331, the downmixer unit 2332, the switch unit 2333, and the DAC unit 2334 correspond to the mixer 233 in Fig. 2.

[0082] This diagram shows the flow of audio signal processing within the receiver 2. The receiver 2 extracts multiple audio assets via the TLV / MMT separation processing unit 22. The audio asset selection unit 231 selects an audio asset to be output from these, and the decoder unit 232 decodes and outputs the audio. Here, the audio asset selected by the audio asset selection unit 231 is input to the decoder unit 232, which decodes it according to its audio mode (audio codec).

[0083] The switch unit 2333 selects an asset suitable for the external AV amplifier from among a plurality of audio assets, and outputs it to the decoder unit 232 and the external output I / F unit 251. The decoder unit 232 decodes the audio asset according to the input audio asset. If the decoded data sequence is an MPEG-H data sequence, that is, if the audio asset is an MPEG-H audio asset, the decoder unit 232 outputs the data sequence to the mixer unit 2331. If the decoded data sequence is a 5.1ch PCM data sequence, the decoder unit 232 outputs the data sequence to the downmixer unit 2331. If the decoded data sequence is a 2ch PCM data sequence, the decoder unit 232 outputs the data sequence to the switch unit 2333. The mixer unit 2331 performs downmixing processing on the input data stream by synthesizing audio for each sound material in the MPEG-H audio asset. The downmixer unit 2331 performs downmixing processing to convert the input data stream into 2-channel PCM data. The downmixed data stream is output to the switch unit 2333. The switch unit 2333 outputs a data string to the DAC unit 2334 or the external output I / F (interface) unit 251 in accordance with an instruction based on the control information from the audio asset selection unit 231. The DAC unit 2334 converts the input data string into an analog audio signal and outputs it to the speaker 234.

[0084] <Audio switching menu> FIG. 11 is a diagram showing an example of the audio switching menu according to the present embodiment. The audio switching menu F81 is an example of an audio switching menu displayed on a receiver 2 that supports MPEG-H. The audio switching menu F82 is an example of an audio switching menu displayed on a receiver 2a that does not support MPEG-H. The audio switching menu F81 is also an example of an audio switching menu displayed on a receiver 2a that supports MPEG-H when MPEG-H audio assets exist (when MPEG-H audio assets and MPEG-4 audio assets exist). The audio switching menu is a menu for selecting one of the audio types of simultaneous audio. When an MPEG-H audio asset is selected, the receiver 2 may separate audio types in different languages ​​from a single asset and display them again in the audio switching menu.

[0085] The audio type F811 is an audio type for selecting MPEG-H audio. The two audio types F812 are audio types for selecting Japanese as the language, MPEG-4, 5.1ch or 2ch audio. The audio type F813 is an audio type for selecting English as the language, MPEG-4, 5.1ch or 2ch audio. Note that a receiver 2 that supports 5.1ch may or may not display the audio type for selecting 2ch. The audio type uses the audio notation described in the text_char field of the MH-audio component descriptor. If multiple languages ​​exist in an MPEG-H audio asset, multiple audio types (audio notation) may be described in the text_char field. If the audio is MPEG-H, the language may not be selected on the menu screen. In this case, for example, audio type F811 is the audio type for selecting MPEG-H audio.

[0086] [MH-Audio Component Descriptor] FIG. 12 is a schematic diagram showing an example of the structure of the MH-audio component descriptor according to this embodiment. The MH-Audio Component Descriptor describes each parameter of an audio elementary stream in an asset and is also used to express the elementary stream in text format. MPEG-4 audio is multiplexed as an audio elementary stream for each audio configuration (e.g., language, number of channels). MPEG-H audio contains various audio configurations in one audio elementary stream.

[0087] In the MH-audio component descriptor, the meaning of each field is as follows: "descriptor_tag" describes a fixed value that indicates that it is an MH-audio component descriptor. "descriptor_length" describes the descriptor length of the MH-audio component descriptor.

[0088] "nga_profile_level" (field F91) indicates whether or not an MPEG-H audio asset exists (whether or not MPEG-H audio exists). If MPEG-H exists, it indicates its profile and level. A profile represents a set of functions defined for a specific purpose. There are two types of profiles: "Basic Profile" and "Low Complexity Profile" (LC). LC is a standard profile. A "Baseline Profile" is an LC profile that omits specific functions. There may be three or more types of profiles. A level represents processing capability and is information corresponding to the processing capability. For example, the level may correspond to the processing load or memory usage, but may also correspond to the number of channels. The pair of profile and level may specify, for example, the performance of the device or the performance required to decode the bitstream.

[0089] "stream_content" is set to a specific value (0x03) for MPEG-4 AAC audio streams, and to another value (0x04) for MPEG-4 ALS audio streams. Note that an even further value (e.g., 0x05) may be set for MPEG-H audio streams.

[0090] "component_type" specifies the type of audio component, with 8 bits (b7-b0) defined as b7: dialogue control, b6-b5: audio for disabled people, and b4-b0: audio mode. Note that "component_type" may have an increased number of bits and a value (e.g., b8) added, and the added value may be defined as MPEG-H. "component_tag" (component tag) is a label for identifying a component stream, and has the same value as the component tag in the MH-stream identification descriptor. "stream_type" specifies a fixed value that indicates the LATM / LOAS stream format.

[0091] The "simulcast_group_tag" is assigned the same number to components that are simulcasting (transmitting the same content in different encoding formats or audio modes). For components that are not simulcasting, a specific value (0xFF) is set. "main_component_flag" takes a specific value when the audio component is the main audio. "quality_indicator" indicates the sound quality mode. "sampling_rate" indicates the sampling frequency. "ISO_639_language_code" indicates the language of the audio component. In ES multilingual mode, it indicates the language of the first audio component. The language code is expressed as a three-letter alphabetic code. Each character is written as 8 bits and inserted in that order into the 24-bit field. "text_char" describes the name of the audio type. If this description is the default character string, this field may be omitted.

[0092] The simulcast_group_tag (simulcast group identification) is used to distinguish between stereo audio transmitted simultaneously with MPEG-H, 22.2ch surround, or 5.1ch surround, and MPEG-4 AAC stereo audio transmitted simultaneously with the ALS encoding method, etc. The same value for simulcast_group_tag is sent for audio transmitted simultaneously.

[0093] <Transmission Operation Rules and Reception Processing Standards> FIG. 13 is a schematic diagram showing an example of the transmission operation rules according to this embodiment. FIG. 14 is a schematic diagram showing an example of a receiving processing standard according to this embodiment. FIG. 13 shows an example of the sending operation rules for the MH-audio component descriptor, and FIG. 14 shows an example of the receiving processing criteria for the MH-audio component descriptor.

[0094] In these examples, "nga_profile_level" is a label for identifying the presence or absence of MPEG-H, the profile, and the level. Of the four bits of "nga_profile_level," the first bit is set to the profile, and the remaining three bits are set to the presence or absence and level of MPEG-H. The field representing the profile is "nga_profile," which is bslbf (bit string, left bit first). A value of "1" for "nga_profile" indicates that the profile is "Basic Profile," and a value of "0" indicates that the profile is "Low Complexity Profile." The field indicating the presence or absence and level of MPEG-H is "nga_level", which is an uimsbf (unsigned integer most significant bit first). A value of "0" for "nga_level" indicates that an MPEG-H audio signal is not present (not MPEG-H audio), and any other value indicates that an MPEG-H audio signal is present (MPEG-H audio). A value of "1" for "nga_level" indicates that the level is "Level 1", a value of "2" for "Level 2", a value of "3" for "Level 3", and a value of "4" for "Level 4". Values ​​"5" through "7" for "nga_level" are unused and can be assigned in the future.

[0095] "stream_content" is set to 0x03 for an MPEG-4 AAC audio stream, 0x04 for an MPEG-4 ALS audio stream, and 0x05 for an MPEG-H audio stream. If the value of "stream_content" is anything other than these values, the receiver 2 invalidates the descriptor.

[0096] Regarding "simulcast_group_tag," when transmitting / receiving MPEG-H audio, the same value of "simulcast_group_tag" is set for the MPEG-H audio asset and one or more MPEG-4 audio assets. In this case, a different value of "simulcast_group_tag" is set for one or more MPEG-4 audio assets. The latter "simulcast_group_tag" is not set for the MPEG-H audio asset. In other words, the latter simulcast audio includes only MPEG-4 audio assets and does not include MPEG-H audio assets. "ISO_639_language_code2" may indicate one or more language names of MPEG-H audio assets. Specifically, the first language (e.g., Japanese) may be written in "ISO_639_language_code" and the second language (e.g., English) may be written in "ISO_639_language_code2." For MPEG-H audio assets, multiple languages ​​(for example, Japanese and English) may be written in "ISO_639_language_code2."

[0097] In the transmission operation of the MH-Audio Component Descriptor, when updating the parameters of the audio stream within the same event, broadcasting station 1 generally changes the contents of the MH-Audio Component Descriptor in the MPT and updates the MPT version. However, there are exceptions to this rule, where broadcasting station 1 performs transmission operations without updating this descriptor. In these cases, the contents of the audio stream and the MH-Audio Component Descriptor temporarily become inconsistent. This can occur, for example, when transitioning from the main program to a commercial or during fluid programming. In these cases, broadcasting station 1 does not update the MPT version, and the receiver continues to play the audio stream with the same component tag value. Such an operation where this descriptor is not updated when updating the parameters of the audio stream is permitted only when the audio encoding method is AAC and the audio mode is switched between audio modes of 5.1ch or less.

[0098] During the reception process of the MH-Audio Component Descriptor, if the MPT version is updated and the number of audio streams or the contents of this descriptor are updated, the receiver 2 will play the audio appropriately according to the contents of this descriptor. As long as the MPT version has not been updated, the receiver 2 will, in principle, continue to play the audio stream with the same component tag value. When switching between audio modes of 5.1ch or lower, the audio stream and the contents of this descriptor may differ. In such cases, the receiver 2 will prioritize decoding the contents of the audio stream.

[0099] [Select Audio Asset] The broadcast will use multiple audio modes (MPEG-H, MPEG-4 AAC2ch, AAC5.1ch, AAC7.1ch, AAC22.2ch, ALS2ch, ALS5.1ch). Mono and dual mono will not be used.

[0100] The receiver 2 has the following audio decoding functions: MPEG-4 AAC2ch playback MPEG-4 AAC 5.1ch to 2ch downmix playback function To meet these requirements, AAC2ch is used in simulcast mode in MPEG-H, MPEG-4 AAC7.1ch, or AAC22.2ch audio modes (AAC5.1ch may also be used in simulcast mode). Also, AAC2ch or AAC5.1ch is used in simulcast mode in MPEG-H audio mode or ALS audio mode.

[0101] When multiple audio assets are in operation, the receiver 2 has the ability to switch and select as follows. When playing back on the receiver 2 itself, the receiver 2 determines the audio mode that can be played on the receiver 2 itself, and switches and plays back assets in ascending order of component tag value. When selecting a channel, the asset with the smallest component tag value that can be played back is played back as the default audio. In a playback environment up to 2ch, if AAC2ch audio is simultaneously operated in conjunction with MPEG-H or AAC5.1ch, receiver 2 will prioritize playback of AAC2ch audio. In a playback environment up to a specific level, if MPEG-H is operated, receiver 2 will prioritize playback of the highest or lowest level of the levels below the specific level. However, if the audio mode is switched without updating the MPT version, the asset that is currently being played will continue to play.

[0102] The receiver 2 determines whether multiple languages ​​are being used by referring to the simulcast group identification, and plays the asset (language) with the smallest component tag value as the default language. For MPEG-H, the receiver 2 may play audio in a predetermined default language. Even if the language is switched, the receiver 2 will return to the default language when reselecting a channel. For MPEG-H, the receiver 2 may have a language fixed mode. In the receiver 2, the selection of a valid audio asset is cyclically switched by using an audio button on a remote control, etc. For example, in the receiver 2, the selection of an MPEG-H audio asset and one or more MPEG-4 audio assets is cyclically switched. In the user interface where the receiver selects audio from a menu, audio information shall be displayed according to the information in the MH-Audio Component Descriptor. Note that the audio type notation given in the text_char field in the MH-Audio Component Descriptor shall take priority. However, for MPEG-H, the receiver 2 may also give priority to a predetermined audio notation.

[0103] When the audio mode is switched within the same audio asset, or when the receiver automatically switches to the audio of a different asset, the receiver 2 switches in a way that does not sound unnatural to the receiver. The switching operation of the receiver 2 is as follows. (1) The receiver 2, which has learned of the audio mode or asset switch due to the MPT update approximately 0.5 seconds ago, mutes the output of the preceding audio after fading. (2) After the receiver 2 has performed the necessary processing for switching, it cancels the mute and resumes outputting the subsequent audio. The time required for the switching process varies depending on whether the audio asset has been switched and the type of audio mode being updated. Generally, the time required for switching the encoding method is the longest. During the switching process, a silent interval is set on the sending side. (3) When the receiver 2 switches from an MPEG-4 audio asset to an MPEG-H audio asset, it displays the MPEG-H audio digital mixer.

[0104] [Audio switching behavior] FIG. 15 is a flowchart showing a detailed example of switching according to this embodiment. This diagram shows the audio switching operation of the receiver 2. The processes of the following steps S101 to S104, S11, S112, and S121, and the control of steps S122 and S123 are performed by the computer of the receiver 2 (CPU 255: control unit).

[0105] (Step S101) The receiver selects a channel using a remote control or the like, or the receiver 2 automatically selects a channel. Then, the process of step S102 is performed. (Step S102) The receiver 2 updates the MPT, and then the process of step S103 is performed. (Step S103) The receiver 2 checks the default asset. Specifically, the default asset is an asset whose "component_tag" has a specific value. Default assets are predetermined for each asset type. If the asset type is "broadcast transmission audio," a specific value "0x0010" is assigned to the default asset. This specific value is set as the initial value of the variable i (i=0x0010). Then, the processing of step S104 is performed.

[0106] (Step S104) The receiver 2 determines whether or not "component_tag" is an asset of broadcast transmission audio. Assets whose asset type is "broadcast transmission audio" are assigned values ​​from "0x0010" to "0x002F". The receiver 2 determines whether or not the asset is an asset of broadcast transmission audio by determining whether or not the value of the variable i is equal to or less than "0x002F". Note that MPEG-H audio assets are assigned a smaller value in "component_tag" than MPEG-4 audio assets. In this case, the playability of the MPEG-H audio assets is determined first. However, MPEG-H audio assets may be assigned a larger value in "component_tag" than MPEG-4 audio assets. If it is determined that the asset is a broadcast transmission audio asset (Yes), the process proceeds to step S111. On the other hand, if it is determined that the asset is not a broadcast transmission audio asset (No), the process proceeds to step S121.

[0107] (Step S111) The receiver 2 determines whether the stream can be played back by the receiver 2 itself. Specifically, the receiver 2 determines whether the stream can be played back by the receiver 2 itself, using "stream_content", "component_type", "stream_type", and "nga_profile_level".

[0108] To determine whether MPEG-H audio can be played back, the receiver 2 makes the following determination.

[0109] The receiver 2 separates "nga_profile" and "nga_level" from "nga_profile_level". The receiver 2 determines whether an MPEG-H audio signal is not present (not MPEG-H audio) by determining whether "nga_level" is "0". In other words, the receiver 2 determines whether an MPEG-H audio signal is present (is MPEG-H audio) by determining whether "nga_level" is not "0".

[0110] If "nga_level" is "0", the MPEG-4 audio stream does not have an MPEG-H audio asset, so the receiver 2 determines whether or not its own device can play back this MPEG-4 audio stream (MPEG-4 audio). If the receiver 2 cannot play back this MPEG-4 audio stream (No), it performs the process of step S104. On the other hand, if the receiver 2 can play back this MPEG-4 audio stream (Yes), it performs the process of step S112.

[0111] If "nga_level" is not "0" (is between "1" and "4", or is "1" or greater), the MPEG-H audio stream has an MPEG-H audio asset, so the receiver 2 determines whether or not its own device can play back this MPEG-H audio stream (MPEG-H audio). If the receiver 2 cannot play back this MPEG-H audio stream (No), the process proceeds to step S104.

[0112] On the other hand, if playback is possible, the receiver 2 determines whether the level indicated by "nga_level" is at a level at which playback is possible on the receiver 2 itself (or whether it is at or below the level at which playback is possible). If the receiver 2 determines that the level is not at a level at which playback is possible (No), the process proceeds to step S104. If the level is playable, the receiver 2 determines whether the profile indicated by "nga_profile" is a profile playable by the receiver 2 itself. If it is not a playable profile (No), the process proceeds to step S104. On the other hand, if it is a playable profile (Yes), the process proceeds to step S112. If it is determined in at least one of the determinations in step S111 that the asset is not playable (No), the variable i is incremented by 1, and the process in step S104 is performed on the asset with the next "component_tag" value.

[0113] (Step S112) The receiver 2 checks whether simulcast is available and the language, etc. Specifically, the receiver 2 performs this process using "simulcast_group_tag", "ES_multi_lingual_flag", "main_component_flag", "ISO_639_language_code", "ISO_639_language_code2", and "text_char". The receiver 2 may use either "ISO_639_language_code" or "ISO_639_language_code2", or a combination thereof, for the language of MPEG-H audio.

[0114] (Step S113) The receiver 2 adds the asset information acquired in the process of S112 to a list in memory (RAM 254 or auxiliary storage device 252). This lists the playable streams. Then, the process of step S104 is performed. Here, the receiver 2 increments the variable i by 1, and performs the process of step S104 for the asset with the next "component_tag" value.

[0115] (Step S121) The receiver 2 selects a stream from the list of playable streams (the list created in the processing of S113). The selection may be made automatically by the receiver 2 or manually by the recipient. Then, the processing of step S122 is performed. (Step S122) The receiver 2 displays the selected language, etc. The selection may be made automatically by the receiver 2 or manually by the recipient. Then, the processing proceeds to step S123. In MPEG-H audio, audio data strings in multiple languages ​​may be stored in one asset. In this case, if an MPEG-H audio stream is selected in step S121, the receiver 2 can select a language without changing (selecting) the asset. (Step S123) The receiver 2 plays back the audio stream selected in steps S121 and S122. Thereafter, the receiver 2 repeats the operation of this flowchart by performing the processing of step S102.

[0116] If no particular specification is made, the receiver 2 selects the stream with the smallest component_tag value (for example, the processing of step S121). The receiver 2 also allows the recipient to select arbitrarily from a list (see FIG. 11). If an arbitrary selection is made, the receiver 2 continues to hold that value (plays the stream). However, if the component_tag value (the stream being played) disappears, or if conditions such as the presence or absence of simulcast or language change, the receiver 2 performs an appropriate switching process and immediately selects and plays the stream with the smallest component_tag value that can be played.

[0117] FIG. 16 is a flowchart showing another detailed example of switching according to this embodiment. In this figure, the same processes as those in FIG. 15 are denoted by the same reference numerals, and the description thereof will be omitted.

[0118] (Step S1111) The receiver 2 determines whether or not an MPEG-H audio signal is present by determining whether "nga_level" is "0." If "nga_level" is "0" (Yes), the receiver 2 performs the process of step S1114. On the other hand, if "nga_level" is not "0" (No), the receiver 2 performs the process of step S1112. (Step S1112) The receiver 2 determines whether the profile indicated by "nga_profile" is a profile that the receiver 2 can play (support). If the receiver 2 determines that the profile is not a playable profile (No), the receiver 2 performs the process of step S104. On the other hand, if the profile is a playable profile (Yes), the receiver 2 performs the process of step S1113.

[0119] (Step S1113) The receiver 2 determines whether the level indicated by "nga_level" is a level that the receiver 2 can play (support) (or whether it is below the playable level). If it is not a playable level (No), the process proceeds to step S104. If it is a playable level (Yes), the process proceeds to step S113. (Step S1114) The receiver 2 determines whether or not the stream is playable by its own device. Specifically, the receiver 2 determines whether or not the stream is playable by its own device using "stream_content", "component_type", and "stream_type". If the stream is not playable (No), the process of step S104 is performed. If the stream is playable (Yes), the process of step S112 is performed. If the processing in step S113 determines that the content is at a reproducible level (Yes), the processing in step S1114 or S112 may be performed.

[0120] [Use Case] Below, an example of a use case will be shown regarding a broadcast (asset) pattern, a receiver, and audio selection in the receiver in this embodiment. Broadcast pattern 1 (for example, program A) is a pattern in which the following audio (asset) 11 to 14 are broadcast. Broadcast pattern 1: (1A) Audio 11: MPEG-H Level 4 (22.2ch (9 / 10 / 3), dialogue Japanese / English, commentary Japanese) (1B) Audio 12: MPEG-4 (7.1ch, Japanese) (1C) Audio 13: MPEG-4 (Stereo Japanese) (1D) Audio 14: MPEG-4 (Stereo English) Audio 11 is MPEG-H audio, level 4, 22.2ch audio (BGM), with Japanese and English dialogue and Japanese commentary. Audio 13 is audio transmitted simultaneously with Audio 12.

[0121] Broadcast pattern 2 (for example, program B other than program A) is a pattern in which the following sounds 11 to 14 are broadcast. Broadcast pattern 2: (2A) Audio 11: MPEG-H Level 3 (11.1ch (4 / 7 / 0), dialogue Japanese / English, commentary Japanese) (2B) Audio 12: MPEG-4 (7.1ch, Japanese) (2C) Audio 13: MPEG-4 (Stereo Japanese) (2D) Audio 14: MPEG-4 (Stereo English) Comparing Broadcast Pattern 1 and Broadcast Pattern 2, the differences are that the MPEG-H (Audio 11) levels are "Level 4" and "Level 3," and the number of channel audio is "22.2ch" and "11.1ch." Audio 13 is audio transmitted simultaneously with Audio 12.

[0122] Receiver 2 has the following capabilities: Receiver A is a receiver that supports (can play) MPEG-H Audio Level 4. Receiver B is a receiver that supports (can play) MPEG-H Audio Level 3. Receiver C is a receiver that does not support (cannot play) MPEG-H Audio.

[0123] The audio selection in each of the receivers A to C is as follows. Receiver A is compatible with MPEG-H level 4 and below, as well as MPEG-4. Therefore, receiver A selects MPEG-H audio 11 for both broadcast patterns 1 and 2. Receiver B is compatible with MPEG-H level 3 and below and MPEG-4, but not with MPEG-H level 4. Therefore, in broadcast pattern 1, receiver B selects either MPEG-4 audio 12, 13, or 14, but in broadcast pattern 2, it selects MPEG-H audio 11. Receiver C selects one of MPEG-4 audio 12, 13, or 14 for both broadcast patterns 1 and 2. Even if the receiver 2 can select (or is able to select) MPEG-H audio, it may select MPEG-4 audio through user operation or settings.

[0124] More specifically, it is as follows. <Receiver B: Supports MPEG-H Level 3> In the case of broadcast pattern 1, receiver B can select either MPEG-4 audio 12, 13, or 14. Here, receiver B selects audio from the set of audio 12 and audio 14, or the set of audio 13 and audio 14, to avoid overlapping selection from simultaneous audio (audio 12 and audio 13). In the case of broadcasting pattern 2, receiver B can select MPEG-H audio 11.

[0125] In the case of broadcasting pattern 2, if receiver B is equipped with, for example, 7.1 speakers, it selects audio 11 and generates 7.1ch audio by rendering (down-converting) the selected audio 11. In this case, receiver B can play the following audio (B11) to (B13). (B11) 7.1ch dialogue Japanese (B12) 7.1ch dialogue English (B13) 7.1ch dialogue Japanese + commentary audio Japanese

[0126] In the case of broadcasting pattern 2, when using stereo headphones, for example, receiver B selects audio 11 and generates stereo audio by rendering (down-converting) the selected audio 11. In this case, receiver B can play the following audio (B21) to (B23). (B21) Stereo Dialogue Japanese (B22) Stereo Dialogue English (B23) Stereo dialogue Japanese + commentary Japanese

[0127] <Receiver A: Supports MPEG-H Level 4> The case where receiver A is equipped with 22.2ch speakers will be described. In the case of broadcast pattern 1, receiver A selects audio 11, and by performing a rendering process on the selected audio 11, can reproduce the following audio (A11) to (B13). (A11) 22.2ch dialogue Japanese (A12) 22.2ch dialogue English (A13) 22.2ch dialogue Japanese + commentary audio Japanese

[0128] In the case of broadcasting pattern 2, receiver A selects audio 11 and can play the following audio (A21) to (B23) by rendering (upconverting) the selected audio 11. (A21) 22.2ch dialogue Japanese (A22) 22.2ch dialogue English (A23) 22.2ch dialogue Japanese + commentary audio Japanese

[0129] As described above, in this embodiment, the broadcasting system Sys broadcasts audio including MPEG-H audio in audio assets (an example of a component). The separator 22 of the receiver 2 acquires nga_profile_level (an example of identification information indicating whether MPEG-H audio is present) from the broadcast airwaves at the multiplexing layer. The CPU 255 (an example of a control unit) selects an audio asset according to the capabilities of the receiver 2 based on nga_profile_level. The audio decoder 232 decodes the audio data of the selected audio asset. This allows the receiver 2 in the broadcasting system Sys to perform appropriate audio reproduction according to its own capabilities.

[0130] In this embodiment, nga_profile_level is placed in an MH-audio component descriptor, which is an MMT (MPEG Media Transport) descriptor that describes parameters related to audio signals among program elements. The separator 22 acquires the MH-audio component descriptor. The CPU 255 selects an audio component according to the capabilities of the receiver 2 based on the nga_profile_level placed in the MH-audio component descriptor.

[0131] This allows the receiver 2 to read the nga_profile_level from the MPEG-H audio component descriptor and perform appropriate audio playback according to its own device's capabilities. Furthermore, the receiver 2 can determine whether MPEG-H audio is present before inputting the audio data to the audio decoder 232 (before decoding). This allows the receiver 2 to avoid decoding MPEG-H audio data when it cannot decode MPEG-H audio or when decoding is not necessary (for example, when the receiver has selected MPEG-4 audio).

[0132] In this embodiment, the broadcast includes MPEG-H audio assets and MPEG-4 audio assets. The audio decoder 232 includes a mixer unit 2331 (an example of a first decoding unit) that decodes audio data from the MPEG-H audio assets, and a downmixer unit 2332 (an example of a second decoding unit) that decodes audio data from the MPEG-4 audio assets. The separator 22 separates the MPEG-H audio from the MPEG-4 audio. If the CPU 255 determines that MPEG-H audio can be played based on the nga_profile_level specified in the MPEG-H audio component descriptor, it selects either the MPEG-H audio or the MPEG-4 audio audio asset. The audio decoder 232 decodes the audio data of the selected audio asset, either the MPEG-H audio or the MPEG-4 audio.

[0133] This allows the receiver 2 to select either MPEG-H audio or MPEG-4 audio, and the receiver 2 can play back either MPEG-H audio or MPEG-4 audio depending on its own capabilities. If the receiver 2 is capable of playing back MPEG-H audio, the recipient can select either MPEG-H audio or MPEG-4 audio.

[0134] In this embodiment, the broadcast also includes MPEG-H audio assets and MPEG-4 audio assets. The audio decoder 232 decodes the audio data of the MPEG-4 audio assets. The separator 22 separates the MPEG-H audio from the MPEG-4 audio. The CPU 255 selects an MPEG-4 audio asset when it determines that the MPEG-H audio cannot be played based on the nga_profile_level specified in the MPEG-H audio component descriptor. The audio decoder 232 decodes the audio data of the selected MPEG-4 audio asset. This allows the receiver 2 to select MPEG-4 audio if its own device (including external amplifiers and external speakers) does not support MPEG-H audio, and the receiver 2 can play MPEG-H audio or MPEG-4 audio depending on the capabilities of its own device.

[0135] In this embodiment, each audio asset is identified by a component_tag. MPEG-4 audio is composed of one or more audio assets (they are multiplexed and transmitted). The separator 22 reads the MH-audio component descriptors (an example of configuration information) of the MPEG-H audio asset and the MPEG-4 audio asset according to the order of the component_tag, and obtains the nga_profile_level from the MH-audio component descriptor. This allows the receiver 2 to determine for each audio asset whether it is MPEG-H audio along with the audio asset information.

[0136] In this embodiment, the MH-audio component descriptor contains nga_profile_level (whether nga_level is 0 or not), and nga_level (an example of a level indicating the processing capability of MPEG-H audio) or nga_profile (an example of a profile indicating a set of functions). The CPU 255 selects an audio component according to the capabilities of the receiver 2 based on nga_level or nga_profile. This allows the receiver 2 to determine whether MPEG-H audio is present and to obtain the level or profile of the MPEG-H audio. The receiver 2 can then perform appropriate audio playback according to the level, profile, and its own device capabilities.

[0137] [Variation: MPEG-H playback capability determination] The receiver 2 may determine whether MPEG-H can be played back as follows: <Case 1> The receiver 2 may determine whether an asset is MPEG-H audio using "component_tag." In this case, a specific value is assigned to the MPEG-H audio asset in "component_tag." For example, MPEG-4 audio assets in Japanese are assigned to 0x0010 to 0x001F, and MPEG-4 audio assets in English are assigned to 0x0020 to 0x0025. MPEG-H audio assets are assigned to 0x0026 to 0x002F. The receiver 2 determines that an asset is an MPEG-H audio asset if "component_tag" is any of 0x0026 to 0x002F. If "component_tag" is not any of 0x0026 to 0x002F, the receiver 2 determines that the asset is not an MPEG-H audio asset. In this case, the receiver 2 may assign "component_tag" so that the MPEG-H profile and level can also be identified. The receiver 2 may assign the MPEG-H profile and level to "nga_profile_level."

[0138] In Case 1 of this modified example, the receiver 2 can determine whether or not an asset is an MPEG-H audio asset in the process of step S104 (without performing the process of step S111). Furthermore, the receiver 2 can determine whether or not an asset is an MPEG-H audio asset without providing a new field. Comparing the above embodiment with this variant 1, in the above embodiment, "nga_profile_level" stores whether it is MPEG-H audio, and the MPEG-H profile and level, so it is possible to operate the contents, restrictions (number, number of bytes, etc.) or rules of the values ​​of other fields (component_tag in this variant) while maintaining the existing operation.

[0139] <Case 2> The receiver 2 may use "component_type" to determine whether the asset is MPEG-H audio. For example, the "component_type" value may be defined as b8:MPEG-H. If "component_type" is b8, the receiver 2 determines that the asset is an MPEG-H audio asset. If "component_type" is not b8, the receiver 2 determines that the asset is not an MPEG-H audio asset. In this case, the receiver 2 may assign "component_type" so that the MPEG-H profile and level can also be identified. The receiver 2 may assign the MPEG-H profile and level to "nga_profile_level." In case 2 of this modified example, the receiver 2 can determine whether or not an asset is an MPEG-H audio asset without providing a new field. Comparing the above embodiment with this variant 2, in the above embodiment, "nga_profile_level" stores whether it is MPEG-H audio, and the MPEG-H profile and level, so it is possible to operate the contents, restrictions (number, number of bytes, etc.) or rules of the values ​​of other fields (component_type in this variant) while maintaining the existing operation.

[0140] <Case 3> The receiver 2 may use "stream_content" to determine whether the asset is MPEG-H audio. As described above, the "stream_content" value is set to a specific value (0x03) for an MPEG-4 AAC audio stream, and to another value (0x04) for an MPEG-4 ALS audio stream. Furthermore, a different value (e.g., 0x05) is set for an MPEG-H audio stream. If "stream_content" is 0x05, the receiver 2 determines that the asset is an MPEG-H audio asset. If "stream_content" is not 0x05, the receiver 2 determines that the asset is not an MPEG-H audio asset. In this case, the receiver 2 may assign "stream_content" so that the MPEG-H profile and level can also be identified. The receiver 2 may assign the MPEG-H profile and level to "nga_profile_level." In case 3 of this modified example, the receiver 2 can also determine whether an asset is an MPEG-H audio asset without providing a new field. Comparing the above embodiment with this variant 3, in the above embodiment, "nga_profile_level" stores whether it is MPEG-H audio or not, and the MPEG-H profile and level, so it is possible to operate the contents, restrictions (number, number of bytes, etc.) or rules of the values ​​of other fields (stream_content in this variant) while maintaining the existing operation.

[0141] Similarly, other fields referenced in steps S111 and S112, such as "stream_type" or "asset_type," may be set with information indicating whether the data is MPEG-H audio or information that can identify the profile and level.

[0142] <Case 4> In this modification, a new descriptor, "MH-MPEG-H audio descriptor," is introduced. The receiver 2 uses the "MH-MPEG-H audio descriptor" to determine whether the content is MPEG-H audio and to identify the profile and level.

[0143] 17 is a schematic diagram showing an example of the structure of an MH-MPEG-H audio descriptor according to this modification. The "MH-MPEG-H Audio Descriptor()" is one of the descriptors of MMT-SI, and is written in MMT-SI in addition to the above-mentioned descriptors such as the "MH-Audio Component Descriptor".

[0144] In the "MH-MPEG-H Audio Descriptor", the meaning of each field is as follows: "descriptor_tag" describes a fixed value that indicates that it is an MH-MPEG-H audio descriptor. "descriptor_length" describes the descriptor length of the MH-MPEG-H audio descriptor.

[0145] "nga_type" indicates the type of next-generation audio. For example, '1' is set for MPEG-H audio. Note that for audio other than MPEG-H (MPEG-4), '0' may be set, or another next-generation audio type may be set. "profile_level" indicates the profile and level. Of the three bits in "profile_level," the first bit indicates the profile, and the remaining two bits indicate the level. For example, if the value of the first bit is "1," it indicates that the profile is "Basic Profile," and if the value is "0," it indicates that the profile is "Low Complexity Profile." For the remaining two bits, if the value is "00," the level is "Level 1," if the value is "01," the level is "Level 2," if the value is "10," the level is "Level 3," and if the value is "11," the level is "Level 4." The profile and the level may be separate fields.

[0146] The "component_tag" has the same value as the component_tag in the "MH-audio component descriptor." Note that the "component_tag" in the "MH-MPEG-H audio descriptor" also has the same value as the component_tag in the MH-stream identification descriptor.

[0147] "preset()" indicates whether or not there are presets and the number of presets. A preset is, for example, a set of pre-defined audio values ​​(values ​​for adjusting the audio) for each of multiple audio materials. Different preset values ​​result in different audio balances between audio materials. Presets include those based on the audio playback environment (speaker type and placement), those for improving accessibility, and those tailored to the viewer's preferences. Presets for these include those that increase (or decrease) the volume of dialogue only, insert audio commentary, or play audio on a second device (such as a nearby speaker). By selecting a preset, the recipient can select and enjoy their preferred audio (sound balance, etc.) without having to make detailed settings for all audio materials.

[0148] "interactive()" indicates whether it is interactive or not and the operation content. The interactive status indicates whether the receiver can adjust each sound material. The operation details indicate the operation details for adjusting each sound material. The operation details include, for example, information about the material (object to be adjusted for audio), the adjustment details (type of strength or effect, type of user interface, conditions such as upper and lower limits), and the adjustment tool (tool name, provider, download location, etc.). The receiver can adjust the MPEG-H audio according to the broadcast program.

[0149] <Transmission Operation Rules and Reception Processing Standards> FIG. 18 is a schematic diagram showing an example of the transmission operation rules according to this modification. FIG. 19 is a schematic diagram showing an example of a reception processing standard according to this modification. FIG. 18 shows an example of the transmission operation rules for the MH-MPEG-H audio descriptor, and FIG. 19 shows an example of the reception processing standards for the MH-MPEG-H audio descriptor.

[0150] The "descriptor_tag" is set to a fixed value of "0x8040." This value is greater than the fixed value of "0x8014" of "descriptor_tag" in the "MH-Audio Component Descriptor." In this case, the receiver 2 treats the "MH-Audio Component Descriptor" as having a higher priority than the "MH-MPEG-H Audio Descriptor," and reads it first. In other words, the "MH-MPEG-H Audio Descriptor" is treated as auxiliary data for the "MH-Audio Component Descriptor." Note that the value of "descriptor_tag" may be a value smaller than "0x8014."

[0151] If the receiver 2 determines that an "MH-MPEG-H audio descriptor" exists in the MPT (see FIG. 7) (a descriptor with a descriptor_tag value of "0x8040"), it determines that an MPEG-H audio asset exists. In this case, the receiver 2 identifies the audio asset with the "component_tag" described in the "MH-MPEG-H Audio Descriptor" as an MPEG-H audio asset. Note that if the "MH-MPEG-H Audio Descriptor" also describes audio assets other than MPEG-H audio assets, the receiver 2 may identify the asset with the "component_tag" corresponding to the "nga_type" value (e.g., a value of "0": MPEG-H) as an MPEG-H audio asset.

[0152] In the process of step S111 in Fig. 15, if the variable i matches the "component_tag" value of the audio asset identified as an MPEG-H audio asset, the receiver 2 determines that an MPEG-H audio signal is present. On the other hand, if there is no match, the receiver 2 determines that an MPEG-H audio signal is not present. If the receiver 2 determines that an MPEG-H audio signal is present, it determines whether its own device can play back this MPEG-H audio stream. If the receiver 2 cannot play back this MPEG-H audio stream (No), it performs the process of step S104. On the other hand, if it can play back, the receiver 2 determines whether the level indicated by "profile_level" is a level that its own device can play back. If the receiver 2 determines that it is not a level that can be played back (No), it performs the process of step S104. If the level is a level that can be played back, the receiver 2 determines whether the profile indicated by "profile_level" is a profile that its own device can play back. If the receiver 2 determines that it is not a profile that can be played back (No), it performs the process of step S104. On the other hand, if the profile is a profile that can be played back (Yes), it performs the process of step S112.

[0153] 15, the receiver 2 adds the asset information acquired in S112 to a list in memory. Here, for MPEG-H audio assets, the receiver 2 also stores the "preset()" and "interactive()" information in the list in memory. In the process of step S122 or S123, the receiver 2 may automatically select a preset based on the information in "preset()," or may display the contents of the presets and let the receiver select a preset. In the process of step S122 or S123, the receiver 2 may display the MPEG-H audio digital mixer based on the information in "interactive()." In these cases, the receiver 2 plays back the audio stream based on the selected preset or the adjustment results by the digital mixer in step S123.

[0154] As described above, in this modification, the broadcast includes MMT (MPEG Media Transport) descriptors, including an MH-audio component descriptor, which is a descriptor describing parameters related to audio signals among program elements, and an MPEG-H audio descriptor, which is a descriptor describing parameters related to MPEG-H audio. The nga_type is placed in the MPEG-H audio descriptor. The separator 22 acquires the MH-audio component descriptor and the MPEG-H audio descriptor. The CPU 255 selects an audio component according to the capabilities of the receiver 2 based on the nga_type placed in the MPEG-H audio descriptor. This allows the receiver 2 to read out nga_type from the MPEG-H audio descriptor and perform appropriate audio playback according to the capabilities of the receiver 2 itself.

[0155] In this modification, preset() and interactive() are placed in the MPEG-H audio descriptor. The separator 22 acquires the MPEG-H audio component descriptor and the MPEG-H audio descriptor. The CPU 255 plays back audio based on preset() or interactive() placed in the MPEG-H audio descriptor. This allows the receiver 2 to play back audio based on preset(). The receiver can select and enjoy the audio of their choice (sound balance, etc.) without having to make detailed settings for all the audio material. The receiver 2 can adjust the audio and play back the audio based on interactive(). The receiver can adjust the MPEG-H-based audio according to the broadcast program.

[0156] In the above embodiment (including the modified examples), asset priority may be as follows: For video and audio assets, if multiple assets of the same asset_type are defined in one MPT, and if multiple component descriptors (MH-audio component descriptors) are placed in the MH-EIT, the receiver 2 prioritizes the assets in ascending order of component_tag values. In other words, the receiver 2 determines that the default asset has the highest priority, and the higher the value, the lower the priority. This priority can be used, for example, when displaying a list of streams on an EPG or in the display order when a stream switching button is pressed. However, for audio simultaneously transmitted using assets with the same simulcast_group_tag value, in a playback environment of up to 2ch, if AAC2ch audio is simultaneously operated in accordance with AAC5.1ch, the receiver 2 prioritizes playback of AAC2ch audio.

[0157] 20 is a schematic diagram showing an example of the operation of simulcast audio according to this embodiment (including modifications), which shows an example of the operation of the simulcast_group_tag value in audio transmitted in a simulcast manner. In this diagram, an asset with an audio mode of "MPEG-H" is an MPEG-H audio asset, and is available in multiple languages ​​(Japanese and English). This audio asset is assigned a component_tag value of "0x0026." This value is higher than the other audio assets (MPEG-4 audio assets: 0x0010-0x0013, 0x0020-0x0023). In other words, in this example, the priority of MPEG-H audio assets is set lower than that of MPEG-4 audio assets.

[0158] This diagram shows an example of operating two simulcast audios. The first simulcast audio has a simulcast_group_tag value of "0x00", and the second simulcast audio has a simulcast_group_tag value of "0x01". The first simul audio is composed of four MPEG-4 audio assets (0x0010, 0x0011, 0x0020, 0x0021) and one MPEG-H audio asset (0x0026). In other words, it is simul audio that includes both MPEG-4 and MPEG-H audio assets. The second simul audio is composed of four MPEG-4 audio assets (0x0012, 0x0013, 0x0022, 0x0023). In other words, it is simul audio consisting of only MPEG-4 audio assets. This diagram shows that the audio is made up of a first simulcast audio and a second simulcast audio.

[0159] 21 is a schematic diagram showing another example of the operation of simulcast audio according to this embodiment (including modifications), which shows an example of the operation of the simulcast_group_tag value in audio transmitted in a simulcast manner. In this diagram, an asset with an audio mode of "MPEG-H" is an MPEG-H audio asset, and is available in multiple languages ​​(Japanese and English). This audio asset is assigned a component_tag value of "0x0010." This value is assigned a smaller value than other audio assets (MPEG-4 audio assets: 0x0010, 0x0012, 0x0020, 0x0021). In other words, in this example, the priority of MPEG-H audio assets is set higher than that of MPEG-4 audio assets. The simulcast_group_tag value in this diagram is "0x00", and it consists of four MPEG-4 audio assets (0x0010, 0x0011, 0x0020, 0x0021) and one MPEG-H audio asset (0x0026). This diagram shows an example of operation with one simulcast audio.

[0160] Note that multiple simulcast audio may exist in Fig. 21, as in Fig. 20. For example, the four MPEG-4 audio assets in Fig. 20 (0x0012, 0x0013, 0x0022, 0x0023, simulcast_group_tag value "0x01") may be added to the audio asset in Fig. 21. Also, in Figure 20, as shown in Figure 21, there may be one simulcast audio (MPEG-H audio assets have a lower priority), and the four audio assets with the simulcast_group_tag value of "0x01" may be excluded.

[0161] In the above embodiment (including the modified examples), the "MH-audio component descriptor" may include preset() or interactive(). In the above embodiment, "voice" may be replaced with "audio." "Asset" may be replaced with "component," and "component" may be replaced with "asset."

[0162] Note that in the above-described embodiment, at least some of the broadcast station 1 (broadcast device 1), receiver 2, broadcast station server 3, and provider server 4, for example, the receiver 2's separator (Demux, TLV / MMT separator) 22, 22a, selector (audio asset selector) 231, audio decoder (decoder) 232, mixer 2331, downmixer 2332, switch 2333, DAC 2334, mixer 233, video decoder 241, presentation processor 242, input / output device 251, CPU 255, and communication chip 256, may be implemented by a computer. In this case, a program for implementing this control function may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into and executed by a computer system. Note that the "computer system" referred to here refers to a computer system built into the broadcast station 1, receiver 2, broadcast station server 3, or provider server 4, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above-mentioned programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system. Furthermore, some or all of the broadcast station 1, receiver 2, broadcast station server 3, and provider server 4 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the broadcast station 1, receiver 2, broadcast station server 3, and provider server 4 may be individually implemented as a processor, or some or all of them may be integrated into a processor. Furthermore, the integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.

[0163] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0164] Broadcasting System Sys Relay station Sa Broadcasting stations, broadcasting equipment 1 MPEG-H Encoder 11 MPEG-4 Encoder 111, 112, 113 Mux (Multiplexer) 12 Receiver 2 Broadcasting station server 3 Operator Server 4 Tuner 211 Demodulator 212 Separator (Demux, TLV / MMT separation section) 22, 22a Selector (audio asset selection section) 231 Audio decoder (decoder part) 232 MPEG-H Audio Core Decoder 232-1 MPEG-4 Decoder 232-2 Mixer section 2331 Downmixer 2332 MPEG-H Audio Renderer 233-1 Mixer 233-2 Switch section 2333 DAC section 2334 Mixer 233 Speaker 234 Video decoder 241 Presentation Processor 242 Display 243 Input / output device (external output I / F) 251 Auxiliary storage 252 ROM 253 RAM 254 CPU 255 Communication chip 256

Claims

1. a receiving unit for receiving a broadcast including MPEG-H audio in the audio component; an acquisition unit that acquires information indicating the presence of MPEG-H audio from the broadcast; a selector for selecting an audio component based on the information; a decoding unit that decodes the audio data of the selected audio component; the broadcast includes a descriptor describing information about the object-based audio; The acquisition unit acquires a descriptor that describes information about the object-based audio. Receiver.

2. The descriptor describing information related to the object-based audio includes a profile indicating a level or a set of functions indicating processing capabilities of MPEG-H audio; The selection unit selects an audio component according to the capabilities of a receiver based on the level or the profile.

2. The receiver of claim 1.

3. A transmitting device for transmitting at least an audio component, When MPEG-H audio is included as the audio component, information indicating the presence of MPEG-H audio; and a descriptor describing information about the object-based audio. Transmitting device.

4. a transmitting device according to claim 3; The receiver according to claim 1, Broadcasting system.

5. A receiving method in a receiver, comprising: The receiver includes: Receiving a broadcast that includes MPEG-H audio in the audio component; obtaining information indicating the presence of MPEG-H audio from the broadcast; selecting an audio component based on said information; Decode the audio data for the selected audio component; the broadcast includes a descriptor describing information about the object-based audio; Obtain a descriptor that describes information about the object-based audio. Receiving method.

6. A transmission method in a transmission device that transmits at least an audio component, comprising: The transmitting device When MPEG-H audio is included as the audio component, information indicating the presence of MPEG-H audio; and a descriptor describing information about the object-based audio. Sending method.

7. 1. A method in a system including a transmitter and a receiver, comprising: Send it with at least an audio component, When MPEG-H audio is included as the audio component, information indicating the presence of MPEG-H audio; a descriptor describing information about the object-based audio; Send it with more included, Receiving a broadcast that includes MPEG-H audio in the audio component; From the above broadcast: information indicating the presence of MPEG-H audio; a descriptor describing information about the object-based audio; Get selecting an audio component based on the information indicating the presence of MPEG-H audio; Decoding the audio data of the selected audio component; method.

Citation Information

Patent Citations

  • Method and system for generating and rendering object-based audio with conditional rendering metadata

    JP2016521380A

  • Receiver

    JP2019110571A

  • Selection of next generation coded audio data for transport

    JP2019504341A

  • Interactive audio metadata handling

    US20170280169A1

  • Method and apparatus for processing of auxiliary media streams embedded in a mpegh 3D audio stream

    US20200395027A1