Receiving device
Patent Information
- Application Number
- JP2022119736
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-09-22
- Estimated Expiration
- 2042-07-27
AI Technical Summary
Existing receivers are unable to appropriately process object-based audio signals such as AC-4, leading to inadequate audio reproduction based on their capabilities.
A broadcasting system that includes a broadcasting device multiplexing identification information for AC-4 audio components, a receiver that separates and decodes audio data based on receiver capabilities using MMT descriptors, and a control unit to select appropriate audio components for decoding.
Enables appropriate audio reproduction on receivers by identifying and processing AC-4 audio components according to their capabilities, ensuring optimal playback.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a broadcasting system, a receiver, a receiving method, and a program. [Background technology]
[0002] In broadcasting, the use of object-based audio signals such as AC-4 (ETSI TS 103 190) is being considered. Patent Document 1 describes that an audio signal of an audio object (an object-based audio signal) is treated as a priority signal and is played back with priority. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2021-124719 A Summary of the Invention [Problem to be solved by the invention]
[0004] In Patent Document 1, an audio encoding device encodes an object-based audio signal, and an audio decoding device performs decoding processing on the resulting bitstream. However, some receivers that receive broadcasts do not have the capability to process object-based audio signals, and it is desirable to be able to perform appropriate audio reproduction according to the capabilities of the receiver.
[0005] The present invention has been made in view of the above circumstances, and provides a broadcasting system, a receiver, a receiving method, and a program capable of performing appropriate audio reproduction according to the capabilities of the receiver. [Means for solving the problem]
[0006] The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is a broadcasting system that broadcasts including an audio component, comprising a broadcasting device that broadcasts broadcast waves in which identification information indicating whether the audio component is an AC-4 audio component is multiplexed into a multiplexing layer, and a receiver that receives the broadcast waves, the receiver comprising a separation unit that acquires the identification information from the broadcast waves at the multiplexing layer, a control unit that selects an audio component according to the capabilities of the receiver based on the identification information, and a decoding unit that decodes the audio data of the selected audio component.
[0007] Another aspect of the present invention is the above-mentioned broadcasting system, in which the identification information is placed in an audio component descriptor, which is an MMT (MPEG Media Transport) descriptor and is a descriptor that describes parameters related to an audio signal among program elements.
[0008] Another aspect of the present invention is the above-mentioned broadcasting system, wherein the broadcast includes a plurality of audio components including an AC-4 audio component and an MPEG-4 audio component, the control unit selects one audio component from the plurality of audio components based on the identification information placed in the audio component descriptor, the separation unit separates the selected audio component, and the decoding unit decodes the audio data of the selected audio component.
[0009] Another aspect of the present invention is the broadcasting system described above, wherein when the audio component is an AC-4 audio component, the broadcasting device broadcasts the broadcast wave in which information indicating the sound material contained in the AC-4 audio component is multiplexed into a multiplexing layer, the separation unit acquires the information indicating the sound material from the broadcast wave at the multiplexing layer, the control unit selects a combination of the sound materials based on the information indicating the sound materials, and the decoding unit decodes the audio data of the selected sound materials.
[0010] Another aspect of the present invention is a receiver comprising a separation unit that acquires identification information indicating whether an audio component is an AC-4 audio component from a broadcast wave at a multiplexing layer, a control unit that selects an audio component according to the capabilities of the device based on the identification information, and a decoding unit that decodes the audio data of the selected audio component.
[0011] Another aspect of the present invention is a receiving method comprising the steps of: acquiring, at a multiplexing layer, identification information from a broadcast wave indicating whether an audio component is an AC-4 audio component; selecting an audio component according to the capabilities of the device based on the identification information; and decoding audio data of the selected audio component.
[0012] Another aspect of the present invention is a program for causing a receiver computer having a separation unit that acquires identification information from a broadcast wave at a multiplexing layer indicating whether an audio component is an AC-4 audio component, and a decoding unit that decodes the audio data of a selected audio component, to function as a control unit that selects an audio component according to the capabilities of the device based on the identification information. Effect of the Invention
[0013] According to the present invention, appropriate audio reproduction can be performed according to the capabilities of the receiver. [Brief description of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an example of a configuration of a broadcasting system Sys according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram illustrating an example of a broadcasting system Sys according to the embodiment. [Diagram 3] FIG. 2 is a diagram showing a comparative example of the broadcasting system Sys according to the embodiment. [Figure 4]FIG. 2 is a diagram showing another example of the broadcasting system Sys according to the embodiment. [Diagram 5] 1 is an explanatory diagram illustrating an overview of a broadcasting system Sys according to the embodiment. FIG. [Figure 6] FIG. 2 is a diagram illustrating an example of a protocol stack structure according to the embodiment. [Figure 7] FIG. 11 is a diagram showing the data structure of an MPT according to the embodiment. [Figure 8] 2 is a schematic diagram showing a hardware configuration of a receiver 2 according to the embodiment. FIG. [Figure 9] 2 is a schematic diagram illustrating an example of a flow of signal processing in a receiver according to the embodiment. FIG. [Figure 10] FIG. 11 is a diagram showing an example of an audio switching menu according to the embodiment. [Figure 11] 11 is a schematic diagram showing an example of the structure of an MH-audio component descriptor according to the embodiment. FIG. [Figure 12] 13 is a table showing example values of nga_type according to the embodiment. [Figure 13] 13 is a table showing example values of nga_level according to the embodiment. [Figure 14] 13 is a schematic diagram showing another example of the structure of the MH-audio component descriptor according to the embodiment. FIG. [Figure 15] 13 is a table showing example values of stream_content according to the embodiment. [Figure 16] 13 is a table showing another example of the value of nga_level according to the embodiment. [Figure 17] 1 is a schematic diagram showing an example of the structure of an MH-AC-4 audio descriptor according to the embodiment. FIG. [Figure 18] FIG. 11 is a schematic diagram showing an example of the structure of presentation() according to the embodiment. [Figure 19] FIG. 13 is a schematic diagram showing another example of the structure of presentation() according to the embodiment. [Figure 20] 11 is a flowchart showing a detailed example of switching according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0016] [System Configuration] FIG. 1 is a diagram showing an example of the configuration of a broadcasting system Sys according to an embodiment of the present invention. The broadcasting system Sys includes a broadcasting device 1 of a broadcasting station (referred to as "broadcasting station 1"), a relay station Sa, a receiver 2, a broadcasting station server 3, and a carrier server 4. The broadcasting is, for example, terrestrial digital broadcasting, but may also be, for example, advanced BS (Broadcasting Satellites) digital broadcasting or advanced wideband CS (Communication Satellites) digital broadcasting. The present invention is not limited to these broadcasting, and the broadcasting may also be broadcasting that does not use the relay station Sa. The broadcasting may also be wired broadcasting such as cable television. The relay station Sa is, for example, a digital relay station, but may also be a broadcasting satellite.
[0017] In the broadcasting system Sys, a broadcasting station 1 transmits digital broadcasting signals, application control information, presentation control information, etc., via airwaves. A service provider provides metadata and video content related to programs from a provider server 4. The application control information notifies receivers compatible with this system of applications that are linked to programs, and also sends commands and control information for starting and stopping them. The control information regarding presentation transmits control information regarding whether or not an application can be presented, and whether or not an application can be superimposed on a single TV screen with a broadcast program. The broadcast station operates a broadcast station server 3 in the broadcast system Sys. The broadcast station server 3 provides metadata such as program title, program ID, program summary, cast, broadcast date and time, etc. The information provided by the broadcast station to the service provider is provided through an API (Application Programming Interface) provided by the broadcast station server 3.
[0018] A service provider is a party that provides services through the broadcasting system Sys, and produces and distributes content and applications for providing the services, and operates the broadcasting station server 3 to realize each individual service. Here, the services include broadcasting and communication integration services that integrate broadcasting and communication. The broadcasting station server 3 "manages and distributes applications" by sending them to the receiver 2. As a "server for each service," the broadcasting station server 3 provides server functions for implementing individual services (MPEG-H service, VOD program recommendation service, multilingual subtitle service, etc.).
[0019] MPEG-H was approved by the ISO / IEC Moving Picture Experts for a digital container standard, a video compression standard, an audio compression standard, and two conformance test standards. It is a series of standards under development by the Moving Picture Experts Group (MPEG). MPEG-H services can include the audio component of AC-4 audio. AC-4 audio, for example, enables object-based audio. In object-based audio, an "object" is each of the sound materials that make up a program, such as music or human voices. In object-based audio, an audio signal is recorded for each sound material, making it possible to control the audio for each material. In addition, when playing back on receiver 2, it is also possible to play back the program according to the actual position of the speakers, based on the playback position information of the material.
[0020] The broadcasting station server 3 not only realizes the functional aspects of these services, but also transmits the content that constitutes the services (AC-4 audio data, VOD content, subtitle data, etc.). As a "repository," the broadcasting station server 3 registers applications for the broadcasting system Sys for distribution, and provides and searches for a list of available applications in response to inquiries from the receiver 2.
[0021] The receiver 2 includes a receiver equipped with a function for realizing broadcast and communication cooperation services in addition to a function for receiving existing digital broadcasts. The receiver 2 has the following functions in addition to a broadband network connection function. - Function to execute applications in response to application control signals from broadcasting - Function to present information by linking broadcasting and communication - Device linking function Here, the terminal includes, for example, a user terminal such as a smartphone, a smart speaker, etc. The terminal link function of the receiver 2 accesses broadcast resources such as program information in response to a request from another terminal, and calls receiver functions such as playback control. An example of the application includes a digital mixer for AC-4 audio. A user (also called a "receiver") can use the digital mixer received from the business server 4 to adjust the volume or effects of the audio signal for each sound material, or adjust the balance between the sound materials. These adjustments can also be made for each speaker.
[0022] More specifically, the receiver 2 has the following functions. The receiver 2 has a "broadcast reception and playback" function that receives broadcast radio waves, selects a specific broadcast service, and synchronously plays back the video, audio, subtitles, and data broadcast that constitute the service. The receiver 2 has a "communication content reception and playback" function that accesses video content stored on a server on the communication network (e.g., the operator server 4), receives it as VOD streaming, and synchronously plays back the video, audio, and subtitles that make up the content. The receiver 2 has an "application control" function that, based on application control information obtained from a server on the communications network or from broadcast signals, acts on the application engine primarily with regard to managed applications, and controls and manages the life cycle and events of each application. The receiver 2 has an "application engine" function that acquires and executes applications. This function is realized by, for example, an HTML5 browser. The receiver 2 has a "presentation synchronization control" function that controls the stream presentation synchronization of video, audio, etc. received by broadcasting and video, audio, etc. received by streaming. The receiver 2 has an "application launcher" function that mainly includes a navigation function for the user to select and launch non-broadcast managed applications.
[0023] FIG. 2 is a diagram showing an example of a broadcasting system Sys according to the present embodiment. The receiver 2 in FIG. 2 is a receiver compatible with AC-4 audio.
[0024] Broadcast station 1 multiplexes the video and audio signals and transmits the multiplexed signal. The multiplexing method used is MMT (MPEG Media Transport)·TLV (Type Length Value). Broadcasting station 1 generates and transmits an audio signal A1 (also referred to as an "advanced audio signal") that includes both MPEG-4 audio (channel-based audio: e.g., audio A12 to audio A14) signals and AC-4 audio (object-based audio: audio A11) signals. In this way, the broadcasting station 1 transmits the advanced audio signal and the MPEG-4 audio signal in parallel.
[0025] More specifically, the broadcasting station 1 treats AC-4 audio as an audio component (asset) and generates an advanced audio signal A1 by multiplexing this audio component with each audio component of MPEG-4 audio. This multiplexing is performed after the audio data string of each audio component is encoded. The broadcasting station 1 transmits a broadcast wave on which this advanced audio signal A1 is multiplexed. In a table describing asset information, the broadcasting station 1 describes in a descriptor whether the asset (audio component) is an AC-4 audio (object-based audio: for example, audio A11) component. Note that the descriptor indicating whether the asset is an AC-4 audio component may be a descriptor indicating whether an advanced audio signal A1 exists, or may indicate whether an AC-4 audio signal is included. In MMT, components such as video and audio are defined as assets. An example of audio A11 in AC-4 is an audio with a maximum channel audio of 11.1ch, dialogue audio in Japanese or English, and commentary audio in Japanese or English.
[0026] The receiver 2 includes a tuner 211, a Demux (demultiplexer) 22, a selector 231, an audio decoder (decoding unit) 232, a mixer 233, and a video decoder 241. The detailed configuration of the receiver 2 will be described later.
[0027] The tuner 211 receives broadcast waves via an antenna, and tunes (selects) a channel selected based on a user operation. The tuned signal is demodulated and input to the Demux 22 as data. The Demux 22 separates the input data into a video data string, an audio data string, a superimposed text data string, a subtitle data string, etc. The separated audio data string is output to the selector 231. The separated video data string is output to the video decoder 241.
[0028] Here, the Demux 22 separates the audio data string into audio data strings of each audio component, that is, the AC-4 audio A11 and the MPEG-4 audio A12, A13, and A14. More specifically, the Demux 22 determines whether an AC-4 audio signal exists based on a descriptor in a table describing asset information. When the Demux 22 determines that an AC-4 audio signal exists, and the receiver 2 has an AC-4 audio decoding capability, the Demux 22 separates the audio A11, A12, A13, and A14 from the data of the advanced audio signal. When the Demux 22 determines that an AC-4 audio signal does not exist, the Demux 22 separates only the audio A12, A13, and A14. Note that the Demux 22 may separate only the audio A11, A12, A13, and A14 that are to be decoded from among the audio A11, A12, A13, and A14.
[0029] The audio data string of each audio component output from the Demux 22 is input to the selector 231. The selector 231 selects the audio data string of the audio component in accordance with a user operation or the capabilities of the receiver 2. The capabilities of the receiver 2 include, for example, the number of channels that can be decoded simultaneously, or the type and capabilities of speakers that can be played back. The selector 231 outputs the selected audio data string to the audio decoder 232. The audio decoder 232 decodes the audio data string of the audio component input from the selector 231 . If the audio data string decoded by the audio decoder 232 is an AC-4 audio data string, the mixer 233 synthesizes the audio for each sound material and performs downmix processing. The downmixed audio data string is converted into audio and output from a speaker. If the audio data string decoded by the audio decoder 232 is an MPEG-4 audio data string, the audio data string is converted into audio and output from a speaker. In other words, audio synthesis for each sound material and downmix processing are not performed on the MPEG-4 audio data string.
[0030] The video data stream output from the Demux 22 is input to a video decoder 241 . The video decoder 241 decodes the input video data string. The decoded video data string undergoes color space conversion processing as necessary and is used to display the video on a display. The superimposed text data string and subtitle data string separated by the Demux 22 are decoded by a superimposed text decoder and subtitle decoder (not shown), respectively, and the decoded character strings are superimposed on the video.
[0031] As described above, the receiver 2 according to this embodiment obtains information indicating that the component is an AC-4 audio in a multiplex layer in a broadcast that includes AC-4 audio (audio) in the audio component. The receiver 2 selects an audio component according to the capabilities of the receiver 2 itself, and therefore can perform appropriate audio playback according to the capabilities of the receiver 2.
[0032] FIG. 3 is a diagram showing a comparative example of the broadcasting system Sys according to the present embodiment. This diagram shows an example in which broadcast station C1 transmits only AC-4 audio, which is not multiplexed with MPEG-4 audio. In this example, AC-4 audio is used as the only audio component. Therefore, the table describing the asset information does not include a descriptor indicating whether or not it is an AC-4 audio component. In this case, Demux C22 can obtain AC-4 audio as an audio component (audio configuration), but cannot determine whether it can be processed (level, etc.). The audio data string output from Demux C22 is decoded by audio decoder C232 and output to mixer C233.
[0033] In contrast to the comparative example in Fig. 3, the broadcasting station 1 in this embodiment transmits AC-4 audio signals and MPEG-4 audio signals in parallel. The receiver 2 first judges whether the signal is an AC-4 audio signal by the descriptor, and selects AC-4 or MPEG-4 audio according to its audio decoding capability. This allows the receiver 2 to play back audio in the appropriate format, either AC-4 or MPEG-4, depending on the capabilities of the receiver 2 itself.
[0034] 4 is a diagram showing another example of the broadcasting system Sys according to the present embodiment, in which a receiver 2a is not compatible with AC-4. The receiver 2a in this figure includes a tuner 211, a Demux 22a, a selector 231a, an audio decoder 232a, and a video decoder 241. In this figure, the same functional units as those in the receiver 2 in Figure 3 are given the same reference numerals, and their description will be omitted.
[0035] The Demux 22a separates the input data into a video data string, an audio data string, a superimposed text data string, a subtitle data string, etc. The separated audio data string is output to the selector 231a. The separated video data string is output to the video decoder 241. Here, the Demux 22a separates the audio data string into MPEG-4 audio A12, A13, and A14, and audio data strings of each audio component. More specifically, the Demux 22a determines whether each audio component is an AC-4 audio component or not by a descriptor in a table describing asset information. The Demux 22a determines that audio A11, which is an AC-4 audio component, cannot be played. Demux 22 separates audio A12, audio A13, and audio A14 from the data of advanced audio signal A1. Note that such signal selection (selection of an MPEG-4 audio signal, or selection of a simulcast of only an MPEG-4 audio signal) may be performed by the selector 231.
[0036] FIG. 5 is an explanatory diagram illustrating an outline of the broadcasting system Sys according to this embodiment. In the broadcasting system Sys, the broadcasting station 1 includes an AC-4 encoder 11, MPEG-4 encoders 111-113, and a Mux (multiplexer) 12. The broadcasting station 1 also includes other functional units necessary for broadcasting. Although the figure shows an example of three MPEG-4 encoders, the broadcasting station 1 may include two or less MPEG-4 encoders, or may include four or more MPEG-4 encoders.
[0037] Receiver 2 includes a Demux 22, a selector 231, an AC-4 decoder 232-1, an MPEG-4 decoder 232-2, an AC-4 renderer 233-1, and a mixer 233-2. In this figure, the same functional units as those in receiver 2 in Figure 2 are denoted by the same reference numerals. Note that AC-4 decoder 232-1 and MPEG-4 decoder 232-2 correspond to audio decoder 232 in Figure 2. AC-4 renderer 233-1 and mixer 233-2 correspond to mixer 233 in Figure 2.
[0038] In the broadcasting station 1, data of background sounds (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), commentary audio (Japanese), and commentary audio (English) are input as AC-4 audio material to an AC-4 encoder 11. In addition, data of 7.1ch audio including Japanese dialogue, stereo audio including Japanese dialogue, and stereo data including English dialogue are input as MPEG-4 audio material to MPEG-4 encoders 111, 112, and 113, respectively.
[0039] The AC-4 encoder 11 outputs an AC-4 audio stream St1 by encoding the input audio. This stream is also called an AC-4 stream, and is one elementary stream including multiple audio objects (background sound, dialogue (Japanese), dialogue (English), audio commentary (Japanese), and audio commentary (English)). MPEG-4 encoders 111, 112, and 113 encode the input audio and output MPEG-4 audio streams St2, St3, and St4, respectively.
[0040] A video stream, SI (Signaling Information), an MPEG-H audio stream St1, and MPEG-4 audio streams St2, St3, and St4 are input to Mux 12. Mux 12 multiplexes these data. The multiplexed data is modulated, and the modulated signal is broadcast as a broadcast wave.
[0041] The broadcast waves received by the receiver 2 are demodulated, and the demodulated data is input to the Demux 22. The Demux 22 separates the input data into a video stream, SI, AC-4 audio stream St1, and MPEG-4 audio streams St2, St3, and St4. The AC-4 audio stream St1, and the MPEG-4 audio streams St2, St3, and St4 are input to a selector 231. A descriptor indicating whether or not an audio component is an AC-4 audio component is extracted from the MPT (MMT Package Table) of the SI.
[0042] The selector 231 determines whether an AC-4 audio signal is present based on the extracted descriptors, and if an AC-4 audio signal is present, the selector 231 outputs the AC-4 audio stream St1 to the AC-4 decoder 232-1. The selector 231 outputs the MPEG-4 audio stream St2, St3 or St4 to the MPEG-4 decoder 232-2 based on the MPT.
[0043] The AC-4 decoder 232-1 extracts data of background sounds (22.2ch / 11.1ch), dialogue (Japanese), dialogue (English), commentary audio (Japanese), and commentary audio (English) by decoding the AC-4 audio stream St1. The AC-4 renderer 233-1 is an audio renderer for AC-4 audio, and performs rendering processing (including down-conversion or up-conversion) on the audio data extracted by the AC-4 decoder 232-1, and outputs the result to the mixer 233. The MPEG-4 decoder 232-2 decodes the MPEG-4 audio streams St2, St3, or St4 to extract 7.1ch audio including Japanese subtitles, stereo audio including Japanese subtitles, and stereo data including English subtitles, and outputs them to the mixer 233. The mixer 233-2 synthesizes the audio of the input data, and the synthesized audio is output from each speaker or headphones, etc.
[0044] [Regarding the control information of the broadcast wave] The broadcast wave according to this embodiment will be described. In the broadcast wave, the control information is superimposed and transmitted by each broadcaster on its broadcast signal, which is a TLV stream. The control information includes TLV-SI (TLV-Signaling Information) related to the TLV multiplexing method and MMT-SI (MMT-Signaling Information) related to MMT, which is a media transport method. Hereinafter, a "component" (of video or audio) will also be referred to as an "asset".
[0045] [The protocol stack structure of the system using MMT] In a system using MMT, an example of the protocol stack structure in which control information is arranged will be described. FIG. 6 is a diagram showing an example of the protocol stack structure according to this embodiment. As shown in this figure, the protocol stack used in the broadcast system is TMCC (Transmission and Multiplexing Configuration It is composed of MMT-SI, time information, encoded video data, encoded audio data, encoded subtitle data, MMT-SI, applications written in HTML5 standard (also simply called apps), EPG (Electronic Program Guide), content download data, etc. The video and audio signals of the broadcast program are encoded in MFU (Media Fragment Unit) / MPU. The MFU / MPU is then loaded onto the MMTP payload and packetized by broadcast station 1 as an MMTP packet, which is then transmitted by broadcast station 1 as an IP packet. For the transmission of data content, data is packetized by broadcast station 1 as an MMTP packet, which is then transmitted by broadcast station 1 as an IP packet. When the IP packet thus configured is broadcast using a broadcast transmission path, it is transmitted by broadcast station 1 in the form of a TLV packet. One IP packet or one header-compressed IP packet is transmitted by broadcast station 1 as one TLV packet.
[0046] Furthermore, the protocol stack used in the broadcasting system provides two types of control information: MMT-SI and TLV-SI. MMT-SI is control information that indicates the configuration of a broadcast program, etc. MMT-SI is in the form of an MMT control message, placed on the MMTP payload by broadcasting station 1, packetized as an MMTP packet, and transmitted by broadcasting station 1 in an IP packet. TLV-SI is control information related to multiplexing of IP packets, and provides information for channel selection and correspondence information between IP addresses and services.
[0047] TMCC is a type of control information that is inserted into transmission frames and transmitted in a hierarchical modulation method that specifies the modulation method and error correction method for each unit (slot) of the signal on the transmission path. HEVC (High Efficiency VIDeo Coding) is a method of encoding video signals. VVC (Versatile Video Codec) may also be used as a method of encoding video signals. AAC (Advanced Audio Coding), ALS (Audio Lossless Coding), and AC-4 are methods of encoding audio signals. UDP / IP (User Datagram Protocol / Internet Protocol) is one of the protocols used in communication. TLV (TYPE LENGTH VALUE) is one of the methods of multiplexing data. TLV consists of three parts for encoding data: data type, length, and value.
[0048] <Message type and identification> MMT-SI includes messages, tables, and descriptors. The messages include a Package Access (PA) message, an M2 section message, a CA message, an M2 short section message, a data transmission message, and a message set by the operator. The MMT-SI messages used in broadcasting are as follows:
[0049] The "PA message" carries the PLT and MPT to indicate the entry point of the service. The "M2 Section Message" transmits the section extension format of MPEG-2 Systems. The "CA message" transmits information regarding the conditional access method. The "M2 Short Section Message" transmits the MPEG-2 Systems section short format. The "data transmission message" transmits a table relating to data transmission.
[0050] The TLV-SI table used in broadcasting is as follows.
[0051] “TLV-NIT(Network Information Table for "TLV" transmits information that associates transmission path information, such as modulation frequency, with broadcast programs in TLV packet transmission. The "AMT (Address Map Table)" transmits information that associates a service identifier, which identifies a broadcast program number, with an IP packet. The MMT-SI table used in broadcasting is as follows. "MPT (MMT Package Table)" provides information that constitutes a package, such as a list of assets and their locations. "PLT (Package List Table)" indicates a list of packet IDs that transmit PA messages including MPTs of services provided as broadcast services. An "ECM (Entity Control Message)" transmits common information consisting of program information (information about the program and a key for descrambling, etc.) and control information (a forced on / off command for the decoder's scrambling function).
[0052] "EMM (Entity Management Message)" transmits individual information including contract information for each subscriber and a work key for decrypting common information. The "CAT (MH) (Conditional Access Table)" specifies the packet identifier of the MMTP packet that transmits individual information among the related information that constitutes the conditional reception broadcasting. The "MH-EIT (MH-Event Information Table)" transmits information related to the program, such as the program name, broadcast date and time, and a description of the content. The "MH-AIT (MH-Application Information Table)" transmits dynamic control information related to an application and additional information required for execution. "MH-BIT (MH-Broadcaster Information Table)" is used to present information about broadcasters present on the network.
[0053] The "MH-SDTT (MH-Software DownLoad TriggerTable)" transmits notification information such as the download service ID, schedule information, and the type of receiver to be updated. The "MH-SDT (MH-Service Description Table)" transmits information about the organized channels, such as the name of the organized channel and the name of the broadcasting company. The "MH-TOT (MH-Time Offset Table)" transmits the current date and time, as well as the difference between the actual time and the time displayed to the human system. The "MH-CDT (MH-Common Data Table)" transmits data that is commonly required for all receivers, such as operator logo marks, and is assumed to be stored in non-volatile memory. The "DDMT (Data Directory Management Table)" provides a directory structure for the files that make up an application.
[0054] "DAMT (Data Asset Management Table)" provides the configuration of MPUs in an asset and version information for each MPU. "DCCT (Data Content Configuration Table)" provides configuration information of files as data contents. "EMT (Event Message Table)" is used to transmit information about event messages.
[0055] <MMTパッケージテーブル> The MPT (MMT Package Table) provides information that constitutes a package, such as a list of assets and their locations on the network. FIG. 7 is a diagram showing the data structure of the MPT according to this embodiment. "table_id" (table identifier) is an 8-bit field that identifies each table. "version" is an area in which the version number of the table is written. "Length" (table length) is an area in which the number of data bytes following this field is written.
[0056] "MMT_package_id_length" indicates the length of the package ID byte in bytes. "MMT_package_id_byte" indicates the package ID. It shall be the same value as the service ID for identifying the service. "MPT_descriptors_length" indicates the length of the MPT descriptor area in bytes. "MPT_descriptors_byte" (MPT descriptor area) is an area that stores the MPT descriptors. If the program is a multi-view program, the MPT descriptor area includes the MH-Component Group Descriptor (MH-Component_Group_Descriptor()). On the other hand, if the program is not a multi-view program, the MPT descriptor area does not include the MH-Component Group Descriptor. "number_of_assets" indicates the number of assets for which this table provides information.
[0057] The MPT has an area that describes one or more assets. This area contains the following fields for each asset: "identifier_type" indicates the ID system of the MMTP packet flow. If the ID system indicates an asset ID, a specific value (0x00) is set. "asset_id_scheme" (asset ID format) indicates the format of the asset ID. For "asset_id", the receiver 2 uses the component_tag value for reception operations. The receiver 2 uses the component_tag value to identify the asset. "asset_id_length" indicates the length of the asset ID byte in bytes. "asset_id_byte" (asset ID byte) indicates the asset ID.
[0058] "asset_type" indicates the type of asset. The asset type may be, for example, hvc1, which indicates video data encoded in HEVC, mp4a, which indicates audio data encoded in MPEG-4 audio, or mha1, mha2, mhm1, or mhm2, which indicate audio data encoded in MPEG-H audio, or ac-4, which indicates audio data encoded in AC-4.
[0059] "asset_clock_relation_flag" (clock information flag) indicates whether or not the asset has a clock information field. "location_count" (number of locations) indicates the number of location information items of the asset. "MMT_general_location_info" (location information) indicates the location information of the asset. "asset_descriptors_length" indicates the total byte length of the subsequent descriptors. "asset_descriptors_byte" (asset descriptor area) is an area for storing asset descriptors.
[0060] <Descriptor types and identification> The TLV-SI descriptors used in broadcasting are as follows. The "Service List Descriptor" is a description of a list of organized channels and their types. "Satellite Delivery System Descriptor" "Satellite Channel Descriptor" is a description of the physical conditions of the satellite transmission path. The "System Management Descriptor" is an identification such as broadcast / non-broadcast. A "Network Name Descriptor" is a description of a network name.
[0061] The MMT-SI descriptors used in broadcasting are as follows: The "remote control key descriptor" uniquely provides a service for assignment to one-touch keys on the receiver remote control (remote controller). The "asset group descriptor" provides the group relationship of assets and their priority within the group. The "MPU timestamp descriptor" provides the presentation time of the MPU. The "access control descriptor" identifies the conditional access method. The "scrambling scheme descriptor" identifies the scrambling subsystem. The "Emergency Information Descriptor (MH)" provides a description of the necessary information and functions as an emergency alert signal. The "MH-Event Group Descriptor" describes grouping information for multiple events. The "MH-Service List Descriptor" describes a list of organized channels and their types. The "MH-Short Event Descriptor" describes the program name and a brief description of the program. The "MH-extended event descriptor" describes detailed information about a program.
[0062] The "video component descriptor" describes parameters, explanations, etc. relating to the video signal among the program element signals. The "MH-stream identification descriptor" is used to identify individual program element signals. The "MH-Content Descriptor" describes the program genre. The "MH-Parental Rate Descriptor" describes the age restriction for permitted viewing. The "MH-audio component descriptor" describes parameters relating to audio signals among program elements. The "MH-Target Area Descriptor" describes the target area. The "MH-series descriptor" describes series information spanning multiple events. The "MH-SI transmission parameter descriptor" describes the parameters of SI transmission (such as the cycle group and retransmission cycle). The "MH-broadcaster name descriptor" describes the broadcaster name. The "MH-service descriptor" describes the name of the channel and the name of its service provider.
[0063] The "MH-Data Encoding Method Descriptor" is used to identify the data encoding method. The "UTC-NPT reference descriptor" conveys the relationship between NPT and UTC. The "Event Message Descriptor" conveys information about the event message in general. The "MH-Local Time Offset Descriptor" describes the difference in time between the actual time and the time displayed in the human system when daylight saving time is in effect. The "MH-logo transmission descriptor" describes the character string for the simple logo, pointing to the logo in CDT format, etc. The "MPU extended timestamp descriptor" provides the decoding time of an access unit in the MPU, etc. The "MPU download content descriptor" describes attribute information of content that is downloaded using the MPU. The "MH-Application Descriptor" describes information about an application. The "MH-transmission protocol descriptor" specifies a transmission protocol and describes location information of an application that depends on the transmission protocol. The "MH-Simple Application Location Descriptor" describes in detail where the application can be obtained.
[0064] The "MH-application boundary authority setting descriptor" describes the application boundary settings and the broadcast resource access authority settings for each area (URL). The "link destination PU descriptor" describes information about the link destination presentation unit. The "application service descriptor" describes entry information of an application related to the service. The "MPU node descriptor" indicates that the MPU corresponds to a directory node defined in the data directory management table. The "PU configuration descriptor" indicates a list of MPUs that make up a presentation unit. The "MH-hierarchical coding descriptor" describes information for identifying a hierarchically coded video stream component.
[0065] The "content copy control descriptor" is placed when indicating control information regarding digital copying for the entire service or when describing the maximum transmission rate. The "content usage control descriptor" is used to describe control information related to storage and output for the program. It is also used to specify whether or not to operate the "number-limited copying" for the program or asset. The "associated broadcaster descriptor" indicates the identification value of the BS / broadband CS digital broadcasting broadcaster and the terrestrial digital broadcasting series required for accessing the NVRAM. The "multimedia service information descriptor" describes detailed information about each content of a multimedia service, such as the presence or absence of data content and the presence or absence of subtitles. The "emergency news descriptor" indicates that an emergency news flash related to safety and security (emergency earthquake alert, special news, super flash news) is being broadcast. The "MH-CA contract information descriptor" describes information for confirming that a service or event can be reserved. The "MH-CA service descriptor" indicates the composition channel of the entity that operates the automatic display message and describes the display control information of the message. The "MH-AC-4 audio descriptor" describes the parameters related to the audio components of AC-4.
[0066] <Arrangement of MH-Audio Component Descriptor and MH-AC-4 Audio Descriptor> The "MH-Audio Component Descriptor" and the "MH-AC-4 Audio Descriptor" are arranged in the following table. ·MPT (Asset Descriptor Area) ·MH-EIT[p / f actual] (MH-EIT[p / f]) ·MH-EIT[schedule actual basic] (MH-EIT[schedule basic])
[0067] "MPT" is stored in the "PA message". "MH-EIT[p / f]" is the time-series information regarding the current and next events, where the former is called present and the latter is called following. "MH-EIT[p / f actual]" and "MH-EIT[schedule actual basic]" are tables that describe events included in the services operating in the self-TLV stream and are stored in the "M2 section message".
[0068] "MH-AIT" is also a table that indicates control information that indicates the life cycle and constraints of an application. "MMT" is also a multiplexing method that enables integrated transmission over multiple transmission paths. "MP4 ACC" is an audio coding method defined by ISO / IEC 14496-3. "MP4 ALS" (ALS: Audio Lossless Coding) is an audio lossless coding method defined by ISO / IEC 14496-3. "MPT" is an abbreviation for MMT Package Table. "MPT" is a table that provides information that constitutes a service (package), such as a list of assets and their locations. It has elements and attributes that indicate specific information. "Tables" are stored in messages and transmitted in MMTP packets. Messages that store "tables" are determined according to the table. "Package" refers to a unit of content in the MMT standard. "Messages" store tables and descriptors. Messages are stored in MMTP payloads and transmitted using MMTP packets.
[0069] "SI information" is also information that describes the contents of the multiplexed information, identification information, etc. Receiver 2 is, for example, a "digital terrestrial broadcast receiver" and has the function of selecting and demodulating the receiving channel from the IF signal, selecting and decoding the desired program, and outputting the baseband signal. However, the receiver 2 may also be an "advanced BS digital broadcast receiver," which in addition to having these functions is a device capable of receiving advanced BS digital broadcasts in the 11.7 GHz to 12.75 GHz frequency band. The receiver 2 is also called an STB or IRD. An "item" is the smallest unit of transmission that constitutes an MPU in application data transmission based on the MMT transmission method. An "item" is equivalent to a file. An "MPU" is a transmission unit that is comprised of a collection of items contained within one component. It is envisioned that an "MPU" will be operated in correspondence with a presentation unit (PU), update unit, or storage control unit.
[0070] A "component" (asset) is a unit that has the same packet ID in one IP data flow. In the MPT, it is referred to as an asset. A "component" is identified by the component_tag, which will be described later. The application set being transmitted is switched depending on the data event. An "asset" is a transmission unit for video, audio, etc. multiplexed using the MMT method. An "asset type" is the type of content being transmitted in each asset. "Simul audio" is the simultaneous transmission of multiple different audio modes within the same event. An "event" is a collection of streams with fixed start and end times within the same service (programming channel), such as news or dramas.
[0071] [Receiver 2 hardware configuration] FIG. 8 is a schematic diagram showing a hardware configuration of the receiver 2 according to this embodiment. The receiver 2 includes a tuner 211, a demodulator 212, a separator 22, a selector 231, an audio decoder 232, a speaker 234, a video decoder 241, a presentation processor 242, a display 243, an input / output device 251, an auxiliary storage device 252, a ROM (Read Only The computer 100 includes a RAM (Random Access Memory) 253, a RAM (Random Access Memory) 254, a CPU (Central Processing Unit) 255, and a communication chip 256. The demodulator 212, the separator 22, the selector 231, the audio decoder 232, and the speaker 234 are also referred to as an audio processing unit M. Note that the configuration for processing data (for example, the separator 22, the selector 231, the audio decoder 232, the video decoder 241, and the presentation processor 242) may be realized by software (arithmetic processing by the CPU 255). 8, the hardware configuration corresponding to each component of the receiver 2 in FIG. 2, FIG. 4, or FIG. 5 is given the same numerals as those given to the components in FIG.
[0072] A digital broadcast signal received by an antenna is input to the receiver 2 via an input terminal, converted into a TLV stream by the tuner 211 and demodulator 212, and separated into video, audio, other assets, and various message tables of MMT through TLV / MMT separation processing by the separator 22. The scrambled assets are decoded by a descrambler using the obtained key after processing the EMM / ECM extracted by the TLV / MMT separation processing in a CAS module (not shown). The video assets are decoded by the video decoder 241, and output after processing to present text and graphics images. The audio assets are decoded by the audio decoder 232 and then output. For video and audio output, the receiver body may be provided with video and audio output means (display 243 and speaker 234), or may be provided with a digital video and audio output that outputs decoded video and audio signals to an external device, or a digital audio output that outputs only audio to an external device. A high-speed digital interface may also be provided. The receiver may also be provided with a storage means such as an auxiliary storage device 252 (HDD, etc.) inside, providing a broadcast storage function. The receiver 2 has memories such as a RAM 254 used by receiver applications such as EPG and multimedia services, an auxiliary storage device 252 (non-volatile memory: NVRAM, etc.) for storing logo data and EPG data for services, and a ROM 253 (which can be substituted with NVRAM) for storing fonts, etc.
[0073] The separator 22 determines whether an AC-4 audio signal is present based on a descriptor in a table describing asset information. The AC-4 audio signal includes AC-4 audio assets (also referred to as "AC-4 audio assets"). If separator 22 determines that an AC-4 audio signal is present, it separates the AC-4 audio assets and each of one or more MPEG-4 audio assets (MPEG-4 assets) if receiver 2 has AC-4 audio decoding capability. If separator 22 determines that an AC-4 audio signal is not present, it separates each of the one or more MPEG-4 audio assets from the data of the MPEG-4 audio signal.
[0074] <Input terminals, tuners, demodulators> Receiver 2 has two terminals for inputting digital broadcast signals: IF input and optical input. However, the receiver 2 has only an IF input and does not necessarily have an optical input. The tuner 211 supports right-hand circular band IF frequencies, left-hand circular band IF frequencies, or both. The demodulator 212 performs front-end signal processing.
[0075] <Separator / Video Decoder> The TLV / MMT separation process by the separator 22 consists of two processes: TLV separation and MMT separation. The receiver 2 in broadcast transmission has the ability to simultaneously process at least 12 assets per service. The receiver 2 may handle a maximum of 22 assets per service. Video assets may be subjected to screen split coding. Furthermore, the receiver 2 may not have built-in video decoding processing, but may be equipped with a function for streaming from a high-speed digital interface. The receiver 2 may output HDR (High Dynamic Range) video to an SDR (Standard Dynamic Range) compatible display.
[0076] <Audio decoder> In the receiver 2 to which downmix processing for an external pseudo surround processor and downmix processing for stereo sound field expansion are added as options, the presentation processor 242 displays the downmix setting status on the display 243. This allows the receiver to understand the setting status. If the receiver 2 is equipped with a digital audio output for MPEG-4 AAC (Advanced Audio Coding) audio streams, it will output in a format that complies with the AAC extension and is multiplexed by the LATM / LOAS (Low-overhead MPEG-4 Audio Transport Multiplex / Low Overhead Audio Stream) broadcast format.If the receiver 2 is equipped with a digital audio output for MPEG-4 ALS audio streams, it will output in a format that complies with the ALS extension and is multiplexed by the LATM / LOAS broadcast format.
[0077] <Output terminal> A digital video / audio output terminal and a digital audio output terminal are described below as output terminals of the receiver 2. However, the receiver 2 may be equipped with a high-speed digital interface instead of these output terminals. In the case of a receiver 2 with a built-in display device, it is not necessary to equip it with a digital video and audio output terminal. In the case of a receiver 2 that does not have a display device (display 243) such as an STB, the receiver 2 is equipped with a digital video and audio output terminal, which is either an HDMI (registered trademark, the same applies below) terminal, a terminal for MHL / superMHL output, or a terminal for wireless digital video and audio output function.
[0078] The receiver may be equipped with an optical digital audio output terminal or a coaxial digital audio output terminal as a digital audio output terminal. It may also be equipped with an HDMI terminal and provide a digital audio output function using the HDMI Audio Return Channel (HDMI-ARC) defined in HDMI 1.4. When outputting an MPEG-4 AAC audio stream to the digital audio output terminal, it must conform to the AAC extension, but the output of 22.2ch multichannel audio may be TBD. When outputting an MPEG-4 ALS audio stream to the digital audio output terminal, it must conform to the ALS extension, but the output of the MPEG-4 ALS stream may be TBD.
[0079] NVRAM is used as memory for downloading receiver software and data common to all receivers, and as memory for downloading data transmitted by the MH-CDT method, such as logo data. NVRAM stores data types, data common to all receivers (genre code table, program characteristic code table, reserved word table), logo data, multimedia services, received e-mail, etc., and for example, an AC-4 audio digital mixer is stored.
[0080] <Signal processing flow in the audio processing unit M> FIG. 9 is a schematic diagram showing an example of the flow of signal processing in the receiver according to this embodiment. This diagram is an example of the audio processing unit M. The audio processing unit M includes a demodulator 212, a TLV / MMT separation unit 22, a selector 231, a decoder unit 232, a mixer unit 2331, a downmixer (DMIX) unit 2332, a switch (SW) unit 2333, a DAC (Digital-Analog Converter) unit 2334, and an external output I / F (interface) unit 251. 9, components corresponding to those of the receiver 2 in Fig. 2, 4, or 5 are given the same numbers as those given to the components in Fig. 2, 4, or 5. Note that the mixer unit 2331, the downmixer unit 2332, the switch unit 2333, and the DAC unit 2334 correspond to the mixer 233 in Fig. 2.
[0081] This diagram shows the flow of audio signal processing within the receiver 2. The receiver 2 extracts multiple audio assets via the TLV / MMT separation processing unit 22. The selector 231 selects an audio asset to be output from these, and the decoder unit 232 decodes and outputs the audio. Here, the audio asset selected by the selector 231 is input to the decoder unit 232, and decoding is performed according to the audio mode (audio codec).
[0082] The switch unit 2333 selects an asset suitable for an external AV amplifier from among a plurality of audio assets, and outputs the asset to the decoder unit 232 and the external output I / F unit 251. The decoder unit 232 decodes the audio asset according to the input audio asset. If the decoded data sequence is an AC-4 data sequence, that is, if the audio asset is an AC-4 audio asset, the decoder unit 232 outputs the data sequence to the mixer unit 2331. If the decoded data sequence is a 5.1ch PCM data sequence, the decoder unit 232 outputs the data sequence to the downmixer unit 2332. If the decoded data sequence is a 2ch PCM data sequence, the decoder unit 232 outputs the data sequence to the switch unit 2333. The mixer unit 2331 performs downmixing processing by synthesizing the audio for each sound material in the AC-4 audio asset for the input data string. The downmixer unit 2331 performs downmixing processing to convert the input data string into 2ch PCM data. The data stream that has undergone the downmix process is output to the switch unit 2333 . The switch unit 2333 outputs a data string to the DAC unit 2334 or the external output I / F (interface) unit 251 in response to an instruction based on the control information from the selector 231. The DAC unit 2334 converts the input data string into an analog audio signal and outputs it to the speaker 234.
[0083] <Audio switching menu> 10 is a diagram showing an example of an audio switching menu according to the present embodiment. The audio switching menu is a menu for selecting any one of the audio types of simulcast audio. The audio switching menu F81 is an example of an audio switching menu displayed on a receiver 2 that supports AC-4. The audio switching menu F82 is an example of an audio switching menu displayed on a receiver 2a that does not support AC-4. These are examples in which one AC-4 audio asset and three MPEG-4 audio assets are transmitted, and the control unit of the receiver 2 generates these audio switching menus F81 and F82 by referring to the MH-audio component descriptor arranged in the MPT or MH-EIT. When creating the audio switching menu F81, the MH-AC-4 audio descriptor described later may also be referred to. This one AC-4 audio asset includes 11.1ch background sounds, dialogue (Japanese), dialogue (English), commentary audio (Japanese), and commentary audio (English). The three MPEG-4 audio assets are 5.1ch Japanese, 2ch (stereo) Japanese, and 2ch English, respectively.
[0084] The audio type F811 is an audio type for selecting background sound (11.1ch) and dialogue (Japanese) included in the AC-4 audio asset. The audio type F812 is an audio type for selecting background sound (11.1ch) and dialogue (English) included in the AC-4 audio asset. The audio type F813 is an audio type for selecting background sound (11.1ch) and commentary audio (Japanese) included in the AC-4 audio asset. The audio type F814 is an audio type for selecting background sound (11.1ch) and commentary audio (English) included in the AC-4 audio asset.
[0085] The audio type F815 is an audio type for selecting 5.1ch Japanese, which is one of the three MPEG-4 audio assets. The audio type F816 is an audio type for selecting 2ch (stereo) Japanese, which is one of the three MPEG-4 audio assets. The audio type F817 is an audio type for selecting 2ch (stereo) English, which is one of the three MPEG-4 audio assets. The audio switching menu F81 shows a state in which the audio type F811 is selected. When any of the audio types F811 to F814 is selected, the audio objects included in the selected audio type are mixed by the mixer 233.
[0086] The audio type F821 is an audio type for selecting 5.1ch Japanese, which is one of three MPEG-4 audio assets. The audio type F822 is an audio type for selecting 2ch (stereo) Japanese, which is one of three MPEG-4 audio assets. The audio type F823 is an audio type for selecting 2ch (stereo) English, which is one of three MPEG-4 audio assets. The audio switching menu F82 shows that the audio type F821 is selected.
[0087] In addition, a receiver 2 that supports 5.1ch may or may not display the audio type for selecting 2ch. Also, the phonetic notation described in the text_char field of the MH-audio component descriptor is used for the audio type. If there are multiple languages in the AC-4 audio asset, multiple audio types (phonetic notations) may be described in the text_char field.
[0088] [MH-Audio Component Descriptor] FIG. 11 is a schematic diagram showing an example of the structure of the MH-audio component descriptor according to this embodiment. The MH-Audio Component Descriptor describes each parameter of an audio elementary stream in an asset and is also used to express the elementary stream in character form. MPEG-4 audio is multiplexed as an audio elementary stream for each audio configuration (e.g., language, number of channels). AC-4 audio includes various audio configurations in one audio elementary stream. When the MH-Audio Component Descriptor is placed in the MPT, it is placed in the asset_descriptors_byte of the corresponding audio asset in the MPT.
[0089] In the MH-audio component descriptor, the meaning of each field is as follows: Note that in Fig. 11 etc., "uimsbf" stands for unsigned integer most significant bit first, and "bslbf" stands for bit string left bit first. "descriptor_tag" describes a fixed value indicating that it is an MH-audio component descriptor. "descriptor_length" describes the descriptor length of the MH-audio component descriptor.
[0090] "nga_type" (field F91) indicates the type of next-generation audio of the audio asset corresponding to this MH-audio component descriptor. Fig. 12 is a table showing examples of values of nga_type according to this embodiment. In the example of Fig. 12, when the value of "nga_type" is 0b0, it indicates that the type of next-generation audio is MPEG-H 3D Audio (Baseline Profile), and when it is 0b1, it indicates that it is AC-4.
[0091] "nga_level" (field F92) indicates the level of the audio asset corresponding to this MH-audio component descriptor. The level is information indicating the processing power required to decode the audio asset. The level is, for example, the processing load or the amount of memory used, but may also correspond to the number of channels. For example, the performance of the device or the performance required to decode the bitstream may be specified by a pair of type and level. FIG. 13 is a table showing an example of the value of nga_level according to this embodiment. In the example shown in FIG. 13, when the value of "nga_level" is 0b000, it indicates that it is not NGA (next generation audio). In this case, the value of "nga_type" has no meaning. Also, when the value of "nga_level" is 0b001, it indicates that the level is 1 (Level 1). Similarly, when it is 0b010, it indicates that the level is 2 (Level 2). When it is 0b011, it indicates that the level is 3 (Level 3). When it is 0b100, it indicates that the level is 4 (Level 4). Additionally, 0b101 to 0b111 are unused. Note that "0b" indicates that the following sequence of numbers is binary.
[0092] "stream_content" is set to a specific value (0x03) for MPEG-4 AAC audio streams, and to another value (0x04) for MPEG-4 ALS audio streams. Note that for AC-4 audio streams, a further value (e.g., 0x07) may be set.
[0093] "component_type" specifies the type of audio component, and 8 bits (b7-b0) are defined as b7: dialogue control, b6-b5: audio for people with disabilities, and b4-b0: audio mode. Note that "component_type" may have an increased number of bits and a value (e.g. b8) may be added, and the added value may be defined as AC-4. "component_tag" is a label for identifying a component stream, and has the same value as the component tag in the MH-stream identification descriptor. "stream_type" describes a fixed value indicating the LATM / LOAS stream format.
[0094] The "simulcast_group_tag" is assigned the same number to components that are simulcasting (transmitting the same content in different encoding formats or audio modes), and is set to a specific value (0xFF) for components that are not simulcasting. "main_component_flag" has a specific value when the audio component is the main audio. "quality_indicator" indicates the sound quality mode. "sampling_rate" indicates the sampling frequency. "ISO_639_language_code" indicates the language of the audio component. In ES multilingual mode, it indicates the language of the first audio component. The language code is expressed as a three-letter alphabetic code. Each character is written as 8 bits and inserted in that order into the 24-bit field. "text_char" describes the name of the audio type. If this description is the default character string, this field may be omitted.
[0095] The simulcast_group_tag (simulcast group identification) is used to distinguish between stereo audio sent simulcast with AC-4, 22.2ch surround, or 5.1ch surround, and MPEG-4 AAC stereo audio sent simulcast with the ALS coding method, etc., on the receiver side. The simulcast_group_tag value is sent with the same value for audio sent simulcast.
[0096] Note that "ISO_639_language_code2" may indicate one or more language names of the AC-4 audio asset. Specifically, a first language (e.g., Japanese) may be described in "ISO_639_language_code" and a second language (e.g., English) may be described in "ISO_639_language_code2". Also, for AC-4 audio assets, multiple languages (e.g., Japanese and English) may be described in "ISO_639_language_code2".
[0097] In the transmission operation of the MH-audio component descriptor, when the broadcasting station 1 updates the parameters of the audio stream within the same event, in principle it changes the contents of the MH-audio component descriptor in the MPT and updates the version of the MPT, but there are exceptions to this rule, where this descriptor is not updated. In this case, the contents of the audio stream and the MH-audio component descriptor temporarily do not match. For example, this may occur when transitioning from the main program to a commercial or during fluid programming. In this case, the broadcasting station 1 does not update the version of the MPT, so the receiver continues to play the audio stream with the same component tag value. Such an operation where this descriptor is not updated when updating parameters of an audio stream is permitted only when the audio encoding method is AAC and the audio mode is switched between audio modes of 5.1ch or less.
[0098] During the reception process of the MH-audio component descriptor, if the MPT version is updated and the number of audio streams or the contents of this descriptor are updated, the receiver 2 will play the audio appropriately according to the contents of this descriptor. As long as the MPT version has not been updated, the receiver 2 will, in principle, continue to play the audio stream with the same component tag value. When switching between audio modes of 5.1ch or lower, the audio stream and the contents of this descriptor may differ. In that case, the receiver 2 will give priority to decoding the contents of the audio stream.
[0099] [Select audio asset] The broadcast operates in multiple audio modes (AC-4, MPEG-H, MPEG-4 AAC2ch, AAC5.1ch, AAC7.1ch, AAC22.2ch, ALS2ch, ALS5.1ch).
[0100] The receiver 2a has the following functions as an audio decoding function. MPEG-4 AAC2ch playback MPEG-4 AAC 5.1ch to 2ch downmix playback function To meet these conditions, AAC2ch is used in simulcast mode in AC-4, MPEG-H, MPEG-4 AAC7.1ch, or AAC22.2ch audio modes (AAC5.1ch may be used in simulcast mode). Also, AAC2ch or AAC5.1ch is used in simulcast mode in AC-4, MPEG-H audio mode, or ALS audio mode.
[0101] The receiver 2 has the ability to switch and select as follows when multiple audio assets are in operation. When playing back on the receiver 2 itself, the receiver 2 determines the audio mode that can be played back on the receiver 2 itself, and switches to and plays back the asset with the smallest component tag value first. When selecting a channel, the asset with the smallest playable component tag value is played back as the default audio. In a playback environment of up to 2ch, the receiver 2 will prioritize AAC2ch audio when AAC2ch audio is simulcast to AC-4, MPEG-H, or AAC5.1ch. In a playback environment of up to a specific level, when AC-4 is in operation, the receiver 2 will prioritize playback of the maximum or minimum level among the levels below the specific level. However, if the audio mode is switched without updating the MPT version, the asset that was being played will continue to be played.
[0102] The receiver 2 determines whether multiple languages are used by referring to the simulcast group identification, and plays the asset (language) with the smaller component tag value as the default language. For AC-4, the receiver 2 may play audio in a predetermined default language. Even if the language has been switched, the receiver 2 returns to the default language when a channel is reselected. The receiver 2 may be provided with a language fixed mode for AC-4. In the receiver 2, the selection of a valid audio asset is cyclically switched by an audio button on a remote control, etc. For example, in the receiver 2, the selection of an AC-4 audio asset and one or more MPEG-4 audio assets is cyclically switched. In the user interface where the receiver selects the audio from the menu, the audio information shall be displayed according to the information in the MH-Audio Component Descriptor. Note that the phonetic notation described in the text_char field in the MH-Audio Component Descriptor shall take precedence over the notation characters for the audio type. However, for AC-4, the receiver 2 may take precedence over a predefined phonetic notation.
[0103] When the audio mode is switched within the same audio asset, or when the receiver automatically switches to the audio of a different asset, the receiver 2 switches in a manner that does not cause the receiver to feel unnatural. The switching operation of the receiver 2 is as follows. (1) When the receiver 2 detects a change in audio mode or asset due to an MPT update approximately 0.5 seconds ago, it mutes the output of the preceding audio after fading. (2) After the receiver 2 has performed the processing required for switching, it cancels the mute and resumes output of the subsequent audio. The time required for the switching process varies depending on whether the audio asset has been switched and the type of audio mode being updated. In general, it takes the longest time to switch the encoding method. During the switching process, a silent section is provided on the sending side. (3) When receiver 2 switches from MPEG-4 audio assets to AC-4 audio assets, it displays the AC-4 audio digital mixer.
[0104] Fig. 14 is a schematic diagram showing another example of the structure of the MH-audio component descriptor according to this embodiment. The example in Fig. 14 differs from the example in Fig. 11 only in that it does not have "nga_type" (field F91) and "nga_level" (field F92) but has a 4-bit "nga_level" (field F93). "nga_level" (field F93) indicates the level of the audio asset corresponding to this MH-audio component descriptor. In the example in Fig. 14, "stream_content" identifies whether or not it is an AC-4 audio asset, and if it is an AC-4 audio asset, "nga_level" identifies the level.
[0105] Fig. 15 is a table showing example values of stream_content according to this embodiment. In the example of Fig. 15, when the value of "stream_content" is 0x3, it indicates that the audio asset is MPEG-4 AAC. Similarly, when it is 0x4, it indicates that the audio asset is MPEG-4 ALS. When it is 0x6, it indicates that the audio asset is MPEG-H 3D Audio (Baseline Profile). When it is 0x7, it indicates that the audio asset is AC-4. 0x5 is unused.
[0106] Fig. 16 is a table showing another example of the value of nga_level according to this embodiment. In the example of Fig. 16, when the value of "nga_level" is 0b0001, it indicates that the level is 1 (Level 1). Similarly, when it is 0b0010, it indicates that the level is 2 (Level 2). When it is 0b0011, it indicates that the level is 3 (Level 3). When it is 0b0100, it indicates that the level is 4 (Level 4). 0b0000 and 0b0101 to 0b1111 are unused.
[0107] 17 is a schematic diagram showing an example of the structure of the MH-AC-4 audio descriptor according to the present embodiment. Like the MH-audio component descriptor, the MH-AC-4 audio descriptor is also placed in the asset_descriptors_byte of the corresponding asset in the MPT.
[0108] "descriptor_tag" describes a fixed value indicating that it is an MH-AC-4 audio descriptor. "descriptor_length" describes the descriptor length of the MH-AC-4 audio descriptor. "nga_type" and "nga_level" are the same as "nga_type" and "nga_level" in Fig. 11. The receiver 2 may refer to "nga_type" and "nga_level" in the MH-AC-4 audio descriptor and not provide "nga_type" and "nga_level" in the MH-audio component descriptor.
[0109] "presentation()" indicates audio material included in the AC-4 audio asset. Furthermore, "presentation()" may indicate a designation of a combination of audio materials to be mixed. For example, each of the combination of background sound (11.1ch) and dialogue (Japanese) of audio type F811 in FIG. 10, the combination of background sound (11.1ch) and dialogue (English) of audio type F812, the combination of background sound (11.1ch) and commentary audio (Japanese) of audio type F813, and the combination of background sound (11.1ch) and commentary audio (English) of audio type F814 may be defined in "presentation()". Note that it may be ac4_presentation_info defined in ETSI TS 103 190-1 and ETSI TS 103 190-2.
[0110] "dialogue_enhancement()" indicates the presence or absence of a dialogue enhancement function in the AC-4 audio asset, and may further indicate information available for dialogue enhancement. Note that "presentation()" and "dialogue_enhancement()" may be included in the MH-AC-4 audio descriptor multiple times (N times) as shown in FIG. 17. When multiple "presentation()" and "dialogue_enhancement()" are included, each may correspond to one of the options included in the audio switching menu.
[0111] Fig. 18 is a schematic diagram showing an example of the structure of presentation() according to this embodiment. The example shown in Fig. 18 is an excerpt from ac4_presentation_info - AC-4 presentation information, section 4.2.3.2 of ETSI TS 103 190-1 V1.3.1 (2018-02) Digital Audio Compression (AC-4) Standard; Part 1: Channel based coding.
[0112] Fig. 19 is a schematic diagram showing another example of the structure of presentation() according to the present embodiment. The example shown in Fig. 19 is an excerpt from ac4_presentation_info, section 6.2.1.2 of ETSI TS 103 190-2 V1.2.1 (2018-02) Digital Audio Compression (AC-4) Standard; Part 2: Immersive and personalized.
[0113] [Audio switching behavior] FIG. 20 is a flowchart showing a detailed example of switching according to this embodiment. This diagram shows the audio switching operation of the receiver 2. The processes of the following steps S101 to S104, S112 to S114, and S121, and the control of steps S122 and S123 are performed by a computer (CPU 255: control unit) of the receiver 2. Note that the example of Fig. 20 is a flowchart in the case where the structure of the MH-audio component descriptor is the example shown in Fig. 11.
[0114] (Step S101) A channel is selected to receive a channel specified by a remote control or the like by the receiver 2, or a channel automatically specified by the receiver 2, that is, settings are made to the tuner 211 and the Demux 22. Then, the process of step S102 is performed. (Step S102) The receiver 2 updates the MPT. Then, the process of step S103 is performed. (Step S103) The receiver 2 checks the default asset. Specifically, the default asset is an asset whose "component_tag" has a specific value. Default assets are determined in advance for each asset type. If the asset type is "broadcast transmission audio", a specific value "0x0010" is assigned to the default asset. This specific value is set as the initial value of the variable i (i=0x0010). Then, the process of step S104 is performed.
[0115] (Step S104) The receiver 2 determines whether or not "component_tag" is an asset of broadcast transmission audio. Assets whose asset type is "broadcast transmission audio" are assigned values from "0x0010" to "0x002F". The receiver 2 determines whether or not the asset is an asset of broadcast transmission audio by determining whether or not the value of the variable i is equal to or less than "0x002F". Note that AC-4 audio assets may be assigned a smaller value in "component_tag" than MPEG-4 audio assets. In this case, the playability of the AC-4 audio assets is determined first. However, AC-4 audio assets may be assigned a larger value in "component_tag" than MPEG-4 audio assets. If it is determined that the asset is a broadcast transmission audio (Yes), the process proceeds to step S1111. On the other hand, if it is determined that the asset is not a broadcast transmission audio (No), the process proceeds to step S121.
[0116] (Step S1111) The receiver 2 determines whether or not "nga_level" of the MH-audio component descriptor is "0". This determines whether or not the corresponding asset is an audio asset of next-generation audio. If "nga_level" is "0" (Yes), that is, if it is not an audio asset of next-generation audio, processing of step S1114 is performed. On the other hand, if "nga_level" is not "0" (No), that is, if it is an audio asset of next-generation audio, processing of step S1112 is performed.
[0117] (Step S1112) The receiver 2 determines whether the type of next-generation audio indicated by "nga_type" is a type that the receiver 2 can play (support). If the type is not playable (No), the receiver 2 increments the variable i and then performs the process of step S104. On the other hand, if the type is playable (Yes), the receiver 2 performs the process of step S1113.
[0118] (Step S1113) The receiver 2 judges whether the level indicated by "nga_level" is a level that the receiver itself can play (support) (or is lower than the playable level). If it is not a playable level (No), the variable i is incremented and the process of step S104 is performed. If it is a playable level (Yes), the process of step S113 is performed. If it is determined in the process of step S1113 that the content is at a reproducible level (Yes), the process of step S114 or S112 may be performed.
[0119] (Step S114) The receiver 2 determines whether or not the stream is playable by the receiver 2 itself. Specifically, the receiver 2 uses "stream_content", "component_type", and "stream_type" to determine whether or not the stream is playable by the receiver 2 itself. If the stream is not playable (No), the variable i is incremented and the process of step S104 is performed. If the stream is playable (Yes), the process of step S112 is performed.
[0120] (Step S112) The receiver 2 checks whether simulcast is available and the language, etc. Specifically, the receiver 2 performs this process using "simulcast_group_tag", "ES_multi_lingual_flag", "main_component_flag", "ISO_639_language_code", "ISO_639_language_code2", and "text_char".
[0121] (Step S113) The receiver 2 adds to a list in memory (RAM 254 or auxiliary storage device 252) the asset information acquired in the process of step S112 or the asset information determined to be at a playable level in the process of step S1113. This lists combinations of playable MPEG-4 audio assets and AC-4 sound materials. Note that the information added to the list is, for example, the information indicated by the presentation() of the MH-AC-4 audio descriptor in the case of an AC-4 audio asset. After that, the receiver 2 increments the variable i and performs the process of step S104 on the asset with the next "component_tag" value.
[0122] (Step S121) The receiver 2 selects a combination of MPEG-4 assets or AC-4 sound materials from the list compiled in the process of step S113. The selection may be automatic by the receiver 2 or manual by the recipient. If the receiver selects manually, the receiver 2 displays an audio switching menu, for example, as shown in FIG. 10, based on the list compiled in the process of step S113 and allows the receiver to select. Then, the process of step S122 is performed. (Step S122) The receiver 2 displays the selected language, etc. The selection may be automatic by the receiver 2 or may be manually selected by the recipient. Then, the process proceeds to step S123. In addition, in AC-4 audio, audio data strings in multiple languages may be stored in one asset. In this case, when an AC-4 audio asset is selected in step S121, the receiver 2 can select a language without changing (selecting) the asset. (Step S123) The receiver 2 plays back the combination of audio assets or sound materials selected in steps S121 and S122. After that, the process of step S102 is performed, whereby the receiver 2 repeats the operation of this flowchart.
[0123] In step S1111, the receiver 2 may determine whether or not the stream_content of the MH-audio component descriptor is 0x3 or 0x4. That is, when the stream_content is 0x3 or 0x4, it may be determined that the audio is not next-generation audio, and the process of step S114 may be performed. Also, when the stream_content is not 0x3 or 0x4, or is 0x6 or 0x7, it may be determined that the audio is next-generation audio, and the process of step S1111 may be performed.
[0124] Furthermore, in step S1112, the receiver 2 may refer to the stream_content of the MH-audio component descriptor to determine whether or not it is compatible. That is, when the stream_content is 0x6, the receiver 2 determines whether or not it is compatible with MPEG-H 3D Audio (Baseline Profile). Also, when the stream_content is 0x7, the receiver 2 determines whether or not it is compatible with AC-4.
[0125] As described above, in this embodiment, the broadcasting system Sys broadcasts audio assets (an example of a component) including AC-4 audio. The broadcasting station 1 broadcasts broadcast waves in which identification information indicating whether an audio component is an AC-4 audio component is multiplexed in the multiplex layer. The separator 22 of the receiver 2 acquires this identification information from the broadcast waves in the multiplex layer. The CPU 255 (an example of a control unit) selects an audio asset according to the capabilities of the receiver 2 based on this identification information. The audio decoder 232 decodes the audio data of the selected audio asset.
[0126] Furthermore, in this embodiment, the broadcasting system Sys broadcasts audio assets (an example of a component) including AC-4 audio. The broadcasting station 1 broadcasts broadcast waves including MMT control information in which identification information indicating whether an audio component is an AC-4 audio component is placed. The separator 22 of the receiver 2 separates the MMT control information in which this identification information is placed from the broadcast waves. The CPU 255 (an example of a control unit) selects an audio asset according to the capabilities of the receiver 2 based on this identification information. The audio decoder 232 decodes the audio data of the selected audio asset.
[0127] In the above embodiment, the multiplexing layer in which the identification information indicating whether the audio component is an AC-4 audio component is multiplexed is the MMT layer, but may be the UDP / IP or TLV layer. The identification information indicating whether the audio component is an AC-4 audio component may be identification information included in the data nga_type or stream_content, i.e., the control information, a descriptor indicating whether the audio component is an AC-4 audio component (MH-audio component descriptor), or control information indicating whether the audio component is an AC-4 audio component (MMT-SI, MPT, MH-EIT). This allows the receiver 2 in the broadcasting system Sys to perform appropriate audio reproduction according to its own capabilities.
[0128] In the above embodiment, "voice" may be replaced with "audio." "Asset" may be replaced with "component," and "component" may be replaced with "asset."
[0129] In addition, in the above-mentioned embodiment, a part of the broadcast station 1 (broadcast device), the receiver 2, the broadcast station server 3, and the provider server 4, for example, at least a part of the separator (Demux, TLV / MMT separator) 22, 22a, the selector 231, the audio decoder (decoder unit) 232, the mixer unit 2331, the downmixer unit 2332, the switch unit 2333, the DAC unit 2334, the mixer 233, the video decoder 241, the presentation processor 242, the input / output device 251, the CPU 255, and the communication chip 256 of the receiver 2 may be realized by a computer. In that case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into and executed by a computer system. In addition, the "computer system" referred to here is a computer system built into the broadcast station 1, the receiver 2, the broadcast station server 3, or the provider server 4, and includes hardware such as an OS and peripheral devices. In addition, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Furthermore, the term "computer-readable recording medium" may also include devices that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and devices that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such cases. Furthermore, the above-mentioned program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system. In addition, the broadcast station 1, the receiver 2, the broadcast station server 3, and the provider server 4 in the above-mentioned embodiment may be realized as an integrated circuit such as a large scale integration (LSI). Each functional block of the broadcast station 1, the receiver 2, the broadcast station server 3, and the provider server 4 may be individually processed, or may be integrated into a processor in part or in whole. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the advancement of semiconductor technology, an integrated circuit based on that technology may be used.
[0130] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0131] Broadcasting System Sys Relay station Sa Broadcasting station (broadcasting equipment) 1 AC-4 Encoder 11 MPEG-4 Encoder 111, 112, 113 Mux (Multiplexer) 12 Receiver 2, 2a Broadcasting Server 3 Business Server 4 Tuner 211 Demodulator 212 Separator (Demux, TLV / MMT separation section) 22, 22a Selector 231 Audio decoder (decoder section) 232 AC-4 Decoder 232-1 MPEG-4 Decoder 232-2 Mixer section 2331 Downmixer section 2332 AC-4 Renderer 233-1 Mixer 233-2 Switch section 2333 DAC section 2334 Mixer 233 Speaker 234 Video Decoder 241 Presentation Processor 242 Display 243 Input / Output Device (External Output I / F) 251 Auxiliary storage 252 ROM 253 RAM 254 CPU 255 Communication chip 256
Claims
1. an acquisition unit that acquires, from a broadcast wave, identification information indicating whether an audio component is an AC-4 audio component at a multiplexing layer; a selection unit that selects an audio component according to the capability of the device itself based on the identification information; a decoding unit for decoding the audio data of the selected audio component; Equipped with Receiving device.
2. The identification information is included in an audio component descriptor, which is an MMT (MPEG Media Transport) descriptor and describes parameters related to an audio signal.
2. The receiving device according to claim 1.
3. the broadcast includes a plurality of audio components, including an AC-4 audio component and an MPEG-4 audio component; the selection unit selects one audio component from the plurality of audio components based on the identification information included in the audio component descriptor; The acquisition unit acquires the selected audio component, The decoding unit decodes the audio data of the selected audio component.
3. The receiving device according to claim 2.
4. the acquisition unit acquires, from the broadcast wave, information indicating sound material included in an audio component of the AC-4 audio at the multiplex layer; the selection unit selects the combination of sound materials based on information indicating the sound materials; the decoding unit decodes the audio data of the sound materials of the selected combination.
2. The receiving device according to claim 1.
5. The broadcast wave includes a descriptor describing information related to object-based audio, The acquisition unit acquires a descriptor that describes information about the object-based audio.
3. The receiving device according to claim 2.
6. The descriptor describing information related to the object-based audio includes a profile representing a level or a set of functions representing processing capabilities of AC-4 audio, The selection unit selects an audio component according to the capability of a receiving device based on the level or the profile.
6. The receiving device according to claim 5.