Receiving device, broadcasting system, receiving method, and program
Patent Information
- Application Number
- JP2022105741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing broadcast systems fail to provide an efficient means for viewers to adjust the characteristics of spoken voices without decoding audio streams, especially in next-generation broadcasting formats like Dolby AC-4, which combines multiple audio materials into a single stream.
A receiving device that separates control information from a broadcast signal to identify adjustable audio components, allowing viewers to adjust volume and volume ratios through an adjustment guide, even before decoding the audio.
Enables efficient adjustment of audio characteristics, enhancing viewer experience by allowing personalized control over spoken voice levels and directions without full audio decoding.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a receiving device, a broadcasting system, a receiving method, and a program. [Background technology]
[0002] It has been proposed to emphasize the dialogue spoken by the characters in broadcast audiovisual content. For example, a function has been proposed to extract the dialogue from the received audio and adjust the volume of the extracted audio separately from other background sounds. The dialogue and speech contained in the content are sometimes called dialogue.
[0003] For example, Patent Document 1 describes a receiving device that detects dialogue voice data to be reproduced on a specific channel from received voice data based on auxiliary information related to the voice of a broadcast program. The receiving device controls the volume of a dialogue-only channel that reproduces dialogue voice among multiple channels and the volume of the other channels, and outputs notification information related to the detected dialogue voice.
[0004] The adoption of the Dolby (registered trademark) AC-4 format (referred to as the "AC-4 format" in this application) is being considered as an audio format for next-generation broadcasting. The AC-4 format defines specifications for sending parameters for extracting the audio components of dialogue from the entire audio of multiple channels. Such parameters can be obtained by analyzing frequency components of audio acquired during content production or existing audio of dialogue. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2016-187136 A Summary of the Invention [Problem to be solved by the invention]
[0006] When a viewer adjusts the characteristics of a speech voice, a system is being considered that detects the presence or absence of an adjustment function for the speech voice at the time the voice is received and displays a screen showing the adjustable items. On the other hand, in next-generation broadcasting, only one audio stream containing a variety of audio materials may be provided. In order to realize a display at the same hierarchical level as the audio switching menu in the conventional method, a means is required to display the presence or absence of an adjustment function and the adjustment range after extracting the parameters decoded from the audio decoder. In other words, it was not possible to determine the presence or absence of an adjustment function without going through the audio decoder. [Means for solving the problem]
[0007] The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is a receiving device that includes a separation unit that separates control information indicating the composition of a broadcast program and at least audio assets of the broadcast program from a broadcast signal, and an audio processing unit that, when the control information includes adjustment information indicating that characteristics of an element component that is a part of the audio asset are adjustable, outputs adjustment guidance information to guide adjustment of the characteristics, and adjusts the characteristics of the element component in accordance with an input. Effect of the Invention
[0008] According to an embodiment of the present invention, it is possible to efficiently provide an adjustment function for some components of the broadcast audio. [Brief description of the drawings]
[0009] [Figure 1] 1 is a schematic block diagram showing an example of the functional configuration of a broadcasting system according to a first embodiment. [Diagram 2] 1 is a schematic block diagram illustrating an example of a functional configuration of a broadcasting device according to a first embodiment. [Diagram 3] 2 is a schematic block diagram illustrating an example of a functional configuration of a receiving device according to the first embodiment. FIG. [Figure 4]FIG. 2 is a schematic block diagram showing an example of a configuration related to generation of a speech enhancement parameter set. [Diagram 5] FIG. 2 is a schematic block diagram showing a first configuration example relating to emphasis of an uttered voice. [Figure 6] FIG. 11 is a schematic block diagram showing a second configuration example relating to emphasis of an uttered voice. [Figure 7] A figure showing an example of the structure of an MH-audio component descriptor. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of an MPT. [Figure 9] A figure showing an example of the structure of an MH-AC-4 audio descriptor. [Figure 10] FIG. 4 is a diagram showing an example of a selection screen according to the first embodiment. [Figure 11] FIG. 4 is a diagram showing an example of a sound adjustment menu screen according to the first embodiment. [Figure 12] 5 is a flowchart showing an example of a voice receiving process according to the first embodiment. [Figure 13] FIG. 11 is a schematic block diagram illustrating an example of a functional configuration of a receiving device according to a second embodiment. [Figure 14] FIG. 11 is a diagram showing a display example of an electronic program guide according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] <First embodiment> Hereinafter, an embodiment of the present invention will be described with reference to the drawings. First, an overview of a broadcasting system 1 according to a first embodiment of the present invention will be described below. Fig. 1 is a schematic block diagram showing an example of the functional configuration of the broadcasting system 1 according to this embodiment. The broadcasting system 1 includes a broadcasting device 10 and a receiving device 20. In the example of Fig. 1, the number of receiving devices 20 is one, but generally there may be a plurality of receiving devices.
[0011] The broadcasting device 10 generates multiplexed data by multiplexing control information indicating the configuration of a broadcasting program with a package of the broadcasting program. The broadcasting device 10 transmits a broadcast signal carrying the generated multiplexed data to a broadcast transmission path BT. A package refers to a unit of content equivalent to a broadcasting program (event) and is associated with a broadcasting service. In this embodiment, a single audio asset is included as a component of one package. One package may also include other types of components, such as a video asset and a data broadcasting asset.
[0012] An audio asset includes one or more components. For example, audio data of multi-channel audio may be included as a component. Audio data of one system of multi-channel audio includes audio signals of multiple channels. In general, audio content is not limited to one audio material, but is produced by mixing audio materials of multiple different sound sources as elemental components. The audio material may include an audio signal indicating a speech voice. In this application, "speech voice" refers to audio intended to transmit linguistic information, and is not limited to audio generated by speech, and may be, for example, audio generated by text-to-speech synthesis processing. In this application, speech voice may be called "lines" or "dialogue". In other words, "lines" in this application is not limited to "script" or linguistic information according to a script, and may also include audio that transmits linguistic information improvisedly. In addition, "dialogue" generally means a dialogue or conversation between multiple people, but in this application, it may also include utterances that are not intended to transmit information to a specific conversation partner, such as monologue and commentary voice. In addition, a channel for audio signals is sometimes called an "audio channel" to distinguish it from a "broadcast channel", which is a channel for carrying broadcast signals.
[0013] The control information includes information about the components that make up the package. In this embodiment, the control information may include adjustable information. The adjustable information is information indicating that the listening characteristics of an elementary component, which is a part of the sound transmitted in the sound asset, are adjustable. The adjustable information indicates that one or both of the volume and the volume ratio between channels are adjustable as characteristics of the speech sound that constitutes the elementary component.
[0014] Depending on the encoding method related to multi-channel audio, the audio stream obtained by encoding may include a parameter set used to relatively emphasize speech components from the audio signals of multiple channels. The relatively emphasized speech components may also be called extraction, dialogue enhancement, etc. In the following description, this parameter set may be called the "speech enhancement parameter set."
[0015] The broadcast transmission path BT is a transmission path that can transmit a broadcast signal unidirectionally to a plurality of unspecified destinations to the receiving device 20. The broadcast transmission path BT is typically formed by broadcast waves of a predetermined frequency band corresponding to a broadcast channel. The broadcast transmission path BT may be configured to include a communication network as a part thereof. Such a communication network may be any type of network, such as the Internet, a public wireless network, an in-house network, or a dedicated line.
[0016] The receiving device 20 receives a broadcast signal transmitted using the broadcast transmission path BT, and separates control information and audio assets of the broadcast program from the multiplexed data carried by the received broadcast signal. The receiving device 20 determines whether the separated control information includes adjustable information, and when it determines that the information includes adjustable information, outputs adjustment guide information for guiding the viewer to adjust the characteristics of the element components. As the adjustment guide information, for example, a menu screen may be presented to guide the viewer to adjust the volume of the speech voice constituting the element components of the multi-channel audio and the volume ratio between the channels. The receiving device 20 accepts an operation input from the viewer, adjusts the characteristics of the element components according to the instructed characteristics, and outputs audio including the adjusted element components.
[0017] Next, an example of the functional configuration of the broadcasting device 10 according to this embodiment will be described. Fig. 2 is a schematic block diagram showing an example of the functional configuration of the broadcasting device 10 according to this embodiment. The broadcasting device 10 includes a content editing unit 120, a multiplexing unit 126, a modulation unit 128, and a transmission unit 130.
[0018] The content organizing unit 120 acquires a plurality of pieces of content that are to become materials as elemental content, and organizes the content of a broadcast program having the acquired elemental content as broadcast content. The organized broadcast content forms one package and is broadcast as a program. The broadcast content typically includes a plurality of components. The plurality of components includes a single audio asset and a video asset. The audio asset includes an audio stream, i.e., audio data that is continuous over time. The broadcast content may further include data broadcast content. The content organizing unit 120 outputs content data indicating the organized broadcast content to the multiplexing unit 126.
[0019] The content organizing unit 120 generates an audio signal that constitutes audio content by mixing audio materials from multiple different sound sources as components according to instructions from a creator. For example, the audio of a drama or movie includes the speech of characters, sounds, music, and the like. Audio signals of multi-channel audio may be adopted in audio assets. When producing audio content corresponding to multi-channel audio, the content organizing unit 120 may adjust the volume of each audio material as well as the volume ratio between audio channels. For audio content corresponding to multi-channel audio, the content organizing unit 120 may set whether or not the volume can be adjusted for a listener and whether or not the volume ratio can be adjusted for each audio material. That is, audio materials that are instructed to be adjustable in either or both of the volume adjustment and the volume ratio are set as element components whose listening characteristics are adjustable.
[0020] The content editing unit 120 includes an audio processing unit 122 and a video processing unit 124 . The audio processing unit 122 acquires an audio signal organized as the audio of a broadcast program, and generates an audio asset forming a single audio stream from the acquired audio signal using a predetermined audio encoding method. For example, AC-4 can be used as the audio encoding method. AC-4 is one of the multi-channel audio encoding methods. AC-4 can encode audio signals with up to 11.1 channels. AC-4 may also be applied to encoding audio signals with a smaller number of channels, for example, 7.1 channels or 5.1 channels. An audio stream encoded by AC-4 may include multiple audio materials as elementary components. Speech may be applied to each elementary component.
[0021] Audio content may include a portion that allows the reproduction characteristics of element components, such as the volume or volume ratio, to be adjusted. In this case, the audio processing unit 122 calculates a speech voice enhancement parameter set for each predetermined audio frequency band based on the audio signal of each channel to be encoded and the audio signal that constitutes the adjustable element component. The audio processing unit 122 includes the speech voice enhancement parameter set in an audio stream and outputs it to the multiplexing unit 126 as an audio asset.
[0022] The video processor 124 acquires video data organized as video of a broadcast program, and generates video assets that form a video stream from the acquired video using a predetermined video encoding method. As the video encoding method, for example, VVC (Versatile Video Coding) can be used.
[0023] The multiplexing unit 126 receives content data related to the content of a broadcast program from the content editing unit 120. The input content data includes the above-mentioned audio assets and video assets. The multiplexing unit 126 multiplexes the input content data and multiplexing information (control information) indicating information constituting the package, using a predetermined multiplexing method, to generate multiplexed data. As the multiplexing method, for example, the MMT-TLV (MPEG Media Transport-Type Length Value) method can be used. In the MMT-TLV method, the multiplexing information is described using MMT-SI (MMT-Signaling Information).
[0024] MMT-SI includes an MMT package table (MPT). The MPT is an information table in which assets constituting a package and attributes of each asset are described. When an audio asset includes speech as an element component with adjustable characteristics of either or both of volume and volume ratio, the multiplexing unit 126 describes in the MPT adjustment information indicating that the audio asset includes an element component with adjustable playback characteristics. The multiplexing unit 126 describes in the MPT the speech enhancement parameter set input from the audio processing unit 122 in association with the element component of the corresponding audio asset.
[0025] In addition, the creator of the audio content may set the initial values for the volume of the audio material that constitutes the element components and the audio channel to which it is distributed. In that case, the multiplexing unit 126 may further describe the set initial values in the MPT in association with the element components of the corresponding audio assets. If there are multiple audio channels to which it is distributed, the volume is set for each audio channel. The volume ratio between the audio channels is specified by the volume of each audio channel. The attributes of the audio asset are described using the MH-Audio Component Description.
[0026] The multiplexing unit 126 generates MMTP packets by dividing the content data and multiplexing information into blocks each having a predetermined amount of information. The multiplexing unit 126 configures a TLV packet that stores the MMTP packets acquired at each predetermined transmission time interval. The multiplexing unit 126 outputs a TLV stream consisting of a series of TLV packets to the modulation unit 128 as multiplexed data.
[0027] The modulation unit 128 modulates the multiplexed data input from the multiplexing unit 126 using a predetermined modulation method and converts it into a broadcast signal. As the predetermined modulation method, for example, 64QAM (Quadrature Amplitude Modulation), 256QAM, etc. can be used. The modulation unit 128 outputs the converted broadcast signal to the transmission unit 130.
[0028] The transmitting unit 130 transmits the broadcast signal input from the modulating unit 128 to the broadcast transmission path BT. The transmitting unit 130 is, for example, a transmitter, and is connected to an antenna. The transmitting unit 130 up-converts the input broadcast signal from a base frequency to a predetermined carrier frequency and supplies the up-converted signal to the antenna as a transmission signal. Broadcast waves carrying the broadcast signal are transmitted from the antenna.
[0029] Next, an example of the functional configuration of the receiving device 20 according to this embodiment will be described. Fig. 3 is a schematic block diagram showing an example of the functional configuration of the receiving device 20 according to this embodiment. The receiving device 20 includes a receiving unit 212, a demodulating unit 214, a separating unit 216, an audio decoding unit 222, a video decoding unit 224, a receiving processing unit 240, a reproducing unit 250, a display unit 260, and an input unit 270.
[0030] The receiving unit 212 receives a broadcast signal transmitted via the broadcast transmission path BT, and outputs the received broadcast signal to the demodulation unit 214. The receiving unit 212 is, for example, a tuner, and is connected to an antenna. The receiving unit 212 down-converts the carrier frequency component of the received signal obtained by receiving the signal via the antenna to a base frequency, and outputs the down-converted signal to the demodulation unit 214 as a broadcast signal. As the carrier frequency, a carrier frequency corresponding to the broadcast channel specified by the reception processing unit 240 is specified.
[0031] The demodulation unit 214 demodulates the broadcast signal input from the receiving unit 212 using a predetermined demodulation method and converts it into multiplexed data. The demodulation unit 214 outputs the converted multiplexed data to the separation unit 216. As the demodulation method, a method corresponding to the modulation method used for modulating the transmitted multiplexed data is used.
[0032] The separation unit 216 separates the multiplexed information and the content from the multiplexed data input from the demodulation unit 214. Separation of the multiplexed information is also called de-multiplexing. When the MMT-TLV method is used as the multiplexing method, the separation unit 216 extracts MMTP packets from TLV packets, which are units of a TLV stream that constitutes the multiplexed data. The separation unit 216 can extract the multiplexed information and the content data from the extracted MMTP packets. The separation unit 216 separates an information table that describes various types of multiplexed information from the multiplexed data, and outputs the separated information table to the reception processing unit 240. The separated information table includes the MPT.
[0033] The separation unit 216 refers to the MPT and identifies audio assets and video assets as assets that are components that make up a package. The separation unit 216 separates the identified audio assets and outputs the separated audio assets to the audio decoding unit 222. The separation unit 216 separates the identified video assets and outputs the separated video assets to the video decoding unit 224. When the content data includes data broadcasting content, the separation unit 216 outputs the data broadcasting content to the reception processing unit 240.
[0034] The audio decoding unit 222 decodes the audio asset input from the separation unit 216 using a predetermined audio decoding method, and outputs the audio signal obtained by decoding to the reception processing unit 240. The audio decoding unit 222 may use, as the audio decoding method, an audio decoding method corresponding to the audio encoding method used for encoding the audio signal. The video decoding unit 224 decodes the video stream forming the video asset input from the separation unit 216 using a predetermined video decoding method, and outputs the video data obtained by the decoding to the reception processing unit 240. The video decoding unit 224 may use, as the video decoding method, a video decoding method corresponding to the video encoding method used to encode the video stream.
[0035] The reception processing unit 240 executes a process for presenting broadcast content received by broadcast. The function of the reception processing unit 240 can be realized by executing a browser as a pre-installed program and executing commands described in the data content on the browser. That is, the function of the reception processing unit 240 can be realized by the computer system of the receiving device 20 analyzing commands described in the data broadcast content input from the separation unit 216 as a process instructed by the command described in the browser, and executing the process instructed by the analyzed command. For example, the data broadcast content can instruct the start and end of presentation of element contents constituting the broadcast content, the display area of the video, and the like. In this application, executing a process indicated by an instruction written in an application program such as a browser or other program may be referred to as "executing a program" or "running a program."
[0036] The reception processing unit 240 controls the reception of broadcast content and the presentation of the received broadcast content based on instructions by operation signals input from the input unit 270 . For example, the reception processing unit 240 instructs the receiving unit 212 to start receiving broadcast content based on an operation signal (start reception). At this time, the reception processing unit 240 starts outputting to the reproduction unit 250 a playback signal representing audio based on the audio signal input from the audio decoding unit 222. The reception processing unit 240 starts outputting to the display unit 260 display data representing video based on the video data input from the video decoding unit 224.
[0037] Furthermore, the reception processing unit 240 instructs the reception unit 212 to stop receiving the broadcast content based on an operation signal input from the input unit 270 (stop reception). At this time, the reception processing unit 240 stops outputting to the reproduction unit 250 a playback signal indicating audio based on the audio signal input from the audio decoding unit 222. The reception processing unit 240 stops outputting to the display unit 260 display data indicating video based on the video data input from the video decoding unit 224. The reception processing unit 240 notifies the receiving unit 212 of the broadcast channel designated by the operation signal, and starts receiving the broadcast content on the notified broadcast channel (channel switching).
[0038] The reception processing unit 240 includes a voice processing unit 242 and a display processing unit 244 . The audio processing unit 242 performs processing for playing back audio represented by the audio asset input from the audio decoding unit 222. The audio processing unit 242 determines whether or not the MPT input from the separation unit 216 includes adjustable information for the audio asset related to the audio signal input from the audio decoding unit 222.
[0039] When the adjustable information is included, the audio processing unit 242 generates an audio adjustment menu screen as a setting screen for guiding adjustment of the playback characteristics of the speech sound that is an element component related to the adjustable information, and outputs it to the display processing unit 244. The display processing unit 244 outputs display data on which the audio adjustment menu screen input from the audio processing unit 242 is superimposed to the display unit 260. The user can visually recognize the adjustment guidance information displayed on the audio adjustment menu screen and know that the received audio includes a speech sound whose characteristics can be adjusted. That is, the user is informed that the volume of the speech sound included in the audio content, or the volume ratio between channels, can be adjusted by operation.
[0040] When the adjustable information is included, the speech asset includes a speech enhancement parameter set for extracting element components. The speech processing unit 242 extracts the speech enhancement parameter set from the input speech asset. The speech processing unit 242 generates a speech signal in which the speech components are enhanced as element components by using the speech enhancement parameter set extracted from the multi-channel speech signal obtained by decoding (speech enhancement). The speech enhancement process is also called dialogue enhancement (DE). The speech processing unit 242 adjusts the volume of the speech according to the volume specified by the operation signal. The speech processing unit 242 adjusts the distribution ratio between the speech channels according to the distribution ratio specified by the operation signal. The distribution ratio between the audio channels determines the volume of the speech generated from the playback sound source corresponding to each audio channel. The direction in which the speech is perceived (localized) by the listener, who is the user, is adjusted based on the volume ratio of the speech between the playback sound sources at different positions. The audio processing unit 242 outputs an audio signal including a speech component in which one or both of the volume and the distribution ratio between the audio channels have been adjusted to the playback unit 250 as a playback signal.
[0041] When the initial values of the volume and the audio channel to which the signals are distributed are set in the MPT, the audio processor 242 generates an audio signal in which the speech component is emphasized before the operation signal is input. The audio processor 242 outputs the audio signal obtained by adjusting the volume of each audio channel for the generated audio signal as a playback signal to the playback unit 250 so that the set volume is distributed to the set audio channel. A specific example of the audio adjustment menu screen will be described later.
[0042] The display processing unit 244 performs processing for playing back the video represented by the video data input from the video decoding unit 224. The display processing unit 244 generates a display screen in which the video represented by the video asset is arranged in a predetermined display area, and outputs display data representing the generated display screen to the display unit 260. The display processing unit 244 outputs display data representing various setting screens or an electronic program guide, etc., to the display unit 260 based on instructions of operation signals input from the input unit 270. The setting screen or the electronic program guide used for setting various parameters may be displayed in parallel with the display screen of the broadcast program.
[0043] The reproduction unit 250 includes a plurality of speakers for reproducing the sound represented by the audio signal input from the reception processing unit 240. The plurality of speakers are expected to be arranged at positions defined by a predetermined reproduction method. Each speaker corresponds to one audio channel, and emits a sound based on that audio channel. The number of speakers included in the reproduction unit 250 may be equal to or greater than the number of channels defined by the reproduction method. The display unit 260 includes a device for presenting various display screens indicated by the display data input from the receiving and processing unit 240. The display unit 260 includes, for example, a display.
[0044] The input unit 270 receives a user's operation and outputs an operation signal corresponding to the received operation to the reception processing unit 240. The input unit 270 may include a general-purpose member such as a mouse or a touch panel, or may include a dedicated member such as a button, a lever, or a knob. A touch sensor used as the input unit 270 and a display used as the display unit 260 may be integrated so as to overlap each other and configured as a touch panel. The input unit 270 may include an operation signal sensor that detects an operation signal from another device (for example, a remote control device, a smartphone, etc.). The operation signal sensor outputs the detected operation signal to the reception processing unit 240.
[0045] Next, a description will be given of a method for generating the speech enhancement parameter set in the content editing unit 120. In the example of Fig. 4, the speech processing unit 122 includes a speech analysis unit 122a and a speech encoding unit 122b. The speech analysis unit 122a receives, in parallel, a multi-channel speech signal (multi-channel speech signal) of multiple channels to be transmitted and a speech signal indicating a speech voice. The speech analysis unit 122a divides the speech signal of each speech channel and the speech signal into band components, which are components of multiple predetermined frequency bands, for each period of a predetermined length (frame, for example, 15 to 100 ms). The speech analysis unit 122a determines a weighting factor for each speech channel so that the difference between a weighted sum, which is the sum of multiplication values obtained by multiplying the band components of each speech channel by a weighting factor, and the band components of the speech voice is reduced (minimized). The speech analysis unit 122a outputs a parameter set obtained by integrating the weighting factor determined for each speech channel between frequency bands as a speech voice enhancement parameter set to the speech encoding unit 122b. The audio encoding unit 122b outputs to the multiplexing unit 126 an audio stream having a code sequence obtained by performing a predetermined audio encoding process on the multi-channel audio signal for each period of a predetermined length, together with the speech enhancement parameter set input from the audio analysis unit 122a.
[0046] Next, a description will be given of a method for emphasizing an uttered voice in the reception processing unit 240. In the example of Fig. 5, the voice processing unit 242 includes an uttered voice emphasis control unit 242a and an uttered voice emphasis processing unit 242b. The speech voice enhancement control unit 242a receives the speech voice enhancement parameter set from the speech decoding unit 222 and specifies the volume instructed by the operation signal by the user operation. The speech voice enhancement control unit 242a outputs the input speech voice enhancement parameter set and the specified volume to the speech voice enhancement processing unit 242b. The volume ratio between the voice channels is expressed using the volume of each voice channel. The volume of each voice channel corresponds to the product of the volume ratio and the volume common to the voice channels. The voice channel that accepts the volume instruction by the operation is called the speech voice channel and is distinguished from other voice channels. When there is no particular instruction for a speech voice channel, one speech voice channel may be predetermined.
[0047] The speech enhancement processing unit 242b receives a multi-channel audio signal obtained by performing audio decoding processing on the audio stream input from the separation unit 216. The speech enhancement processing unit 242b divides the audio signal of each audio channel into bands to calculate band components. In a certain frequency band, a weighted sum, which is the sum of multiplication values across audio channels obtained by multiplying band components for each audio channel by a weighting coefficient, is estimated as a speech component (hereinafter referred to as an "estimated speech component"). The above multiplication value is estimated as a speech component contained in each audio channel (hereinafter referred to as an "audio-channel-specific speech component"). Therefore, a residual component obtained by subtracting the audio-channel-specific speech component from the band component of each audio channel is estimated as a component other than the speech component related to that audio channel (hereinafter referred to as an "estimated residual component").
[0048] Therefore, the speech enhancement processing unit 242b calculates, for each voice channel, a band component obtained by multiplying the estimated residual component by a first gain (hereinafter referred to as an "attenuated residual component") and adding a multiplication value obtained by multiplying the estimated speech component by a second gain corresponding to the volume specified for that voice channel (hereinafter referred to as a "gain-adjusted speech component") as the band component after speech enhancement. The first gain is a parameter indicating the magnitude of the estimated residual component. The first gain is a predetermined real number that is greater than 0 and less than or equal to 1. The second gain is a parameter indicating the magnitude of the estimated speech component to be redistributed to each voice channel. When the number of voice channels to be redistributed is one, the second gain is a real number greater than the first gain. The speech enhancement processing unit 242b synthesizes the calculated band components between frequency bands for each voice channel to generate a voice signal after speech enhancement. The speech voice enhancement processing unit 242b outputs the generated speech signal as a playback signal to the playback unit 250. In this embodiment, the method described in JP-T-2017-534904 can be used as a method for generating a speech voice enhancement parameter set and a method for generating a speech signal after speech enhancement.
[0049] A multi-channel audio reproduction method is applied to the audio signal decoded from the audio stream generated using AC-4. AC-4 allows encoding of audio signals of up to 11.1 channels. It also specifies the speaker arrangement used for multi-channel audio reproduction. Twelve speakers are used for 11.1 channel reproduction. Seven of the twelve speakers are arranged at the same height as the listening position. The directions of the seven speakers are as follows, with the direction in front of the listening position being 0 degrees: 0 degrees in front (C), 22-30 degrees in front left (FL), 22-30 degrees in front right (FR), 90-110 degrees in rear left side (L), 90-110 degrees in rear right side (R), 135-150 degrees in rear left (RL), and 135-150 degrees in rear right (RR). Of the twelve speakers, four speakers are arranged at a higher position than the listening position. The elevation angles of the four speakers are all between 30 and 55 degrees, and they are located at the rear left (CRL), front left (CFL), front right (CFR), and rear right (CRR) of the listening position, respectively. The four speakers are usually installed on the ceiling of the room. The remaining speaker is a subwoofer dedicated to reproducing low-frequency components mainly below 200Hz. The subwoofer is installed in front of the listener to the left, but does not contribute to directional perception. Due to this speaker arrangement, the maximum number of audio channels applicable to AC-4 is sometimes expressed as 7.1.4ch.
[0050] AC-4 is sometimes used to encode audio signals with fewer channels. In reproducing 7.1-channel audio signals, eight speakers installed on the floor are used instead of the four speakers installed on the ceiling. In reproducing 5.1-channel audio signals, six of the eight speakers installed on the floor are used. The six speakers are positioned in front of the listening position, at the front left, front right, rear left, rear right, and at the subwoofer.
[0051] Of the seven speakers at the same height as the listening position, the directions of three speakers, 0 degrees in front (C), 22-30 degrees in front left (FL), and 22-30 degrees in front right (FR), are distributed across the screen panel constituting the display unit 260. Therefore, these three speakers may be used to ensure the localization of the voice of a person projected on the screen panel. In AC-4, each of the three speakers or an audio channel corresponding to the front speaker is specified as a dialogue-only channel to which speech voice is distributed. In other words, any of the speakers in front, left, or right, or a channel corresponding to some or all of these speakers, may be applied as an audio channel that allows the volume of speech voice to be adjusted. The dialogue-only channel corresponds to the above-mentioned speech voice channel.
[0052] Generally, the perceived direction of a sound is adjusted by changing the volume ratio between audio channels. For example, when a speech sound is emitted at the same volume from a speaker in front of the front and a speaker in front of the left, a listener at a listening position perceives the speech sound in a direction between the speakers. By adjusting the volume ratio so that the volume from the speaker in front of the front is relatively louder than the speaker in front of the left, the perceived direction of the sound moves from the front to the front. Conversely, by adjusting the volume ratio so that the volume from the speaker in front of the left is relatively louder than the speaker in front of the front, the perceived direction of the sound moves from the front to the front to the left. As the relationship between the target direction in which sound is perceived and the volume ratio between audio channels, for example, the sine law, the tangent law, the VBAP (Vector Based Amplitude Panning) method, etc. can be used in adjusting the volume ratio.
[0053] In the above description, the method for encoding audio in a broadcast program is mainly AC-4, but the present invention is not limited to this. The audio processing unit 122 of the broadcasting device 10 may generate an audio stream obtained by encoding using another encoding method other than AC-4 as an audio asset. The multiplexing unit 126 may further multiplex the generated audio assets to form multiplexed data. The formed multiplexed data is transmitted to the receiving device 20 using the broadcast transmission path BT.
[0054] On the other hand, in the receiving device 20, the audio decoding unit 222 decodes audio streams of a plurality of decoding methods and converts them into audio signals of one or more channels. As illustrated in FIG. 6, the receiving device 20 includes a selection unit 226 between the separation unit 216 and the audio decoding unit 222. The selection unit 226 receives audio streams of a plurality of decoding methods from the separation unit 216, selects an audio stream of a decoding method corresponding to an audio mode instructed by the audio processing unit 242, and outputs the selected audio stream to the audio decoding unit 222. The selection unit 226 may be configured to include a dedicated selector. When audio streams of a plurality of decoding methods are transmitted, the audio processing unit 242 generates a selection screen as a setting screen for guiding the selection of an audio mode, and causes the display unit 260 to display the generated selection screen using the display processing unit 244.
[0055] The audio processing unit 242 can detect audio streams of multiple decoding methods by referring to the audio mode described for each audio asset described in the MPT. When adjustable information is included for an audio asset related to AC-4, information for guiding adjustment of the playback characteristics of the speech voice may be written in the display field of AC-4 as the decoding method on the setting screen. In that case, the user can select an audio mode taking into consideration that the playback characteristics of the speech voice are adjustable. In this way, in this embodiment, the adjustable information is described in the MPT, so that a clue for selecting an audio mode can be obtained before decoding each audio stream.
[0056] Next, an example of the MH-audio component descriptor (MH-Audio_Component_Descriptor) will be described. The MH-audio component descriptor is provided for each audio asset in the asset descriptor area (asset_descriptors_byte) of the MPT (FIG. 8), and is used to describe parameters set for the audio asset. In this embodiment, as illustrated in FIG. 7, the MH-audio component descriptor sets the type and level of the audio mode using a new parameter nga_profile_level in a 4-bit free area that was previously used as a free area (reserved_future_use). In other words, the audio processing unit 242 can identify the type and level of the audio mode by the setting value of the parameter nga_profile_level described in the MH-audio component descriptor set for the audio asset in the MPT.
[0057] More specifically, one bit of the four bits assigned to the parameter nga_profile_level represents the type of Next Generation Audio (NGA). NGA means audio provided by the next generation television broadcasting service. Either AC-4 or MPEG-H 3DA (3D Audio) is represented as the type of NGA. The value of one bit is "1" indicates AC-4, and the value of 0 indicates MPEG-H 3DA. Three bits of the four bits represent the presence or absence of NGA and the level. Therefore, the audio processing unit 242 can determine whether the provided audio mode, that is, the encoding method, is AC-4 or not, depending on whether the value of the parameter nga_profile_level is "1". Among the three-bit integers, 0 indicates non-NGA, and 1 to 5 indicate levels 1 to 5, respectively. The level is an index of required hardware resources, such as processing load and memory usage. The level represents, for example, a combination of the sampling frequency and maximum number of elements of each audio channel. The maximum number of elements is the maximum value of the total number of audio channels and objects. Here, an object refers to an object that is a single sound source. In this embodiment, a human speaker can also be an object.
[0058] As another example, in the MH-audio component descriptor, a new audio mode may be indicated as a value that has not been used in the past among the 4-bit values assigned to stream_content. Conventionally, MPEG-4 AAC has been assigned to 0x03, and MPEG-4 ALS to 0x04, but in this embodiment, MPEG-H 3DA is indicated to 0x06, and AC-4 is assigned to 0x07.
[0059] For audio assets in which a value indicating AC-4 is described as the audio mode, more detailed information is described in the asset descriptor area of the MPT (Figure 8) using the MH-AC-4 Audio_Descriptor. The following parameters are described in the MH-AC-4 audio descriptor illustrated in FIG. 9. The descriptor tag (descriptor_tag) indicates that the descriptor is an MH-AC-4 audio descriptor. The descriptor length (descriptor_length) indicates the amount of information described in the descriptor. The next-generation audio type (nga_type) indicates AC-4 as the type of next-generation audio. In the MH-AC-4 audio descriptor, the type of next-generation audio does not necessarily need to be described. The profile level (Profile_level) indicates the profile level. The profile level indicates a set of a profile and a level. A profile indicates a set of functions defined by purpose or use. A profile is, for example, a predetermined parameter set related to the functions of AC-4. In other words, the profile level indicates the functions and scale required for the execution of AC-4. These depend on the number of channels that make up the audio stream, the number of preset audio, and the sampling frequency of the audio per channel.
[0060] The item of preset voice (presentation) describes whether or not a preset voice is included in the voice asset, and the number of preset voices when a preset voice is included. The item of preset voice may describe contents similar to "ac4_presentation" specified in ETSI TS103 190-2. Some or all of the preset voices set may be the target of dialogue emphasis. Dialogue emphasis (dialogue_enhancement) is set for each preset voice. N illustrated in FIG. 9 corresponds to the number of preset voices. The item of dialogue emphasis describes whether or not dialogue emphasis is performed for each preset voice, and the adjustment contents when dialogue emphasis is performed. Items that can be adjusted by operation, such as the volume and volume ratio described above, are set as adjustment contents. The same contents as "dialog_enhancement" specified in ETSI TS (European Telecommunications Standards Institute Technical Specification) 103 190-1 or 190-2 may be described. In this embodiment, a parameter equivalent to at least "b_de_data_present" of "dialog_enhancement" may be set in the preset voice item. "b_de_data_present" is a 1-bit value indicating the presence or absence of dialogue enhancement data. A value of "1" for "b_de_data_present" indicates the presence of dialogue enhancement, and a value of "0" for "b_de_data_present" indicates the absence of dialogue enhancement. In this embodiment, the value indicating the presence or absence of dialogue enhancement can be used to express the possibility of adjusting the speech voice.
[0061] Therefore, for an audio asset whose encoding method is determined to be AC-4, the audio processing unit 242 can refer to the MH-AC-4 audio descriptor from the MPT and determine whether or not the audio asset includes adjustable information based on the presence or absence of dialogue emphasis settings in the items of preset audio (presentation) and dialogue emphasis (dialog_enhancement) or the setting value thereof. The MH-AC-4 audio descriptor may be set to include an independent speech speech emphasis parameter set for each line or section. The set speech speech emphasis parameter set is used to adjust each line or the line in the corresponding section. In that case, the audio stream obtained by encoding may not include the speech speech emphasis parameter set. The audio processing unit 242 may determine whether or not the audio asset includes adjustable information based on the presence or absence of a speech speech emphasis parameter set in the item of preset audio.
[0062] Next, a selection screen according to this embodiment will be described. FIG. 10 is a diagram showing an example of the selection screen according to this embodiment. The selection screen shown in FIG. 10 is a screen for guiding the selection of one audio service to be used for playback from a set of broadcast services provided by four audio streams transmitted by a selected broadcast channel. In the example of FIG. 10, any one of the four audio streams can be selected by operation. The first of the four audio streams is a stream of 11.1ch sound related to AC-4, and includes Japanese audio and English audio, and either one of them can be selected by operation. The Japanese audio and English audio each include dialogue that allows the playback characteristics to be adjusted, and whether or not dialogue emphasis is required can be selected. That is, four services are provided from the first system, such as a language of Japanese with or without dialogue emphasis, and a second system, a language of English with or without dialogue emphasis. The selection screen guides the selection of any one of the services. In FIG. 10, the filled-in items represent selected items. That is, 11.1ch sound is selected as the audio mode, Japanese is selected as the language, and dialogue emphasis is required (dialogue emphasis is on) is selected as the dialogue emphasis.
[0063] The second stream shown in the figure is a 5.1ch audio stream related to MPEG-4 and includes Japanese audio. The third stream is a 2ch stereo stream related to MPEG-4 and includes Japanese audio. The fourth stream is a 2ch stereo stream related to MPEG-4 and includes English audio. None of the streams in the second to fourth streams include audio material that enables dialogue emphasis.
[0064] Next, an example of the audio adjustment menu screen according to this embodiment will be described. FIG. 11 is a diagram showing an example of the audio adjustment menu screen according to this embodiment. The audio adjustment menu screen exemplified in FIG. 11 is displayed when the dialogue emphasis requirement is selected for the Japanese dialogue voice. The display of "Japanese dialogue" on the audio adjustment menu screen indicates that the Japanese dialogue voice is selected as the voice to be played back. The display of "Emphasis ON" indicates that the dialogue emphasis requirement is selected. The dial with the character string "Dialogue volume" written thereon is used to adjust the volume of the dialogue voice selected by operation. The slider with the character string "Balance" written thereon is used to indicate the direction in which the dialogue voice selected by operation is perceived by adjusting the volume ratio between the audio channels.
[0065] In this example, the direction of the dialogue voice can be adjusted from the front left to the front right with respect to a predetermined listening position. The voice processing unit 242 determines a volume ratio between the voice channels of the left front speaker and the front speaker, or between the voice channels of the front speaker and the right front speaker, for the estimated speech voice components that make up the selected dialogue voice. The voice processing unit 242 reduces the volume of the estimated residual components other than the dialogue voice for each voice channel relatively to the volume for the dialogue voice.
[0066] The timing at which the exemplified setting screen is generated and displayed may be, for example, each time a channel is selected, or at any time when an operation signal is input from the input unit 270. In addition, when a speech voice is provided for a multi-channel audio signal in a past broadcast program, the audio processing unit 242 may apply the same setting as the speech voice provided in the past to the speech voice provided in the currently received broadcast program. In that case, it is not necessarily necessary to display the setting screen.
[0067] Next, an example of the voice reception process according to the present embodiment will be described below with reference to a flowchart shown in FIG 12. (Step S202) The demodulation unit 214 of the receiving device 20 demodulates the broadcast signal received by the receiving unit 212 to obtain multiplexed data. (Step S204) The demultiplexer 216 demultiplexes the MPT from the acquired multiplexed data. The demultiplexer 216 refers to the demultiplexed MPT to identify components that constitute a package of a broadcast program, and demultiplexes the identified audio assets from the multiplexed data.
[0068] (Step S206) The audio processing unit 242 refers to the MH-audio component descriptor written in the MPT to identify the audio mode of the separated audio asset. (Step S208) The audio processing unit 242 judges whether the acquired audio assets include an audio asset whose audio mode is AC-4. If it is judged that the acquired audio assets include an audio asset whose audio mode is AC-4 (step S208 YES), the process proceeds to step S210. If it is judged that the acquired audio assets do not include an audio asset (step S208 NO), the process proceeds to step S212.
[0069] (Step S210) The audio processing unit 242 refers to the MPT and determines whether or not there is adjustment possible information for an audio asset whose audio mode is AC-4. Then, the process proceeds to step S212. (Step S212) The audio processing unit 242 configures a selection screen for guiding the selection of a service that can be provided from the acquired audio assets based on other parameters transmitted by the MPT, such as the identified audio mode, adjustable information, language information, etc. The audio processing unit 242 outputs the configured selection screen to the display processing unit 244, and causes the display unit 260 to display it. (Step S214) The audio processing unit 242 determines whether or not a service with adjustable playback characteristics has been selected based on an operation signal. If it is determined that the service has been selected (step S214, YES), the process proceeds to step S216. If it is determined that the service has not been selected (step S214, NO), the audio signal related to the selected service is output to the playback unit 250.
[0070] (Step S216) The audio processing unit 242 refers to the MPT, identifies adjustable items for the selected service, and generates an audio adjustment menu screen for adjusting the identified items by operation. The audio processing unit 242 outputs the audio adjustment menu screen to the display processing unit 244, and causes it to be output to the display unit 260. (Step S218) The audio processing unit 242 determines whether or not an operation signal instructing an adjustment parameter has been input. If it is determined that an operation signal has been output (step S218 YES), the process proceeds to step S220. If it is determined that an operation signal has not been output (step S218 NO), the process of step S218 is repeated. (Step S220) The voice processing unit 242 adjusts the playback characteristics of the lines, which are the spoken voice, based on the acquired adjustment parameters. The voice processing unit 242 outputs a voice signal indicating the spoken voice whose playback characteristics have been adjusted, as a playback signal, to the playback unit 250. Then, the process of FIG. 12 ends.
[0071] <Second embodiment> Next, a second embodiment of the present invention will be described. In the following description, differences from the first embodiment will be mainly described, and unless otherwise specified, the same reference numerals will be used to refer to the same items as the first embodiment.
[0072] In the broadcasting system 1 according to the present embodiment, the multiplexing unit 126 (FIG. 2) of the broadcasting device 10 generates multiplexed data using the MMT-TLV multiplexing method. The MMT-SI constituting the multiplexed data further includes an MH-Event Information Table (MH-EIT). The MH-EIT is an information table in which information about individual broadcast programs is described. The MH-EIT includes information about broadcast programs that will be broadcast in the future or in the past, as well as the current broadcast programs. The multiplexing unit 126 acquires information about individual broadcast programs in advance and temporarily stores it. As information about broadcast programs, information about names (program names (titles)), broadcast dates and times, broadcast channels, and information about audio assets and video assets, which are components of broadcast programs, are described in the MH-EIT. Information about audio assets is described using the MH-audio component descriptor, as in the MPT. Therefore, whether or not the audio includes an element component whose playback characteristics can be adjusted is transmitted for each broadcast program. In the following description, information about individual broadcast programs may be called "event information."
[0073] Next, a functional configuration example of the receiving device 20 according to this embodiment will be described. Fig. 13 is a schematic block diagram showing a functional configuration example of the receiving device 20 according to this embodiment. The receiving device 20 according to this embodiment includes a reception processing unit 240 including an audio processing unit 242, a display processing unit 244, and a reservation processing unit 246. The display processing unit 244 constructs an electronic program guide based on the MH-EIT input from the separation unit 216. The display processing unit 244 refers to the MH-EIT, identifies the broadcast time and the broadcast channel from the event information for each broadcast program, and extracts a predetermined type of element information from the event information. The element information extracted from the MH-EIT includes, for example, the broadcast date and time, the title, the audio mode, and the like. The display processing unit 244 includes information on the presence or absence of dialogue audio whose playback characteristics can be adjusted in the electronic program depending on the presence or absence of adjustable information for each broadcast program. The display processing unit 244 can construct an electronic program guide by arranging a predetermined type of display item from the extracted element information in the order of broadcast time for each column corresponding to the broadcast channel. The display processing unit 244 outputs display data including the constructed electronic program guide to the display unit 260. The display unit 260 displays the electronic program guide based on the display data input from the display processing unit 244.
[0074] FIG. 14 is a diagram showing a display example of an electronic program guide according to this embodiment. In the illustrated electronic program guide, the broadcast start time, title, audio mode, and adjustable information are displayed for each broadcast program. The adjustable information is expressed using the character "se". The notation "se" is an abbreviation of "dialogue". This display indicates that there is a dialogue whose playback characteristics are adjustable. In the illustrated example, the display mode of the display column of the broadcast program in which the adjustable information is displayed is different from the display mode of the display column of the broadcast program in which the adjustable information is not displayed in that the background brightness is lower. Due to the difference in the display mode, the user can immediately find the broadcast program in which the adjustable information is displayed. In addition, the audio mode set for the broadcast program in which the adjustable information is not displayed is AC-4. However, even if the audio mode is AC-4 for a broadcast program, the adjustable information may not be displayed. The absence of adjustable information indicates that there is no dialogue whose playback characteristics are adjustable.
[0075] Returning to FIG. 13, the reservation processing unit 246 executes processing related to reservation (playback reservation) for presentation of the contents of a broadcast program or reservation (recording reservation) for recording. The reservation processing unit 246 specifies the broadcast channel and broadcast time designated by the operation signal input from the input unit 270. The reservation processing unit 246 refers to the MH-EIT and specifies a broadcast program to be broadcast on the specified broadcast channel at the specified broadcast time. The reservation processing unit 246 may specify the broadcast channel and broadcast time specified in the operation signal, or may specify the broadcast channel and broadcast time of a specific broadcast program for which pressing is detected among the display columns of multiple broadcast programs arranged in the electronic program guide. Here, "pressing" includes not only actual pressing, but also the meaning of acquiring an operation signal indicating a position within the range of the display area. Then, the reservation processing unit 246 determines whether the function designated by the operation signal is a playback reservation or a recording reservation.
[0076] When a playback reservation is instructed by an operation signal, the reservation processing unit 246 notifies the receiving unit 212 of the specified broadcast channel when the current time reaches the start time of the broadcast time, and starts receiving the broadcast signal transmitted on that broadcast channel. The receiving processing unit 240 performs processing related to the presentation of the contents of the broadcast program based on the broadcast signal received as described above. Therefore, a broadcast program in which the playback characteristics for the element components can be adjusted is presented during the broadcast time.
[0077] When a recording reservation is instructed by an operation signal, the reservation processing unit 246 notifies the receiving unit 212 of the instructed broadcast channel when the current time reaches the start time of the broadcast period, and starts receiving the broadcast signal transmitted on that broadcast channel. The reception processing unit 240 records recording data in which components of a broadcast program including audio assets separated by the separation unit 216 are associated with control information including an MPT. In response to an instruction from the operation signal, the reservation processing unit 246 starts processing related to the presentation of a broadcast program based on the recorded recording data (recording and playback). Thus, a broadcast program in which the playback characteristics for element components can be adjusted is presented at any time after recording.
[0078] As described above, the receiving device 20 according to the above embodiment includes a separation unit 216 that separates control information (e.g., MPT) indicating the composition of a broadcast program from a broadcast signal and at least the audio assets of the broadcast program, and an audio processing unit 242 that outputs adjustment guidance information (e.g., an audio adjustment menu screen) to guide the user in adjusting the characteristics when the control information includes adjustment capability information indicating that the characteristics (e.g., volume, volume ratio between audio channels) of an element component (e.g., speech sound), which is a part of the audio assets, are adjustable, and adjusts the characteristics of the element component in accordance with an input (e.g., an operation signal). With this configuration, the user is informed that the characteristics of the element components of the transmitted voice can be adjusted without decoding or analyzing the voice, and the characteristics of the element components can be adjusted to desired characteristics by inputting. Therefore, it is possible to efficiently guide and execute the adjustment of the characteristics of the element components to each user.
[0079] Additionally, the audio processor 242 may adjust the volume as a characteristic of the element component in response to the input. With this configuration, the user can arbitrarily adjust the volume of the element components by inputting an input, and by listening to the element components having the desired volume, the user can increase his / her interest in the broadcast program.
[0080] Also, the audio assets may include audio signals of multiple audio channels, and the audio processing unit 242 may adjust the volume ratio between the audio channels to distribute the element components as a characteristic of the element components in response to the input. With this configuration, the user can arbitrarily adjust the balance between the sound channels of the element components by inputting an input, and since the user can perceive the sound of the element components from the direction corresponding to the desired volume ratio, the user can increase his / her interest in the broadcast program.
[0081] Furthermore, the receiving device 20 may include a display processing unit 244. The separation unit 216 may separate event information (e.g., MH-EIT) related to a program provided by a broadcast service from the control information. The display processing unit 244 may determine whether or not the characteristics of the element components are adjustable based on the presence or absence of adjustable information based on the separated event information, and output to the display unit 260 program guide information indicating the presence or absence of an element component whose characteristics are adjustable for each program. With this configuration, program guide information indicating the presence or absence of an element component whose characteristic is adjustable for each program is displayed, so that a user can easily select a program to watch by visually checking the program guide information and based on the presence or absence of an element component whose characteristic is adjustable.
[0082] The receiving device 20 may also include a reservation processing unit 246 that identifies a program designated by input based on the event information, and reserves the presentation of content including the audio of the program or the recording of the data of the content. With this configuration, the program designated by input from the displayed program guide information is played back or recorded, so that the user can easily instruct to watch or record the selected program scheduled to be broadcast.
[0083] In addition, when the audio processing unit 242 determines that the audio mode of an audio asset is AC-4 based on an MMT package table (MPT) that describes control information, it may determine whether or not a value indicating adjustable information is described in the MH-AC audio descriptor for that audio asset. With this configuration, it is possible to determine whether or not the playback characteristics of audio transmitted using AC-4 as the audio encoding method can be adjusted, without decoding or analyzing the transmitted audio.
[0084] Although the embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above embodiment, and the present invention also includes designs that do not deviate from the gist of the present invention. The configurations described in the above embodiment can be combined in any combination.
[0085] For example, the audio material that is an element component included in the audio asset is not necessarily limited to speech, but may be the sound of a musical instrument, a sound, or the like. The number of audio materials that are element components is not limited to one in one broadcast program, but may be two or more. The adjustable characteristics of the element components are not limited to the volume or the volume ratio between audio channels, but may include other types of characteristics such as frequency characteristics. The positions, sizes, and positions of the display information of various display screens may be set arbitrarily. The audio processing unit 242 may also determine whether or not adjustment based on the adjustable information acquired from the control information can be realized by the processing capabilities of the audio decoding unit 222 or the audio processing unit 242. The audio processing unit 242 may discard the adjustable information that is determined to be impossible to realize, and may not include it in the setting screen. A parameter set that is a premise for adjusting each element component may be transmitted to the receiving device 20, and the audio processing unit 242 of the receiving device 20 may estimate the element components from the transmitted parameter set and the multi-channel audio signal. In that case, even when an audio stream generated using an encoding method other than AC-4 is transmitted, the audio processing unit 242 may output adjustment guidance information for guiding adjustment of the playback characteristics of the element components, and adjust the playback characteristics based on an operation signal.
[0086] In addition, some of the components of the receiving device 20 may be omitted, or other components may be added. For example, in the receiving device 20, any one of the reproduction unit 250, the display unit 260, and the input unit 270, or any combination thereof, may be omitted as long as they can be connected to other functional units of the receiving device 20 so as to be capable of inputting and outputting data therefrom. In addition, a program for implementing some or all of the functions of the above-mentioned receiving device 20, for example, the separation unit 216, the audio decoding unit 222, the video decoding unit 224, and the receiving processing unit 240, may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into a computer system and executed. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer system" may also include multiple computer devices connected via a network including the Internet, a WAN, a LAN, a dedicated line, or other communication lines.
[0087] Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Thus, the recording medium storing the program may be a non-transient recording medium such as a CD-ROM. The recording medium also includes internal or external recording media accessible from a distribution server to distribute the program. The program code stored in the recording medium of the distribution server may be different from the program code in a format executable by a terminal device. In other words, the format in which the program is stored in the distribution server does not matter as long as it can be downloaded from the distribution server and installed in a format executable by a terminal device.
[0088] In addition, the program may be divided into a plurality of parts, each of which may be downloaded at a different timing and then integrated by a terminal device, or a distribution server for distributing each of the divided programs may be different. The programs for realizing the functions of each part may be individually configured. For example, the function of the reception processing unit 240 may be realized by a computer system executing a browser as a program. Here, the computer system may perform a process related to the browser by parsing a command written in an application that is carried as part of the content or separately from the content in a broadcast signal, and execute a process instructed by the specified command, thereby realizing a part of the function of the reception processing unit 240. The "computer-readable recording medium" includes a memory that holds a program for a certain period of time, such as a volatile memory (for example, RAM) inside a computer system that becomes a server or a client when a program is transmitted via a network. The above program may be a program for realizing a part of the above-mentioned functions. Furthermore, the above-mentioned functions may be realized in combination with a program already recorded in the computer system, that is, a so-called difference file (difference program).
[0089] 1...broadcast system, 10...broadcast device, 20...receiving device, 120...content compilation section, 122...audio processing section, 124...video processing section, 126...multiplexing section, 128...modulation section, 130...transmission section, 212...receiving section, 214...demodulation section, 216...separation section, 222...audio decoding section, 224...video decoding section, 226...selection section, 240...reception processing section, 242...audio processing section, 244...display processing section, 246...schedule processing section, 250...playback section, 260...display section, 270...input section
Claims
1. An acquisition unit that acquires control information indicating the configuration of a broadcast program and an audio asset of the broadcast program from a broadcast signal, When the control information includes adjustment possible information indicating that characteristics of an element component, which is a part of the audio asset, are adjustable, an audio processing unit that outputs adjustment guidance information for guiding adjustment of the characteristics and adjusts the characteristics of the element component according to an input, A receiving apparatus comprising the same.
2. The audio processing unit adjusts the volume of the element component as the characteristic according to the input. The receiving apparatus according to claim 1.
3. The audio asset includes audio signals of a plurality of audio channels, The audio processing unit adjusts a volume ratio between audio channels that distribute the element component as the characteristic according to the input. The receiving apparatus according to claim 2.
4. Comprising a display processing unit, The acquisition unit acquires event information regarding a program provided by a broadcast service from the control information, The display processing unit determines whether adjustment of the characteristics is possible based on the presence or absence of the adjustment possible information based on the event information, and outputs program expression information indicating the presence or absence of an element component for which the characteristics can be adjusted for each program to a display unit. The receiving apparatus according to claim 1.
5. Based on the event information, comprising a reservation processing unit that identifies a program instructed by an input and reserves presentation of content including the audio of the program or recording of data of the content. The receiving apparatus according to claim 4.
6. When the audio processing unit determines that the audio mode of the audio asset is AC-4 based on the control information, the audio processing unit determines whether a value indicating the adjustment possible information is described in a descriptor for the audio asset. The receiving apparatus according to claim 1.
7. A program for causing a computer to execute an acquisition step of acquiring control information indicating the configuration of a broadcast program and an audio asset of the broadcast program from a broadcast signal, and an audio processing step of outputting adjustment guidance information for guiding adjustment of the characteristics and adjusting the characteristics of the element component according to an input when the control information includes adjustment possible information indicating that characteristics of an element component, which is a part of the audio asset, are adjustable.
8. A receiving method in a receiving apparatus, wherein the receiving apparatus executes an acquisition step of acquiring control information indicating the configuration of a broadcast program and an audio asset of the broadcast program from a broadcast signal. When the control information includes adjustable information indicating that characteristics of an element component, which is a part of the voice asset, are adjustable, execute a voice processing step of outputting adjustment guidance information for guiding adjustment of the characteristics and adjusting the characteristics of the element component according to an input. Receiving method.