Transmitting device, receiving device and receiving method

The solution allows for effective sound pressure adjustment of object content in 3D audio by encoding and inserting allowable ranges, addressing hearing challenges in 3D audio systems.

JP7768328B2Active Publication Date: 2025-11-12SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024228460
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-06-17
Filing Date
2024-12-25
Publication Date
2025-11-12
Estimated Expiration
2036-06-13

AI Technical Summary

Technical Problem

Existing 3D audio technologies face challenges in adjusting the sound pressure of object content effectively, leading to difficulties in hearing dialogue language due to background sounds or viewing environments.

Method used

An audio encoding unit generates an audio stream with encoded data of object contents, and an information insertion unit inserts information on the allowable range of sound pressure increase or decrease for each object content into the audio stream or container, allowing the receiving side to adjust sound pressure within specified limits.

Benefits of technology

Enables precise adjustment of sound pressure for each object content, ensuring clear audio reproduction by increasing or decreasing sound pressure as needed, while maintaining overall sound pressure consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007768328000001
    Figure 0007768328000001
  • Figure 0007768328000002
    Figure 0007768328000002
  • Figure 0007768328000003
    Figure 0007768328000003
Patent Text Reader

Abstract

To enable good sound pressure adjustment of an object content at a receiving end.SOLUTION: An audio stream having encoded data of a predetermined number of object contents is generated, and a container of a predetermined format containing this audio stream is transmitted, then information indicating a permissible range of increase or decrease of sound pressure for each object content is inserted into a layer of the audio stream and / or a layer of the container. Based on this information, On a receiving side, sound pressure for each object content increases and decreases within the allowable range.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This technology is a transmitting device , receiving Receiving device and receiving method Regarding do. [Background technology]

[0002] BACKGROUND ART Conventionally, a technique has been proposed as a stereoscopic (3D) audio technique in which encoded sample data is mapped to speakers located at arbitrary positions based on metadata and then rendered (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication No. 2014-520491 Summary of the Invention [Problem to be solved by the invention]

[0004] It is conceivable that various types of object content coded data consisting of coded sample data and metadata can be transmitted together with channel coded data such as 5.1 channels and 7.1 channels, enabling audio reproduction with enhanced realism at the receiving end. For example, object content such as dialogue language can be difficult to hear depending on the background sound or viewing environment.

[0005] An object of the present technology is to enable the receiving side to properly adjust the sound pressure of object content. [Means for solving the problem]

[0006] The concept of this technology is: an audio encoding unit that generates an audio stream having encoded data of a predetermined number of object contents; a transmitter for transmitting a container in a predetermined format including the audio stream; an information inserting unit that inserts information indicating an allowable range of increase or decrease in sound pressure for each object content into the layer of the audio stream and / or the layer of the container; Located in the transmitting device.

[0007] In this technology, an audio encoding unit generates an audio stream having encoded data of a predetermined number of object contents, and an information insertion unit inserts information indicating an allowable range of increase or decrease in sound pressure for each object content into a layer of the audio stream and / or a layer of a container.

[0008] For example, the information indicating the allowable range of increase or decrease in sound pressure for each object content may be information on the upper and lower limits of sound pressure. Also, for example, the coding format of the audio stream may be MPEG-H 3D Audio, and the information inserting unit may include, in the audio frame, an extension element having information indicating the allowable range of increase or decrease in sound pressure for each object content.

[0009] In this way, in the present technology, information indicating the allowable range of increase or decrease in sound pressure for each object content is inserted into the audio stream layer and / or the container layer, which makes it easy for the receiving side to adjust the increase or decrease in sound pressure for each object content within the allowable range by using this inserted information.

[0010] In the present technology, for example, each of a predetermined number of object contents may belong to one of a predetermined number of content groups, and the information insertion unit may insert information indicating the allowable range of increase or decrease in sound pressure for each content group into the audio stream layer and / or the container layer. In this case, it is sufficient to send information indicating the allowable range of increase or decrease in sound pressure as many times as the number of content groups, making it possible to efficiently transmit information indicating the allowable range of increase or decrease in sound pressure for each object content.

[0011] Furthermore, in the present technology, for example, factor type information indicating which of a plurality of factor types to apply may be added to the information indicating the allowable range of increase or decrease in sound pressure for each object content, making it possible to apply an appropriate factor type for each object content.

[0012] Another concept of the present technology is a receiving unit for receiving a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents; A control unit is provided to control the sound pressure increase / decrease process for increasing / decreasing the sound pressure for the object content selected by the user. It is in the receiving device.

[0013] In this technology, a receiving unit receives a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents, and a control unit controls sound pressure increase / decrease processing for increasing / decreasing sound pressure for the object contents selected by a user.

[0014] In this way, the present technology increases or decreases the sound pressure of object content selected by the user, which makes it possible to, for example, increase the sound pressure of a specific object content and decrease the sound pressure of other object content, thereby effectively adjusting the sound pressure of a specific number of object contents.

[0015] In the present technology, for example, information indicating an allowable range of increase or decrease in sound pressure for each object content may be inserted into the audio stream layer and / or the container layer, and the control unit may further control an information extraction process that extracts the information indicating the allowable range of increase or decrease in sound pressure for each object content from the audio stream layer and / or the container layer, and the sound pressure increase or decrease process may increase or decrease the sound pressure for the object content selected by the user based on the extracted information. In this case, it is easy to adjust the sound pressure of each object content within the allowable range.

[0016] In addition, in the present technology, for example, in the sound pressure increase / decrease process, when the sound pressure for the object content selected by the user is increased, the sound pressure for the other object contents may be decreased, and when the sound pressure for the object content selected by the user is decreased, the sound pressure for the other object contents may be increased. In this case, it is possible to keep the sound pressure of the entire object content constant without requiring the user to perform any operation.

[0017] In addition, in the present technology, for example, the control unit may further control a display process for displaying a user interface screen showing the sound pressure state of the object content whose sound pressure is increased or decreased by the sound pressure increasing / decreasing process. In this case, the user can easily check the sound pressure state of each object content and easily set the sound pressure. [Effects of the Invention]

[0018] According to the present technology, it is possible to effectively adjust the sound pressure of object content on the receiving side. Note that the effects described in this specification are merely examples and are not intended to be limiting, and additional effects may also be provided. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a block diagram showing an example of the configuration of a transmission / reception system according to an embodiment; [Figure 2] FIG. 1 is a diagram illustrating an example of the structure of transmission data of MPEG-H 3D Audio. [Figure 3] FIG. 1 is a diagram illustrating an example of the structure of an audio frame in transmission data of MPEG-H 3D Audio. [Figure 4] FIG. 10 is a diagram showing the correspondence between the type (ExElementType) of an extension element and its value (Value). [Figure 5] FIG. 10 is a diagram showing an example of the structure of a content enhancement frame that includes, as an extension element, information indicating the allowable range of increase or decrease in sound pressure for each content group. [Figure 6] This figure shows the contents of the main information in an example structure of a content enhancement frame. [Figure 7] FIG. 10 is a diagram showing an example of a sound pressure value (factor value) indicated by information indicating an allowable range of increase or decrease in sound pressure. [Figure 8] FIG. 10 is a diagram illustrating an example of the structure of an audio content enhancement descriptor. [Figure 9] 10 is a block diagram showing an example of the configuration of a stream generating unit included in the service transmitter. FIG. [Figure 10] FIG. 1 is a diagram illustrating an example of the structure of a transport stream TS. [Figure 11] FIG. 2 is a block diagram illustrating an example of the configuration of a service receiver. [Figure 12] FIG. 2 is a block diagram showing an example of the configuration of an audio decoding unit. [Figure 13] FIG. 10 is a diagram showing an example of a user interface screen showing the current sound pressure state of each object content. [Figure 14] 10 is a flowchart showing an example of a sound pressure increase / decrease process in the object enhancer in response to a unit operation by a user. [Figure 15] 10A and 10B are diagrams for explaining examples of sound pressure adjustment of object content and the effects thereof. [Figure 16]10 is a diagram showing another example of sound pressure values ​​(factor values) indicated by information indicating the allowable range of increase or decrease in sound pressure. FIG. [Figure 17] FIG. 10 is a diagram showing another example structure of a content enhancement frame that includes, as an extension element, information indicating the allowable range of increase or decrease in sound pressure for each content group. [Figure 18] This figure shows the contents of the main information in an example structure of a content enhancement frame. [Figure 19] FIG. 10 is a diagram illustrating another example structure of an audio content enhancement descriptor. [Figure 20] 10 is a flowchart showing another example of the sound pressure increasing / decreasing process in the object enhancer in response to a unit operation by the user. [Figure 21] FIG. 10 is a diagram illustrating an example of the structure of an MMT stream. DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, modes for carrying out the invention (hereinafter referred to as "embodiments") will be described. The description will be made in the following order. 1. Embodiment 2. Variations

[0021] <1. Embodiment> [Example of transmission and reception system configuration] 1 shows an example of the configuration of a transmission / reception system 10 according to an embodiment. This transmission / reception system 10 is made up of a service transmitter 100 and a service receiver 200. The service transmitter 100 transmits a transport stream TS via broadcast waves or network packets.

[0022] The transport stream TS includes an audio stream, or a video stream and an audio stream. The audio stream includes channel-encoded data as well as encoded data of a predetermined number of object contents (object-encoded data). In this embodiment, the encoding format of the audio stream is MPEG-H 3D Audio.

[0023] The service transmitter 100 inserts information (information on upper and lower limits) indicating the allowable range of increase or decrease in sound pressure for each object content into the audio stream layer and / or the transport stream TS layer as a container. For example, each of a predetermined number of object contents belongs to one of a predetermined number of content groups, and the service transmitter 200 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group into the audio stream layer and / or the container layer.

[0024] Figure 2 shows an example of the structure of MPEG-H 3D Audio transmission data. This example structure consists of one channel coded data and six object coded data. The one channel coded data is 5.1 channel coded data (CD) and consists of coded sample data for SCE1, CPE1.1, CPE1.2, and LFE1.

[0025] Of the six object codings, the first three belong to the Dialog Language Object Content Group Codings (DOD): these three object codings are the Object for dialog language codings for the first, second, and third languages, respectively.

[0026] The coded data for the dialogue language objects corresponding to the first, second, and third languages ​​consists of coded sample data SCE2, SCE3, and SCE4, respectively, and metadata (Object metadata) for mapping and rendering the data to speakers located at any position.

[0027] Of the six coded objects, the remaining three belong to the coded data of the sound effect object content group (SEO). These three coded objects are coded data of sound effect objects (Object for sound effect) corresponding to the first, second, and third sound effects, respectively.

[0028] The coded data for the sound effect objects corresponding to the first, second, and third sound effects consist of coded sample data SCE5, SCE6, and SCE7, respectively, and metadata (Object metadata) for mapping and rendering the data to speakers located at any position.

[0029] Encoded data is categorized by type into groups. In this configuration example, the 5.1 channel encoded data is group 1. The encoded data for dialogue language objects corresponding to the first, second, and third languages ​​are group 2, group 3, and group 4, respectively. The encoded data for sound effect objects corresponding to the first, second, and third sound effects are group 5, group 6, and group 7, respectively.

[0030] Furthermore, those that can be selected between groups on the receiving side are registered in a switch group (SW Group) and encoded. In this configuration example, groups 2, 3, and 4, which belong to the content group of the dialogue language object, are set to switch group 1 (SW Group 1). Furthermore, groups 5, 6, and 7, which belong to the content group of the sound effect object, are set to switch group 2 (SW Group 2).

[0031] Figure 3 shows an example of the structure of an audio frame in MPEG-H 3D Audio transmission data. This audio frame consists of multiple MPEG Audio Stream Packets. Each MPEG Audio Stream Packet consists of a header and a payload.

[0032] The header contains information such as packet type, packet label, and packet length. The payload contains information defined by the packet type in the header. This payload information contains "SYNC," which is the synchronization start code, "Frame," which is the actual data for 3D audio transmission, and "Config," which indicates the configuration of this "Frame."

[0033] A "Frame" contains channel-encoded data and object-encoded data that make up the 3D audio transmission data. Here, the channel-encoded data consists of encoded sample data such as SCE (Single Channel Element), CPE (Channel Pair Element), and LFE (Low Frequency Element). Also, the object-encoded data consists of encoded sample data of SCE (Single Channel Element) and metadata for mapping and rendering it to speakers located at any position. This metadata is included as an extension element (Ext_element).

[0034] In this embodiment, an element (Ext_content_enhancement) that contains information indicating the allowable range of increase or decrease in sound pressure for each content group is newly defined as an extension element (Ext_element). Accordingly, configuration information for this element (content_enhancement config) is newly defined in "Config."

[0035] 4 shows the correspondence between the type (ExElementType) of the extension element (Ext_element) and its value (Value). For example, 128 is newly defined as the value of the type "ID_EXT_ELE_content_enhancement".

[0036] Figure 5 shows an example of the syntax of a content enhancement frame (Content_Enhancement_frame()) that includes information indicating the allowable range of increase or decrease in sound pressure for each content group as an extension element. Figure 6 shows the semantics of the main information in this example.

[0037] The 8-bit field "num_of_content_groups" indicates the number of content groups. The 8-bit fields "content_group_id", "content_type", "content_enhancement_plus_factor", and "content_enhancement_minus_factor" are repeated for each content group.

[0038] The "content_group_id" field indicates the ID (identification) of the content group. The "content_type" field indicates the type of the content group. For example, "0" indicates "dialog language", "1" indicates "sound effect", "2" indicates "BGM", and "3" indicates "spoken subtitles".

[0039] The "content_enhancement_plus_factor" field indicates the upper limit for increasing or decreasing the sound pressure. For example, as shown in the table of FIG. 7, "0x00" indicates 1 (0 dB), "0x01" indicates 1.4 (+3 dB), ..., "0xFF" indicates infinite (+infinite dB). The "content_enhancement_minus_factor" field indicates the lower limit for increasing or decreasing the sound pressure. For example, as shown in the table of FIG. 7, "0x00" indicates 1 (0 dB), "0x01" indicates 0.7 (-3 dB), ..., "0xFF" indicates 0.00 (-infinite dB). The table of FIG. 7 is shared by service receiver 200.

[0040] In this embodiment, a new Audio_Content_Enhancement descriptor is defined, which contains information indicating the allowable range of sound pressure increase or decrease for each content group. This descriptor is then inserted into the audio elementary stream loop under the Program Map Table (PMT).

[0041] Figure 8 shows an example of the structure (Syntax) of an audio content enhancement descriptor. The 8-bit field "descriptor_tag" indicates the descriptor type. In this case, it indicates that it is an audio content enhancement descriptor. The 8-bit field "descriptor_length" indicates the length (size) of the descriptor, and indicates the number of subsequent bytes as the length of the descriptor.

[0042] The 8-bit field "num_of_content_groups" indicates the number of content groups. The 8-bit fields "content_group_id", "content_type", "content_enhancement_plus_factor", and "content_enhancement_minus_factor" are repeated for each content group. The information in each field is the same as that described in the Content Enhancement Frame (see Figure 5).

[0043] Returning to Fig. 1, the service receiver 200 receives a transport stream TS transmitted by broadcast waves or network packets from the service transmitter 100. This transport stream TS includes an audio stream in addition to a video stream. The audio stream includes channel-encoded data and encoded data of a predetermined number of object contents (object-encoded data), which constitute the transmission data of 3D audio.

[0044] Information indicating the allowable range of increase or decrease in sound pressure for each object content is inserted into the audio stream layer and / or the transport stream TS layer as a container. For example, information indicating the allowable range of increase or decrease in sound pressure for a predetermined number of content groups is inserted. Here, one or more object contents belong to one content group.

[0045] The service receiver 200 decodes the video stream to obtain video data, and also decodes the audio stream to obtain 3D audio data.

[0046] The service receiver 200 processes the increase / decrease in sound pressure for the object content selected by the user, and limits the range of increase / decrease in sound pressure based on the allowable range of increase / decrease in sound pressure for each object content inserted in the layer of the audio stream and / or the layer of the transport stream TS as a container.

[0047] [Stream generation part of the service transmitter] 9 shows an example of the configuration of the stream generation unit 110 included in the service transmitter 100. The stream generation unit 110 includes a control unit 111, a video encoder 112, an audio encoder 113, and a multiplexer 114.

[0048] The video encoder 112 receives video data SV and encodes it to generate a video stream (video elementary stream). The audio encoder 113 receives object data of a predetermined number of content groups as audio data SA, along with channel data. Each content group contains one or more object contents.

[0049] The audio encoder 113 encodes the audio data SA to obtain 3D audio transmission data, and generates an audio stream (audio elementary stream) including this 3D audio transmission data. The 3D audio transmission data includes channel-encoded data as well as object-encoded data of a predetermined number of content groups.

[0050] For example, as shown in the configuration example of Figure 2, it includes channel coded data (CD), coded data of a content group of dialogue language objects (DOD), and coded data of a content group of sound effect objects (SEO).

[0051] The audio encoder 113 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group into the audio stream under the control of the control unit 111. In this embodiment, a newly defined element (Ext_content_enhancement) having information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted as an extension element (Ext_element) into the audio frame (see FIGS. 3 and 5).

[0052] The multiplexer 114 converts the video stream output from the video encoder 112 and a predetermined number of audio streams output from the audio encoder 113 into PES packets, and then converts them into transport packets and multiplexes them to obtain a transport stream TS as a multiplexed stream.

[0053] The multiplexer 114 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group into the transport stream TS as a container under the control of the control unit 111. In this embodiment, a newly defined audio content enhancement descriptor (Audio_Content_Enhancement descriptor) having information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into the audio elementary stream loop existing under the PMT (see FIG. 8).

[0054] The operation of the stream generation unit 110 shown in Fig. 9 will now be briefly described. Video data is supplied to a video encoder 112. This video encoder 112 encodes the video data SV, generating a video stream including the encoded video data. This video stream is supplied to a multiplexer 114.

[0055] The audio data SA is supplied to the audio encoder 113. This audio data SA includes channel data as well as object data of a predetermined number of content groups, where each content group includes one or more object contents.

[0056] The audio encoder 113 encodes the audio data SA to obtain 3D audio transmission data. This 3D audio transmission data includes channel-encoded data as well as object-encoded data of a predetermined number of content groups. The audio encoder 113 then generates an audio stream including this 3D audio transmission data.

[0057] At this time, the audio encoder 113 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group into the audio stream under the control of the control unit 111. That is, a newly defined element (Ext_content_enhancement) having information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted as an extension element (Ext_element) into the audio frame (see FIGS. 3 and 5).

[0058] The video stream generated by the video encoder 112 is supplied to the multiplexer 114. Also, the audio stream generated by the audio encoder 113 is supplied to the multiplexer 114. In the multiplexer 114, the streams supplied from the respective encoders are PES packetized, and further transport packetized and multiplexed, thereby obtaining a transport stream TS as a multiplexed stream.

[0059] At this time, the multiplexer 114 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group into the transport stream TS as a container under the control of the control unit 111. That is, a newly defined audio content enhancement descriptor (Audio_Content_Enhancement descriptor) having information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into the audio elementary stream loop existing under the PMT (see FIG. 8).

[0060] [Transport Stream TS Structure] Fig. 10 shows an example structure of a transport stream TS. In this example structure, there is a PES packet "video PES" for a video stream identified by PID1, and there is also a PES packet "audio PES" for an audio stream identified by PID2. A PES packet consists of a PES header (PES_header) and a PES payload (PES_payload). DTS and PTS timestamps are inserted in the PES header.

[0061] An audio stream (Audio coded stream) is inserted into the PES payload of the PES packet of the audio stream. A content enhancement frame (Content_Enhancement_frame()) containing information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into the audio frame of this audio stream.

[0062] The transport stream TS also contains a PMT (Program Map Table) as PSI (Program Specific Information). The PSI is information that describes which program each elementary stream in the transport stream belongs to. The PMT contains a program loop that describes information related to the entire program.

[0063] The PMT also contains elementary stream loops that contain information related to each elementary stream. In this configuration example, there is a video elementary stream loop (video ES loop) corresponding to the video stream, and an audio elementary stream loop (audio ES loop) corresponding to the audio stream.

[0064] The video elementary stream loop (video ES loop) contains information such as the stream type and PID (packet identifier) ​​for each video stream, as well as descriptors describing information related to the video stream. The "Stream_type" value for this video stream is set to "0x24," and the PID information indicates PID1, which is assigned to the PES packet "video PES" of the video stream, as described above. One of the descriptors included is the HEVC descriptor.

[0065] The audio elementary stream loop (audio ES loop) contains information such as the stream type and PID (packet identifier) ​​for each audio stream, as well as a descriptor describing information related to that audio stream. The "Stream_type" value for this audio stream is set to "0x2C," and the PID information indicates the PID2 assigned to the audio stream's PES packet, "audio PES," as described above. One of the descriptors is the audio content enhancement descriptor (Audio_Content_Enhancement descriptor), which contains information indicating the allowable range for increasing or decreasing the sound pressure for each content group.

[0066] [Example of service receiver configuration] 11 shows an example of the configuration of a service receiver 200. This service receiver 200 has a receiving unit 201, a demultiplexer 202, a video decoding unit 203, a video processing circuit 204, a panel driving circuit 205, and a display panel 206. This service receiver 200 also has an audio decoding unit 214, an audio output circuit 215, and a speaker system 216. This service receiver 200 also has a CPU 221, a flash ROM 222, a DRAM 223, an internal bus 224, a remote control receiving unit 225, and a remote control transmitter 226.

[0067] The CPU 221 controls the operation of each part of the service receiver 200. The flash ROM 222 stores control software and data. The DRAM 223 constitutes a work area for the CPU 221. The CPU 221 loads software and data read from the flash ROM 222 onto the DRAM 223, starts up the software, and controls each part of the service receiver 200.

[0068] The remote control receiving unit 225 receives a remote control signal (remote control code) transmitted from the remote control transmitter 226 and supplies it to the CPU 221. The CPU 221 controls each unit of the service receiver 200 based on the remote control code. The CPU 221, the flash ROM 222, and the DRAM 223 are connected to an internal bus 224.

[0069] The receiving unit 201 receives a transport stream TS transmitted by broadcast waves or network packets from the service transmitter 100. This transport stream TS includes an audio stream in addition to a video stream. The audio stream includes channel-encoded data and encoded data of a predetermined number of object contents (object-encoded data), which constitute the transmission data of 3D audio.

[0070] Information indicating the allowable range of increase or decrease in sound pressure for a predetermined number of content groups is inserted into the audio stream layer and / or the container layer of the transport stream TS. Note that one content group contains one or more object contents.

[0071] Here, a newly defined element (Ext_content_enhancement) that contains information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted as an extension element (Ext_element) into the audio frame (see Figures 3 and 5). Also, a newly defined audio content enhancement descriptor (Audio_Content_Enhancement descriptor) that contains information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into the audio elementary stream loop under the PMT (see Figure 8).

[0072] The demultiplexer 202 extracts the video stream from the transport stream TS and sends it to the video decoding unit 203. The video decoding unit 203 performs a decoding process on the video stream to obtain uncompressed video data.

[0073] Video processing circuit 204 performs scaling processing, image quality adjustment processing, etc. on the video data obtained by video decoding unit 203 to obtain video data for display. Panel driving circuit 205 drives display panel 206 based on the display image data obtained by video processing circuit 204. Display panel 206 is configured, for example, with an LCD (Liquid Crystal Display), an organic EL display (organic electroluminescence display), or the like.

[0074] The demultiplexer 202 also extracts various information such as descriptor information from the transport stream TS and sends it to the CPU 221. This information also includes an audio content enhancement descriptor that contains information indicating the allowable range of increase or decrease in sound pressure for each content group described above. The CPU 221 can use this descriptor to recognize the allowable range (upper and lower limits) of increase or decrease in sound pressure for each content group.

[0075] The demultiplexer 202 also extracts an audio stream from the transport stream TS and sends it to an audio decoding unit 214. The audio decoding unit 214 performs a decoding process on the audio stream to obtain audio data for driving each speaker that constitutes a speaker system 216.

[0076] In this case, the audio decoding unit 214, under the control of the CPU 221, decodes only the encoded data of one of the object contents selected by the user out of the encoded data of a predetermined number of object contents contained in the audio stream, for the encoded data of the multiple object contents that constitute a switch group.

[0077] The audio decoding unit 214 also extracts various pieces of information inserted into the audio stream and transmits them to the CPU 221. This information also includes elements that contain information indicating the allowable range of increase or decrease in sound pressure for each of the content groups described above. The CPU 221 can recognize the allowable range (upper limit and lower limit) of increase or decrease in sound pressure for each of the content groups using these elements.

[0078] The audio decoding unit 214 also processes the increase / decrease of sound pressure for object content selected by the user under the control of the CPU 221. At this time, the audio decoding unit 214 limits the range of increase / decrease of sound pressure based on the allowable range (upper limit, lower limit) of increase / decrease of sound pressure for each object content inserted in the layer of the audio stream and / or the layer of the transport stream TS as a container. The audio decoding unit 214 will be described in detail later.

[0079] The audio output processing circuit 215 performs necessary processing such as D / A conversion and amplification on the audio data for driving each speaker obtained by the audio decoding unit 214, and supplies the audio data to the speaker system 216. The speaker system 216 includes multiple channels, such as 2 channels, 5.1 channels, 7.1 channels, 22.2 channels, etc.

[0080] "Example of audio decoding section configuration" 12 shows an example of the configuration of the audio decoding unit 214. The audio decoding unit 214 includes a decoder 231, an object enhancer 232, an object renderer 233, and a mixer 234.

[0081] The decoder 231 decodes the audio stream extracted by the demultiplexer 202 to obtain object data of a predetermined number of object contents together with the channel data. This decoder 213 performs processing that is essentially the reverse of that of the audio encoder 113 in the stream generation unit 110 in Fig. 9. Note that, with respect to the multiple object contents that make up a switch group, under the control of the CPU 221, only the object data of one of the object contents selected by the user is obtained.

[0082] The decoder 231 also extracts various pieces of information inserted into the audio stream and transmits them to the CPU 221. This information also includes elements that contain information indicating the allowable range of increase or decrease in sound pressure for each content group. This element allows the CPU 221 to recognize the allowable range (upper and lower limits) of increase or decrease in sound pressure for each content group.

[0083] The object enhancer 232 increases or decreases the sound pressure of object content selected by the user from among a predetermined number of object data obtained by the decoder 231. When increasing or decreasing the sound pressure, the object enhancer 232 is provided with a target content (target_content) indicating the object content for which sound pressure should be increased or decreased, a command indicating whether to increase or decrease the sound pressure, and an allowable range (upper limit, lower limit) for increasing or decreasing the sound pressure for the target content, in response to a user operation.

[0084] For each user operation, the object enhancer 232 changes the sound pressure of the object content of the target content (target_content) by a predetermined amount in the direction (increase or decrease) indicated by the command (command). In this case, if the sound pressure is already within the limit indicated by the allowable range (upper limit, lower limit), the sound pressure is left unchanged.

[0085] Furthermore, the object enhancer 232 determines the change range (predetermined range) of the sound pressure by, for example, referring to the table in Fig. 7. For example, if the current state is 1 (0 dB) and the user's unit operation is an increase, the state is changed to 1.4 (+3 dB). Also, for example, if the current state is 1.4 (+3 dB) and the user's unit operation is an increase, the state is changed to 1.9 (+6 dB).

[0086] For example, if the current state is 1 (0 dB) and the user's unit operation is a decrease, the state is changed to 0.7 (-3 dB). For example, if the current state is 0.7 (-3 dB) and the user's unit operation is an increase, the state is changed to 0.5 (-6 dB).

[0087] Furthermore, when increasing or decreasing the sound pressure, the object enhancer 232 sends information indicating the sound pressure state of each object data to the CPU 221. Based on this information, the CPU 221 displays a user interface screen indicating the current sound pressure state of each object content on a display unit, for example, the display panel 206, to facilitate the user's sound pressure setting.

[0088] Figure 13 shows an example of a user interface screen that displays the sound pressure status. This example shows the case where there are two object contents: a Dialogue Language Object (DOD) and a Sound Effect Object (SEO) (see Figure 2). The hatched mark indicates the current sound pressure status. Note that "plus_i" indicates the upper limit, and "minus_i" indicates the lower limit.

[0089] 14 shows an example of the process of increasing or decreasing sound pressure in the object enhancer 232 in response to a unit operation by the user. The object enhancer 232 starts the process in step ST1. Thereafter, the object enhancer 232 proceeds to the process in step ST2.

[0090] In step ST2, the object enhancer 232 determines whether the command is an increase command. If it is an increase command, the object enhancer 232 proceeds to processing in step ST3. In step ST3, the object enhancer 232 increases the sound pressure of the object content of the target content (target_content) by a predetermined amount if it is not at the upper limit value. After processing in step ST3, the object enhancer 232 ends the processing in step ST4.

[0091] If the command in step ST2 is not an increase command, i.e., if it is a decrease command, the object enhancer 232 proceeds to processing in step ST5. In this step ST5, the object enhancer 232 decreases the sound pressure of the object content of the target content (target_content) by a predetermined amount if it is not at the lower limit value. After processing in step ST5, the object enhancer 232 ends the processing in step ST4.

[0092] 12, the object renderer 233 performs rendering processing on the object data of a predetermined number of object contents obtained through the object enhancer 232 to obtain channel data of the predetermined number of object contents. Here, the object data is composed of audio data of an object sound source and position information of this object sound source. The object renderer 233 obtains channel data by mapping the audio data of the object sound source to an arbitrary speaker position based on the position information of the object sound source.

[0093] The mixer 234 combines the channel data obtained by the decoder 231 with the channel data of each object content obtained by the object renderer 233 to obtain audio data (channel data) for driving each speaker that constitutes the speaker system 216.

[0094] The operation of the service receiver 200 shown in Fig. 11 will now be briefly described. The receiver 201 receives a transport stream TS transmitted by broadcast waves or network packets from the service transmitter 100. This transport stream TS includes an audio stream in addition to a video stream.

[0095] The audio stream contains channel-encoded data and a predetermined number of object content encoded data (object-encoded data), which constitute 3D audio transmission data. Each of the predetermined number of object contents belongs to one of a predetermined number of content groups. In other words, one or more object contents belong to one content group.

[0096] This transport stream TS is supplied to a demultiplexer 202. In the demultiplexer 202, a video stream is extracted from the transport stream TS and supplied to a video decoding unit 203. In the video decoding unit 203, a decoding process is performed on the video stream to obtain uncompressed video data. This video data is supplied to a video processing circuit 204.

[0097] The video processing circuit 204 performs scaling, image quality adjustment, and other processes on the video data to obtain video data for display. This video data for display is supplied to a panel driving circuit 205. The panel driving circuit 205 drives a display panel 206 based on the video data for display. As a result, an image corresponding to the video data for display is displayed on the display panel 206.

[0098] The demultiplexer 202 also extracts various types of information, such as descriptor information, from the transport stream TS and sends it to the CPU 221. This information also includes an audio content enhancement descriptor that contains information indicating the allowable range of increase or decrease in sound pressure for each content group. The CPU 221 uses this descriptor to recognize the allowable range (upper and lower limits) of increase or decrease in sound pressure for each content group.

[0099] Furthermore, the demultiplexer 202 extracts an audio stream from the transport stream TS and sends it to an audio decoding unit 214. The audio decoding unit 214 decodes the audio stream to obtain audio data for driving each speaker that constitutes a speaker system 216.

[0100] In this case, in the audio decoding unit 214, among the encoded data of a predetermined number of object contents contained in the audio stream, with regard to the encoded data of the multiple object contents that constitute a switch group, under the control of the CPU 221, only the encoded data of any one of the object contents selected by the user is to be decoded.

[0101] The audio decoding unit 214 also extracts various pieces of information inserted into the audio stream and transmits them to the CPU 221. This information also includes an element having information indicating the allowable range of increase or decrease in sound pressure for each of the content groups described above. The CPU 221 recognizes the allowable range (upper limit and lower limit) of increase or decrease in sound pressure for each of the content groups using this element.

[0102] Furthermore, the audio decoding unit 214 performs processing to increase or decrease the sound pressure for the object content selected by the user under the control of the CPU 221. At this time, the audio decoding unit 214 limits the range of increase or decrease in sound pressure based on the allowable range (upper limit value, lower limit value) of increase or decrease in sound pressure for each object content.

[0103] In other words, in this case, in response to user operation, the CPU 221 provides the audio decoding unit 214 with target content (target_content) indicating the object content for which sound pressure should be increased or decreased, a command indicating whether to increase or decrease, and an allowable range (upper limit, lower limit) for increasing or decreasing the sound pressure for the target content.

[0104] Then, in the audio decoding unit 214, for each unit operation by the user, the sound pressure of the object data belonging to the content group of the target content (target_content) is changed by a predetermined amount in the direction (increase or decrease) indicated by the command (command). In this case, if the sound pressure is already within the limit value indicated by the allowable range (upper limit value, lower limit value), the sound pressure is left as it is without being changed.

[0105] The audio data for driving each speaker obtained by the audio decoding unit 214 is supplied to an audio output processing circuit 215. The audio output processing circuit 215 performs necessary processing such as D / A conversion and amplification on the audio data. The processed audio data is then supplied to a speaker system 216. As a result, the speaker system 216 produces an audio output corresponding to the image displayed on the display panel 206.

[0106] As described above, in the transmission / reception system 10 shown in Fig. 1, the service receiver 200 increases or decreases the sound pressure of the object content selected by the user. Therefore, for example, it is possible to increase the sound pressure of a specific object content and decrease the sound pressure of other object content, thereby effectively adjusting the sound pressure of a specific number of object contents.

[0107] Figure 15(a) shows a schematic representation of the waveform of the audio data of the dialogue language object content, and Figure 15(b) shows a schematic representation of the waveform of the audio data of other object content. Figure 15(c) shows a schematic representation of the waveform of these pieces of audio data combined. In this case, the amplitude of the waveform of the dialogue language audio data is greater than the amplitude of the audio data of the other object content, so the sound of the dialogue language is masked by the sounds of the other object content, making it very difficult to hear.

[0108] Figure 15(d) shows a schematic waveform of the audio data of the dialogue language object content with increased sound pressure, Figure 15(e) shows a schematic waveform of the audio data of the other object content with decreased sound pressure, and Figure 15(f) shows a schematic waveform of the combined audio data.

[0109] In this case, the amplitude of the waveform of the dialogue language audio data becomes larger than the amplitude of the waveform of the audio data of the other object contents, so the sound of the dialogue language becomes easier to hear without being masked by the sounds of the other object contents. Also, in this case, the sound pressure of the dialogue language object content is increased, but the sound pressure of the other object contents is decreased, so the overall sound pressure of the object contents is kept constant.

[0110] 1, the service transmitter 100 inserts information indicating the allowable range of increase or decrease in sound pressure for each object content into the audio stream layer and / or the transport stream TS layer as a container. Therefore, by using this inserted information, the receiving side can easily adjust the increase or decrease in sound pressure of each object content within the allowable range.

[0111] 1, the service transmitter 100 inserts information indicating the allowable range of increase or decrease in sound pressure for each content group to which a predetermined number of object contents belong into the transport stream TS as a layer and / or container of the audio stream. Therefore, it is only necessary to send information indicating the allowable range of increase or decrease in sound pressure as many times as the number of content groups, making it possible to efficiently transmit information indicating the allowable range of increase or decrease in sound pressure for each object content.

[0112] <2. Modifications> In the above embodiment, an example was shown in which there is one factor type for the information indicating the allowable range of increase or decrease in sound pressure for each object content, and therefore for each content group (see FIG. 7). However, it is also possible to make it possible to select from multiple factor types for the information indicating the allowable range of increase or decrease in sound pressure for each object content.

[0113] 16 shows an example of a table in which the factor type of information indicating the allowable range of increase or decrease in sound pressure for each content group can be selected from multiple types. In this example, there are two factor types: "factor_1" and "factor_2."

[0114] In this case, for a content group for which "factor_1" is specified, the receiving side refers to the "factor_1" section of the table to recognize the upper and lower limits of the sound pressure, as well as the range of change in the increase / decrease of the sound pressure. Similarly, for a content group for which "factor_2" is specified, the receiving side refers to the "factor_2" section of the table to recognize the upper and lower limits of the sound pressure, as well as the range of change in the increase / decrease of the sound pressure.

[0115] For example, even if the "content_enhancement_plus_factor" is the same and is "0x02," if "factor_1" is specified, the upper limit is recognized as 1.9 (+6 dB), and if "factor_2" is specified, the upper limit is recognized as 3.9 (+12 dB). Also, if an increase command is issued from a state of 1 (0 dB), if "factor_1" is specified, it will change to a state of 1.4 (+3 dB), and if "factor_2" is specified, it will change to a state of 1.9 (+6 dB). Also, for either factor, if the specified value is "0x00," both the upper and lower limits are 0 dB, which means that the sound pressure cannot be changed for the target content group.

[0116] Figure 17 shows an example of the syntax of a content enhancement frame (Content_Enhancement_frame()) in which the factor type of information indicating the allowable range of increase or decrease in sound pressure for each content group can be selected from multiple types. Figure 18 shows the semantics of the main information in this example.

[0117] The 8-bit field "num_of_content_groups" indicates the number of content groups. The 8-bit fields "content_group_id", "content_type", "factor_type", "content_enhancement_plus_factor", and "content_enhancement_minus_factor" are repeated for each content group.

[0118] The "content_group_id" field indicates the ID (identification) of the content group. The "content_type" field indicates the type of content group. For example, "0" indicates "dialog language", "1" indicates "sound effect", "2" indicates "BGM", and "3" indicates "spoken subtitles". The "factor_type" field indicates the applied factor type. For example, "0" indicates "factor_1", and "1" indicates "factor_2".

[0119] The "content_enhancement_plus_factor" field indicates the upper limit for increasing or decreasing sound pressure. For example, as shown in the table in Fig. 16, when the applied factor type is "factor_1," "0x00" indicates 1 (0 dB), "0x01" indicates 1.4 (+3 dB), ..., "0xFF" indicates infinite (+infinite dB), and when the applied factor type is "factor_2," "0x00" indicates 1 (0 dB), "0x01" indicates 1.9 (+6 dB), ..., "0x7F" indicates infinite (+infinite dB).

[0120] The "content_enhancement_minus_factor" field indicates the lower limit of the increase or decrease in sound pressure. For example, as shown in the table in Fig. 16, when the applied factor type is "factor_1," "0x00" indicates 1 (0 dB), "0x01" indicates 0.7 (-3 dB), ..., "0xFF" indicates 0.00 (-infinite dB), and when the applied factor type is "factor_2," "0x00" indicates 1 (0 dB), "0x01" indicates 0.5 (-6 dB), ..., "0x7F" indicates 0.00 (-infinite dB).

[0121] FIG. 19 shows an example of the syntax of an Audio_Content_Enhancement descriptor in which the factor type of information indicating the allowable range of increase or decrease in sound pressure for each content group can be selected from multiple types.

[0122] The 8-bit field "descriptor_tag" indicates the descriptor type. In this case, it indicates that it is an audio content enhancement descriptor. The 8-bit field "descriptor_length" indicates the length (size) of the descriptor, and indicates the number of subsequent bytes as the length of the descriptor.

[0123] The 8-bit field "num_of_content_groups" indicates the number of content groups. The 8-bit fields "content_group_id", "content_type", "factor_type", "content_enhancement_plus_factor", and "content_enhancement_minus_factor" are repeated for each content group. The information in each field is the same as that described in the Content Enhancement Frame (see Figure 17).

[0124] In the above embodiment, an example has been shown in which the sound pressure of the object content of the target content (target_content) selected by the user is changed by a predetermined amount in the direction (increase or decrease) indicated by the command in the service receiver 200. However, when increasing or decreasing the sound pressure of the object content of the target content (target_content), it is also possible to automatically increase or decrease the sound pressure of other object content in the opposite direction.

[0125] In this way, for example, the user can cause the service receiver 200 to execute the processes in Figures 15(d) and (e) simply by performing an operation to increase the object content of the dialogue language.

[0126] The flowchart in Fig. 20 shows an example of the sound pressure increase / decrease process in the object enhancer 232 (see Fig. 12) in response to the user's unit operation in this case. The object enhancer 232 starts the process in step ST11. Thereafter, the object enhancer 232 proceeds to the process in step ST12.

[0127] In step ST12, the object enhancer 232 determines whether the command is an increase command. If it is an increase command, the object enhancer 232 proceeds to the processing of step ST13. In step ST13, the object enhancer 232 increases the sound pressure of the object content of the target content (target_content) by a predetermined amount if it is not at the upper limit value.

[0128] Next, in step ST14, the object enhancer 232 reduces the sound pressure of the object content other than the target content (target_content) to keep the overall sound pressure of the object content constant. In this case, the sound pressure is reduced by an amount corresponding to the increase in sound pressure of the object content of the target content (target_content). In this case, the other object content to be reduced in sound pressure may be one or more. After processing in step ST14, the object enhancer 232 ends the processing in step ST15.

[0129] If the command in step ST12 is not an increase command, i.e., if it is a decrease command, the object enhancer 232 proceeds to the processing of step ST16. In this step ST16, the object enhancer 232 decreases the sound pressure of the object content of the target content (target_content) by a predetermined amount if it is not at the lower limit value.

[0130] Next, in step ST17, the object enhancer 232 increases the sound pressure of the object content other than the target content (target_content) to keep the overall sound pressure of the object content constant. In this case, the sound pressure is reduced by an amount corresponding to the increase in sound pressure of the object content of the target content (target_content). In this case, the other object content to be reduced in sound pressure may be one or more. After processing in step ST17, the object enhancer 232 ends the processing in step ST15.

[0131] In the above embodiment, an example was shown in which information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into both the audio stream layer and the transport stream TS layer as a container. However, it is also possible to insert this information only into the audio stream layer or only into the transport stream TS layer as a container.

[0132] In the above-described embodiment, an example was shown in which the container was a transport stream (MPEG-2 TS). However, the present technology can be similarly applied to systems that deliver data in containers of MP4 or other formats. For example, this includes an MPEG-DASH-based stream delivery system or a transmission / reception system that handles an MMT (MPEG Media Transport) structured transmission stream.

[0133] 21 shows an example of the structure of an MMT stream. In an MMT stream, there is an MMT packet for each asset such as video and audio. In this example structure, there is an MMT packet for an audio asset identified by ID2 along with an MMT packet for a video asset identified by ID1.

[0134] A content enhancement frame (Content_Enhancement_frame()) containing information indicating the allowable range of increase or decrease in sound pressure for each content group is inserted into the audio frame of the audio asset (audio stream).

[0135] Additionally, MMT streams contain message packets such as PA (Packet Access) message packets. PA message packets contain tables such as the MMT Package Table. MP tables contain information for each asset. Corresponding to audio assets (audio streams), Audio_Content_Enhancement descriptors are placed, which contain information indicating the allowable range for increasing or decreasing the sound pressure for each content group.

[0136] The present technology can also be configured as follows. (1) an audio encoding unit for generating an audio stream having encoded data of a predetermined number of object contents; a transmitter for transmitting a container in a predetermined format including the audio stream; an information inserting unit that inserts information indicating an allowable range of increase or decrease in sound pressure for each object content into the layer of the audio stream and / or the layer of the container; Transmitting device. (2) each of the predetermined number of object contents belongs to one of a predetermined number of content groups; The information inserting unit inserts information indicating an allowable range of increase or decrease in sound pressure for each content group into a layer of the audio stream and / or a layer of the container. The transmitting device according to (1) above. (3) The encoding format of the audio stream is MPEG-H 3D Audio, The information inserting unit includes an extension element having information indicating an allowable range of increase or decrease in sound pressure for each object content in the audio frame. The transmitting device according to (1) or (2). (4) The information indicating the allowable range of increase or decrease in sound pressure for each object content is accompanied by factor selection information indicating one of a plurality of factors. The transmitting device according to any one of (1) to (3). (5) an audio encoding step for generating an audio stream having encoded data of a predetermined number of object contents; a transmitting step of transmitting a container in a predetermined format including the audio stream by a transmitting unit; an information inserting step of inserting information indicating an allowable range of increase or decrease in sound pressure for each object content into the layer of the audio stream and / or the layer of the container; Sending method. (6) a receiving unit for receiving a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents; A processing unit for increasing or decreasing the sound pressure of the object content selected by the user is provided. Receiving device. (7) Information indicating the allowable range of increase or decrease in sound pressure for each object content is inserted into the layer of the audio stream and / or the layer of the container; an information extracting unit for extracting information indicating an allowable range of increase or decrease in sound pressure for each object content from the audio stream layer and / or the container layer; The processing unit processes the increase / decrease of the sound pressure for the object content related to the user selection based on the extracted information. The receiving device according to (6) above. (8) The processing unit When the sound pressure for the object content related to the user selection is increased, the sound pressure for the other object contents is decreased, and when the sound pressure for the object content related to the user selection is decreased, the sound pressure for the other object contents is increased. The receiving device according to (6) or (7). (9) A display control unit is further provided that displays a UI screen showing the sound pressure state of the object content whose sound pressure is increased or decreased by the processing unit. The receiving device according to any one of (6) to (8). (10) a receiving step of receiving, by a receiving unit, a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents; A processing step for increasing or decreasing the sound pressure for the object content selected by the user is included. Receiving method.

[0137] The main feature of this technology is that it inserts information indicating the allowable range of sound pressure increase / decrease for each object content into the audio stream layer and / or container layer, enabling the receiving side to appropriately adjust the increase / decrease in sound pressure for each object content within the allowable range (see Figures 9 and 10). [Explanation of symbols]

[0138] 10. Transmitting and receiving system 100···Service Transmitter 110 Stream generation unit 111 Control unit 112...Video Encoder 113 Audio Encoder 114 Multiplexer 200···Service receiver 201 Receiving unit 202 Demultiplexer 203 Video decoding unit 204...Video processing circuit 205 Panel drive circuit 206···Display panel 214 Audio Decoder 215 Audio output processing circuit 216···Speaker System 221 CPU 222···Flash ROM 223 DRAM 224 Internal Bus 225 Remote control receiver 226···Remote control transmitter 231... decoder 232···Object Enhancer 233···Object Renderer 234···Mixer

Claims

1. an audio encoding unit that generates an audio stream having encoded data of a predetermined number of object contents; a transmitter for transmitting a container in a predetermined format including the audio stream; an information inserting unit that inserts information indicating an allowable range of increase or decrease in sound pressure for each object content into a layer of the audio stream and / or a layer of the container; The encoding format of the audio stream is MPEG-H 3D Audio. Transmitting device.

2. Each of the predetermined number of object contents belongs to one of a predetermined number of content groups; The information inserting unit inserts information indicating an allowable range of increase or decrease in sound pressure for each content group into a layer of the audio stream and / or a layer of the container. The transmitting device according to claim 1 .

3. The information indicating the allowable range of increase or decrease in sound pressure for each object content is added with factor type information indicating which of a plurality of factor types is to be applied. The transmitting device according to claim 1 .

4. a receiving unit for receiving a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents; a control unit for controlling a sound pressure increase / decrease process for increasing / decreasing a sound pressure for an object content selected by a user; The encoding format of the audio stream is MPEG-H 3D Audio. Receiving device.

5. information indicating an allowable range of increase or decrease in sound pressure for each object content is inserted into a layer of the audio stream and / or a layer of the container; the control unit further controls an information extraction process for extracting information indicating an allowable range of increase or decrease in sound pressure for each object content from the audio stream layer and / or the container layer; In the sound pressure increase / decrease process, sound pressure for the object content selected by the user is increased or decreased based on the extracted information.

5. The receiving device according to claim 4.

6. In the above sound pressure increase / decrease process, When the sound pressure for the object content related to the user selection is increased, the sound pressure for the other object contents is decreased, and when the sound pressure for the object content related to the user selection is decreased, the sound pressure for the other object contents is increased.

5. The receiving device according to claim 4.

7. The control unit further controls a display process for displaying a user interface screen showing the sound pressure state of the object content whose sound pressure is increased or decreased by the sound pressure increase / decrease process.

5. The receiving device according to claim 4.

8. a receiving step of receiving, by a receiving unit, a container in a predetermined format including an audio stream having encoded data of a predetermined number of object contents; a sound pressure increase / decrease processing step for increasing / decreasing sound pressure for object content selected by a user; The encoding format of the audio stream is MPEG-H 3D Audio. Receiving method.

Citation Information

Patent Citations

  • Drug capable of inhibiting melanization for external use

    JP1988008311A

  • Stream playback device

    JP2009151926A

  • Apparatus and method for generating an audio output signal using object-based metadata.

    JP2011528200A

  • Systems and tools for improved 3D audio creation and expression

    JP2014520491A

  • Encoding and playback of 3D audio soundtracks

    JP2014525048A