Method and apparatus for improving dialogue intelligibility during playback of audio data

By adjusting the mix of dialogue audio data with music and effect audio data based on the volume mixing ratio and sound pressure level on the playback device, the problem of insufficient playback volume of the dialogue audio track is solved and the intelligibility of the dialogue is improved.

CN115668372BActive Publication Date: 2025-09-09DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180035484.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-15
Filing Date
2021-05-12
Publication Date
2025-09-09
Estimated Expiration
2041-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot flexibly adjust the playback volume of dialogue audio tracks, resulting in insufficient dialogue intelligibility in different environments, especially when music and effect audio are too loud.

Method used

By determining the volume mixing ratio based on the playback volume value, the mixing ratio of dialogue audio data to music and effect audio data is automatically adjusted, the sound pressure level is used as a function to optimize the volume mixing ratio, and dynamic adjustments are made based on the user-defined volume setting and the ambient sound pressure level.

Benefits of technology

It improves the comprehensibility of conversation audio data in different environments, ensures that the conversation audio can be flexibly adjusted according to actual needs on the playback device, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668372B_ABST
    Figure CN115668372B_ABST
Patent Text Reader

Abstract

This document describes a method for improving dialogue intelligibility during playback of audio data on a playback device, wherein the audio data includes dialogue audio data and at least one of music and effects audio data, the method comprising the steps of: determining a volume mixing ratio based on a playback volume value; mixing the dialogue audio data with the at least one of the music and effects audio data based on the volume mixing ratio; and outputting the mixed audio data for playback. A corresponding playback device and a corresponding computer program product are further described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 025,479 filed on May 15, 2020 (reference number: D18144USP1) and European Application No. 20174974.4 filed on May 15, 2020 (reference number: D18144EP), which are incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to a method for improving dialog intelligibility during playback of audio data on a playback device, and more particularly to mixing audio data including dialog audio data and at least one of music and effects audio data based on a volume mixing ratio determined based on a playback volume value. The present disclosure further relates to a playback device and a corresponding computer program product implementing a media playback system for improving dialog intelligibility during playback of audio data.

[0004] Although some embodiments will be described herein with particular reference to the disclosure, it will be understood that the disclosure is not limited to this field of use and has applicability in a wider context. Background Art

[0005] Any discussion of the background art throughout the disclosure should in no way be considered as an admission that such art is widely known or forms part of the common general knowledge in the art.

[0006] Media content such as movies, TV series, sports broadcasts, entertainment programs, and news typically consists of video data and associated audio data. Depending on the type of media content, this audio data may include dialogue audio data as well as music and / or effects audio data. This audio data is typically distributed in the form of a premixed sound. However, this premixed sound cannot currently be modified on the consumer side, which can result in poor sound reproduction. For example, in the case of movies, if the premixed sound is generated with cinematic environments in mind rather than private environments, the dialogue track may be perceived as being quite unbalanced, with the music and effects tracks being too loud.

[0007] In view of the foregoing, there exists an existing need for a method and apparatus that allows flexible and separate adjustment of the playback volume of a dialog audio track relative to music and effects audio tracks. In particular, it is desirable to be able to automatically adjust the gain of the dialog audio track independently of the music and effects audio tracks. Summary of the Invention

[0008] According to a first aspect of the present disclosure, a method for improving the intelligibility of dialogue during playback of audio data on a playback device is provided, wherein the audio data may include dialogue audio data, and at least one of music and effect audio data. The method may include the following steps: (a) determining a volume mixing ratio as a function of the sound pressure level based on the playback volume value by mapping the playback volume value to the sound pressure level, wherein the volume mixing ratio refers to the ratio of the volume of the dialogue audio data to the volume of at least one of the music and effect audio data. The method may further include the following steps: (b) mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio. And, the method may include the following steps: (c) outputting the mixed audio data for playback.

[0009] As configured above, the described method allows for automatically adjusting the playback volume of a dialog track in media content based on the current playback volume of the device playing the media content. In this context, the volume mix ratio of the dialog audio data to the music and effects audio data can decrease as the (absolute) playback volume increases.

[0010] In some embodiments, in step (b), mixing the dialog audio data with at least one of the music and effect audio data based on the volume mixing ratio may include applying a gain to at least the dialog audio data.

[0011] In some embodiments, in step (a), the playback volume value may be mapped to a sound pressure level, and the volume mixing ratio may be determined as a function of the sound pressure level. The relationship between the volume mixing ratio and the sound pressure level (SPL) may be linear, wherein the volume mixing ratio may decrease linearly as the sound pressure level increases.

[0012] In some embodiments, in step (a), the playback volume value can be set based on the volume value of the playback device. This configuration allows for flexible and individual adjustment of dialogue intelligibility based on absolute playback volume.

[0013] In some embodiments, the volume value setting may be a user-defined value. Such a configuration allows for adjusting dialogue intelligibility relative to the user's preferred current absolute playback volume.

[0014] In some embodiments, in step (a), the volume mixing ratio may be further determined based on the ambient sound pressure level. This configuration allows further consideration of the individual environmental conditions of the playback device, such as the effects of a noisy background, obstacles in the room, or the overall design of the room (environmental awareness / room compensation).

[0015] In some embodiments, the ambient sound pressure level may be determined based on measurements from one or more microphones.

[0016] In some embodiments, before step (a), the method may further comprise:

[0017] (i) receiving a bit stream comprising compressed audio data; and

[0018] (ii) core decoding the compressed audio data by a core decoder and providing the dialogue audio data, and at least one of the music and effect audio data.

[0019] In some embodiments, in step (i), the received bitstream may further include information about the audio content type, and in step (a), the volume mix ratio may be further determined based on the audio content type. This configuration enables further adjustment of dialogue intelligibility by taking into account that the mix may differ for different content types. For example, in the case of a movie, the dialogue track may not be as prominent as in a sports broadcast or news program. Advantageously, the audio content type can be signaled via metadata included in the bitstream received by the playback device.

[0020] In some embodiments, the audio content type may include one or more of audio content of a movie, audio content of a news program, audio content of a sports broadcast, and audio content of an episode.

[0021] In some embodiments, the method may further include analyzing the compressed audio data to provide the dialog audio data and at least one of the music and effects audio data. In such a configuration, the analysis may include not only distinguishing between the dialog, music, and effects audio data, but also determining the content type as described above.

[0022] According to a second aspect of the present disclosure, a playback device implementing a media playback system is provided, the playback device being configured to improve dialogue intelligibility during playback of audio data, the audio data comprising dialogue audio data and at least one of music and effect audio data. The media playback system may include (a) an audio processor configured to determine a volume mixing ratio as a function of the sound pressure level based on the playback volume value by mapping a playback volume value to a sound pressure level, wherein the volume mixing ratio refers to a ratio of the volume of the dialogue audio data to the volume of at least one of the music and effect audio data. The media playback system may further include (b) a mixer configured to mix the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio. Furthermore, the media playback system may include (c) a controller configured to output the mixed audio data for playback.

[0023] In some embodiments, the audio mixer may be further configured to apply a gain to at least the conversational audio data.

[0024] In some embodiments, the audio processor may be configured to map the playback volume value to a sound pressure level to determine the volume mixing ratio as a function of the sound pressure level.

[0025] In some embodiments, the playback device may further include a user interface for receiving a volume value setting of a user, and the playback volume value may be based on the volume value setting.

[0026] In some embodiments, the playback device may further include one or more microphones for determining an ambient sound pressure level, and the audio processor may be configured to determine the volume mixing ratio further based on the ambient sound pressure level.

[0027] In some embodiments, the playback device may further include (i) a receiver configured to receive a bitstream comprising compressed audio data, and (ii) a core decoder configured to perform core decoding on the compressed audio data and provide the dialog audio data and at least one of the music and effect audio data.

[0028] In some embodiments, the core decoder may be further configured to analyze the compressed audio data to provide at least one of the dialog audio data, and the music and effects audio data.

[0029] According to a third aspect of the present disclosure, there is provided a computer program product having instructions adapted to cause a device having processing capabilities to perform a method for improving dialogue intelligibility during playback of audio data on a playback device. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Example embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0031] Figure 1 An example of a method for improving dialog intelligibility during playback of audio data on a playback device is illustrated, where the audio data may include dialog audio data, and at least one of music and effects audio data.

[0032] Figure 2 An example of a playback device implementing a media playback system to improve dialog intelligibility during playback of audio data including at least one of dialog audio data and music and effects (M&E) audio data is illustrated.

[0033] Figure 3 Further examples of playback devices implementing a media playback system to improve dialog intelligibility during playback of audio data including at least one of dialog audio data and music and effects (M&E) audio data are illustrated.

[0034] Figure 4 An example of correlation of the ratio of dialog audio data to the sum of music and effect audio data with the sound pressure level is illustrated.

[0035] Figure 5 An example of a device having processing capabilities is illustrated. DETAILED DESCRIPTION

[0036] Dialogue intelligibility during playback of audio data

[0037] A common problem observed during playback of audio data on a playback device is that sound premixing is often insufficient to achieve good sound reproduction. For example, if a movie is played back on a television, but the mix was created for a cinematic environment, dialogue intelligibility may be poor because the music and effects are perceived as too loud. Similarly, if the user is hearing-impaired, dialogue intelligibility may be particularly affected. The design of the playback environment may also contribute to poor dialogue perception during playback.

[0038] The method and apparatus described herein enable improving the intelligibility of conversational audio data during playback of audio data based on a volume mixing ratio of the conversational audio data to music and effect audio data, the audio data further comprising at least one of music and effect audio data. Specifically, the gain of the conversational audio data can be automatically adjusted based on the volume mixing ratio.

[0039] Method for improving dialogue intelligibility during playback of audio data

[0040] refer to Figure 1 , a method for improving dialogue intelligibility during playback of audio data on a playback device is illustrated, wherein the audio data may include dialogue audio data and at least one of music and effect audio data. In step S101, a volume mixing ratio is determined based on a playback volume value.

[0041] In an embodiment, the playback volume value may be mapped to a sound pressure level, and a volume mixing ratio may be determined as a function of the sound pressure level. The volume mixing ratio may refer to a ratio of the volume of the dialogue audio data to the volume (sum) of at least one of the music and effect audio data.

[0042] like Figure 4 As shown in the example of , the volume mixing ratio of the dialogue audio data and the music and effect audio data (M&E) 201 may follow a linear relationship 203 and may decrease linearly as the sound pressure level 202 increases. This relationship may be used as a calibration curve for determining the volume mixing ratio as a function of the sound pressure level. For this purpose, the corresponding values ​​of the volume mixing ratio and the sound pressure level may be stored in one or more lookup tables. Figure 2 or Figure 3 As shown in the example of , the audio processor of the media playback system may then access one or more lookup tables to determine a corresponding volume mixing ratio depending on the current sound pressure level.

[0043] Further, in an embodiment, in step S101, the playback volume value can be based on the volume value setting of the playback device. The volume mixing ratio can be directly determined based on the volume value indicated by the volume value setting. Alternatively or additionally, the volume value indicated by the volume value setting can be mapped to the sound pressure level as described above. Although the volume value setting can be a predetermined setting, in an embodiment, the volume value setting can be a user-defined value. Therefore, the user of the playback device can select the current playback volume value that is suitable for his or her needs, and the volume mixing ratio can be determined based on the user-defined volume value directly and / or by mapping this value to the sound pressure level.

[0044] Alternatively or additionally, in step S101, the volume mixing ratio can be further determined based on the ambient sound pressure level. The ambient sound pressure level can generally refer to any ambient sound in the environment of the playback device. The ambient sound can include, but is not limited to, any background noise and / or any condition in the playback environment that affects the perception of the audio data played back. Although the ambient sound pressure level can be determined in any possible way, in an embodiment, the ambient sound pressure level can be determined based on the measurement results of one or more microphones. The one or more microphones can be implemented in the playback device or can be an external microphone connected to the playback device.

[0045] Prior to step S101, in an embodiment, the method may further include: (i) receiving a bitstream including compressed audio data; and (ii) core decoding the compressed audio data by a core decoder to provide the dialog audio data and at least one of the music and effects audio data. This illustrates the fact that the method can be applied to both uncompressed and compressed audio data. In an embodiment, the compressed audio data can be analyzed to provide dialog audio data and at least one of the music and effects audio data. This can also allow for dialog detection to be performed on the decoder side. To provide the dialog audio data and at least one of the music and effects audio data, the dialog audio data can be further extracted before or after core decoding the compressed audio data. Alternatively or additionally, if the compressed audio data is in a format such as AC-4, the dialog audio data may not be extracted by the decoder. Instead, metadata may be included in the received bitstream and extracted from the received bitstream. The metadata can then be used to improve dialog intelligibility based on the determined volume mixing ratio. In other words, the metadata can be used to enhance certain frequency bands in the transmitted "complete" audio data.

[0046] In a further embodiment, in step (i) above, the received bitstream may further include information about the audio content type, wherein, in step S101, the volume mixing ratio may be further determined based on the audio content type. The audio content type may also be signaled in the bitstream via corresponding metadata. For example, since dialogues in news programs or sports broadcasts may be more pronounced or more obvious than in movies, for example, taking the audio content type into account may allow different sound mixes to be considered to improve dialogue intelligibility. In an embodiment, the audio content type may include one or more of audio content of a movie, audio content of a news program, audio content of a sports broadcast, and audio content of an interlude.

[0047] Reference again Figure 1In an example, in step S102, the conversation audio data is mixed with at least one of the music and the effect audio data based on a volume mixing ratio. In an embodiment, in step S102, mixing the conversation audio data with at least one of the music and the effect audio data based on the volume mixing ratio may include applying a gain to at least the conversation audio data. To further balance the volume of the conversation audio data relative to the volume of at least one of the music and the effect audio data during playback, a corresponding gain may also be applied to at least one of the music and the effect audio data.

[0048] In step S103, the mixed audio data is then output for playback. The playback of the audio data may be facilitated in any conceivable manner and is not limited. However, outputting the mixed audio data for playback may also include rendering the mixed audio data. Figure 1 As shown by the arrow connecting step S103 and step S101 in FIG, the sequence of steps of the method may be repeated to continuously adjust the dialogue intelligibility whenever the playback volume may change.

[0049] Playback devices that implement a media playback system

[0050] refer to Figure 2 and Figure 3 , a playback device 100 is shown that implements a media playback system 101 to improve dialogue intelligibility during playback of audio data, the audio data including dialogue audio data and at least one of music and effect audio data. The media playback system 101 includes an audio processor 103 for determining a volume mixing ratio based on a playback volume value.

[0051] In an embodiment, the audio processor 103 may be configured to map the playback volume value to a sound pressure level (SPL) in the volume value to SPL mapping unit 105 to determine a volume mixing ratio as a function of the sound pressure level.

[0052] like Figure 2 As shown in the example of , in an embodiment, the playback device 100 may further include a user interface 106 for receiving a volume value setting by a user. The volume value set by the user may be input into the volume value to SPL mapping unit 105 to determine the volume mixing ratio as a function of the sound pressure level. Alternatively or additionally, the volume value to SPL mapping unit 105 may be bypassed and the volume mixing ratio may be directly determined based on the volume value set by the user.

[0053] In an embodiment, the playback device 100 may further include one or more microphones (not shown) for determining an ambient sound pressure level. In this case, the audio processor 103 may be configured to further determine the volume mixing ratio based on the ambient sound pressure level.

[0054] Reference again Figure 2 In the example of , the media playback system 101 further includes a mixer 102 for mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio determined as described above.

[0055] In an embodiment, the audio mixer 102 may be further configured to apply a gain to at least the conversational audio data.

[0056] The media playback system 101 further includes a controller for outputting the mixed audio data for playback 108. For example, in the case of a 3.0 channel mix, the center channel may only include dialogue. In this case, only the center channel volume may be changed due to the final sound mix output by the mixer 102.

[0057] In an embodiment, the playback device 100 may further include a receiver for receiving a bitstream including compressed audio data 107. The media playback system 101 may then include a core decoder 104 for core decoding the compressed audio data and providing at least one of dialogue audio data, and music and effects audio data. Figure 2 In an embodiment, the core decoder 104 may be further configured to analyze the compressed audio data to provide dialogue audio data and at least one of music and effect audio data. This may also allow dialogue detection to be performed on the decoder side. In order to provide dialogue audio data and at least one of music and effect audio data, the core decoder 104 may be further configured to extract the dialogue audio data before or after core decoding the compressed audio data. Now referring to Figure 3 For example, alternatively or additionally, if the compressed audio data is in a format such as AC-4, for example, the dialog audio data may not be extracted by the core decoder 104, but metadata may be included in the received bitstream, and the core decoder 104 may be configured to extract the metadata from the received bitstream in addition to core decoding the compressed audio data. The mixer 102 may then be configured to use the metadata to improve dialog intelligibility based on the determined volume mixing ratio. In other words, the mixer 102 may be configured to use the metadata to boost certain frequency bands in the transmitted "complete" audio data. It should be noted that although in Figure 3The core decoder 104 and the audio mixer 102 are described as separate entities of the media playback system, but the audio mixer 102 and the core decoder 104 may also be part of a decoder, such that the methods described herein may be performed by a decoder implemented in a corresponding playback device.

[0058] Although the method described herein may be performed by a playback device implementing a media playback system as described above, it should be noted that, alternatively or additionally, the method may also be implemented as a computer program product having instructions adapted to cause a device 300 having processing capabilities 301, 302 to perform the method. Figure 5 Such a device is illustrated exemplarily.

[0059] explain

[0060] Unless otherwise specifically stated, it will be apparent from the following discussion that it should be understood that throughout the disclosed discussion, terms such as "process," "calculate," "determine," "analyze," etc. are utilized to refer to the actions and / or processes of a computer or computing system or similar electronic device that manipulates and / or transforms data represented as physical (e.g., electronic) quantities into other data similarly represented as physical quantities.

[0061] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data to transform that electronic data into other electronic data.A "computer" or "computing machine" or "computing platform" may include one or more processors.

[0062] As described above, the methods described herein can be implemented as a computer program product having instructions adapted to cause a device with processing capabilities to perform the methods. This includes any processor capable of executing a set of instructions (sequential or otherwise) specifying an action to be taken. Thus, an example may be a typical processing system that may include one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, a tensor processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem comprising main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may further be a distributed processing system in which processors are coupled together via a network. If the processing system requires a display, it may include such a display, such as a liquid crystal display (LCD), any type of light-emitting diode display (LED), including, for example, an OLED (organic light-emitting diode) display, or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include input devices, such as one or more of an alphanumeric input unit (e.g., a keyboard), a pointing control device (e.g., a mouse), and the like. The processing system may also encompass storage systems such as disk drives. The processing system may include a sound output device, such as one or more speakers or a headphone port, and a network interface device.

[0063] For example, a computer program product can be software. Software can be implemented in a variety of ways. Software can be sent or received over a network via a network interface device, or distributed via carrier media. Carrier media can include, but are not limited to, non-volatile media, volatile media, and transmission media. For example, non-volatile media can include optical disks, magnetic disks, and magneto-optical disks. Volatile media can include dynamic memory, such as main memory. Transmission media can include coaxial cables, copper wire, and optical fiber, including the wires that comprise a bus subsystem. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should be taken to include, but are not limited to, solid-state memory, computer products embodied in optical and magnetic media; media carrying propagated signals detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, implement a method; and transmission media in a network carrying propagated signals detectable by at least one of one or more processors and representing the set of instructions.

[0064] It should be noted that when the method to be performed comprises several elements (eg several steps), no order of these elements is implied unless specifically stated.

[0065] It will be understood that, in an example embodiment, the steps of the method discussed are performed by an appropriate processor (or multiple processors) in a processing (e.g., computer) system executing instructions (computer-readable code) stored in a storage device. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure can be implemented using any suitable technique for implementing the functionality described herein. The present disclosure is not limited to any particular programming language or operating system.

[0066] Reference throughout this disclosure to "one embodiment," "some embodiments," or "an embodiment" means that a particular feature described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" throughout this disclosure are not necessarily all referring to the same embodiment. Furthermore, in one or more embodiments, the particular features may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art in light of this disclosure.

[0067] In the claims below and in the description herein, any of the terms comprising, including, or which comprises, is an open term, which means including at least the elements / features that follow, but not excluding other elements / features. Thus, when the term "comprising" is used in a claim, the term should not be interpreted as being limited to the devices or elements or steps listed thereafter. As used herein, the term including, or any of which includes, or that includes, is also an open term, which means including at least the elements / features that follow the term, but not excluding other elements / features. Thus, including is synonymous with comprising and means comprising.

[0068] It should be understood that in the above description of example embodiments of the present disclosure, various features of the present disclosure are sometimes grouped together in a single example embodiment, figure, or description thereof in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this approach to the present disclosure should not be interpreted as reflecting an intention that the claims require more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, each inventive aspect lies in less than all the features of a single, previously disclosed example embodiment. Therefore, the claims following the specification are hereby expressly incorporated into this specification, with each claim standing on its own as a separate example embodiment of the present disclosure.

[0069] In addition, although some example embodiments described herein include some features included in other example embodiments and do not include other features included in other example embodiments, as will be understood by those skilled in the art, combinations of features from different example embodiments are intended to be within the scope of this disclosure and to form different example embodiments. For example, in the appended claims, any of the example embodiments claimed for protection may be used in any combination.

[0070] In the description provided herein, numerous specific details are set forth. However, it should be understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, device structures, and techniques are not shown in detail to avoid obscuring the understanding of this specification.

[0071] Therefore, although there has been described what is believed to be the best mode of the present disclosure, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the present disclosure, and it is intended to claim all such changes and modifications that fall within the scope of the present disclosure. For example, steps may be added or deleted to the methods described within the scope of the present disclosure.

[0072] Various aspects of the present invention can be understood from the following enumerated example embodiment (EEE):

[0073] EEE1. A method for improving dialog intelligibility during playback of audio data on a playback device, wherein the audio data comprises dialog audio data and at least one of music and effects audio data, the method comprising the steps of:

[0074] (a) determining a volume mixing ratio based on the playback volume value;

[0075] (b) mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio; and

[0076] (c) Output the mixed audio data for playback.

[0077] EEE2. The method according to EEE 1, wherein, in step (b), mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio includes: applying a gain to at least the dialogue audio data.

[0078] EEE3. The method according to EEE 1 or 2, wherein, in step (a), the playback volume value is mapped to a sound pressure level, and the volume mixing ratio is determined as a function of the sound pressure level.

[0079] EEE4. The method according to any one of EEEs 1 to 3, wherein, in step (a), the playback volume value is set based on the volume value of the playback device; and optionally

[0080] Wherein, the volume value setting is a user-defined value.

[0081] EEE5. The method according to any one of EEEs 1 to 4, wherein in step (a), the volume mixing ratio is further determined based on the ambient sound pressure level; and optionally

[0082] The ambient sound pressure level is determined based on measurement results of one or more microphones.

[0083] EEE6. The method according to any one of EEEs 1 to 5, wherein, before step (a), the method further comprises:

[0084] (i) receiving a bit stream comprising compressed audio data; and

[0085] (ii) core decoding the compressed audio data by a core decoder and providing the dialogue audio data, and at least one of the music and effect audio data.

[0086] EEE7. The method according to EEE 6, wherein in step (i), the received bit stream further includes information about an audio content type, and wherein in step (a), the volume mixing ratio is further determined based on the audio content type; and optionally

[0087] The audio content type includes one or more of movie audio content, news program audio content, sports broadcast audio content, and interlude audio content.

[0088] EEE8. The method according to EEE 6 or 7, wherein the method further comprises analyzing the compressed audio data to provide the dialogue audio data, and at least one of the music and effects audio data. EEE

[0089] EEE9. A playback device implementing a media playback system to improve dialog intelligibility during playback of audio data, the audio data comprising dialog audio data and at least one of music and effects audio data, the media playback system comprising:

[0090] (a) an audio processor, the audio processor being configured to determine a volume mixing ratio based on a playback volume value;

[0091] (b) a sound mixer for mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio; and

[0092] (c) A controller configured to output the mixed audio data for playback.

[0093] EEE10. The playback device according to EEE 9, wherein the mixer is further configured to apply a gain to at least the dialogue audio data; and / or

[0094] The audio processor is configured to map the playback volume value to a sound pressure level to determine the volume mixing ratio as a function of the sound pressure level.

[0095] EEE11. The playback device according to EEE 9 or 10, wherein the playback device further comprises a user interface for receiving a volume value setting of a user, and wherein the playback volume value is based on the volume value setting. EEE12.

[0096] EEE12. The playback device according to any one of EEEs 9 to 11, wherein the playback device further includes one or more microphones for determining an ambient sound pressure level, and wherein the audio processor is configured to further determine the volume mixing ratio based on the ambient sound pressure level. EEE13. The playback device according to any one of EEEs 9 to 11, wherein the playback device further includes one or more microphones for determining an ambient sound pressure level, and wherein the audio processor is configured to further determine the volume mixing ratio based on the ambient sound pressure level.

[0097] EEE13. The playback device according to any one of EEEs 9 to 12, wherein the playback device further comprises:

[0098] (i) a receiver for receiving a bit stream comprising compressed audio data; and

[0099] (ii) a core decoder for core-decoding the compressed audio data and providing the dialog audio data, and at least one of the music and effect audio data.

[0100] EEE14. The playback device according to EEE 13, wherein the core decoder is further configured to analyze the compressed audio data to provide the dialogue audio data, and at least one of the music and effects audio data. EEE14.

[0101] EEE15. A computer program product comprising instructions adapted to cause a device having processing capabilities to perform a method according to any one of EEEs 1 to 8. EEE15.

Claims

1. A method for improving the intelligibility of dialogue during playback of audio data on a playback device, wherein: The audio data includes dialogue audio data, and at least one of music and effect audio data, and the method includes the following steps: determining a volume mixing ratio as a function of the sound pressure level based on the playback volume value by mapping the playback volume value to the sound pressure level, wherein when the audio data includes one of music and effect audio data, the volume mixing ratio refers to a ratio of the volume of the dialogue audio data to the volume of the one of the music and effect audio data, and when the audio data includes both music and effect audio data, the volume mixing ratio refers to a ratio of the volume of the dialogue audio data to the sum of the volumes of the music and effect audio data; mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio; and Outputs the mixed audio data for playback.

2. The method according to claim 1, wherein Mixing the dialog audio data with at least one of the music and effect audio data based on the volume mixing ratio includes applying at least a gain to the dialog audio data.

3. The method according to claim 1 or 2, wherein The playback volume value is set based on the volume value of the playback device.

4. The method according to claim 3, wherein: The volume value setting is a user-defined value.

5. The method according to claim 1 or 2, wherein: The volume mixing ratio is further determined based on an ambient sound pressure level.

6. The method according to claim 5, wherein: The ambient sound pressure level is determined based on measurements of one or more microphones.

7. The method according to claim 1 or 2, wherein: Before determining the volume mixing ratio, the method further includes: receiving a bitstream comprising compressed audio data; and The compressed audio data is core-decoded by a core decoder and provides the dialog audio data, and at least one of the music and effect audio data.

8. The method according to claim 7, wherein: The received bitstream further includes information on an audio content type, and wherein the volume mixing ratio is further determined based on the audio content type.

9. The method according to claim 8, wherein The audio content type includes one or more of movie audio content, news program audio content, sports broadcast audio content, and episode audio content.

10. The method according to claim 7, wherein: The method further includes analyzing the compressed audio data to provide the dialog audio data, and at least one of the music and effects audio data.

11. A playback device implementing a media playback system to improve dialog intelligibility during playback of audio data, the audio data comprising dialog audio data and at least one of music and effects audio data, the media playback system comprising: an audio processor for determining a volume mixing ratio as a function of the sound pressure level based on the playback volume value by mapping the playback volume value to the sound pressure level, wherein when the audio data includes one of music and effect audio data, the volume mixing ratio refers to a ratio of the volume of the dialogue audio data to the volume of the one of the music and effect audio data, and when the audio data includes both music and effect audio data, the volume mixing ratio refers to a ratio of the volume of the dialogue audio data to the sum of the volumes of the music and effect audio data; a sound mixer for mixing the dialogue audio data with at least one of the music and effect audio data based on the volume mixing ratio; as well as A controller is configured to output the mixed audio data for playback.

12. The playback device according to claim 11, wherein The audio mixer is further configured to apply a gain to at least the conversational audio data.

13. The playback device according to claim 11 or 12, wherein: The audio processor is configured to map the playback volume value to a sound pressure level to determine the volume mixing ratio as a function of the sound pressure level.

14. The playback device according to claim 11 or 12, wherein: The playback device further comprises a user interface for receiving a volume value setting of a user, and wherein the playback volume value is based on the volume value setting.

15. The playback device according to claim 11 or 12, wherein: The playback device further comprises one or more microphones for determining an ambient sound pressure level, and wherein the audio processor is configured to determine the volume mixing ratio further based on the ambient sound pressure level.

16. The playback device according to claim 11 or 12, wherein: The playback device further comprises: a receiver for receiving a bit stream comprising compressed audio data; and A core decoder is configured to perform core decoding on the compressed audio data and provide the dialog audio data and at least one of the music and effect audio data.

17. The playback device according to claim 16, wherein: The core decoder is further configured to analyze the compressed audio data to provide the dialog audio data, and at least one of the music and effects audio data.

18. A computer program product having instructions adapted to cause a device having processing capabilities to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • System, method and storage medium for loudness-based audio-signal compensation

    CN108337606A

  • Object-based audio signal balancing

    CN108432130A