Conversation intelligibility enhancement method and system

By separating and adjusting the dialogue and non-dialogue components in the audio signal, the problem of insufficient dialogue intelligibility in mixed audio tracks for consumers is solved, thereby improving dialogue intelligibility and the viewing experience.

CN120660137APending Publication Date: 2025-09-16DTS INC(US)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480011474.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-16
Filing Date
2024-02-07
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Consumers often have trouble understanding dialogue when watching content with mixed soundtracks because dialogue intelligibility is disrupted, leading to frequent volume adjustments and a detrimental viewing experience.

Method used

By separating the audio signal into dialogue and non-dialogue components, processing and adjusting their loudness separately, ensuring that the dialogue component reaches the minimum intelligibility level, and appropriately adjusting the loudness of the non-dialogue component to combine them into an audio signal with improved dialogue intelligibility.

Benefits of technology

This improves dialogue intelligibility, reduces the need to adjust dialogue and non-dialogue volume, and enhances the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660137A_ABST
    Figure CN120660137A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate to a method and system for enhancing conversation intelligibility in a raw audio signal including a conversation component and a non-conversation component. The method comprises providing a dialogue component of the original audio signal in a first individual audio signal, providing a non-dialogue component of the original audio signal in a second individual audio signal, separately processing the first individual audio signal and the second individual audio signal, wherein processing the first and second individual audio signals comprises processing the loudness of the first individual audio signal and / or the second individual audio signal, and combining the processed first and second individual audio signals to provide a processed audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications and priority claims

[0002] This application is related to and claims priority to U.S. Provisional Application No. 63 / 483,737, filed on February 7, 2023, and entitled “DIALOG ENHANCEMENT ECOSYSTEM FOR STREAMING AND BROADCASTING MEDIA,” and is related to and claims priority to U.S. Provisional Application No. 63 / 508,811, filed on June 16, 2023, and entitled “DIALOG ENHANCEMENT ECOSYSTEM FOR STREAMING AND BROADCASTING MEDIA,” which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to enhancing dialogue intelligibility in an audio signal that includes both dialogue and non-dialogue components. For example, an audio soundtrack of video content that may be played back on a media device such as a set-top box, television, laptop, or the like. The mixed soundtrack may be composed of narrative dialogue and non-dialogue audio components. The non-dialogue components may include, for example, ambient or environmental sounds, music, and sound effects. Background Art

[0004] Consumers often have trouble understanding dialogue in mixed soundtracks because it's being played back through the sound reproduction system in their playback environment. This can be difficult for consumers to understand due to a number of factors that can reduce the intelligibility of spoken words. This often forces consumers to constantly adjust content volume levels, turning it down if the music and sound effects are too loud and up if the dialogue is too quiet. This disconnects them from the content viewing experience and can cause frustration. Furthermore, simply increasing the device's master volume level doesn't resolve intelligibility issues because it increases the volume of both the dialogue and distracting non-dialogue tracks.

[0005] Accordingly, there is a need to improve the intelligibility of the dialogue component in the audio signal. Summary of the Invention

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0007] One aspect of the present invention provides a method for enhancing the intelligibility of dialogue in an original audio signal, the original audio signal comprising a dialogue component and a non-dialogue component. The method comprises providing the dialogue component of the original audio signal in a first separate audio signal, providing the non-dialogue component of the original audio signal in a second separate audio signal, separately processing the first separate audio signal and the second separate audio signal, wherein processing the first and second separate audio signals comprises processing the loudness of the first separate audio signal and / or the second separate audio signal, and combining the processed first and second separate audio signals to provide a processed audio signal.

[0008] Thus, aspects of the present invention are based on the concept of independently analyzing and processing the dialogue and non-dialogue components of an audio signal. This allows for separate processing of the dialogue and non-dialogue components, thereby providing improved intelligibility of the dialogue components. In particular, the dialogue components can be adjusted or equalized differently than the non-dialogue components. For example, the dialogue components can receive loudness normalization and, optionally, spectral enhancement, while the non-dialogue components can receive dynamic range compression, as discussed below.

[0009] In the meaning of the present invention, the dialogue component is the component related to the spoken language (including the silent intervals between the spoken words), wherein the non-dialogue component is related to other components of the audio signal, such as music and sound effects. The dialogue component can also be called foreground speech, wherein the non-dialogue component can also be called background sound.

[0010] It is noted that the method steps are not necessarily performed by the same entity. For example, the steps of providing the dialogue component of the original audio signal in a first separate audio signal and providing the non-dialogue component of the original audio signal in a second separate audio signal can be performed in the head-end system or in the cloud. The steps of separately processing the first separate audio signal and the second separate audio signal and combining the processed signals can be performed on the consumer device. In another embodiment, dialogue separation is performed in a higher-power device at the customer site (such as a set-top box or TV), while processing the first and second separate audio signals is provided by another customer device (such as a consumer device). However, in other embodiments, all steps are implemented in the same device (such as a consumer device).

[0011] In an embodiment, providing the dialogue component in the first separate audio signal and providing the non-dialogue component in the second separate audio signal comprises receiving the first and second separate audio signals from a source from which the first and second separate audio signals are separately obtainable. Accordingly, if separate dialogue-only and non-dialogue signals are already available at the production stage, they may be used directly. For example, object-based audio (such as DTS: Dolby or ) to provide discrete conversation flows.

[0012] In another embodiment, providing a dialogue component in a first separate audio signal and providing a non-dialogue component in a second separate audio signal includes separating the dialogue component from the non-dialogue component in the original audio signal. Separating the dialogue component from the non-dialogue component can be achieved by a variety of methods. For example, dialogue separation can be achieved by deep learning models (such as convolutional neural networks and recurrent neural networks), which have the ability to isolate different sources, including dialogue. There are currently commercially available neural network-based dialogue separation products, such as RX Dialogue Isolate from iZotope, Inc. Another approach relies on analyzing object-based audio, as discussed in J. Paulus et al., "Source Separation for Enabling Dialogue Enhancement in Object-Based Broadcast with MPEG-H," J. Audio Eng. Soc., Vol. 67, No. 7 / 8, July / August 2019.

[0013] In an embodiment, processing the first individual audio signal comprises determining a short-term loudness level of the first individual audio signal, and determining whether the determined short-term loudness level is less than a predefined minimum dialog loudness level DLL MIN When the determined short-term loudness level is less than the predefined minimum dialogue loudness level DLL MIN In the case of MIN If the determined short-term loudness level is not less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is not modified.

[0014] In this aspect of the invention, the parameter "minimum dialogue loudness level" (DLL MIN ) defines the target short-term average loudness level for the dialog component. If the measured dialog level is less than the target DLL MIN , then the first individual audio signal (dialogue signal) is amplified towards the target minimum level. It should be noted that if the dialogue loudness is already above the DLL MIN , then no signal modifications are applied. DLL MIN Typical default values ​​for will match industry recommendations for digital dialogue loudness levels. This value is typically in the range of -22 LUFS to -27 LUFS. Since most program content will adhere to these recommendations, it may not be necessary to significantly modify the dialogue loudness to achieve this goal.

[0015] In a further embodiment, the processed first individual audio signal is spectrally enhanced before being combined with the processed second individual audio signal.Such spectral enhancement is optional and may include applying a specific filter to the dialogue component.

[0016] In yet further embodiments, the method further comprises determining voice activity in the first individual audio signal, and moving the first individual audio signal towards the minimum dialogue loudness level DLL only if voice activity has been determined. MIN This embodiment is based on the idea that the dialog loudness should be boosted only when there is speech activity. Otherwise, very low-level dialog components in the dialog signal (such as background dialog or noise / artifacts) will be subjected to undesirably high gain to match the DLL. MIN This can result in undesirable loudness spikes during transitions from quiet segments to segments with narrative dialogue, as the normalization ballistics need time to adjust to the rapid loudness changes.

[0017] One example of determining voice activity includes determining whether a short-term loudness level of the first individual audio signal is above a threshold dialogue loudness level DLL THRESH , wherein the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL only if the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL THRESH DLL towards the minimum dialogue loudness level MIN The first separate audio signal is amplified. In this embodiment, the parameter "threshold dialogue loudness level (DLL) THRESH )” is used as a voice activity detector (VAD), below which the dialogue loudness level is not boosted. In addition, DLL THRESH Helps avoid amplifying low-level processing artifacts from the previous dialog separation process.

[0018] It is noted that this aspect of the invention is not limited to specific implementations using VAD as the threshold parameter. It can also include other voice activity detection implementations, such as those using output masks from a dialogue separation process or machine learning algorithms designed specifically for voice activity detection.

[0019] In an embodiment, amplifying the first separate audio signal includes using a dynamic range processor that applies gain using a modifiable curve defined by a plurality of control points. Such a modifiable curve may be defined by five control points (x / y coordinates) and allows the processor to function as a compressor, expander, loudness adjuster, or a hybrid of these modes. A smoothing parameter may also be incorporated to ensure seamless transitions between operating zones.

[0020] In an embodiment, processing the second individual audio signal comprises determining a short-term loudness level of the first individual audio signal or obtaining a predefined minimum dialogue loudness level DLL of the first individual audio signal. MIN , determining a short-term loudness level of the second individual audio signal, and determining a difference between the short-term loudness level of the first individual audio signal and the short-term loudness level of the second individual audio signal or a minimum dialogue loudness level DLL MIN Is the difference between the short-term loudness level of the first and second separate audio signals less than a predefined minimum dialogue-to-non-dialogue ratio D2ND? MIN If it is less than, then the loudness level of the second separate audio signal is reduced so that the difference approaches the minimum dialogue to non-dialogue ratio D2ND MIN If not, then the second separate audio signal is not modified.

[0021] In this embodiment, the parameter "minimum dialogue to non-dialogue ratio" (D2ND MIN ) represents the minimum difference between the short-term dialogue loudness level and the short-term non-dialogue loudness level. If the measured levels have a loudness difference less than this value, the non-dialogue signal is compressed until the average difference between the dialogue loudness level and the non-dialogue loudness level approaches D2ND MIN It is important to note that non-dialogue levels should only be lowered when necessary.

[0022] In an embodiment, reducing the loudness level of the second individual audio signal comprises compressing the dynamic range of the second individual audio signal. This may be achieved by using a dynamic range processor that applies a gain using a modifiable curve determined by a plurality of control points, the control points allowing the processor to act as a compressor and / or loudness adjuster. For example, if the difference mentioned above is below D2ND MIN value, then a certain compression ratio (such as 2:1) can be achieved so that the difference is close to D2ND MIN value.

[0023] In yet another embodiment, a short-term loudness level (of a first separate audio signal including a dialog component or a second separate audio signal including a non-dialog component) is determined for a continuous window of a predefined length, wherein the loudness level is determined according to an industry standard. The window may be in the range of 10 ms to 100 ms. For example, the window may have a length of 20 ms. The industry standard according to which the loudness level is determined may be the ITU-R BS.1770 standard, in which loudness is expressed in LKFS (K-weighted loudness relative to full scale) or its synonym LUFS (Loudness Units relative to Full Scale) introduced in EBU R128, which is a standard loudness measurement unit used for audio normalization in broadcast television systems and other video and music streaming services. In particular, the first iteration of this standard, ITU-R BS.1770-1, may be used for determining loudness, as it is particularly well-suited to handling instantaneous loudness fluctuations through continuous, short-term measurements.

[0024] In another embodiment, the first individual audio signal and the second individual audio signal are processed in multiple processing paths, the processing paths including a universal processing path in which the processed audio signals are provided to any number of listeners, and at least one individualized processing path in which the processed audio signals are provided to an individual listener, wherein processing the first individual audio signal and / or processing the second individual audio signal includes using parameters personalized for the individual listener during processing. The personalized parameters may include a personal hearing profile and subjective listening preferences specific to the listener. This embodiment addresses the situation where not everyone in the listening space desires to hear the common audio output from the dialogue enhancement system. Thus, alternative levels of dialogue processing can be implemented based on the needs of one or more listeners.

[0025] In an embodiment, the original audio signal is an audio soundtrack, i.e., the sound that accompanies and is synchronized with the images of a movie, television program, video game, radio program, etc. The original soundtrack can be in the form of a digital audio file. However, the present invention is not limited to this embodiment. For example, the original audio signal can be a live audio signal.

[0026] In another embodiment, the original audio signal is a stereo or multi-channel signal. Provision can be made for each channel of the stereo or multi-channel signal to include a dialog component in a first separate audio signal and a non-dialog component in a second separate audio signal, wherein the first and second separate audio signals are processed separately and then combined. Accordingly, in this embodiment, the number of channels at the input is maintained at the output.

[0027] In another embodiment, the original audio signal is a stereo signal, wherein the stereo signal is upmixed into a three-channel signal comprising a center channel, a left channel, and a right channel, wherein the signal component of the stereo signal originally panned to the center is extracted to the center channel. Furthermore, provision is made for providing a dialogue component only for the center channel in a first separate audio signal and a non-dialogue component in a second separate audio signal. The first separate audio signal and the second separate audio signal are processed according to the present invention, wherein the second separate audio channel is combined with the left and right channels for loudness processing. This embodiment requires that the channel dialogue separation be performed only on a single channel, thereby reducing complexity.

[0028] In another embodiment, the original audio signal is a multi-channel signal including a center channel and a plurality of additional channels, wherein, for the center channel only, a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal, assuming that dialogue is most prevalent in the center channel. The first separate audio signal and the second separate audio signal are processed according to the present invention, wherein the second separate audio channel is combined with the additional channels for loudness processing. This embodiment requires that channel dialogue separation be performed only on a single channel of the multi-channel signal, thereby reducing complexity.

[0029] In another embodiment, the original audio signal is a multi-channel signal comprising a center channel, a left channel, a right channel, and further channels, wherein the center channel, the left channel, and the right channel are downmixed to two channels, wherein for each of the two downmixed channels, a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal. This embodiment assumes that dialogue is primarily present in the front (center, left, right) channels. The first separate audio signal and the second separate audio signal are processed according to the present invention, wherein the second separate audio signal is combined with the further channels for loudness processing. This embodiment requires that channel dialogue separation be performed only on a single channel of the multi-channel signal, thereby reducing complexity.

[0030] In another embodiment, the processed first and second individual audio signals are further processed by applying spatial audio processing and / or specific algorithms before the processed first and second individual audio signals are combined. This embodiment is based on the recognition that the first individual audio signal having a dialog component and the second individual audio signal having a non-dialog component can be kept separate before combining for further downstream processing. For example, the further processing can include applying algorithms that include spatial audio processing for headphones and speakers. Other examples of downstream processing include algorithms that are better applied to only the non-dialog components of the input signal, such as bass boost.

[0031] According to another aspect of the present invention, a method for enhancing dialogue in an original audio signal including a dialogue component and a non-dialogue component is provided. The method includes receiving the dialogue component of the original audio signal in a first separate audio signal, receiving the non-dialogue component of the original audio signal in a second separate audio signal, and separately processing the first separate audio signal and the second separate audio signal, wherein processing the first and second separate audio signals includes processing the loudness of the first separate audio signal and / or the second separate audio signal.

[0032] This aspect of the invention focuses on separate processing of the first individual audio signal and the second individual audio signal.The method may be implemented at a consumer site in a consumer device such as a television, laptop, smartphone or headset.

[0033] According to another aspect of the present invention, a system for enhancing dialogue in an original audio signal comprising a dialogue component and a non-dialogue component is provided. The system includes a dialogue separation unit configured to provide the dialogue component of the original audio signal in a first separate audio signal and to provide the non-dialogue component of the original audio signal in a second separate audio signal. The system also includes a loudness processing unit configured to separately process the first separate audio signal and the second separate audio signal, wherein processing the first and second separate audio signals includes processing the loudness of the first separate audio signal and / or the second separate audio signal. Furthermore, an audio mixer is provided, configured to combine the processed first and second separate audio signals to provide a processed audio signal.

[0034] It should be noted that the dialog separation unit, loudness processing unit, and audio mixer are not necessarily included in the same entity. For example, the dialog separation unit can be implemented in the headend or cloud. The loudness processing unit and audio mixer can be implemented on the consumer device. In another embodiment, the dialog separation unit is implemented in a higher-power device at the customer site, such as a set-top box or television, while the loudness processing unit and audio mixer are implemented in another consumer device, such as a consumer device.

[0035] Embodiments of the system correspond to the embodiments of the method discussed above. For example, the dialogue separation unit can be configured to receive a dialogue component and a non-dialogue component from a source from which the first and second separate audio signals can be separately obtained. Alternatively, the dialogue separation unit can be configured to provide the dialogue component and the non-dialogue component by separating the dialogue component from the non-dialogue component in the original audio signal.

[0036] According to yet another aspect of the present invention, a non-transitory computer-readable medium having executable instructions stored thereon is provided, wherein when the instructions are executed by a processor, the following operations are performed: providing a dialogue component of an original audio signal in a first separate audio signal, providing a non-dialogue component of the original audio signal in a second separate audio signal, separately processing the first separate audio signal and the second separate audio signal, wherein processing the first and second separate audio signals includes processing loudness of the first separate audio signal and / or the second separate audio signal, and combining the processed first and second separate audio signals to provide a processed audio signal.

[0037] Embodiments of the non-transitory computer readable medium correspond to the embodiments of the method discussed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Throughout the drawings, reference numerals are repeated to indicate correspondence between referenced elements.The drawings are provided to illustrate embodiments of the invention described herein and not to limit its scope.

[0039] Figure 1 is an overall system architecture of a system for enhancing dialogue in an original audio signal, wherein the system architecture includes a dialogue separation unit, a loudness processing unit, and an audio mixing unit;

[0040] Figure 2 Shown Figure 1 An embodiment of a conversation separation unit;

[0041] Figure 3 is a flow chart of a method for enhancing dialogue in an original audio signal;

[0042] Figure 4 is a flow chart of a method for processing a first separate audio signal comprising a dialog component of an original audio signal;

[0043] Figure 5 is a flow chart of a method for processing a second separate audio signal comprising a non-dialogue component of an original audio signal;

[0044] Figure 6 yes Figure 1 An embodiment of a loudness processing unit;

[0045] Figure 7 is an embodiment of a dynamic range processor;

[0046] Figure 8 The figure shows the loudness measurement of multichannel audio signals according to ITU-R BS.1770-1;

[0047] Figure 9 is an example of a dialogue loudness modification curve;

[0048] Figure 10 This is an example of a non-dialogue loudness modification curve;

[0049] Figure 11 is a flow chart of an example method for processing a first separate audio signal comprising a dialog component of an original audio signal;

[0050] Figure 12 is a flow chart of an example method for processing a second separate audio signal comprising a non-dialogue component of an original audio signal;

[0051] Figure 13 A system for implementing centralized conversation separation and distributed personalized conversation enhancement is illustrated;

[0052] Figure 14 An additional system for implementing centralized conversation separation and distributed personalized conversation enhancement is illustrated, wherein two-channel wireless transmission and endpoint-specific post-processing are provided;

[0053] Figure 15 yes Figure 1 A stereo input / output implementation method for a general system;

[0054] Figure 16 yes Figure 1 an alternative stereo implementation of the general system of claim 1 wherein a separate dialogue channel is retained for additional post-processing;

[0055] Figure 17 yes Figure 1 An alternative stereo implementation of the general system of claim 1, wherein mono dialogue separation is achieved;

[0056] Figure 18 An embodiment of a stereo shuffler is shown;

[0057] Figure 19 yes Figure 1 An alternative stereo implementation of the general system of claim 1, wherein mono dialogue separation is achieved using a stereo shuffler;

[0058] Figure 20 yes Figure 1 A multi-channel implementation of a general system wherein mono-channel dialogue separation is achieved;

[0059] Figure 21 yes Figure 1 Another multi-channel embodiment of the general system of claim 1, wherein stereo dialogue separation is implemented; and

[0060] Figure 22 yes Figure 1 Another multi-channel implementation of the general system of FIG. 1 , in which 3-channel dialogue separation is implemented. DETAILED DESCRIPTION

[0061] The following description describes various embodiments of methods and systems for enhancing dialogue in an original audio signal that includes both dialogue and non-dialogue components.

[0062] Figure 1 A general system architecture is shown. The system includes a dialogue separation unit 1, a loudness processing unit 2, and an audio mixing unit 3. The dialogue separation unit 1 receives an original audio signal A and separates the dialogue component from the other components of the signal. The dialogue component is provided in a first separate audio signal 11, and the non-dialogue component is provided in a second separate audio signal 12. Once the dialogue component is separated, it can be analyzed and processed separately from the non-dialogue component. The dialogue component includes spoken language. The non-dialogue component represents any signal component that is not a narrative dialogue signal. These can include music and sound effects.

[0063] The original audio signal A can be a digital audio signal. It can be stored in a file or streamed. For example, the original audio signal is the audio track of video content that can be played back on a media device such as a set-top box, a television, a laptop, etc. In another example, the original audio signal is a live radio transmission.

[0064] The loudness processing unit 2 receives a first individual audio signal 11 including a dialogue component and a second individual audio signal 12 including a non-dialogue component, and separately processes the first individual audio signal 11 and the second individual audio signal 12. In particular, the audio signals 11 and 12 are analyzed and processed separately for loudness, as will be described in detail in relation to the embodiment of the present invention. Figure 4 、 Figure 5 、 Figure 11 and Figure 12 As discussed in the embodiments of FIG. , the loudness processing unit 2 is configured to ensure that the dialogue loudness level never falls below the listener's desired level and also to adapt the non-dialogue loudness level so that a consistent minimum dialogue-to-non-dialogue ratio exists. The audio mixer 3 receives the processed first individual audio signal 110 and the processed second individual audio signal 120 and combines them to provide a processed audio signal B with dialogue enhancement as output. For example, the audio mixer 3 outputs a dialogue-enhanced audio track.

[0065] It is important to note that system components 1, 2, and 3 can be split between devices. For example, the conversation separation unit 1 can reside in the headend or cloud, while the conversation processing unit 21 can reside on the consumer device. In another example, one higher-power device (e.g., a set-top box or TV) includes the conversation separation unit 1, while another device includes the conversation processing unit 2 and audio mixer 3.

[0066] Figure 2 Shown Figure 1 An example embodiment of a dialogue separation unit 1 is shown. The purpose of the dialogue separation unit 1 is to separate an incoming raw audio signal A into two separate signals 11 and 12, each having a dialogue component and a non-dialogue component. In the depicted embodiment, but not necessarily, a machine learning network is trained to perform this task. However, other audio source separation techniques may also be used.

[0067] according to Figure 2 In an embodiment, an original audio mixture is converted to the STFT domain in a short-time Fourier transform (STFT) unit 101, and relevant audio features are extracted from the STFT data in a feature extraction unit 102. For example, the data can be converted to a logarithmic magnitude representation, and the linear frequency bands can be grouped into a smaller number of key perceptual bands. The input features are taken as the input layer of a machine learning network 103 (e.g., a network based on the U-NET architecture), with the goal of outputting a set of features. In this case, the output features represent a set of frequency-domain weights that represent the filtering required to isolate speech and / or non-speech signal components. The output features are processed so that a linear STFT domain representation can be derived. These per-frequency-bin weights are applied to a delayed version of the original STFT domain signal (delayed by a delay unit 105) in a speech filter unit 106. The delay is used to compensate for the processing delays involved in feature extraction, network inference, and reconstruction. The resulting filtered STFT data is then converted back to the time domain using an inverse STFT process in an STFT unit 107. The extracted dialogue component is output in a first individual audio signal 11 and the non-dialogue component is output in a second individual audio signal 12 .

[0068] The machine learning network may have been trained on a large database of separated examples of conversational and non-conversational signals (which represent desired outputs) and their corresponding mixtures (provided at the input).

[0069] If the first and second separate audio signals 11 , 12 are readily available from a source such as a production stage source, then no dialog separation is required, so that the dialog separation unit 2 can be bypassed or simplified to a unit that only receives the first and second separate audio signals 11 , 12 .

[0070] Figure 3 FIG. 1 is a flow chart of a method for enhancing dialogue in an original audio signal. In steps 301 and 302, the dialogue component and the non-dialogue component of the original audio signal are provided in first and second separate audio signals. The dialogue component and the non-dialogue component can be separated by a dialogue separation technique (such as a Figure 2In step 303, the first and second individual audio signals are processed separately, wherein the loudness of the first and / or second individual audio signals is adjusted to improve the intelligibility of the dialogue. This can be done in Figure 1 The loudness processing unit 2 is implemented, and its more specific embodiment is about Figure 6 Discussion proceeds. In step 304, the processed first and second individual audio signals are combined to provide a processed audio signal.

[0071] Figure 4 and Figure 5 An example of how to process the first and second separate audio signals is provided in . Figure 4 In a first step 401, a short-term dialogue loudness level (STDLL) of a first individual audio signal is determined. This loudness level determination can be based on a standard loudness measurement, as will be discussed below. The measurement of short-term loudness means measuring the loudness over a window. For example, the short-term loudness is measured over a small window of 20 ms, with a look-ahead of 10 ms so that the window is centered around the current sample, but it should be understood that this is merely an example and other window lengths and look-aheads can also be implemented.

[0072] In step 402, it is determined whether the determined STDLL is less than a predefined minimum dialogue loudness level (DLL MIN If this is the case, then in step 403 the first individual audio signal is passed towards the DLL MIN If not, then the first separate audio signal is not modified. Accordingly, the loudness of the first separate audio signal with the dialogue input is normalized to a predefined minimum dialogue loudness level DLL MIN .

[0073] It is important to note that this approach differs significantly from prior art methods. Prior art methods implement volume normalization to ensure that the level of quiet passages is increased to a more audible level, and implement dynamic range compression to ensure that the level of excessively loud passages is reduced. However, the combination of these processes can produce audible artifacts when applied to the original soundtrack mix. For example, with volume normalization enabled, quiet non-dialogue passages (e.g., a non-dialogue forest tree landscape) will become unnaturally loud. Similarly, using dynamic range compression on the louder non-dialogue soundtrack components will also affect the louder dialogue, making it more difficult to hear the dialogue in the presence of non-dialogue sounds. In addition, these solutions often apply a high-frequency spectral boost to the signal to increase dialogue clarity. Such filters are typically applied to the non-dialogue components in the mix as well as the dialogue components, affecting the overall spectral balance of the soundtrack.

[0074] On the other hand, when the dialogue component and the non-dialogue component are first separated into two separate signals, the short-term loudness of the separated dialogue component can be analyzed independently of the non-dialogue component. This allows loudness normalization to be applied only to the dialogue signal.

[0075] Figure 5 This is an example of how a second, separate audio signal including a non-dialogue component can be processed. In step 501, the short-term non-dialogue loudness level (STNDLL) of the second, separate audio signal including the non-dialogue component is determined. Again, this is achieved by windowing the second, separate audio signal. Similar to the processing of the first, separate audio signal, the short-term loudness can be measured over a small window of 20 ms, with a look-ahead of 10 ms. In general, the window size used to determine the short-term loudness of the first, separate audio signal can be the same as the window size used to determine the short-term loudness of the second, separate audio signal, or it can be different. In this regard, it should be noted that the technique of applying different processing to the dialogue and non-dialogue streams is associated with the additional advantage that different loudness windows can be applied to each stream.

[0076] In step 502, the minimum dialogue loudness level DLL is determined MIN (with about Figure 3 The same DLLs already discussed MIN The difference between the STNDLL level) and the STNDLL level determined in step 501 is less than a predefined minimum dialogue to non-dialogue ratio D2ND. MIN If it is less than, then the loudness level of the second individual audio signal is reduced so that the difference approaches the ratio D2ND MIN (Step 503). If not, then the second individual audio signal is not modified (Step 504). Reducing the loudness level of the second individual audio signal may be achieved by compressing the dynamic range of the second individual audio signal.

[0077] Loudness level reduction is only applied to non-dialogue signals.

[0078] Alternatively, instead of determining the minimum dialogue loudness level DLL MIN Instead, it is determined and analyzed whether the difference between the short-term loudness level (STDLL) of the first individual audio signal and the STNDLL level is less than a predefined minimum dialogue to non-dialogue ratio D2ND. MIN .

[0079] Figure 6 Shown Figure 1 The loudness processing unit 2 comprises a unit 201 for determining the short-term loudness of a first individual audio signal 11. The unit 201 receives information about Figure 4 Predefined minimum dialogue loudness level DLL for discussionMIN As input parameters. The loudness processing unit 2 further comprises an amplification unit 202, which is configured to amplify the first individual audio signal 11 according to the control signal received by the unit 201. In addition, optionally, a spectrum enhancement unit 203 for the first individual audio signal 11 is provided. With respect to the second individual audio signal 12, a unit 204 for determining the short-term loudness of the second individual audio signal 12 is provided. The unit 204 receives information about Figure 5 The predefined minimum conversation-to-non-conversation ratio D2ND discussed MIN The loudness processing unit 2 further comprises a compression unit 205 configured to compress the second individual audio signal 12 according to the control signal provided by the unit 204. The loudness processing unit 2 outputs a processed first individual audio signal 110 and a processed second individual audio signal 120.

[0080] Parameter DLL MIN and D2ND MIN Can be set by the consumer or system integrator.

[0081] Units 201, 202 and 204, 205 can use Figure 7 4. The DSP can be adapted to handle both dialog and non-dialog input streams. The signal received at input 401 is replicated and delayed in a look-ahead delay unit 402 in a first path 408. For example, a 10ms look-ahead delay is incorporated to prepare the DRP to actively respond to incoming loudness transients. In a second path 409, a loudness measurement unit 403, a gain calculator 404, and a gain smoother 405 are provided. The determined gain is applied to the signal in path 408 in unit 406 and sent to output 407.

[0082] In the loudness measurement unit 403, the loudness of the signal is measured. In particular, the short-term loudness of the first individual audio signal 11 or the short-term loudness of the second individual audio signal 12 can be measured. The loudness measurement is performed according to industry standards. In an embodiment, the loudness is estimated using the industry standard ITU-R BS.1770-1 and is measured with a specific window size (such as 20ms) to ensure both accuracy and responsiveness. In the standard ITU-R BS.1770, loudness is expressed in LKFS (K-weighted loudness relative to full scale) or its synonym LUFS (Loudness Unit relative to Full Scale) introduced in EBU R128, which is a standard loudness measurement unit used for audio normalization in broadcast television systems and other video and music streaming services. In particular, the first iteration of this standard, ITU-R BS.1770-1, can be used, which is particularly suitable for handling instantaneous loudness fluctuations through continuous short-term measurements.

[0083] The ITU-R BS.1770-1 standard processes each audio channel by first applying a pair of second-order IIR filters for pre-filtering, followed by RLB (Revised Low-B Curve) filtering to simulate the frequency response of the human ear. The mean square energy of the filtered signal is then calculated over the measurement interval T, yielding the value z for each channel i. The mean square energy is determined as follows:

[0084]

[0085] After the mean square calculation, the channel-specific weights G are applied i , and finally get the total loudness value:

[0086]

[0087] Channel weights can be assigned as follows: Left (G L ): 1.0; right (G R ): 1.0; Center (G C ): 1.0; Left surround (G Ls ): 1.41; right surround (G Rs ): 1.41. Figure 8 Shown in.

[0088] Reference again Figure 7 , the gain computer 404 calculates the gain using a modifiable curve determined by five control points (x / y coordinates). This allows the processor to be used as a compressor, expander, loudness adjuster, or a hybrid of these modes. A smoothing parameter can also be incorporated into the gain smoother 405 to ensure seamless transitions between operating zones. The gain smoother can be a branching attack and release smoother with settings for fast / slow attack and release times, which allows the processor to quickly adapt to significant changes in input loudness.

[0089] Figure 9 and Figure 10 Shown are examples of a dialogue loudness modification curve 51 and a non-dialogue loudness modification curve 52. As mentioned, the gain is calculated using a modifiable curve determined by five control points, which provides flexibility in using it for both dialogue (upward gain) and non-dialogue (attenuation) processing. Figure 9 and Figure 10 Example curves for dialogue and non-dialogue are shown in . In these examples, the following parameter values ​​are used: DLL MIN =-20dBFS; D2ND MIN =10dB; and DLL TRESH = -80dB. In practice, these curves are smoothed to ensure seamless transitions between operating regions.

[0090] Parameter DLL MIN and D2ND MIN This has been discussed previously. TRESH Defines the threshold dialogue loudness level. This parameter is used as a Voice Activity Detector (VAD), below which dialogue loudness is not boosted. If there is no DLL TRESH , then very low level dialog components (e.g. background dialog or noise / artifacts on the dialog channel) will experience undesirably high gains to match the DLL MIN This can lead to undesirable loudness spikes during the transition from quiet segments to segments with narrative dialogue, as the normalized trajectory needs time to adapt to the rapid loudness change. THRESH Helps avoid amplifying low-level processing artifacts from the previous dialog separation process.

[0091] exist Figure 9 In the example shown, the dialogue loudness modification curve 51 consists of four loudness operating zones. In the first zone, from -inf dB to -80 dB, the loudness is attenuated by 10 dB. In the given example, -80 dB is the threshold dialogue loudness level DLL. TRESH As mentioned earlier, in the DLL TRESH Below this level, the dialogue loudness does not increase. In this example, it is even attenuated by 10dB. In the second zone, from -80dB to -70dB, there is a gradual transition from compression to normalization. The third zone, from -70dB to -20dB, is where the loudness is normalized to DLL. MIN In the fourth zone, from -20 dB to +inf dB, the signal remains unchanged.

[0092] exist Figure 10 In the embodiment of the present invention, the non-dialogue loudness modification curve 52 consists of two loudness operating regions. In the first region, from -inf dB to -30 dB, the signal remains unchanged. In the second region, from -30 dB to +inf dB, the signal is compressed with a compression ratio of 2:1. In other embodiments, the compression ratio can be different.

[0093] Figure 11 Instructions by Figure 6 The method implemented by the loudness processing unit 2 of FIG. 1 for processing a first separate audio signal including a dialogue component. The method is based on Figure 4 In step 111, a block of the dialog stream is input, where the block size is defined by a window. In an embodiment, the window size may be 20 ms. In step 112, the short-term loudness level STDLL of the dialog component is measured, such as with respect to Figure 7 and Figure 8In step 113, it is determined whether the short-term loudness level STDLL is greater than a predefined threshold dialogue loudness level DLL TRESH If not, then in step 116 the unmodified block of the dialog stream is output. If yes, then in step 114 it is further determined whether the short term loudness level STDLL is less than a predefined minimum dialog loudness level DLL MIN , the minimum dialogue loudness level DLL MIN is set according to industry recommendations and may be in the range of -22 LUFS to -27 LUFS (LUFS = "Loudness Unit Full Scale"). If not, then in step 116 the unmodified block of the dialog stream is output. If yes, then in step 115 the level of the dialog component is amplified so that the short term loudness level STDLL is close to or equal to the minimum dialog loudness level DLL MIN ,like Figure 9 The order of steps 113 and 114 can be interchanged.

[0094] Figure 12 Instructions by Figure 6 The method implemented by the loudness processing unit 2 of FIG. 1 for processing a second separate audio signal including a non-dialogue component. The method is based on Figure 5 In step 121, a block of the non-dialogue stream is input, where the block size is defined by a window. In an embodiment, the window size may be 20 ms. In step 122, the short-term loudness level STNDLL of the non-dialogue component is measured, such as with respect to Figure 7 and Figure 8 In step 123, the minimum dialogue loudness level DLL is determined MIN Is the difference between the short-term loudness level STNDLL measured in step 122 and the short-term loudness level STNDLL less than a predefined minimum dialogue to non-dialogue ratio D2ND? MIN If this is not the case, then the unmodified blocks of the non-dialog stream are output in step 125. If this is the case, then the non-dialog signal is compressed so that the mentioned difference DLL MIN -STNDLL approaches the minimum dialogue to non-dialogue ratio D2ND MIN The compression may include dynamic range compression. Figure 10 An example is given in .

[0095] In the above embodiments, it is assumed that everyone in the listening space will hear the common audio output from the dialogue enhancement system. In this case, the algorithm applies dialogue processing based on the preferences of the individual listeners, and everyone in the listening environment will hear it. Figure 13An alternative system is shown in which alternative levels of dialog processing can be provided to individual listeners based on their needs (eg, individuals with more severe hearing loss).

[0096] More specifically, in Figure 13 In the dialogue separation unit, the original audio signal A is split into dialogue and non-dialogue components, as shown in Figure 1 Alternatively, if the dialogue and non-dialogue components are already available separately, they can simply be received. The first and second audio signals 11, 12 are then provided to a series of processing paths. The first processing path is a universal processing path, in which the processed audio signal B is provided to any number of listeners who hear the same processed mix (e.g., through television speakers). The separated signals 11, 12 are processed in a universal dialogue enhancement unit 61, which is connected to the Figure 1 The loudness processing unit 2 corresponds to the audio mixing unit 3.

[0097] In addition, one or more optional individualized processing paths are provided, wherein individualized audio outputs B1, B2 are provided, since personalized parameters of the individual listener are taken into account when processing the first individual audio signal 11 and / or the second individual audio signal 12. For example, a listener-specific personal hearing profile or subjective listening preferences can be implemented in the individualized dialogue enhancement units 62, 63 and applied to the various processing blocks. The individualized audio outputs B1, B2 can be reproduced via headphones or hearing aids.

[0098] The personalized mix can be directed to in-ear monitors or headphones using a wired connection or using low-latency wireless technology such as Bluetooth or Ultra-Wideband Audio (UWB). Figure 14 As shown in . Figure 14 In the example, if the number of channels of audio input A is greater than two, the channels are downmixed to stereo in downmixing unit 64 and wirelessly transmitted to upmixing units 65, 66 associated with the respective users. After upmixing, personalized dialogue enhancement is achieved in units 62, 63. The outputs of dialogue enhancement units 61, 62, 63 can be further enhanced by 3D audio processors 67, 68, 69, which can be embedded in or attached to headphones used by the respective listeners.

[0099] In this regard, a variety of variations can be implemented. In one embodiment, individual listener hearing devices can include noise cancellation features to minimize interference from generic versions of the soundtrack played through the speakers. Multiple personalized mixes can be generated at a hub (e.g., a TV or set-top box) and transmitted simultaneously from that source. Additionally, in some embodiments, the multi-channel audio output content can be downmixed to stereo before transmission. In some embodiments, the multi-channel individual audio output content can be sent to multi-channel headphone virtualization technology, such as DTS Headphone:X, before wireless transmission. This virtualization algorithm can be applied on the transmission device or in the headphones. In some embodiments, the unmixed dialogue and non-dialogue audio streams are wirelessly transmitted to one or more headphone sets with the necessary processing capabilities, and personalized dialogue processing is applied in the headphones. In some embodiments, the dialogue and non-dialogue audio streams are downmixed or encoded (spatially or otherwise) in a manner that allows for lower bandwidth transmission and reception of the original audio channels. For example, the original dialogue and background channels can be stereo downmixed so that the dialogue is center-panned, and the multi-channel non-dialogue signals can be spatially encoded into stereo using an algorithm such as the DTS Neural Surround Downmixer. This stereo signal can then be "upmixed" or decoded back into discrete dialogue and non-dialogue streams on the receiving headphone device. In some embodiments, the original input signal is wirelessly transmitted to headphones with onboard processing capabilities (including a machine learning inference engine). The original audio track is transmitted to the headphones, and dialogue separation and individualized dialogue processing are applied on a processor attached to or embedded in the headphones.

[0100] The following will be about Figures 15 to 22 Different implementation topologies of the systems and methods are discussed.

[0101] exist Figure 15 In Figure 2, input audio signal A is a stereo signal (indicated as "2.0"). Separation into a first independent audio stereo signal 11 for the dialogue component and a second independent audio stereo signal 12 for the non-dialogue component is achieved by a dialogue separation unit 1, which has been trained to separate the dialogue and non-dialogue components from the stereo signal. The separated stereo dialogue and non-dialogue components are then passed to a stereo loudness processing unit 2, and the processed components are remixed. In addition, an output limiter 7 may be provided to ensure that the processed signal does not saturate downstream.

[0102] exist Figure 16In the embodiment of , the input audio signal A is a stereo signal. The output of the loudness processing unit 2 is kept separate for further downstream processing in an additional post-processing unit 30, in which the audio mixing unit is integrated. The post-processing unit 30 may include spatial audio processing algorithms for headphones and loudspeakers. Other examples of downstream processing include algorithms that are better suited for only the non-dialogue components of the input signal, such as bass enhancement. The basic concept of retaining separate dialogue and non-dialogue outputs from the loudness processing unit 2 can be applied to any of the topologies described below. Figure 15 As in , an output limiter 7 is also provided.

[0103] exist Figure 17 In the embodiment of , the input audio signal A is also a stereo signal. However, dialogue separation is only provided for mono. It is assumed that most of the narrative content dialogue is center panned. More specifically, the stereo input A with L, R channels is upmixed to 3 channels (L, C, R) in unit 81 so that the signal components that were initially panned to the center are extracted to a discrete center channel C. The resulting signal now takes two independent paths. The C component is directed to the mono dialogue separation unit 1. Since the main dialogue is assumed to be represented in the extracted center channel, it can be assumed that the residual (L, R) channels represent the non-dialogue components. These residual channels are delayed in delay 82 to compensate for the dialogue separation processing delay and redirected to the loudness processing unit 2 (because they are taken into account when processing the non-dialogue signal components in the loudness processing unit 2). The loudness processing unit 2 assesses the relative loudness of the separated dialogue and the residual C and (L, R) non-dialogue components and applies gain and attenuation to each signal component accordingly, with the (L, R) channels being amplified / compressed in the amplifier / compressor unit 83. The signals are then downmixed to a single stereo pair in the audio mixer 3 and in this case directed to the output limiter 7. Figure 17 In the embodiment of FIG. 1 , the loudness processing unit 2 includes multiple inputs. The first input is the extracted dialogue channel. One or more additional inputs are extracted non-dialogue channels. Furthermore, it is assumed that the L and R channels from the upmix unit 81 contain only non-dialogue.

[0104] Figure 18 and Figure 19 Regarding an alternative embodiment of the system, where the input audio signal A is a stereo signal, where dialogue separation and loudness processing are provided only for a single channel, this embodiment relates to situations where the implementer cannot use an active 2-3 upmixer. In this case, it is possible to apply passive upmixing using a stereo shuffler configuration. Figure 18 A typical stereo shuffler configuration 84 is shown in . The sum and difference of the left and right input channels L and R are formed twice for each channel. The output is the same as the input of a typical stereo shuffler.

[0105] Assuming that narrative dialogue is generally center-panned, the sum of the stereo input channels, L+R, will contain most of the center-panned signal components, while the difference of the input channels, LR, will contain little or no dialogue. Figure 19 As shown in FIG, most of the dialogue can be extracted from the sum component L+R, and this sum component receives dialogue separation and dialogue separation unit 1. A stereo non-dialogue signal is then synthesized using the original difference signal LR and the non-dialogue components of the original sum signal L+R. The recreated stereo non-dialogue signal components and the mono sum signal are then analyzed by loudness processing unit 2 and remixed into an enhanced stereo signal in audio mixer 3. As in the other examples, the resulting output signal is directed to output limiter 7.

[0106] exist Figure 20 In the embodiment of FIG. 4 , the input audio signal A is a multi-channel signal (5.1 in the depicted embodiment). It is assumed that dialogue already dominates the center channel of the multi-channel audio stream. Conversely, it is assumed that all other channels can be non-dialogue. Therefore, the center channel is simply redirected to a mono (1.0) dialogue separation processor 1, and a loudness processing unit 2 considers the loudness of the separated dialogue relative to the residual non-dialogue channels (including the residual from the center channel). The loudness processing sets appropriate gains and delays to each signal component and recombines each signal to match the input multi-channel format (5.1 in this embodiment). Finally, a multi-channel limiter 7 is applied to the resulting 5.1 channel output.

[0107] exist Figure 21 In an embodiment of the invention, the dialogue separation model has been trained to separate the dialogue and dialogue components from the stereo signal, such as Figure 15 . In this case, it can be assumed that most of the dialogue will be present in the front (L, C, R) channel mix. After the channel splitting in unit 85, these channels are downmixed to a stereo signal in downmixer 87 and directed to the stereo dialogue separation unit 1. In this case, the other channels (2.1: LS, RS, LFE) are assumed to contain no narrative dialogue. The loudness of the separated stereo dialogue is compared with the loudness of the residual stereo non-dialogue channels and the (LS, RS, LFE) channels in loudness processing unit 2. Appropriate gains are calculated and applied to all channel signal components. The separated dialogue and non-dialogue outputs (originally obtained from (L, C, R)) are mixed together in audio mixer 3 and further upmixed to their original 3-channel layout using 2-3 channel upmixer 88, and the resulting signal is again combined with the original surround and LFE channel components in channel combiner 86 and reconstructed to match the original input format. As before, a multi-channel limiter can finally be applied to the resulting 5.1 channel output.

[0108] exist Figure 22 In the embodiment of FIG, a dialogue separation unit 1 is used which has been trained on a three-channel input signal. Therefore, there is no need to down-mix or up-mix the (L, C, R) channels.

[0109] The above-described embodiments and implementation topologies are susceptible to various adaptations.

[0110] In some embodiments, the individualized processed audio output, or portions thereof, is directed to specific individuals using a beamforming speaker array.

[0111] In some embodiments, the wireless receiver may be an assistive listening device (hearing aid). In this case, care must be taken to ensure that the selection of the dialogue handling preference is consistent with the built-in assistive listening technology.

[0112] In some embodiments, the common audio output can also be broadcast to multiple wireless receivers without the need for speaker outputs. This will minimize acoustic crosstalk for all listeners.

[0113] In some embodiments, only the dialogue channel is used for individual processing and wireless transmission. This may be the case if the listener only requires enhanced dialogue signals. This can be achieved using bone conduction headphones, near-field speakers, or open-back headphones.

[0114] In some embodiments, the system may use an image sensor that can identify the presence and location of listener(s). This can influence the algorithm parameters used. For example, a preference may be used only when a specific person is in the room. Alternatively, a weighted average of the parameters detected for each person in the room may be used. Listener location can be useful when accounting for ambient noise or when beamforming a conversation to a specific individual.

[0115] In some embodiments, different user preferences may be applied to different types of content. For example, a user may prefer different sets of loudness processing parameters for drama and news. The content type may be obtained from content metadata or may be determined using algorithmic classification.

[0116] In some embodiments, when the described process is applied to a self-contained wearable device (e.g., a hearable, an augmented reality headset), it can be used in environments outside the home (e.g., a movie theater or theater). By default, the loudness adaptation algorithm is based on digital loudness levels (relative to digital full scale). When only the acoustic signal captured by the microphone is available, some amount of SPL to digital level calibration must be performed to ensure a certain degree of equivalence.

[0117] In some embodiments, automatic closed captions may be displayed on an augmented reality display or glasses.

[0118] Alternative Embodiments and Exemplary Operating Environments

[0119] As can be seen herein, there are many other variations besides those described herein. For example, depending on the embodiment, certain actions, events, or functions of any method and algorithm described herein may be performed in a different order, may be added, merged, or omitted entirely, such that not all described actions or events are necessary for the practice of the method and algorithm. Moreover, in some embodiments, actions or events may be performed concurrently, such as by multithreading, interrupt handling, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. Furthermore, different tasks or processes may be performed by different machines and computing systems that may operate together.

[0120] The various illustrative logical blocks, modules, methods, and algorithmic processes and sequences described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and process actions have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. For each specific application, the described functionality may be implemented in varying ways, but such implementation decisions should not be interpreted as causing a departure from the scope of this document.

[0121] The various illustrative logical blocks and modules described in conjunction with the embodiments disclosed herein may be implemented or executed by a machine, such as a general-purpose processor, a processing device, a computing device having one or more processing devices, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices designed to perform the functions described herein, discrete gate or transistor logic, discrete hardware components, or any combination thereof. General-purpose processors and processing devices may be microprocessors, but in alternative embodiments, the processor may be a controller, a microcontroller, or a state machine, a combination thereof, or the like. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration.

[0122] Embodiments of the systems and methods described herein can operate in a variety of general-purpose or special-purpose computing system environments or configurations. In general, a computing environment can include any type of computer system, including but not limited to computer systems based on one or more microprocessors, mainframe computers, digital signal processors, portable computing devices, personal organizers, device controllers, computing engines in devices, mobile phones, desktop computers, mobile computers, tablet computers, smartphones, and devices with embedded computers, among others.

[0123] Such computing devices can typically be found in devices with at least some minimal computing power, including but not limited to personal computers, server computers, handheld computing devices, laptop or mobile computers, communication devices such as cellular phones and PDAs, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, audio or video media players, etc. In some embodiments, the computing device will include one or more processors. Each processor can be a special-purpose microprocessor, such as a digital signal processor (DSP), a very long instruction word (VLIW), or other microcontroller, or can be a conventional central processing unit (CPU) having one or more processing cores, including a core based on a special-purpose graphics processing unit (GPU) in a multi-core CPU.

[0124] The process actions or operations of the methods, processes, or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in any combination of the two. The software module may be contained in a computer-readable medium that can be accessed by a computing device. Computer-readable media include volatile and non-volatile media, which may be removable, non-removable, or some combination thereof. Computer-readable media are used to store information, such as computer-readable or computer-executable instructions, data structures, program modules, or other data. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0125] Computer storage media include, but are not limited to, computer or machine readable media or storage devices, such as Blu-ray Discs (BD), Digital Versatile Discs (DVD), Compact Discs (CD), floppy disks, tape drives, hard drives, optical drives, solid-state memory devices, RAM memory, ROM memory, EPROM memory, EEPROM memory, flash memory or other memory technology, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other device that can be used to store the desired information and can be accessed by one or more computing devices.

[0126] The software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of non-transitory computer readable storage medium, medium, or physical computer storage device known in the art. An exemplary storage medium may be coupled to a processor so that the processor can read information from the storage medium and write information to it. In an alternative embodiment, the storage medium may be an integral part of the processor. The processor and storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a user terminal. Alternatively, the processor and storage medium may reside in a user terminal as discrete components.

[0127] As used in this document, the phrase "non-transitory" means "persistent or long-lived." The phrase "non-transitory computer-readable medium" includes any and all computer-readable media, with the sole exception of transitory propagating signals. By way of example and not limitation, this includes non-transitory computer-readable media such as register memory, processor cache, and random access memory (RAM).

[0128] The phrase "audio signal" is a signal representing physical sound.

[0129] The retention of information such as computer-readable or computer-executable instructions, data structures, program modules, etc. can also be achieved by using a variety of communication media to encode one or more modulated data signals, electromagnetic waves (such as carrier waves) or other transmission mechanisms or communication protocols, and includes any wired or wireless information transmission mechanism. Generally speaking, these communication media refer to signals whose one or more characteristics are set or changed in such a way that information or instructions are encoded in the signal. For example, communication media include wired media (such as a wired network or a direct line connection carrying one or more modulated data signals), and wireless media (such as acoustic, radio frequency RF, infrared, laser, and other wireless media for transmitting, receiving, or both one or more modulated data signals or electromagnetic waves). Any combination of the above should also be included within the scope of communication media.

[0130] Additionally, one or any combination of software, programs, computer program products, or portions thereof that implement some or all of the various embodiments of the systems and methods described herein may be stored, received, transmitted, or read from any desired combination of computer- or machine-readable media or storage devices and communication media in the form of computer-executable instructions or other data structures.

[0131] Embodiments of the systems and methods described herein can be further described in the general context of computer-executable instructions (such as program modules) executed by a computing device. Generally speaking, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments described herein can also be practiced in a distributed computing environment in which tasks are performed by one or more remote processing devices, or in a cloud of one or more devices linked by one or more communication networks. In a distributed computing environment, program modules can be located in local and remote computer storage media including media storage devices. Further, the above instructions can be implemented in part or in whole as hardware logic circuits, which may or may not include a processor.

[0132] Unless otherwise specifically stated or otherwise understood in context as used, conditional language used herein (such as, among others, "can," "might," "could," "for example," etc.) is generally intended to convey that some embodiments include, while other embodiments do not, certain features, elements, and / or states. Thus, such conditional language is generally not intended to imply that features, elements, and / or states are in any way required for one or more embodiments or that one or more embodiments must include logic for determining, with or without author input or prompting, that such features, elements, and / or states are included in or to be performed in any particular embodiment. The terms "comprise," "include," "have," and the like are synonymous and are used inclusively in an open manner and do not exclude additional elements, features, actions, operations, and the like. Moreover, the term "or" is used in its inclusive sense rather than its exclusive sense, such that when used, for example, to connect a list of elements, the term "or" refers to one, some, or all of the elements in the list.

[0133] While the foregoing detailed description has shown, described, and pointed out novel features as applied to various embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms shown may be made without departing from the scope of the present disclosure. As will be appreciated, certain embodiments of the invention described herein may be embodied in a form that does not provide the features and advantages set forth herein, as some features may be used or practiced separately from other features.

Claims

1. A method for enhancing the intelligibility of dialogue in an original audio signal, the original audio signal comprising a dialogue component and a non-dialogue component, the method comprising: providing a dialog component of the original audio signal in a first separate audio signal; providing a non-dialogue component of the original audio signal in a second separate audio signal; separately processing the first individual audio signal and the second individual audio signal, wherein processing the first and second individual audio signals comprises processing the loudness of the first individual audio signal and / or the second individual audio signal, and The processed first and second separate audio signals are combined to provide a processed audio signal.

2. The method of claim 1 , wherein providing a dialogue component in a first individual audio signal and providing a non-dialogue component in a second individual audio signal comprises receiving the first and second individual audio signals from a source from which the first and second individual audio signals are separately obtainable.

3. The method of claim 1 , wherein providing the dialogue component in the first separate audio signal and providing the non-dialogue component in the second separate audio signal comprises separating the dialogue component from the non-dialogue component in the original audio signal.

4. The method of any one of the preceding claims, wherein processing the first individual audio signal comprises: determining a short-term loudness level of the first individual audio signal; Determine whether the determined short-term loudness level is less than a predefined minimum dialogue loudness level DLL MIN , If the determined short-term loudness level is less than the predefined minimum dialogue loudness level DLL MIN , then the first separate audio signal is directed toward the minimum dialogue loudness level DLL MIN enlarge, If the determined short-term loudness level is not less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is not modified.

5. The method of claim 4, wherein the processed first individual audio signal is spectrally enhanced before being combined with the processed second individual audio signal.

6. The method according to claim 4 or 5, wherein the method further comprises: determining speech activity in the first individual audio signal; The first separate audio signal is directed towards the minimum dialogue loudness level DLL only if voice activity has been determined. MIN enlarge.

7. The method of claim 6 , wherein determining the voice activity comprises determining whether the short-term loudness level of the first individual audio signal is above a threshold dialogue loudness level (DLL). THRESH , wherein the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL only if the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL THRESH DLL towards the minimum dialogue loudness level MIN The first individual audio signal is amplified.

8. The method of any one of claims 4 to 7, wherein amplifying the first individual audio signal comprises using a dynamic range processor that applies a gain using a modifiable curve determined by a plurality of control points.

9. The method of any one of the preceding claims, wherein processing the second separate audio signal comprises: Determine the short-term loudness level of the first individual audio signal or obtain a predefined minimum dialogue loudness level DLL of the first individual audio signal MIN ; determining a short-term loudness level of the second separate audio signal; Determining a difference or minimum dialogue loudness level DLL between the short-term loudness level of the first individual audio signal and the short-term loudness level of the second individual audio signal MIN Is the difference between the short-term loudness level of the first and second separate audio signals less than a predefined minimum dialogue-to-non-dialogue ratio D2ND? MIN ; If it is less than, then the loudness level of the second separate audio signal is reduced so that the difference approaches the minimum dialogue to non-dialogue ratio D2ND MIN , If not, then the second separate audio signal is not modified.

10. The method of claim 9, wherein reducing the loudness level of the second individual audio signal comprises compressing the dynamic range of the second individual audio signal.

11. The method of claim 9 or 10, wherein reducing the loudness level of the second individual audio signal comprises using a dynamic range processor that applies a gain using a modifiable curve determined by a plurality of control points.

12. A method as claimed in any one of claims 4 to 11, wherein the short term loudness level is determined for consecutive windows of a predefined length, wherein the loudness level is determined according to an industry standard.

13. The method of any one of the preceding claims, wherein the first individual audio signal and the second individual audio signal are processed in a plurality of processing paths, the processing paths comprising: a universal processing path, in which the processed audio signal is provided to any number of listeners; At least one individualized processing path, wherein a processed audio signal is provided to an individual listener, wherein processing the first individual audio signal and / or processing the second individual audio signal comprises using parameters personalized for the individual listener during processing.

14. The method of claim 13, wherein the parameters personalized for the individual listener include at least one of a personal hearing profile and subjective listening preferences specific to the listener.

15. A method as claimed in any preceding claim, wherein the original audio signal is an audio track.

16. The method of any of the preceding claims, wherein the original audio signal is a stereo signal or a multi-channel signal, wherein for each channel of the stereo signal or multi-channel signal a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal.

17. The method of claim 1 , wherein the original audio signal is a stereo signal, wherein the stereo signal is upmixed to a three-channel signal comprising a center channel, a left channel, and a right channel, wherein a signal component of the stereo signal initially panned to the center is extracted to the center channel, and wherein for the center channel only a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal, and wherein the second separate audio channel is combined with the left and right channels for loudness processing.

18. The method of claim 1 , wherein the original audio signal is a multi-channel signal comprising a center channel and a plurality of further channels, wherein for the center channel only, a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal, and wherein the second separate audio signal is combined with the further channels for loudness processing.

19. The method of claim 1 , wherein the original audio signal is a multi-channel signal comprising a center channel, a left channel, a right channel, and a further channel, wherein the center channel, the left channel, and the right channel are downmixed to two channels, wherein for each of the two downmixed channels, a dialogue component is provided in a first separate audio signal and a non-dialogue component is provided in a second separate audio signal, and wherein the second separate audio signal is combined with the further channel for loudness processing.

20. The method according to any of the preceding claims, wherein the processed first individual audio signal and the processed second individual audio signal are further processed by applying spatial audio processing and / or specific algorithms before combining the processed first and second individual audio signals.

21. A method for enhancing the intelligibility of dialogue in an original audio signal, the original audio signal comprising a dialogue component and a non-dialogue component, the method comprising: receiving a dialog component of an original audio signal in a first separate audio signal; receiving a non-dialogue component of the original audio signal in a second separate audio signal; as well as The first individual audio signal and the second individual audio signal are processed separately, wherein processing the first and second individual audio signals includes processing loudness of the first individual audio signal and / or the second individual audio signal.

22. A system for enhancing the intelligibility of dialogue in an original audio signal, the original audio signal comprising a dialogue component and a non-dialogue component, the system comprising: a dialogue separation unit configured to provide a dialogue component of the original audio signal in a first separate audio signal and a non-dialogue component of the original audio signal in a second separate audio signal; a loudness processing unit configured to separately process the first individual audio signal and the second individual audio signal, wherein processing the first and second individual audio signals comprises processing the loudness of the first individual audio signal and / or the second individual audio signal, and An audio mixer is configured to combine the processed first and second separate audio signals to provide a processed audio signal.

23. The system of claim 22, wherein the dialogue separation unit is configured to provide a dialogue component in a first separate audio signal and a non-dialogue component in a second separate audio signal, including receiving the first and second separate audio signals from a source from which the first and second separate audio signals are separately obtainable.

24. The system of claim 22, wherein the dialogue separation unit is configured to provide the dialogue component in the first separate audio signal and the non-dialogue component in the second separate audio signal by separating the dialogue component from the non-dialogue component in the original audio signal.

25. The system of any one of claims 22 to 24, wherein the loudness processing unit is configured to process the first individual audio signal by: determining a short-term loudness level of the first individual audio signal; Determine whether the determined short-term loudness level is less than a predefined minimum dialogue loudness level DLL MIN , If the determined short-term loudness level is less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is directed towards the predefined minimum dialogue loudness level DLL MIN enlarge, If the determined short-term loudness level is not less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is not modified.

26. The system of claim 25, wherein the loudness processing unit is further configured to perform spectral enhancement on the processed first individual audio signal before combining the processed first individual audio signal with the processed second individual audio signal.

27. The system of claim 25 or 26, wherein the loudness processing unit is further configured to: determining speech activity in the first individual audio signal; The first separate audio signal is directed towards the minimum dialogue loudness level DLL only if voice activity has been determined. MIN enlarge.

28. The system of claim 27, wherein the loudness processing unit is configured to process the first individual audio signal by determining whether the short-term loudness level of the first individual audio signal is higher than a threshold dialogue loudness level DLL. THRESH To determine speech activity, wherein the short-term loudness level is higher than the threshold dialogue loudness level DLL only if the determined short-term loudness level is higher than the threshold dialogue loudness level DLL THRESH DLL towards the minimum dialogue loudness level MIN The first individual audio signal is amplified.

29. The system of any one of claims 25 to 28, wherein the loudness processing unit comprises a dynamic range processor configured to amplify the first individual audio signal by applying a gain, wherein for applying the gain the dynamic range processor is configured to use a modifiable curve determined by a plurality of control points.

30. The system of any one of claims 22 to 29, wherein the loudness processing unit is configured to process the second separate audio signal by: Determine the short-term loudness level of the first individual audio signal or obtain a predefined minimum dialogue loudness level DLL of the first individual audio signal MIN ; determining a short-term loudness level of the second separate audio signal; Determining a difference or minimum dialogue loudness level DLL between the short-term loudness level of the first individual audio signal and the short-term loudness level of the second individual audio signal MIN Is the difference between the short-term loudness level of the first and second separate audio signals less than a predefined minimum dialogue-to-non-dialogue ratio D2ND? MIN ; If it is less than, then the loudness level of the second separate audio signal is reduced so that the difference approaches the minimum dialogue to non-dialogue ratio D2ND MIN , If not, then the second separate audio signal is not modified.

31. The system of claim 30, wherein the loudness processing unit is configured to reduce the loudness level of the second individual audio signal by compressing the dynamic range of the second individual audio signal.

32. The system of claim 30 or 31 , the loudness processing unit comprising a dynamic range processor configured to reduce the loudness level of the second individual audio signal by applying a gain, wherein for applying the gain the dynamic range processor is configured to use a modifiable curve determined by a plurality of control points.

33. The system of any one of claims 22 to 32, wherein the loudness processing unit is configured to determine a short-term loudness level for consecutive windows of a predetermined length, wherein the loudness level is determined according to an industry standard.

34. The system of any one of claims 22 to 33, wherein the first individual audio signal and the second individual audio signal are processed in a plurality of processing paths, the processing paths comprising: a universal processing path including a universal loudness processing unit, wherein the processed audio signal is provided to any number of listeners; At least one individualized processing path, the at least one individualized processing path comprising an individualized loudness processing unit, wherein the processed audio signal is provided to an individual listener, wherein processing the first individual audio signal and / or processing the second individual audio signal comprises using parameters personalized for the individual listener during processing.

35. The system of claim 34, wherein the personalized loudness processing unit is configured to process the first and / or second individual audio signals using personalized parameters comprising at least one of a personal hearing profile and subjective hearing preferences specific to a listener.

36. A system as claimed in any one of claims 22 to 35, wherein the original audio signal is an audio soundtrack.

37. The system of any one of claims 22 to 36, wherein the original audio signal is a stereo signal or a multi-channel signal, and wherein the dialogue separation unit is configured to provide, for each channel of the stereo signal or the multi-channel signal, a dialogue component in a first separate audio signal and a non-dialogue component in a second separate audio signal.

38. The system of any one of claims 22 to 36, wherein the original audio signal is a stereo signal, wherein the system further comprises an upmixer configured to upmix the stereo signal into a three-channel signal comprising a center channel, a left channel, and a right channel, wherein a signal component of the stereo signal initially panned to the center is extracted to the center channel, and wherein the dialog separation unit is configured to provide, for only the center channel, a dialog component in the first separate audio signal and a non-dialog component in the second separate audio signal, and wherein the loudness processing unit is configured to combine the second separate audio channel with the left and right channels for loudness processing.

39. The system of any one of claims 22 to 36, wherein the original audio signal is a multi-channel signal comprising a center channel and a plurality of further channels, wherein the dialog separation unit is configured to provide, for the center channel only, a dialog component in the first separate audio signal and a non-dialog component in the second separate audio signal, and wherein the loudness processing unit is configured to combine the second separate audio signal with the further channels for loudness processing.

40. The system of any one of claims 22 to 36, wherein the original audio signal is a multi-channel signal comprising a center channel, a left channel, a right channel, and an additional channel, wherein the system further comprises a downmixer configured to downmix the center channel, the left channel, and the right channel into two channels, wherein the dialogue separation unit is configured to provide, for each of the two downmixed channels, a dialogue component in a first separate audio signal and a non-dialogue component in a second separate audio signal, and wherein the loudness processing unit is configured to combine the second separate audio signal with the additional channel for loudness processing.

41. The system of any one of claims 22 to 40, further comprising a post-processing unit configured to further process the processed first separate audio signal and the processed second separate audio signal by applying spatial audio processing and / or a specific algorithm before combining the processed first and second separate audio signals.

42. A non-transitory computer-readable medium having stored thereon executable instructions that, when executed by a processor, perform the following operations: providing a dialog component of the original audio signal in a first separate audio signal; providing a non-dialogue component of the original audio signal in a second separate audio signal; separately processing the first individual audio signal and the second individual audio signal, wherein processing the first and second individual audio signals includes processing loudness of the first individual audio signal and / or the second individual audio signal; and The processed first and second separate audio signals are combined to provide a processed audio signal.

43. The non-transitory computer-readable medium of claim 42, wherein processing the first individual audio signal comprises: determining a short-term loudness level of the first individual audio signal; Determine whether the determined short-term loudness level is less than a predefined minimum dialogue loudness level DLL MIN , If the determined short-term loudness level is less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is directed towards the predefined minimum dialogue loudness level DLL MIN enlarge, If the determined short-term loudness level is not less than the minimum dialogue loudness level DLL MIN , then the first separate audio signal is not modified.

44. The non-transitory computer readable medium of claim 43, further performing the following operations: determining speech activity in the first individual audio signal; The first separate audio signal is directed towards the minimum dialogue loudness level DLL only if voice activity has been determined. MIN enlarge.

45. The non-transitory computer-readable medium of claim 44, wherein determining voice activity comprises determining whether a short-term loudness level of the first individual audio signal is above a threshold dialogue loudness level (DLL). THRESH , wherein the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL only if the short-term loudness level is determined to be higher than the threshold dialog loudness level DLL THRESH DLL towards the minimum dialogue loudness level MIN The first individual audio signal is amplified.

46. ​​The non-transitory computer-readable medium of any one of claims 42 to 45, wherein processing the second separate audio signal comprises: Determine the short-term loudness level of the first individual audio signal or obtain a predefined minimum dialogue loudness level DLL of the first individual audio signal MIN ; determining a short-term loudness level of the second separate audio signal; Determining a difference or minimum dialogue loudness level DLL between the short-term loudness level of the first individual audio signal and the short-term loudness level of the second individual audio signal MIN Is the difference between the short-term loudness level of the first and second separate audio signals less than a predefined minimum dialogue-to-non-dialogue ratio D2ND? MIN ; If it is less than, then the loudness level of the second separate audio signal is reduced so that the difference approaches the minimum dialogue to non-dialogue ratio D2ND MIN , If not, then the second separate audio signal is not modified.