Methods and systems for improving dialogue intelligibility
Separating and processing dialogue and non-dialogue components in audio signals addresses the issue of reduced intelligibility by maintaining dialogue clarity and reducing non-dialogue interference, improving the listening experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2026-03-17
AI Technical Summary
Consumers struggle to understand dialogue in audio signals due to reduced intelligibility caused by non-dialogue components, leading to frustration and the need to constantly adjust volume levels.
Separate dialogue and non-dialogue components in audio signals, process them independently to adjust loudness levels, and combine them to enhance dialogue intelligibility.
Improves dialogue clarity by ensuring dialogue components maintain desired loudness levels while reducing non-dialogue components, enhancing overall listening experience without affecting the soundtrack's spectral balance.
Smart Images

Figure 2026509119000001_ABST
Abstract
Description
[Technical Field]
[0001] (Related applications and claims for priority) This application relates to and claims priority to U.S. Patent Provisional Application No. 63 / 483,737, filed on 7 February 2023, for the invention titled "DIALOG ENHANCEMENT ECOSYSTEM FOR STREAMING AND BROADCASTING MEDIA," and further relates to and claims priority to U.S. Patent Provisional Application No. 63 / 508,811, filed on 16 June 2023, both of which are incorporated herein by reference in their entirety.
[0002] This disclosure relates to improving dialogue intelligibility in audio signals that include dialogue and non-dialogue components. For example, an audio soundtrack for video content that can be played on media devices such as set-top boxes, televisions, and laptops. A mixed soundtrack may consist of narrative dialogue and non-dialogue audio components. Non-dialogue components may include, for example, ambient or environmental sounds, music, and sound effects. [Background technology]
[0003] In many cases, consumers are unable to understand dialogue from mixed soundtracks played through their sound playback system in their listening environment. Consumers may be unable to understand dialogue due to numerous factors that can reduce the intelligibility of spoken language. This often forces consumers to constantly adjust the volume levels of their content, lowering the volume if the music and effects are too loud and raising the volume if the dialogue is too quiet. This can detach consumers from their content viewing experience and cause frustration. Furthermore, simply increasing the device's master volume level does not solve the intelligibility problem, as it increases the volume of both the dialogue and any interfering non-dialogue soundtracks.
[0004] Therefore, there is a need to improve the intelligibility of dialogue components in audio signals. [Overview of the project]
[0005] This summary of the invention is provided in a simplified form to introduce selected concepts from those further described below in the modes for carrying out the invention. This summary is not intended to confirm the main or basic features of the subject matter described in the claims, nor is it intended to limit the scope of the subject matter described in the claims.
[0006] One aspect of the present invention provides a method for improving dialogue intelligibility in an original audio signal that includes dialogue and non-dialogue components. The method involves providing the dialogue components of the original audio signal in a first separated audio signal, providing the non-dialogue components of the original audio signal in a second separated audio signal, and processing the first separated audio signal and the second separated audio signal separately, wherein processing the first and second separated audio signals includes processing the loudness of the first separated audio signal and / or the second separated audio signal, and combining the processed first and second separated audio signals to provide a processed audio signal.
[0007] Aspects of the present invention are based on the idea of independently analyzing and processing the dialogue and non-dialogue components of an audio signal. This makes it possible to process the dialogue and non-dialogue components separately, improving the intelligibility of the dialogue components. In particular, the dialogue components can be adjusted or equalized differently from the non-dialogue components. For example, the dialogue components can receive loudness normalization and optionally spectral enhancement, while the non-dialogue components can receive dynamic range compression, as described later.
[0008] Within the scope of the intent of this invention, a dialogue component is a component that considers spoken language (including the silences between spoken words), and a non-dialogue component is a component that considers other components of the audio signal, such as music and sound effects. Furthermore, a dialogue component may also be called foreground speech, and a non-dialogue component may also be called background sound.
[0009] It is pointed out that the method steps do not necessarily have to be executed by the same entity. For example, the step of providing the dialog component of the original audio signal with the first separated audio signal and the step of providing the non-dialog component of the original audio signal with the second separated audio signal may be implemented in the head-end system or the cloud. The step of separately processing the first separated audio signal and the second separated audio signal and the step of combining the processed signals may be implemented on the consumer device. In another embodiment, the dialog separation is implemented in a high-power device at the customer site such as a set-top box or a TV, while the processing of the first separated audio signal and the second separated audio signal is provided by another customer device such as a consumer device. However, in other embodiments, all steps are implemented on the same device such as a consumer-oriented device.
[0010] In one embodiment, providing the dialog component with the first separated audio signal and the non-dialog component with the second separated audio signal includes receiving the first and second separated audio signals from a source where the first and second separated audio signals are separately available. Thus, if individual dialog-only signals and non-dialog signals are already available from the production stage, they can be used directly. For example, an individual dialog stream can be utilized using object-based audio such as DTS:X (registered trademark), Dolby Atmos (registered trademark) or MPEG-H (registered trademark).
[0011] In another embodiment, providing the dialog component with a first separated audio signal and the non-dialog component with a second separated audio signal includes separating the dialog component from the non-dialog component in the original audio signal. Separating the dialog component from the non-dialog component can be implemented by a plurality of methods. For example, dialog separation can be implemented by a deep learning model such as a convolutional neural network and a recurrent neural network that enable the ability to separate different sources containing dialog. There are commercial products for dialog separation based on neural networks such as RX Dialogue Isolate by iZotope. Another method relies on object-based audio analysis as considered in J. Paulus et al.: “Source Separation for Enabling Dialogue Enhancement in Object-Based Broadcast with MPEG-H”, J. Audio Eng. Soc., Vol. 67, No. 7 / 8, 2019 July / August.
[0012] In one embodiment, processing the first separated audio signal includes determining the short-term loudness level of the first separated audio signal and determining whether the determined short-term loudness level is less than a predefined minimum dialog loudness level DLL
[0013] If the determined short-term loudness level is less than the predefined minimum dialog loudness level DLL MIN the first separated audio signal is amplified towards the predefined minimum dialog loudness level DLL MIN If the determined short-term loudness level is not less than the minimum dialog loudness level DLL MIN the first separated audio signal is not changed.
[0013] In this aspect of the invention, the parameter "minimum dialog loudness level" (DLL MIN) determines the target short-term average loudness level of the dialog component. If the measured dialog level is less than the target DLL MIN the first separated audio signal (dialog signal) is amplified towards the target minimum level. If the dialog loudness is already above the DLL MIN no signal correction is applied. The typical default value of the DLL MIN matches the industry recommended value for digital dialog loudness levels. This is generally in the range of -22 LUFS to -27 LUFS. Since most program content follows this recommendation, it may not be necessary to significantly change the dialog loudness to achieve this target.
[0014] In a further embodiment, the processed first separated audio signal is spectrally enhanced before being combined with the processed second separated audio signal. Such spectral enhancement is optionally performed and can include the application of specific filters to the dialog component.
[0015] In a further embodiment, the method further includes determining the voice activity in the first separated audio signal and amplifying the first separated audio signal towards the minimum dialog loudness level DLL MIN only if the voice activity is determined. In this embodiment, it is based on the idea that dialog loudness should only be boosted when there is voice activity. Otherwise, extremely low-level dialog components such as background dialog or noise / artifacts in the dialog signal will receive an undesirably high gain to match the DLL MIN level. This can cause an undesired loudness spike when transitioning from a quiet segment to a segment with story dialog because the normalization ballistic requires adjustment time for rapid loudness changes.
[0016] One example of determining speech activity is when the short-term loudness level of the first isolated audio signal is the threshold dialogue loudness level DLL. THRESH The first separated audio signal includes determining whether the determined short-term loudness level is higher than the threshold dialog loudness level DLL. THRESH Minimum dialog loudness level DLL only if it is higher than [value]. MIN It is amplified toward. In this embodiment, the parameter "Threshold Dialog Loudness Level (DLL)" THRESH ) acts as a Voice Activity Detector (VAD), and below this level, dialogue loudness is not boosted. Furthermore, DLL THRESH This helps to avoid amplification of low-level processing artifacts from the preceding dialogue isolation process.
[0017] It should be noted that this aspect of the present invention is not limited to a specific implementation of VAD as a threshold parameter. It may also include other implementations of speech activity detection, such as those using an output mask from a dialogue separation process or using a machine learning algorithm designed for speech activity detection.
[0018] In one embodiment, amplifying a first isolated audio signal involves using a dynamic range processor that applies gain using a variable curve determined by a plurality of control points. Such a variable curve can be determined by five control points (x / y coordinates), allowing the processor to function as a compressor, expander, loudness leveler, or a hybrid of these modes. Smoothing parameters may also be incorporated to ensure seamless transitions between operating zones.
[0019] In one embodiment, processing a second isolated audio signal involves determining the short-term loudness level of the first isolated audio signal or a predefined minimum dialog loudness level DLL of the first isolated audio signal. MINThe system includes obtaining the difference between the short-term loudness level of the first isolated audio signal and the short-term loudness level of the second isolated audio signal, or the minimum dialog loudness level DLL. MIN The difference between this and the short-term loudness level of the second separated audio signal is the predefined minimum dialogue / non-dialogue ratio D2ND. MIN Determine whether it is less than the minimum dialog-to-non-dialog ratio (D2ND). If it is less than the minimum, the difference is the minimum dialog-to-non-dialog ratio (D2ND). MIN The loudness level of the second isolated audio signal is reduced to approach the target value. If it is not less than the target value, the second isolated audio signal remains unchanged.
[0020] In this embodiment, the parameter "Minimum Dialogue to Non-Dialogue Ratio (D2ND)" is used. MIN ) represents the minimum difference between short-term dialog and short-term non-dialog loudness levels. If the measured level has a loudness difference less than this value, the non-dialog signal is considered to have a mean difference between the dialog loudness level and the non-dialog loudness level. MIN It is compressed until it approaches [a certain value]. It is noted that the non-dialogue level is reduced only when necessary.
[0021] In one embodiment, reducing the loudness level of a second isolated audio signal involves compressing the dynamic range of the second isolated audio signal. This can be accomplished by using a dynamic range processor that applies gain using a variable curve determined by a plurality of control points, where the control points cause the processor to function as a compressor and / or loudness leveler. For example, the difference mentioned is D2ND MIN If the value is less than or equal to the value, the difference is D2ND MIN To get closer to the desired value, a specific compression ratio, such as 2:1, can be implemented.
[0022] In a further embodiment, the short-term loudness level (of a first isolated audio signal including dialogue components, or a second isolated audio signal including non-dialogue components) is determined over a continuous window of a predefined length, where the loudness level is determined according to an industry standard. The window is within the range of 10 ms and 100 ms. For example, the window has a length of 20 ms. The industry standard on which the loudness level is determined may be the ITU-R BS.1770 standard, where loudness is expressed in LKFS (Loudness, K-weighted, relative to Full Scale) or the synonymous term LUFS (Loudness units relative to full scale) introduced in EBU R128, and is the standard loudness measurement unit used for audio normalization in broadcast television systems and other video and music streaming services. In particular, ITU-R BS.1770-1, the first iteration of this standard, can be used to determine loudness because the standard is particularly well-suited to handling immediate loudness fluctuations through continuous short-term measurements.
[0023] In a further embodiment, the first and second isolated audio signals are processed in a plurality of processing paths, the processing paths including a general processing path that provides the processed audio signals to any number of listeners, and at least one personalized processing path that provides the processed audio signals to individual listeners, wherein the processing of the first isolated audio signal and / or the processing of the second isolated audio signal includes using personalized parameters for individual listeners during processing. Personalized parameters may include listener-specific personal auditory profiles and subjective listening preferences. This embodiment addresses the situation where not everyone in the listening space wants to hear a common audio output from the dialogue enhancement system. Thus, alternative degrees of dialogue processing can be implemented according to the needs of one or more people.
[0024] In one embodiment, the original audio signal is an audio soundtrack, i.e., sound accompanying and synchronized with images from a movie, television program, video game, radio program, etc. The original soundtrack may be in the form of a digital audio file. However, the present invention is not limited to such embodiments. For example, the original audio signal may be a live audio signal.
[0025] In a further embodiment, the original audio signal is a stereo signal or a multi-channel signal. For each channel of the stereo signal or multi-channel signal, the dialogue component is provided by a first separated audio signal, the non-dialogue component is provided by a second separated audio signal, and the first and second separated audio signals are processed separately and then combined. Thus, in this embodiment, the number of channels in the input is maintained in the output.
[0026] In a further embodiment, the original audio signal is a stereo signal, which is upmixed into a three-channel signal including a center channel, a left channel, and a right channel, with the signal component of the stereo signal, which is initially panned to the center, being extracted to the center channel. Furthermore, for the center channel only, the dialogue component is provided by a first separated audio signal, and the non-dialogue component is provided by a second separated audio signal. The first and second separated audio signals are processed according to the present invention, and the second separate audio channel is coupled with the left and right channels for loudness processing. In this embodiment, channel dialogue separation of only a single channel is required, thereby reducing complexity.
[0027] In a further embodiment, the original audio signal is a multi-channel signal including a center channel and several further channels, and since it is assumed that the most dialogue is present in the center channel, the dialogue component is provided to the center channel only in the first isolated audio signal, and the non-dialogue component is provided in the second isolated audio signal. The first and second isolated audio signals are processed according to the present invention, and the second isolated audio channel is coupled with further channels for loudness processing. In this embodiment, channel dialogue separation of only one channel is required in the multi-channel signal, thereby reducing complexity.
[0028] In a further embodiment, the original audio signal is a multi-channel signal including a center channel, left channel, right channel, and further channels, where the center channel, left channel, and right channel are downmixed to two channels, and for each of the two downmixed channels, the dialogue component is provided in a first isolated audio signal and the non-dialogue component is provided in a second isolated audio signal. In this embodiment, it is assumed that the majority of the dialogue resides in the front (center, left, right) channels. The first isolated audio signal and the second isolated audio signal are processed according to the present invention, and the second isolated audio signal is coupled with further channels for loudness processing. In this embodiment, channel dialogue separation of only one channel is required in the multi-channel signal, thereby reducing complexity.
[0029] In a further embodiment, the processed first isolated audio signal and the processed second isolated audio signal are further processed by applying spatial audio processing and / or specific algorithms before the processed first isolated audio signal and the second isolated audio signal are combined. In this embodiment, it is based on the understanding that the first isolated audio signal having a dialogue component and the second isolated audio signal having a non-dialogue component may remain separated for further downstream processing before being combined. For example, the further processing may include the application of algorithms that include spatial audio processing for headphones and speakers. Other examples of downstream processing include algorithms that are well applied only to the non-dialogue component of the input signal, such as bass boosting.
[0030] A further aspect of the present invention provides a method for emphasizing dialogue in an original audio signal that includes dialogue and non-dialogue components. This method includes receiving the dialogue components of the original audio signal with a first isolated audio signal, receiving the non-dialogue components of the original audio signal with a second isolated audio signal, and processing the first isolated audio signal and the second isolated audio signal separately, wherein processing the first and second isolated audio signals includes processing the loudness of the first isolated audio signal and / or the second isolated audio signal.
[0031] This aspect of the present invention focuses on the separate processing of a first separated audio signal and a second separated audio signal. This method can be implemented at consumer sites on consumer devices such as televisions, laptops, smartphones, or headphones.
[0032] A further aspect of the present invention provides a system for emphasizing dialogue in an original audio signal, which includes dialogue and non-dialogue components. The system comprises a dialogue separation unit that provides the dialogue components of the original audio signal in a first separated audio signal and the non-dialogue components of the original audio signal in a second separated audio signal. The audio system further comprises a loudness processing unit configured to process the first separated audio signal and the second separated audio signal separately, wherein processing the first separated audio signal and the second separated audio signal includes processing the loudness of the first separated audio signal and / or the loudness of the second separated audio signal. An audio mixer is further provided that is configured to combine the processed first separated audio signal and the second separated audio signal to provide a processed audio signal.
[0033] It should be noted that the dialogue isolation unit, loudness processing unit, and audio mixer do not necessarily belong to the same entity. For example, the dialogue isolation unit may be implemented in the headend or the cloud. The loudness processing unit and audio mixer may be implemented in a consumer device. In another embodiment, the dialogue isolation unit is implemented in a high-power device at the customer site, such as a set-top box or television, while the loudness processing unit and audio mixer are implemented in another customer device, such as a consumer device.
[0034] The system embodiments correspond to the embodiments of the method described above. For example, the dialogue separation unit may be configured to receive dialogue and non-dialogue components from a source where the first and second separated audio signals are separately available. Alternatively, the dialogue separation unit may be configured to provide dialogue and non-dialogue components by separating the dialogue component from the non-dialogue component in the original audio signal.
[0035] According to yet another aspect of the present invention, a non-temporary computer-readable medium storing executable instructions is provided, and when an instruction is executed by a processor, the following operations are performed: an operation to provide the dialogue component of the original audio signal in a first isolated audio signal; an operation to provide the non-dialogue component of the original audio signal in a second isolated audio signal; an operation to process the first isolated audio signal and the second isolated audio signal separately, wherein processing the first and second isolated audio signals includes an operation to process the loudness of the first isolated audio signal and / or the second isolated audio signal; and an operation to combine the processed first and second isolated audio signals to provide a processed audio signal.
[0036] Embodiments of non-temporary computer-readable media correspond to embodiments of the method described above.
[0037] Throughout the drawings, reference numerals are reused to indicate the correspondence between the referenced elements. The drawings are provided to illustrate, and not to limit, embodiments of the invention as described herein. [Brief explanation of the drawing]
[0038] [Figure 1] This is a general system architecture for a system that enhances dialogue in an original audio signal, and the system architecture comprises a dialogue separation unit, a loudness processing unit, and an audio mixing unit.
[0039] [Figure 2] This figure shows one embodiment of the dialog separation unit shown in Figure 1.
[0040] [Figure 3] This is a flowchart showing how to emphasize dialogue in the original audio signal.
[0041] [Figure 4] This is a flowchart of a method for processing a first separated audio signal that includes the dialogue component of the original audio signal.
[0042] [Figure 5] This is a flowchart of a method for processing a second, separated audio signal that contains non-dialogue components of the original audio signal.
[0043] [Figure 6] Figure 1 shows one embodiment of the loudness processing unit.
[0044] [Figure 7] This is one embodiment of a dynamic range processor.
[0045] [Figure 8] This figure shows the loudness measurement of a multi-channel audio signal in accordance with ITU-R BS.1770-1.
[0046] [Figure 9] This is an example of a dialogue loudness correction curve.
[0047] [Figure 10] This is an example of a non-dialogue loudness correction curve.
[0048] [Figure 11] This is a flowchart illustrating an exemplary method for processing a first separated audio signal containing the dialogue component of the original audio signal.
[0049] [Figure 12] This is a flowchart illustrating an exemplary method for processing a second, separated audio signal containing non-dialogue components of the original audio signal.
[0050] [Figure 13]This diagram shows a system that implements centralized dialogue isolation with personalized distributed dialogue emphasis.
[0051] [Figure 14] This figure shows a further system that implements centralized dialogue isolation with personalized distributed dialogue enhancement, which provides 2-channel wireless transmission and endpoint-specific post-processing.
[0052] [Figure 15] Figure 1 shows a typical stereo in / out implementation for a system.
[0053] [Figure 16] A further stereo implementation of the typical system in Figure 1, where the separated dialogue channels are retained for additional post-processing.
[0054] [Figure 17] This is a further stereo implementation of the typical system shown in Figure 1, with single-channel dialogue separation implemented.
[0055] [Figure 18] This is an embodiment of a stereo shuffler.
[0056] [Figure 19] A further stereo implementation of the typical system in Figure 1, where single-channel dialogue separation is implemented using a stereo shuffler.
[0057] [Figure 20] This is a multi-channel implementation of a typical system shown in Figure 1, where single-channel dialogue isolation is implemented.
[0058] [Figure 21] This is a further multi-channel implementation of the typical system shown in Figure 1, in which stereo dialogue separation is implemented.
[0059] [Figure 22] This is a further multi-channel implementation of the typical system shown in Figure 1, where 3-channel dialogue separation is implemented. [Modes for carrying out the invention]
[0060] The following description details various embodiments of methods and systems for enhancing dialogue in an original audio signal that includes both dialogue and non-dialogue components.
[0061] Figure 1 shows a typical system architecture. This system comprises a dialogue separation unit 1, a loudness processing unit 2, and an audio mixing unit 3. The dialogue separation unit 1 receives the original audio signal A and separates the dialogue component from the other components of the signal. The dialogue component is provided by a first separated audio signal 11, and the non-dialogue component is provided by a second separated audio signal 12. Once the dialogue component is separated, it can be analyzed and processed separately from the non-dialogue component. The dialogue component includes the spoken language. The non-dialogue component represents any signal component that is not part of the narrative dialogue signal. These can include music and sound effects.
[0062] The original audio signal A can be a digital audio signal. This may be stored in a file or it may be streamed. For example, the original audio signal may be the audio soundtrack of video content played on a media device such as a set-top box, television, or laptop. In another example, the original audio signal may be a live radio broadcast.
[0063] The loudness processing unit 2 receives a first isolated audio signal 11 containing dialogue components and a second isolated audio signal 12 containing non-dialogue components, and processes the first isolated audio signal 11 and the second isolated audio signal 12 separately. In particular, the audio signals 11 and 12 are separately analyzed and processed for loudness, as illustrated in embodiments with respect to Figures 4, 5, 11, and 12. The loudness processing unit 2 plays a role in ensuring that the dialogue loudness level does not fall below the listener's desired level and in adjusting the non-dialogue loudness level so that the minimum dialogue-to-non-dialogue ratio remains constant. The audio mixer 3 receives the processed first isolated audio signal 110 and the processed second isolated audio signal 120, combines them, and provides a processed audio signal B with enhanced dialogue as an output. For example, a dialogue-enhanced soundtrack is output to the audio mixer 3.
[0064] It is noted that system components 1, 2, and 3 can be separated across devices. For example, the dialogue isolation unit 1 may occur at the headend or in the cloud, while the dialogue processing unit 21 may occur on a consumer device. In another example, one high-power device (e.g., a set-top box or television) may include the dialogue isolation unit 1, while another device may include the dialogue processing unit 2 and the audio mixer 3.
[0065] Figure 2 shows an exemplary embodiment of the dialogue separation unit 1 of Figure 1. The purpose of the dialogue separation unit 1 is to separate the incoming original audio signal A into two separated signals 11 and 12, each having a dialogue component and a non-dialogue component, respectively. In the illustrated embodiment, a machine learning network is trained to perform this task, although this is not necessarily the case. However, it is also possible to use other audio source separation techniques.
[0066] According to the embodiment shown in Figure 2, the original audio mix is transformed into a Short-Time Fourier Transform (STFT) domain in the STFT unit 101, and appropriate audio features are extracted from the STFT data in the feature extraction unit 102. For example, the data can be transformed into a log-magnitude representation, and linear frequency bands can be grouped into fewer important perceptual bands. The input features are taken up as the input layer of the machine learning network 103 (e.g., based on a U-NET architecture), targeting a set of output features. In this case, the output features represent a set of frequency-domain weights that represent filters necessary to separate dialogue and / or non-dialogue signal components. The output features are processed so that the linear STFT domain representation is a derivative. These weights per frequency bin are applied in the dialogue filtering unit 106 to a delayed version of the original STFT domain signal (delayed by the delay unit 105). This delay is used to compensate for processing delays involved in feature extraction, network inference, and reconstruction. As a result, the filtered STFT data is backed into the time domain using inverse STFT processing in the ISTFT unit 107. The extracted dialogue components are output to the first separated audio signal 11, and the non-dialogue components are output to the second separated audio signal 12.
[0067] The machine learning network may be trained on a large database of isolated dialog and non-dialog signal examples (representing the desired output) and their corresponding mixtures (provided in the input).
[0068] If the first and second separated audio signals 11 and 12 are readily available from a source such as a production source, there is no need for dialogue separation such that the dialogue separation unit 2 can then be bypassed or reduced to a unit that simply receives the first and second separated audio signals 11 and 12.
[0069] Figure 3 is a flowchart of a method for enhancing dialogue in the original audio signal. In steps 301 and 302, the dialogue and non-dialogue components of the original audio signal are provided in first and second separated audio signals. The dialogue and non-dialogue components may be provided by dialogue separation techniques as described with respect to Figure 2, or simply received if separately available. In step 303, the first separated audio signal and the second separated audio signal are processed separately, and the loudness of the first separated audio signal and / or the loudness of the second separated audio signal are adjusted to improve dialogue intelligibility. This can be carried out in the loudness processing unit 2 of Figure 1, and a more specific embodiment is described with respect to Figure 6. In step 304, the processed first separated audio signal and the second separated audio signal are combined to provide a processed audio signal.
[0070] Examples of how the first and second isolated audio signals are processed are provided in Figures 4 and 5. Please refer to Figures 4 and 5. According to Figure 4, in the first step 401, the short-term dialogue loudness level (STDLL) of the first isolated audio signal is determined. Such determination of loudness levels can be based on standard loudness measurements, as will be discussed later. Short-term loudness measurement means that the loudness is measured on a window. For example, short-term loudness is measured on a small window of 20 ms with a 10 ms look-ahead, such that the window is centered on the current sample.
[0071] In step 402, the determined STDLL is set to a predefined minimum dialog loudness level (DLL). MIN It is determined whether the value is less than DLL. If so, in step 403, the first separated audio signal is determined to be DLL MINIt is amplified toward the first isolated audio signal. Otherwise, the first isolated audio signal remains unchanged. Thus, the loudness of the first isolated audio signal into which the dialogue is input is set to the predefined minimum dialogue loudness level DLL. MIN It is normalized to this.
[0072] It should be noted that this approach differs substantially from conventional approaches. Conventional approaches employ volume normalization to ensure that the levels of quiet passages are increased to a more audible level, and dynamic range compression to ensure that the levels of excessively loud passages are reduced. However, when applied to the original soundtrack mix, the combination of these processes can introduce audible artifacts. For example, enabling volume normalization can make quiet, non-dialogue passages (e.g., non-dialogue scenes of trees in a forest) unnaturally louder. Similarly, using dynamic range compression on louder non-dialogue soundtrack components can also affect louder dialogue, potentially making dialogue less audible in the presence of non-dialogue sounds. Furthermore, these solutions often apply high-frequency spectral boosts to the signal to improve the intelligibility of dialogue. This filter is typically applied not only to the dialogue components of the mix but also to the non-dialogue components, affecting the spectral balance of the entire soundtrack.
[0073] On the other hand, if the dialogue and non-dialogue components are initially separated into two separate signals, the short-term loudness of the separated dialogue component can be analyzed independently of the non-dialogue component. This allows loudness normalization to be applied only to the dialogue signal.
[0074] Figure 5 shows an example of how a second isolated audio signal containing non-dialogue components can be processed. In step 501, the short-term non-dialogue loudness level (STNDLL) of the second isolated audio signal containing non-dialogue components is determined. This is also done by windowing the second isolated audio signal. Similar to the processing of the first isolated audio signal, the short-term loudness can be measured in a small 20ms window with a 10ms look-ahead. In general, the window size for determining the short-term loudness of the first isolated audio signal may be the same as, or different from, the window size for determining the short-term loudness of the second isolated audio signal. In this regard, it should be noted that the technique of applying different processing to dialogue streams and non-dialogue streams has the further advantage of being able to apply different loudness windows to each stream.
[0075] Step 502 involves the minimum dialog loudness level DLL. MIN (The same DLL as explained in Figure 3) MIN The difference between the level and the STNDLL level determined in step 501 is the predefined minimum dialog-to-non-dialog ratio D2ND. MIN It is determined whether it is less than the ratio D2ND. If so, the loudness level of the second separated audio signal is the ratio D2ND. MIN The second isolated audio signal is reduced to approach this value (step 503). Otherwise, the second isolated audio signal is not modified (step 504). Reducing the loudness level of the second isolated audio signal can be done by compressing the dynamic range of the second isolated audio signal.
[0076] The reduction in loudness level is effective only for non-dialogue signals.
[0077] Alternatively, the minimum dialog loudness level DLL MINInstead of determining the difference between the first isolated audio signal and the STDLL level, the difference between the STDLL level and the first isolated audio signal is determined, and the predefined minimum dialogue / non-dialogue ratio D2ND is determined. MIN Analyze whether it is less than or equal to.
[0078] Figure 6 shows an embodiment of the loudness processing unit 2 of Figure 1. The loudness processing unit 2 includes a unit 201 for determining the short-term loudness of the first separated audio signal 11. Unit 201 is a predefined minimum dialog loudness level DLL as described in relation to Figure 4. MIN The unit receives the following as an input parameter. The loudness processing unit 2 further comprises an amplification unit 202 configured to amplify the first isolated audio signal 11 according to the control signal received by unit 201. Furthermore, optionally, a spectral enhancement unit 203 for the first isolated audio signal 11 is provided. With respect to the second isolated audio signal 12, a unit 204 is provided for determining the short-term loudness of the second isolated audio signal 12. Unit 204 receives the following as an input parameter: the predefined minimum dialogue / non-dialogue ratio D2ND, as described with reference to Figure 5. MIN The loudness processing unit 2 further comprises a compression unit 205 configured to compress the second isolated audio signal 12 according to a control signal supplied from unit 204. The loudness processing unit 2 outputs the processed first isolated audio signal 110 and the processed second isolated audio signal 120.
[0079] Parameter DLL MIN and D2ND MIN This can be configured by the consumer or the system integrator.
[0080] Units 201, 202 and 204, 205 can be implemented using a general-purpose dynamic range processor (DRP), as shown in Figure 7. The DSP can be adapted to handle both dialog and non-dialog input streams. The signal received at input 401 is duplicated and delayed in the look-ahead delay unit 402 of the first path 408. For example, a 10ms look-ahead delay is incorporated to prepare the DRP to proactively respond to loudness transients into which it is input. In the second path 409, a loudness measurement unit 403, a gain computer 404, and a gain smoother 405 are provided. The determined gain is applied to the signal in path 408 by unit 406 and sent to output 407.
[0081] The loudness measurement unit 403 measures the loudness of a signal. In particular, it can measure the short-term loudness of the first isolated audio signal 11 or the short-term loudness of the second isolated audio signal 12. The loudness measurement is performed according to industry standards. In embodiments, loudness is estimated using the industry standard ITU-R BS.1770-1 and measured with a specific window size, such as 20 ms, to ensure both accuracy and responsiveness. In the standard ITU-R BS.1770, loudness is expressed in LKFS (Loudness, K-weighted, relative to Full Scale) or the synonym LUFS (Loudness units relative to full scale), introduced in EBU R128, which is the standard loudness measurement unit used for audio normalization in broadcast television systems and other video and music streaming services. In particular, ITU-R BS.1770-1, the first iteration of this standard, is particularly well-suited to handling immediate loudness fluctuations through continuous short-term measurements.
[0082] The ITU-R BS.1770-1 standard first processes each audio channel by applying pair filtering of second-order IIR filters and RLB (Revised Low-Frequency B-curve) filtering to emulate the frequency response of the human ear. Then, the mean square energy of the filtered signal over a measurement interval T is calculated to obtain the value zi of i for each channel. The mean square energy is determined as follows: JPEG2026509119000002.jpg11150
[0083] After mean square calculation, channel-specific weights Gi are applied, and the aggregated loudness value is obtained: JPEG2026509119000003.jpg13157
[0084] The channel weights are assigned as follows: Left (GL): 1.0, Right (GR): 1.0, Center (GC): Left Surround (GLs): 1.41, Right Surround (GRs): 1.41. This procedure is shown in Figure 8.
[0085] Referring again to Figure 7, the gain computer 404 calculates the gain using a variable curve determined by five control points (x / y coordinates). This allows the processor to function as a compressor, expander, loudness leveler, or a hybrid of these modes. Smoothing parameters can also be incorporated into the gain smoother 405 to ensure seamless transitions between operating zones. The gain smoother may also be a branch attack and release smoother, and the fast / slow attack and release time settings allow the processor to quickly adapt to large changes in input loudness.
[0086] Figures 9 and 10 show examples of dialog loudness correction curves 51 and non-dialog loudness correction curves 52. As previously mentioned, the gain is calculated using a modifiable curve determined by five control points, which provides flexibility for use in both dialog (upward gain) and non-dialog (attenuation) processing. Examples of curves for dialog and non-dialog are shown in Figures 9 and 10. The following parameter values are used in these examples. DLL MIN =-20dBFS;D2ND MIN =10dB; and DLL TRESH = -80dB. In practice, these curves are smoothed to ensure a seamless transition between operational zones.
[0087] Parameter DLL MIN and D2ND MIN This was mentioned earlier. Parameter DLL TRESH This determines the threshold dialog loudness level. This parameter functions as a VAD (Voice Activity Detector). DLL TRESH Without it, extremely low levels of dialog components (such as background dialogs and noise or artifacts in dialog channels) will be present in the DLL. MIN Undesirable high gain will be applied to match this. In this case, because the normalized ballistics need time to adapt to the abrupt change in loudness, undesirable loudness spikes may occur when transitioning from a quiet segment to a segment with narrative dialogue. Furthermore, DLL THRESH This helps avoid amplification of low-level processing artifacts caused by the preceding dialogue isolation process.
[0088] In Figure 9, the dialogue loudness correction curve 51 is made up of four loudness operating zones. In the first zone, from -infdB to -80dB, the loudness is attenuated by 10dB. In this example, -80dB is the threshold dialogue loudness level DLL.TRESH As mentioned above, DLL TRESH Below this level, dialogue loudness is not boosted. In this example, it is attenuated by 10dB. In the second zone from 80dB to -70dB, there is a gradual transition from the compression zone to the normalization zone. In the third zone from 70dB to -20dB, the loudness is DLL MIN This is the zone where the value is normalized. In the fourth zone, from 20dB to +infdB, the signal does not change.
[0089] In Figure 10, the non-dialogue loudness correction curve 52 is made up of two loudness operating zones. In the first zone from -infdB to -30dB, the signal remains unchanged. In the second zone from 30dB to +infdB, the signal is compressed with a compression ratio of 2:1. In other embodiments, the compression ratio may be different.
[0090] Figure 11 shows how the processing of a first separated audio signal containing a dialogue component is performed by the loudness processing unit 2 in Figure 6. This method is based on the method in Figure 4 but includes additional details. In step 111, a block of the dialogue stream is input, and the block size is determined by a window. In an embodiment, the window size may be 20 milliseconds. In step 112, the short-term loudness level STDLL of the dialogue component is measured, as described with respect to Figures 7 and 8. The short-term loudness level STDLL of the dialogue component is measured, as discussed with respect to Figures 7 and 8. In step 113, the short-term loudness level STDLL is set to a predefined threshold dialogue loudness level DLL. TRESH It is determined whether it is greater than or equal to. If not, in step 116, the unmodified block of the dialog stream is output. If it is, in step 114, the short-term loudness level STDLL is set according to industry recommendations and can be in the range of -22 LUFS to -27 LUFS (LUFS = "full scale of loudness units"), which is a predefined minimum dialog loudness level DLL. MINIt is further determined whether it is smaller than DLL. If not, the unmodified block of the dialog stream is output in step 116. If so, the short-term loudness level STDLL is less than the minimum dialog loudness level DLL, as indicated by the indicator in the third zone of Figure 9. MIN In step 115, the level of the dialogue component is amplified so that it approaches or equals the value. The order of steps 113 and 114 may be reversed.
[0091] Figure 12 shows how the loudness processing unit 2 of Figure 6 performs processing of a second separated audio signal containing non-dialogue components. This method is based on the method of Figure 5 but includes additional details. In step 121, a block of the non-dialogue stream is input, with the block size determined by a window. In an embodiment, the window size may be 20 milliseconds. In step 122, the short-term loudness level STNDLL of the non-dialogue components is measured, as described with respect to Figures 7 and 8. Measurements are performed as discussed with respect to Figures 7 and 8. In step 123, the minimum dialogue loudness level DLL is measured. MIN The difference between this and the short-term loudness level STNDLL measured in step 122 is the predefined minimum dialogue-to-non-dialogue ratio D2ND. MIN It is determined whether it is less than [a certain value]. If not, in step 125, the unmodified block of the non-dialog stream is output. If it is, the non-dialog signal is the mentioned difference DLL. MIN -STNDLL is the minimum dialog to non-dialog ratio D2ND MIN It is compressed to approximate the original size. Compression can include dynamic range compression. An example is shown in Figure 10.
[0092] In the above embodiment, it is assumed that everyone in the listening space hears a common audio output from the dialogue enhancement system. In this case, the algorithm applies dialogue processing according to the preferences of one listener, which is then heard by everyone in the listening environment. Figure 13 shows an alternative system that can provide a degree of alternative dialogue processing according to the needs of individual listeners (e.g., individuals with more pronounced hearing loss).
[0093] More specifically, in Figure 13, the original audio signal A is split into a dialogue component and a non-dialogue component in the dialogue separation unit, as described with respect to Figure 1. Alternatively, if the dialogue component and non-dialogue component are already available separately, they are simply received. The first and second audio signals 11 and 12 are then fed into a series of processing paths. The first processing path is a general processing path, and the processed audio signal B is provided to any number of listeners (for example, through television speakers) listening to the same processed mix. The separated signals 11 and 12 are processed in a generalized dialogue enhancement unit 61, which corresponds to the loudness processing unit 2 and audio mixing unit 3 in Figure 1.
[0094] Furthermore, one or more optionally personalized processing paths are provided, and personalized audio outputs B1, B2 are provided so that individual listener personal parameters are taken into consideration when processing the first isolated audio signal 11 and / or the second isolated audio signal 12. For example, a listener-specific personal auditory profile or subjective listening preference can be implemented in personalized dialogue enhancement units 62, 63 and applied to individual processing blocks. Personalized audio outputs B1, B2 can be played back via headphones or an auditory assistance device.
[0095] The personalized mix can be directed to in-ear monitors or headphones using a wired connection or low-latency wireless technology such as Bluetooth or ultra-wideband audio (UWB), as shown in Figure 14. In Figure 14, if the number of channels in audio input A is greater than 2, the channels are downmixed to stereo in downmix unit 64, and the downmixed channels are wirelessly transmitted to upmix units 65, 66 associated with individual users. After upmixing, personalized dialogue enhancement is performed in units 62, 63. The output of dialogue enhancement units 61, 62, 63 may be further enhanced by 3D audio processors 67, 68, 69 which are embedded with or attached to the headphones used by the individual listener.
[0096] In this regard, several modifications can be implemented. In one embodiment, the auditory devices of individual listeners may include noise-canceling features to minimize interference from a generalized version of the soundtrack played on a loudspeaker. Multiple individual mixes can be generated in a hub (e.g., a television or set-top box) and transmitted simultaneously from that source. Furthermore, in some embodiments, multi-channel audio output content may be downmixed to stereo before transmission. In some embodiments, multi-channel individual audio output content may be transmitted to a multi-channel headphone virtualization technology such as DTS Headphone.X. Such a virtualization algorithm can be applied on the transmitting device or within the headphones. In some embodiments, unmixed dialogue and non-dialogue audio streams are wirelessly transmitted to one or more headset sets with the necessary processing power, and personalized dialogue processing is applied in the headphones. In some embodiments, dialogue and non-dialogue audio streams are downmixed or encoded (spatially or otherwise) in a manner that allows for low-bandwidth transmission and reception of the original audio channels. For example, a stereo downmix of the original dialogue and background channels can be performed such that the dialogue is center-panned and the multi-channel non-dialogue signals are spatially encoded into stereo using an algorithm such as the DTS Neural Surround downmixer. This stereo signal can then be "upmixed" or decoded on the receiving headphone device to return to the separate dialogue and non-dialogue streams. In some embodiments, the original input signal is transmitted wirelessly to headphones having onboard processing capabilities, including a machine learning inference engine.The original audio soundtrack is transmitted to the headphones, and dialogue isolation and personalized dialogue processing are applied on a processor attached to or embedded within the headphones.
[0097] Different implementation topologies of the system and method are described below with reference to Figures 15 to 22.
[0098] In Figure 15, the input audio signal A is a stereo signal (indicated as "2.0"). A dialogue separation unit, trained to separate dialogue and non-dialogue components, separates the audio into a first separated audio stereo signal 11 for the dialogue components and a second separated audio stereo signal 12 for the non-dialogue components. The separated stereo dialogue and non-dialogue components are then passed to a stereo loudness processing unit 2, where the processed components are mixed again. Furthermore, an output limiter 7 can be provided to prevent the processed signal from saturating downstream.
[0099] In the embodiment shown in Figure 16, the input audio signal A is a stereo signal. The output of the loudness processing unit 2 remains isolated for further downstream processing in an additional post-processing unit 30 that integrates the audio mixing unit. The post-processing unit 30 may include algorithms for spatial audio processing for headphones and speakers. Other examples of downstream processing include algorithms that are better applied only to the non-dialogue components of the input signal, such as bass boosting. The basic concept of retaining isolated dialogue and non-dialogue outputs from the loudness processing unit 2 can be applied to any of the topologies described below. As in Figure 15, an output limiter 7 is also provided.
[0100] In the embodiment shown in Figure 17, the input audio signal A is a stereo signal. However, dialogue separation is performed on only one channel. Furthermore, it is assumed that the majority of the dialogue in the narrative content is center-panned. More specifically, the stereo input A, which has L and R channels, is upmixed into 3 channels (L, C, R) by unit 81, and the signal component that is initially panned to the center is extracted into a discrete center channel C. As a result, the signal splits into two paths. Component C is led to a single-channel dialogue separation unit 1. Since it is assumed that the main dialogue is represented by the extracted center channel, it can be assumed that the remaining (L, R) channels represent non-dialogue components. These residual channels are delayed in the delay processing unit 82 to compensate for the delay of the dialogue separation process and redirected to the loudness processing unit 2 (because they are taken into consideration when processing the non-dialogue signal components in the loudness processing unit 2). The loudness processing unit 2 evaluates the relative loudness of the isolated dialogue and the loudness of the remaining C and (L,R) non-dialogue components, and applies gain and attenuation to each signal component accordingly, and the (L,R) channels are amplified / compressed by the amplifier / compressor unit 83. The signals are then downmixed by the audio mixer 3 into a single stereo pair, which in this case is led to the output limiter 7. In the embodiment of Figure 17, the loudness processing unit 2 has multiple inputs. The first input is the extracted dialogue channel. One or more further inputs are the extracted non-dialogue channels. Furthermore, the L,R channels from the upmix unit 81 contain only non-dialogue.
[0101] Figures 18 and 19 relate to another embodiment. Figures 18 and 19 relate to an alternative embodiment of a system in which the input audio signal A is a stereo signal and dialogue separation and loudness processing are provided to only one channel. In this embodiment, an active 2-3 upmixer cannot be used. In such cases, it is possible to apply a passive upmix using a stereo shuffler configuration. A typical stereo shuffler configuration 84 is shown in Figure 18. The sum and difference of the left and right input channels L and R are formed twice for each channel. The output is the same as the input of a typical stereo shuffler.
[0102] Assuming that dialogue is generally center-panned, the sum of the L+R signals of the stereo input channel contains a large portion of the center-panned signal component, while the difference between the L and R input channels contains little to no dialogue. Therefore, as shown in Figure 19, the majority of the dialogue can be extracted from the sum component L+R, and this sum component becomes the dialogue separator and the receiver for dialogue separator unit 1. The original difference signal L+R and the non-dialogue component of the original sum signal L+R are then used to synthesize a stereo non-dialogue signal. The reconstructed stereo non-dialogue signal component and the mono sum signal are then analyzed by loudness processing unit 2 and remixed into an augmented stereo signal by audio mixer 3. As in other embodiments, the resulting output signal is led to the output limiter 7.
[0103] In the embodiment shown in Figure 20, the input audio signal A is a multi-channel signal (5.1 in the illustrated embodiment). The center channel of the multi-channel audio stream is assumed to already contain the most dialogue. Conversely, all other channels are assumed to be non-dialogue. Therefore, the center channel is simply redirected to a single-channel (1.0) dialogue separation processor 1, and the loudness processing unit 2 considers the relative loudness of the separated dialogue to the remaining non-dialogue channels (including those remaining from the center channel). The loudness processing gives each signal component appropriate gain and delay, and each signal is recombined to match the input multi-channel format (5.1 in the embodiment). Finally, a multi-channel limiter 7 is applied to the 5.1ch output.
[0104] In the embodiment shown in Figure 21, the dialogue separation model is trained to separate dialogue and dialogue components from a stereo signal, as shown in Figure 15. In this case, it can be assumed that the majority of the dialogue resides in the front (L,C,R) channel mix. After channel splitting in unit 85, these channels are downmixed into a stereo signal in downmixer 87 and led to stereo dialogue separation unit 1. The other channels (2.1: LS, RS, LFE) are assumed in this case to not contain narrative dialogue. The loudness of the separated stereo dialogue is compared in loudness processing unit 2 to the loudness of the stereo non-dialogue channels remaining with the (LS, RS, LFE) channels. An appropriate gain is calculated and applied to all channel signal components. The separated dialogue and non-dialogue outputs, which are originally derivatives from (L,C,R), are mixed together in audio mixer 3 and further upmixed to the original 3-channel layout using 2-3 channel upmixer 88. The resulting signal is then combined again with the original surround and LFE channel components in channel combiner 86 and reconfigured to match the original input format. As before, a multi-channel limiter can ultimately be applied to the resulting 5.1-channel output.
[0105] In the embodiment shown in Figure 22, a dialogue separation unit 1 with training capabilities is used for a 3-channel input signal. As a result, there is no need to downmix or upmix the (L, C, R) channels.
[0106] The embodiments and implementation topologies described above can receive multiple adaptations.
[0107] In some embodiments, the personalized audio output or a portion thereof is directed to a specific individual using a beamforming loudspeaker array.
[0108] In some embodiments, the wireless receiver may be an auditory assistance device (hearing aid). In this case, care must be taken to select a dialogue processing setting that takes into account the built-in auditory assistance technology.
[0109] In some embodiments, a general audio output can be broadcast to multiple wireless receivers without speaker output. This minimizes acoustic crosstalk among all listeners.
[0110] In some embodiments, only the dialogue channel is used for individualized processing and wireless transmission. This is the case when the listener only needs to augment the dialogue signal. This can be done using bone conduction headphones, near-field speakers, or open-ear headphones.
[0111] In some embodiments, the system may use an imaging sensor capable of identifying the presence and location of the listener. This can influence the algorithm parameters used. For example, a particular person's preferences may only be used if that person is in the room. Alternatively, a weighted average of the parameters of everyone detected in the room may be used. The listener's location is useful when considering ambient noise or when beamforming the dialogue to a specific individual.
[0112] In some embodiments, different user preferences can be applied to different types of content. For example, dramas may prefer a different set of loudness processing parameters than news. The content type may be obtained from content metadata or determined using algorithmic classification.
[0113] In some embodiments, when the described processing is applied to self-contained wearable devices (e.g., hearables, augmented reality headsets), it may be used in non-home environments (e.g., movie theaters or theaters). By default, the loudness adaptive algorithm is based on digital loudness levels (relative to digital full scale). If only microphone capture of the acoustic signal is available, SPL calibration against the digital level must be performed to ensure some degree of equivalence.
[0114] In some embodiments, automatic closed captions can be displayed on an augmented reality display or glasses.
[0115] Alternative embodiments and exemplary operating environments Many other variations not described herein will be apparent from this specification. For example, depending on the embodiment, certain operations, events, or functions of any method and algorithm described herein can be performed in a different order, added, integrated, or omitted entirely (thus, not all operations or events described herein are necessary for the implementation of the method and algorithm). Furthermore, in certain embodiments, operations or events can be performed simultaneously rather than sequentially, for example, by multithreading, interrupt handling, or by a multiprocessor or processor core, or on other parallel architectures. In addition, various tasks or processes can be performed by different machines and computing systems that can work together.
[0116] The processing and sequence of various exemplary logic blocks, modules, methods, and algorithms described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this hardware- and software compatibility, the operation of various exemplary components, blocks, modules, and processing is generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. The described functionality can be implemented in different ways for each specific application, but such decisions should not be construed as resulting in a departure from the scope of this specification.
[0117] Various exemplary logic blocks and modules described in relation to embodiments disclosed herein can be implemented or run by machines such as general-purpose processors, processing devices, computing devices having one or more processing devices, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic circuits, discrete hardware components, or any combination thereof designed to perform the functions described herein. General-purpose processors and processing devices may be microprocessors, but in alternative forms, the processor may be a controller, microcontroller, state machine, a combination thereof, or similar. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors working in conjunction with a DSP core, or any other such configuration.
[0118] Embodiments of the systems and methods described herein can operate within many types of general-purpose or dedicated computing system environments or configurations. Generally, computing environments can include any type of computer system, including, but are not limited to, one or more microprocessors, mainframe computers, digital signal processors, portable computing devices, personal organizers, device controllers, computing engines inside electrical products, mobile phones, desktop computers, mobile computers, tablet computers, smartphones, and computer systems based on electrical products with embedded computers, to name a few examples.
[0119] Such computing devices can typically be found in devices with at least some minimum computing power, including, but are not limited to, personal computers, server computers, handheld computing devices, laptops or mobile computers, communication devices such as mobile phones and PDAs, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and audio or video media players. In some embodiments, a computing device will include one or more processors. Each processor may be a specialized microprocessor such as a digital signal processor (DSP), a very long instruction word (VLIW), or another microcontroller, or a conventional central processing unit (CPU) having one or more processing cores, including a specialized graphics processing unit (GPU) based core within a multicore CPU.
[0120] The processing operations of methods, processes, or algorithms described in relation to embodiments disclosed herein can be embodied directly in hardware, in software modules executed by a processor, or in any combination of the two. The software may be contained in a computer-readable medium accessible to a computing device. The computer-readable medium includes both volatile and non-volatile media, which are either removable or non-removable, or any combination thereof. The computer-readable medium is used to store information such as computer-readable instructions or computer-executable instructions, data structures, program modules, or other data. The computer-readable medium may include, but is not limited to, computer storage media and communication media.
[0121] Computer storage media include, but are not limited to, computer-readable or machine-readable media or storage devices such as Blu-ray® discs (BD), digital multipurpose discs (DVD), compact discs (CD), floppy disks, tape drives, hard drives, optical drives, solid-state memory devices, RAM memory, ROM memory, EPROM memory, EEPROM memory, flash memory, or other memory technologies, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or any other devices that can be used to store desired information and are accessible by one or more computing devices.
[0122] Software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or non-temporary computer-readable storage media, multiple media, or any other form of physical computer storage known in this art. An exemplary storage medium may be coupled to a processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The storage medium may be volatile or non-volatile. The processor and storage medium may reside in an ASIC. The ASIC may be embedded in a user terminal. Alternatively, the processor and storage medium may exist as separate components within a user terminal.
[0123] As used herein, the term “non-transient” means “permanent or long-lived.” The term “non-transient computer-readable medium” includes any and all computer-readable medium, with the sole exception being transient propagating signals. This term includes, but is not limited to, non-transient computer-readable medium such as register memory, processor cache, and random access memory (RAM).
[0124] An "audio signal" is a signal that represents physical sound.
[0125] The retention of information such as computer-readable or computer-executable instructions, data structures, and program modules can also be achieved by using various communication media to encode one or more modulated data signals, electromagnetic waves such as carrier waves, or other transport mechanisms or communication protocols, including any wired or wireless information distribution mechanism. Generally, these communication media refer to signals in which one or more of their characteristics are set or modified in a manner that encodes information or instructions within the signal. For example, communication media include wired media such as wired networks or direct wired connections that transmit one or more modulated data signals, and wireless media such as acoustics, radio frequency (RF), infrared, and lasers that transmit, receive, or both transmit one or more modulated data signals or electromagnetic waves. Any combination of the above is also included in the scope of communication media.
[0126] Furthermore, one or any combination of software, programs, or computer program products, or any part thereof, that embody some or all of the various embodiments of the encoding and decoding systems and methods described herein, can be stored, received, transmitted, or read from any desired combination of computer-readable media or machine-readable media or storage devices and communication media in the form of computer-executable instructions or other data structures.
[0127] Embodiments of the systems and methods described herein can be further described in the general context of computer executable instructions, such as program modules, executed by computing devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. Embodiments described herein can also be implemented in a distributed computing environment in which tasks are executed by one or more remote processing devices or in a cloud consisting of one or more devices, linked via one or more communication networks. In a distributed computing environment, program modules can reside in both local and remote computer storage media, including media storage devices. Furthermore, the instructions described above can be implemented in part or as hardware logic circuits, which may or may not include a processor.
[0128] As used herein, conditional words, in particular "can," "might," "may," "eg," and similar terms, unless otherwise explicitly stated or understood in the context in which they are used, are generally intended to convey that a particular embodiment includes certain features, elements, and / or states, while other embodiments do not. Thus, such conditional words are not generally intended to suggest that features, elements, and / or states are necessarily required in one or more embodiments, nor are they intended to suggest that one or more embodiments necessarily include logic for determining whether these features, elements, and / or states are included or performed in any particular embodiment, with or without author input or instructions. The terms "comprising," "including," "having," and similar terms are synonymous and are used in a comprehensive, open-ended manner, not excluding any additional elements, features, actions, or behaviors. Furthermore, the term "or" is used in its inclusive sense (rather than its exclusive sense), and therefore, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list.
[0129] While the above detailed description illustrates, explains, and points out novel features applicable to various embodiments, it will be understood that various omissions, substitutions, and modifications can be implemented in the form and details of the illustrated devices or algorithms without departing from the spirit of this disclosure. As will be recognized, some features can be used or implemented separately from others, so that specific embodiments of the invention described herein can be embodied in forms that do not necessarily provide all of the features and benefits set forth herein.
[0130] Furthermore, while this subject matter is described using terminology specific to structural features and methodological behavior, it should be understood that the subject matter defined in the attached claims is not necessarily limited to the specific features or behaviors described above. Rather, the specific features and behaviors described above are disclosed as exemplary forms for carrying out the claims.
Claims
1. A method for improving the dialogue intelligibility of an original audio signal including dialogue and non-dialogue components, The dialogue component of the original audio signal is provided as a first separated audio signal. The non-dialogue component of the original audio signal is provided in a second separated audio signal. Processing the first isolated audio signal and the second isolated audio signal separately, wherein processing the first isolated audio signal and the second isolated audio signal includes processing the loudness of the first isolated audio signal and / or the second isolated audio signal. The processed first separated audio signal and the second separated audio signal are combined to provide a processed audio signal. Methods that include...
2. Providing the dialogue component in a first isolated audio signal and providing the non-dialogue component in a second isolated audio signal includes receiving the first and second isolated audio signals from sources from which the first and second isolated audio signals are separately available. The method according to claim 1.
3. Providing the dialogue component in a first separated audio signal and providing the non-dialogue component in a second separated audio signal includes separating the dialogue component from the non-dialogue component in the original audio signal. The method according to claim 1.
4. Processing the first separated audio signal is The first separated audio signal is to be determined to have a short-term loudness level, The determined short-term loudness level is a predefined minimum dialog loudness level DLL. MIN Determine whether it is less than, The determined short-term loudness level is the predefined minimum dialog loudness level DLL. MIN If less than the minimum dialog loudness level DLL MIN The first separated audio signal is amplified toward, The short-term loudness level determined above is the small dialog loudness level DLL. MIN If it is not less than, the aforementioned first separated audio signal will not be modified, The method according to any one of claims 1 to 3, including
5. The processed first separated audio signal is spectrally enhanced before being combined with the processed second separated audio signal. The method according to claim 4.
6. The above method further, Determining the audio activity in the first separated audio signal, Only when the aforementioned audio activity is determined, the first separated audio signal is transmitted to the minimum dialogue loudness level DLL. MIN To amplify toward, The method according to claim 4 or 5, including the method described in claim 4 or 5.
7. Determining the aforementioned audio activity is that the short-term loudness level of the first separated audio signal is a threshold dialogue loudness level DLL. THRESH This includes determining whether the determined short-term loudness level is higher than the threshold dialog loudness level DLL. THRESH Only if it is higher than the minimum dialogue loudness level DLL, the first separated audio signal is the minimum dialogue loudness level DLL. MIN It is amplified toward The method according to claim 6.
8. Amplifying the first separated audio signal includes using a dynamic range processor that applies gain using a modifiable curve determined by a plurality of control points. The method according to any one of claims 4 to 7.
9. Processing the second separated audio signal is Determining the short-term loudness level of the first separated audio signal or obtaining a predefined minimum dialogue loudness level DLL of the first separated audio signal MIN and The short-term loudness level of the second separated audio signal is determined, The difference between the short-term loudness level of the first isolated audio signal and the short-term loudness level of the second isolated audio signal, or the minimum dialogue loudness level DLL MIN The difference between this and the short-term loudness level of the second separated audio signal is the predefined minimum dialogue / non-dialogue ratio D2ND MIN Determine whether it is less than, If the difference is less than the minimum dialog / non-dialog ratio D2ND MIN To approach this, the loudness level of the second separated audio signal is reduced, If it is not less than, the second separated audio signal described above will not be changed. The method according to any one of claims 4 to 7, including
10. Reducing the loudness level of the second isolated audio signal includes compressing the dynamic range of the second isolated audio signal. The method according to claim 9.
11. Reducing the loudness level of the second isolated audio signal includes using a dynamic range processor that applies gain using a variable curve determined by multiple level controllers. The method according to claim 9 or 10.
12. A short-term loudness level is determined for a series of windows of a predetermined length, and the loudness level is determined according to industry standards. The method according to any one of claims 4 to 11.
13. The first separated audio signal and the second separated audio signal are processed in multiple processing paths. The aforementioned processing path is A general processing path, wherein the processed audio signal is provided to any number of listeners, At least one individualized processing path, wherein the processed audio signal is provided to individual listeners, Includes, Processing the first isolated audio signal and / or the second isolated audio signal includes using parameters personalized for each individual listener during the processing, The method according to any one of claims 1 to 12.
14. The parameters personalized for each individual listener include at least one of the listener's unique personal auditory profile and subjective listening preferences. The method according to claim 13.
15. The aforementioned original audio signal is an audio soundtrack. The method according to any one of claims 1 to 14.
16. The original audio signal is a stereo signal or a multi-channel signal, and for each channel of the stereo signal or multi-channel signal, the dialogue component is provided by a first separated audio signal, and the non-dialogue component is provided by a second separated audio signal. The method according to any one of claims 1 to 15.
17. The original audio signal is a stereo signal, the stereo signal is upmixed into a three-channel signal including a center channel, a left channel, and a right channel, the signal component of the stereo signal initially panned to the center is extracted to the center channel, the dialogue component is provided to the center channel only by a first separated audio signal, the non-dialogue component is provided by a second separated audio signal, and the second separate audio channel is combined with the left channel and the right channel for loudness processing. The method according to any one of claims 1 to 15.
18. The original audio signal is a multi-channel signal including a center channel and several further channels, wherein the dialogue component is provided to the center channel only by a first separated audio signal, the non-dialogue component is provided by a second separated audio signal, and the second separated audio signal is coupled with the further channels for loudness processing. The method according to any one of claims 1 to 15.
19. The original audio signal is a multi-channel signal including a center channel, a left channel, a right channel, and further channels, the center channel, the left channel, and the right channel are downmixed to two channels, and for each of the two downmixed channels, the dialogue component is provided by a first separated audio signal, the non-dialogue component is provided by a second separated audio signal, and the second separated audio signal is coupled with the further channels for loudness processing. The method according to any one of claims 1 to 15.
20. The processed first separated audio signal and the processed second separated audio signal are further processed by applying spatial audio processing and / or a specific algorithm before the processed first separated audio signal and the second separated audio signal are combined. The method according to any one of claims 1 to 19.
21. A method for improving the dialogue intelligibility of an original audio signal including dialogue and non-dialogue components, Receiving the dialogue component of the original audio signal in the first separated audio signal, Receiving the non-dialogue component of the original audio signal in the second separated audio signal, Processing the first separated audio signal and the second separated audio signal separately, wherein processing the first separated audio signal and the second separated audio signal includes processing the loudness of the first separated audio signal and / or the second separated audio signal. Methods that include...
22. A system for improving dialogue intelligibility in an original audio signal that includes dialogue and non-dialogue components, A dialogue separation unit configured to provide the dialogue component of the original audio signal as a first separated audio signal and the non-dialogue component of the original audio signal as a second separated audio signal, A loudness processing unit configured to process the first isolated audio signal and the second isolated audio signal separately, wherein processing the first isolated audio signal and the second isolated audio signal includes processing the loudness of the first isolated audio signal and / or the second isolated audio signal, An audio mixer configured to combine the processed first separated audio signal and the second separated audio signal to provide a processed audio signal, A system that includes these features.
23. The dialogue separation unit is configured to provide the dialogue component in a first separated audio signal and the non-dialogue component in a second separated audio signal, and includes receiving the first and second separated audio signals from sources from which the first and second separated audio signals are separately available. The audio system according to claim 22.
24. The dialogue separation unit is configured to separate the dialogue component from the non-dialogue component in the original audio signal, thereby providing the dialogue component as a first separated audio signal and the non-dialogue component as a second separated audio signal. The audio system according to claim 22.
25. The loudness processing unit is The first separated audio signal is to be determined to have a short-term loudness level, The determined short-term loudness level is a predefined minimum dialog loudness level DLL. MIN Determine whether it is less than, The determined short-term loudness level is the minimum dialogue loudness level (DLL). MIN If less than the specified value, the first isolated audio signal is treated as a predefined minimum dialogue loudness level DLL. MIN It amplifies toward The determined short-term loudness level is the minimum dialogue loudness level (DLL). MIN If it is not less than, the first separated audio signal will not be modified. The first separated audio signal is configured to be processed by The system according to any one of claims 22 to 24.
26. The loudness processing unit is further configured to spectrally enhance the processed first isolated audio signal before coupling it with the processed second isolated audio signal. The system according to claim 25.
27. The loudness processing unit further, Determine the audio activity in the first separated audio signal described above. Only when vocal activity is determined, the first isolated audio signal is set to the minimum dialogue loudness level DLL. MIN It amplifies toward The system according to claim 25 or 26, configured as follows.
28. The loudness processing unit determines that the short-term loudness level of the first separated audio signal is equal to the threshold dialog loudness level DLL. THRESH It is configured to determine vocal activity by determining whether it is higher than or equal to The first separated audio signal is determined by the threshold dialogue loudness level DLL. THRESH Only if it is higher than the minimum dialogue loudness level DLL MIN It is amplified toward The system according to claim 27.
29. The loudness processing unit comprises a dynamic range processor configured to amplify the first isolated audio signal by applying gain, wherein the dynamic range processor is configured to use a modifiable curve determined by a plurality of control points for applying the gain. The system according to any one of claims 25 to 28.
30. The loudness processing unit is Determining the short-term loudness level of the first isolated audio signal, or a predefined minimum dialogue loudness level DLL of the first isolated audio signal. MIN To obtain, The short-term loudness level of the second separated audio signal is determined, The difference between the short-term loudness level of the first isolated audio signal and the short-term loudness level of the second isolated audio signal, or the minimum dialogue loudness level DLL MIN The difference between this and the short-term loudness level of the second separated audio signal is the predefined minimum dialogue / non-dialogue ratio D2ND MIN Determine whether it is less than, If the difference is less than the minimum dialog / non-dialog ratio D2ND MIN To approach this, the loudness level of the second separated audio signal is reduced, If it is not less than, the second separated audio signal described above will not be changed. The second separated audio signal is configured to be processed by The system according to any one of claims 22 to 29.
31. The loudness processing unit is configured to reduce the loudness level of the second isolated audio signal by compressing the dynamic range of the second isolated audio signal. The audio system according to claim 30.
32. The loudness processing unit includes a dynamic range processor configured to reduce the loudness level of the second isolated audio signal by applying gain, and for applying gain, the dynamic range processor is configured to use a modifiable curve determined by a plurality of control points. The system according to claim 30 or 31.
33. The loudness processing unit is configured to determine a short-term loudness level for a continuous window of a predetermined length, and the loudness level is determined according to industry standards. The system according to any one of claims 22 to 32.
34. The first separated audio signal and the second separated audio signal are processed in multiple processing paths. The aforementioned processing path is A general processing path comprising a general loudness processing unit, wherein the processed audio signal is provided to any number of listeners; A personalized processing path including a personalized loudness processing unit, Equipped with, The processed audio signal is provided to individual listeners, and processing the first isolated audio signal and / or the second isolated audio signal includes using parameters personalized to the individual listeners during the processing. The system according to any one of claims 22 to 33.
35. The personalized loudness processing unit is configured to process the first and / or second isolated audio signals using personalized parameters that include at least one of the listener's unique personal auditory profile and subjective listening preferences. The audio system according to claim 34.
36. The aforementioned original audio signal is an audio soundtrack. An audio system according to any one of claims 22 to 35.
37. The original audio signal is a stereo signal or a multi-channel signal, and the dialogue separation unit is configured to provide the dialogue component to each channel of the stereo signal or multi-channel signal using a first separated audio signal, and the non-dialogue component using a second separated audio signal. The system according to any one of claims 22 to 36.
38. The original audio signal is a stereo signal. The system further comprises an upmixer configured to upmix the stereo signal into a three-channel signal including a center channel, a left channel, and a right channel, wherein the signal component of the stereo signal, initially panned to the center, is extracted to the center channel, the dialogue separation unit is configured to provide only the dialogue component to the center channel as a first separated audio signal and the non-dialogue component as a second separated audio signal, and the loudness processing unit is configured to combine the second separated audio channel with the left channel and the right channel for loudness processing. The system according to any one of claims 22 to 36.
39. The original audio signal is a multi-channel signal including a center channel and a plurality of further channels, the dialogue separation unit is configured to provide the dialogue component to the center channel only with a first separated audio signal and the non-dialogue component with a second separated audio signal, and the loudness processing unit is configured to combine the second separated audio signal with the further channels and perform loudness processing. An audio system according to any one of claims 22 to 36.
40. The original audio signal is a multi-channel signal including a center channel, a left channel, a right channel, and further channels, and the system further comprises a downmixer configured to downmix the center channel, the left channel, and the right channel to two channels. The dialogue separation unit is configured to provide the dialogue component to each of the two downmixed channels as a first separated audio signal and the non-dialogue component as a second separated audio signal, and the loudness processing unit is configured to combine the second separated audio signal with further channels and perform loudness processing. The system according to any one of claims 22 to 36.
41. The system further comprises a post-processing unit configured to further process the processed first and second separated audio signals by applying spatial audio processing and / or a specific algorithm before the processed first and second separated audio signals are combined. The system according to any one of claims 22 to 40.
42. When executed by the processor, The operation of providing the dialogue component of the original audio signal in a first separated audio signal, The operation of providing the non-dialogue component of the original audio signal in a second separated audio signal, An operation for processing the first isolated audio signal and the second isolated audio signal separately, wherein the operation for processing the first isolated audio signal and the second isolated audio signal includes an operation for processing the loudness of the first isolated audio signal and / or the second isolated audio signal, The operation of combining the processed first separated audio signal and the second separated audio signal to provide a processed audio signal, A non-temporary, computer-readable medium containing executable instructions for performing a task.
43. Processing the first separated audio signal is The first separated audio signal is to be determined to have a short-term loudness level, The determined short-term loudness level is a predefined minimum dialog loudness level DLL. MIN Determine whether it is less than, The determined short-term loudness level is the minimum dialogue loudness level DLL. MIN If less than the specified value, the first isolated audio signal is treated as a predefined minimum dialogue loudness level DLL. MIN It amplifies toward, The determined short-term loudness level is the minimum dialogue loudness level DLL. MIN If it is not less than, the first separated audio signal will not be modified. A non-temporary computer-readable medium according to claim 42, including the following:
44. The operation of determining the audio activity in the first separated audio signal, Only when voice activity is determined, the first isolated audio signal is set to the minimum dialogue loudness level DLL. MIN An action that amplifies toward, A non-temporary computer-readable medium according to claim 43, further performing the following:
45. Determining speech activity is the short-term loudness level of the first separated audio signal, which is the threshold dialogue loudness level DLL. THRESH This includes determining whether the determined short-term loudness level is higher than the threshold dialog loudness level DLL. THRESH Only if it is higher than the minimum dialogue loudness level DLL, the first separated audio signal is the minimum dialogue loudness level DLL. MIN It is amplified toward The non-temporary computer-readable medium according to claim 44.
46. Processing the second separated audio signal is Determining the short-term loudness level of the first isolated audio signal, or a predefined minimum dialogue loudness level DLL of the first isolated audio signal. MIN To obtain, The short-term loudness level of the second separated audio signal is determined, The difference between the short-term loudness level of the first isolated audio signal and the short-term loudness level of the second isolated audio signal, or the minimum dialogue loudness level DLL MIN The difference between this and the short-term loudness level of the second separated audio signal is the predefined minimum dialogue / non-dialogue ratio D2ND MIN Determine whether it is less than, If the difference is less than the minimum dialog / non-dialog ratio D2ND MIN To approach this, the loudness level of the second separated audio signal is reduced, If it is not less than, the second separated audio signal described above will not be changed. A non-temporary computer-readable medium according to any one of claims 42 to 45, including the following: