System and method for providing personalized audio streaming and rendering
By processing audio signals through a dialogue clarity engine and a cross-gradient renderer to generate the most accessible mix, the problem of the inability to personalize audio streaming in existing technologies is solved, enabling an audio experience where dialogues can be heard clearly in different noise environments.
Patent Information
- Application Number
- CN202510499812.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-12
- Filing Date
- 2025-04-21
- Publication Date
- 2025-10-24
AI Technical Summary
Existing audio streaming solutions cannot provide a personalized audio experience, cannot automatically adjust audio quality according to ambient noise, and switching audio versions is cumbersome, making it difficult for users to hear conversations clearly in different noisy environments.
The audio signal is processed through the Dialogue Clarity Engine (DCE), including dialogue isolation, enhancement, dynamic range compression, and audio loudness normalization, to generate a Max-Acc-Mix. The audio ratio is adjusted in real time using a cross-gradient renderer, allowing multiple users to listen to different audio versions simultaneously.
It automatically adjusts audio quality based on ambient noise, providing a personalized audio experience. It supports multiple users listening to different audio versions simultaneously, avoiding the tedious process of switching streams and ensuring clear dialogue.
Smart Images

Figure CN120835168A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 636,400, filed April 19, 2024, and U.S. Provisional Patent Application No. 63 / 694,107, filed September 12, 2024, both of which are incorporated herein by reference to the fullest extent permitted by applicable law. BACKGROUND
[0002] Current audio streaming solutions typically rely on a single audio bitstream and decoder per streaming application, which limits the ability of content owners and streaming service providers to personalize the audio experience they provide to end users.
[0003] In particular, existing streaming services provide separate streams for different versions of dialog enhancement, e.g., English dialog enhancement high, English dialog enhancement medium, etc. In this case, the user must select which version they want to listen to. Moreover, in order to change to another audio version, e.g., because the background ambient noise changed, the user must make a request and retrieve the new version from the appropriate content server. Furthermore, most television devices apply post-processing to the decoded audio bitstream (such as AI sound enhancement) to reduce noise or enhance dialog, which does not preserve the integrity of the original sound mix nor the creative intent of the content owner.
[0004] Switching audio streams can be cumbersome, slow, and cause digital streaming file buffering / load delays during audio delivery from a content delivery network (CDN) to an end user. Moreover, conventional systems only provide a predefined number of dialog enhancement versions (e.g., 3-5 versions), without guaranteeing that a given selected version can be ideal for the noise environment of the listener / user. Such an approach can make it difficult for the user to hear the dialog of the audio content when the background noise becomes louder than the expected background noise of the selected stream. For example, if the user selects an “English - low dialog enhancement” audio track, but experiences higher ambient noise than expected, the selected audio stream can not be able to provide full dialog intelligibility to the end user, it would likely be preferable to have “English - high dialog enhancement.”
[0005] Furthermore, existing streaming services require each user in a room or listening environment to consume the exact same audio mix.
[0006] Accordingly, it would be desirable to have a system and method that overcomes the shortcomings of current approaches and enhances the audio experience of users without having to switch streams and always going back to the original server or CDN of the content provider, where all predefined streams are available. BRIEF DESCRIPTION OF DRAWINGS
[0007] FIG. 1 is a top level block diagram of components on the content provider (CP) side of a system for providing personalized audio streaming and rendering that creates a maximum version of an accessible audio mix according to embodiments of the present disclosure.
[0008] FIG. 2A is a top level block diagram of components on the user / listener (or user / recipient) side of a system for providing personalized audio streaming and rendering for a smart television with a cross-fade audio mix controlled by a remote control or user device according to embodiments of the present disclosure.
[0009] FIG. 2B is a top level block diagram of components on the user / listener (or user / recipient) side of a system for providing personalized audio streaming and rendering for a single user device according to embodiments of the present disclosure.
[0010] FIG. 2C is a top level block diagram of components on the user / listener (or user / recipient) side of a system for providing personalized audio streaming and rendering for a smart television with multiple user devices each with a personalized cross-fade audio mix according to embodiments of the present disclosure.
[0011] FIG. 2D is a top level block diagram of components on the user / listener (or user / recipient) side of a system for providing personalized audio streaming and rendering for a smart television with multiple user devices each with a personalized cross-fade audio mix and content selection according to embodiments of the present disclosure.
[0012] FIG. 3A 、 FIG. 3B 、 FIG. 3C 、 FIG. 3D and FIG. 3E illustrate block diagrams of alternative embodiments of a dialog clarity engine (DCE) of FIG. 1 including a combination of dialog isolation (DI), dialog enhancement (DE), dynamic range compression (DRC), enhanced dialog insertion (EDI), and audio loudness normalization (ALN) in various signal processing locations, and corresponding effects on dynamic range according to embodiments of the present disclosure.
[0013] FIG. 4 illustrate two graphs of dynamic range adjustment for dialog enhancement (DE) with integrated loudness (IL) of dialog (Dx) audio the same as non-dialog (or music and effects or M&E or MNE) audio in the first graph and different from non-dialog audio (or music and effects M&E or MNE) in the second graph according to embodiments of the present disclosure.
[0014] FIG. 5Ais a diagram for providing interlaced Cin-Mix and Max-Acc-Mix content audio streams of accessible audio according to embodiments of the present disclosure.
[0015] FIG. 5B is a diagram for providing interlaced Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix content audio streams of accessible audio according to embodiments of the present disclosure.
[0016] FIG. 5C is a diagram for providing interlaced Cin-Mix and Max-Acc-Mix content audio streams of accessible audio with two languages for voice-over audio according to embodiments of the present disclosure.
[0017] FIG. 5D is a block diagram showing a non-encoded content server with separately saved non-encoded Cin-Mix and Max-Acc-Mix, and Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix audio content streams as inputs to an audio encoder, and an encoded content server with interlaced Cin-Mix / Max-Acc-Mix and Cin-Mix / Mid-Acc-Mix / Max-Acc-Mix audio content as outputs from the audio encoder according to embodiments of the present disclosure.
[0018] FIG. 5E is a table showing various encoding and decoding options for a given stream type and user desired configuration according to embodiments of the present disclosure.
[0019] FIG. 6 is a block diagram of various components of a system for providing personalized audio streaming and rendering via a network connection according to embodiments of the present disclosure.
[0020] FIG. 7 shows a block diagram of dialog enhancement (DE) logic of FIG. 3A-FIG. 3E according to embodiments of the present disclosure.
[0021] FIG. 8 shows a block diagram of dynamic range compression (DRC), enhanced dialog insertion (EDI), and audio loudness normalization (ALN) of FIG. 3A-FIG. 3E according to embodiments of the present disclosure.
[0022] FIG. 9 is a plot showing a variable gain function that can be used by the DRC of FIG. 3A-FIG. 3E to vary the DRC gain to provide range compression according to embodiments of the present disclosure.
[0023] FIG. 10A is a block diagram of a cross-fade renderer (X-Fade) for stereo (L / R) input audio, and a corresponding plot showing the X-Fade power hold gain curve, according to embodiments of the present disclosure.
[0024] FIG. 10B is a block diagram of a cross-fade renderer (X-Fade) for stereo (L / R) input audio, and a corresponding plot showing the X-Fade power hold gain curve for a position range of an X-Fade slider adjustment, according to embodiments of the present disclosure.
[0025] FIG. 10C is a block diagram of a cross-fade renderer (X-Fade) for stereo (L / R) input audio with three inputs, and a corresponding plot showing the X-Fade power hold gain curve for a position range of an X-Fade slider adjustment, according to embodiments of the present disclosure.
[0026] FIG. 11 is a block diagram of components of a personalized equalizer (Pers. EQ) logic, according to embodiments of the present disclosure.
[0027] FIG. 12A is a screen shot of a graphical user interface (UI) for DCE gain adjustment and Max-Acc-Mix monitoring used by a CP user / administrator of a content provider, having the ability to set / adjust DCE gain and noise floor and listen to results across a cross-fade range, according to embodiments of the present disclosure.
[0028] FIG. 12B is a diagram of a smart TV remote control for a cross-fade (X-Fade) renderer for manual adjustment FIG. 2A , FIG. 2C and FIG. 2D , according to embodiments of the present disclosure.
[0029] FIG. 12C is a screen shot of a graphical user interface (UI) of an Acc app used by a user, having the ability to set / adjust cross-fade and enable certain features, according to embodiments of the present disclosure.
[0030] FIG. 12D is a screen shot of a graphical user interface (UI) for listening tests to determine parameters of a personal equalizer, according to embodiments of the present disclosure.
[0031] FIG. 13A shows a flowchart of portal / DCE logic FIG. 1 , according to embodiments of the present disclosure.
[0032] FIG. 13BFIG. 1 illustrates a flow diagram of the portal UI logic of FIG. 1 FIG. 2 illustrates a flow diagram of the hub accessibility (Acc) application logic of
[0033] FIG. 13C FIG. 3 illustrates a flow diagram of the Acc application UI logic of FIG. 2A , FIG. 2C , FIG. 2D FIG. 4 illustrates a block diagram of various embodiments of different types of audio input to be streamed and adjustably rendered by the crossfade renderer of
[0034] FIG. 13D FIG. 5 illustrates a block diagram of various embodiments of different types of audio input to be streamed and adjustably rendered by the crossfade renderer of FIG. 2A , FIG. 2C , FIG. 2D FIG. 6 illustrates a block diagram of various embodiments of different types of audio input to be streamed and adjustably rendered by the crossfade renderer of
[0035] FIG. 14A , FIG. 14B , FIG. 14C , FIG. 14D and FIG. 14E FIG. 7 illustrates a block diagram of various embodiments of different types of audio input to be streamed and adjustably rendered by the crossfade renderer of FIG. 10A , FIG. 10B or FIG. 10C for video on demand (VOD) and live sports / events (live) applications.
[0036] FIG. 15A and FIG. 15B FIG. 8 illustrates a block diagram of various embodiments of different types of audio input to be streamed and adjustably rendered by the crossfade renderer of FIG. 10A , FIG. 10B or FIG. 10C for video on demand (VOD) and live sports / events (live) applications. DETAILED DESCRIPTION
[0037] As discussed in more detail below, in some embodiments, the present disclosure relates to systems and methods for providing personalized audio streaming and rendering, including crossfade content delivery, according to the rules, requirements, and needs of the content owner.
[0038] The present disclosure provides a personalized audio streaming experience to the end user by providing an accessible audio mix or soundtrack created from the original studio movie mix, which enables the user to understand the story-telling dialogue (Dx) across a wide range of user devices and environmental background noise levels. The accessible mix is created by providing enhanced audio with dialogue intelligibility and / or dynamic range compression that is adaptive to the environmental background noise, as well as simultaneous multi-dialogue audio streams (for languages, sports commentary, etc.) across single or multiple devices.
[0039] The system and method of personalized audio streaming and rendering of the present disclosure provides various applications and configurations, including: for video on demand (VOD) applications, it provides accessible audio streaming, including enhanced dialogue, audio description (or narration), and voice over (VO), and for live sports / event audio streaming applications, it provides options, such as home and away sports team commentary and sports team commentary focus or stadium sound focus, as discussed herein.
[0040] The present disclosure provides rendering aspects of the decoded stream based on what the end user wants to hear to provide an optimized personal listening experience or personalized audio experience. It also provides audio with multiple streams to enhance the audio experience (with accessibility, localization, and synchronization), and can perform single or multiple audio software decoding within the streaming application itself. As described in the present disclosure, no existing streaming service is capable of playing and rendering at least two different audio mixes simultaneously.
[0041] Current commercial solutions for providing enhanced listening experiences have been standardized (as shown in the standard ATSC-3.0, which supports both Dolby AC-4 and MPEG-H codec audio decoders) and are designed to provide personalized listening and immersive audio experiences with enhanced dialogue intelligibility. However, these existing commercial solutions rely on audio production formats that have not been accepted by the entertainment industry, and require multi-dimensional audio renderers, which are complex and difficult to implement for smart TV manufacturers (OEM) partners and for streaming service.
[0042] One of the advantages of the system and method of the present disclosure is the ability to provide single or multiple device synchronized audiovisual (AV) experiences that are also accessible, personalized, and localized. The solution is “accessible” because native and adaptive dialogue enhancement solutions are provided by the present disclosure. In addition, the solution is “localized” because of the nature of multi-dialogue audio support with alternative language rendering of the native / enhanced (or accessible) dialogue audio tracks, and support for simultaneous dominant and pipe (non-dominant) dialogue audio tracks. This applies to alternative language audio tracks, narration (audio description or AD) audio tracks, voice over (VO), live sports commentary (home / away team commentary), and so on. The solution is “synchronized” because there is only one streaming application that performs audio decoding, which allows multiple users watching different streams (e.g., in different languages) of the same program to be synchronized. Traditional solutions for co-viewing using multiple streaming applications on different devices do not provide perfect synchronization.
[0043] While current AC-4 and MPEG-H compatible products provide some dialog enhancement features, they do not use the ambient noise of the listening environment to automatically adjust the ratio (or mix ratio) between the (dialog and non-dialog) original audio movie mix and the dialog enhancement (or accessible) audio track(s), which is done by the crossfade renderer of the present disclosure. They also do not provide the ability for the user to manually adjust the ratio between the (dialog and non-dialog) original audio movie mix and the dialog enhancement (or accessible) audio track(s) in real-time, which also provides the ability to fine tune without changing (or switching) streams. In addition, conventional products do not support the ability to reproduce multiple and independent audio tracks from a single streaming application, such as a movie / accessible mix including ED (enhanced dialog), AD (audio description), and VO (voice-over), and a sports / personalized mix including home / away commentary and stadium mix, as described herein.
[0044] Reference FIG. 1 FIG. 1 shows a top level block diagram of components on the content provider (CP) side of a system for providing personalized audio streaming and rendering that creates a maximum version of an accessible audio mix (Max-Acc-Mix or MaxAcc mix) in accordance with embodiments of the present disclosure. In particular, the system 10 includes a content provider computer 20 that receives instructions (e.g., content selection) from a content provider (CP) user / administrator 21 and has portal / DCE logic 28 running on the CP computer 20 that receives an unencoded (or unencoded) audio portion of the selected content (e.g., original movie theater near field audio mix (or Cin-Mix or CinMix)) from a content selection API (or content API) 30, as indicated by line 96, which can be obtained from an unencoded content server 36 on line 87. “Near field audio mix” or “near field mix” in a theater audio refers to an audio mix created for home entertainment (such as streaming or DVD) that uses a smaller near field speaker setup to simulate a home theater environment. Such a mix is typically done after the standard theater mix has been approved and finalized by the studio. Near field mixes are designed to be listened to in a more private setup, where the speakers are closer to the listener, unlike the large reverberant space of a movie theater. The present disclosure will also work with original movie mixes that are standard theater mixes designed for movie theaters and the like. The content provider computer 20 can be a smartphone, computer, laptop, tablet, or other computer-based device.
[0045] The Portal / DCE logic 28 provides a Max-Acc-Mix audio soundtrack on line 92 based on the original Cin-Mix content selected by the CP user / administrator 21 and audio signal processing performed by the dialog intelligibility engine 62 having DCE parameters or DCE parameters (e.g., DCE gain, noise floor, etc.) set by default from previously stored values or set (or adjusted) by the CP user / administrator 21 (discussed below). The Portal / DCE logic 28 provides the Max-Acc-Mix to a known audio encoder 26 (such as those discussed below) on line 92, such as the Dolby® Audio Encoders 26 utilizing Dolby® Audio Encoding Technology described herein. FIG. 5E The Portal / DCE logic 28 provides a Max-Acc-Mix audio soundtrack on line 92 based on the original Cin-Mix content selected by the CP user / administrator 21 and audio signal processing performed by the dialog intelligibility engine 62 having DCE parameters or DCE parameters (e.g., DCE gain, noise floor, etc.) set by default from previously stored values or set (or adjusted) by the CP user / administrator 21 (discussed below). The Portal / DCE logic 28 provides the Max-Acc-Mix to a known audio encoder 26 (such as those discussed below) on line 92, such as the Dolby® Audio Encoders 26 utilizing Dolby® Audio Encoding Technology described herein. FIG. 5A 、 FIG. 5B 、 FIG. 5C The Portal / DCE logic 28 provides a Max-Acc-Mix audio soundtrack on line 92 based on the original Cin-Mix content selected by the CP user / administrator 21 and audio signal processing performed by the dialog intelligibility engine 62 having DCE parameters or DCE parameters (e.g., DCE gain, noise floor, etc.) set by default from previously stored values or set (or adjusted) by the CP user / administrator 21 (discussed below). The Portal / DCE logic 28 provides the Max-Acc-Mix to a known audio encoder 26 (such as those discussed below) on line 92, such as the Dolby® Audio Encoders 26 utilizing Dolby® Audio Encoding Technology described herein.
[0046] The Portal / DCE logic 28 can include components such as the Portal UI logic 60 and the dialog intelligibility engine 62 that work together to provide the Max-Acc-Mix and the optional Mid-Acc-Mix.
[0047] In particular, the Portal UI logic 60 provides a Portal user interface (or Portal UI) for the CP user / administrator 21 that can have a content UI portion provided on line 82 to the user / administrator 21 (shown by line 74) via the display 22 that can display a list of available audio content (or audio files) (see FIG. 12A) to create an accessible audio mix that can be obtained from the content API on line 96. The portal UI logic 60 also receives content selections from the CP user / administrator 21 on line 78 via user interaction with the portal UI display 22 (or other input devices connected to the CP computer 20, such as a keyboard, mouse, or other devices). The content selections on line 78 can also be provided to the DCE 62 to allow the DCE to select content as needed, which can be provided directly or from the portal UI logic 60. The UI for content selection can be shown on the display 22 as a "content audio file" as shown in FIG. 1, as indicated by line 74, which is discussed more below. FIG. 12A
[0048] The portal UI logic 60 also requests and receives the raw Cin Mix audio from the non-encoded content server 36 via the content API 30 on line 96. In addition, the portal UI logic 60 provides DCE parameters, such as DCE gain / model type and noise floor (NF), to the dialog clarity engine 62 on line 90, which uses the DCE parameters to perform audio signal processing on the selected Cin-Mix audio tracks, which is discussed more below. The portal UI logic 60 can also communicate with the DCS parameter server 38 on line 85 to retrieve or save DCE parameters. The portal UI logic 60 can receive the DCE parameters from the user via the portal UI (see FIG. 1), for example, or from the DCE parameter server associated with the selected content, or from metadata embedded in the digital audio content if previously determined for the content. The portal UI logic 60 provides the maximum accessible audio mix (Max-Acc-Mix) on line 92, which is also provided (or fed back) to the portal UI logic 60 to allow the CP user / administrator 21 to adjust the accessible mix (Max-Acc-Mix) through the portal UI as desired, as described herein. FIG. 12A
[0049] The dialog clarity engine 62 also receives the DCE parameters (or DCE parameters) from the portal UI logic 60 on line 90, such as NF, DCE gain, and DCE model type, and receives the Max-Acc-Mix = OK signal from the portal UI logic 60 on line 88, and also receives the Cin-Mix audio from the encoded content server 36 via the content API 30 on line 96.
[0050] In some embodiments, the DCE 62 or the portal UI logic 60 can retrieve the default or initial values of the DCE parameters (e.g., gain, model type, NF, and any other required parameters) associated with the selected content from the DCE parameter server 38 on line 83. In addition, the DCE 62 or the portal UI logic 60 saves the values of the DCE parameters provided by the CP user / administrator 21 for the selected content to the DCE parameter server 38.
[0051] The output of the DCE logic 62 is the Max-Acc-Mix audio track provided on line 92 to the audio encoder 26, which is discussed below. The DCE 62 can also store the Max-Acc-Mix audio track on the non-encoded content server 36. Another input to the encoder 26 is the original Cin-Mix audio track on line 96. In some embodiments, the DCE can also provide a second output Mid-Acc-Mix on line 93, which is an audio track tapped from a predetermined mid (or mid) position within the DCE, as discussed below. The output of the encoder 26 is an encoded and interleaved audio stream with both Cin-Mix and Max-Acc-Mix in a single bitstream, such as shown in FIG. 5A 、 FIG. 5B 、 FIG. 5C which is stored on the encoded content server 40 via the content API 30 (shown by lines 94, 95), as discussed below.
[0052] Referring to FIG. 2A 、 FIG. 2B 、 FIG. 2C 、 FIG. 2D , respectively, show several top-level block diagrams 200A, 200B, 200C, 200D of different configurations or applications of the user / listener (or user / recipient) side of the systems and methods of the present disclosure that access the Cin-Mix / Max-Acc-Mix encoded and interleaved content from the encoded content server 40 and decode the digital stream back into separate audio tracks Cin-Mix and Max-Acc-Mix (or Cin-Mix and Max-Pers-Mix) for use by the cross-fade renderer (X-Fade) 100 (discussed more below) in order to provide one or more user / listener’s personalized accessible cross-fade mix, as described herein. The user / listener side of the systems and methods of the present disclosure will be described in detail below.
[0053] Referring back to FIG. 1 and to FIG. 3A-FIG. 3E , show the FIG. 1FIG. 1 shows a block diagram of an alternative embodiment of a dialog clarity engine (DCE) and corresponding impact on audio dynamic range or audio loudness range. In particular, the DCE includes several components or logic, including dialog isolation (DI) 302A, dialog enhancement (DE) 304A, dynamic range compression (DRC) 306A, enhanced dialog insertion (EDI) 308A, and audio loudness normalization (ALN) 310A, which can be configured in several different ways to provide a maximum accessible mix audio (Max-Acc-Mix) on line 92. The DCE model configuration depends on the selected content from the content provider CP user / administrator 21 and the desired accessible mix audio result. In addition, the DCE can also receive DCE parameters from the DCE parameter server 38 on line 83 (discussed below).
[0054] Reference is made to FIG. 3A FIG. 2 shows a first DCE model 300A (DCE model 1) configuration with a component order from left to right of DI-DE-DRC-EDI-ALN 302A, 304A, 306A, 308A, 310A, and corresponding loudness range (LRA) bar graphs 302B, 304B, 306B, 308B, 310B, respectively, showing the impact of each component on an input cinema mix (Cin-Mix) LRA bar graph 301, which is discussed more below.
[0055] The LRA bar graphs 301B, 302B, 304B, 306B, 308B, 310B are plotted on a graph 311 with a vertical axis 312 of loudness measured in LKFS (Loudness K-Weighted Full Scale) and a horizontal time axis (or chronology). The vertical axis 312 also shows a noise floor (NF) range 316, and a noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix.
[0056] The LRA bar graphs 301B, 302B, 304B, 306B, 308B, 310B have an audio loudness range (LRA) for the entire program (or program loudness PL), and a loudness range (LRA) for the dialog (Dx) portion of the audio (or dialog LRA or Dx LRA) 301C, 302C, 304C, 306C, 308C, 310C and a loudness range (LRA) for the non-dialog (or music and effects or M&E or MNE) portion of the audio (or non-dialog LRA or M&E LRA) 301D, 302D, 304D, 306D, 308D, 310D.
[0057] The LRA bar graphs 301B, 302B, 304B, 306B, 308B, 310B can also include the average loudness or integrated loudness (IL) (or dialog LRA or Dx LRA) 301E, 302E, 304E, 306E, 308E, 310E of the dialog (Dx) portion of the audio and the average loudness or integrated loudness (IL) (or non-dialog LRA or M&E LRA) 301F, 302F, 304F, 306F, 308F, 310F of the non-dialog (or music and effects (M&E)) portion of the audio, and the average loudness or integrated loudness (IL) (or program loudness (PL)) 310G of the entire program.
[0058] In FIG. 3A the DCE model 1, the input Cin-Mix is processed by dialog isolation (DI) logic 302A, which uses known tools to separate the audio dialog (Dx) from the non-dialog (or M&E) portion of the audio and provide two audio tracks: a Dx track on line 332A and an M&E track on line 332B. Tools that can be used for the DI logic 302A include AI machine learning (ML) models (open source and commercial) for extracting dialog from video streams, such as (Deezer Research), (Meta™), (AudioShake), etc. The ML models for DI will be trained using existing datasets for extracting dialog from videos across a wide range of content. Any other method or tool or model for isolating (or extracting or separating) Dx from M&E can be used if desired, as long as it provides the functionality and performance described herein.
[0059] In this example, the loudness range LRA of the two tracks (which in this case is dominated by the M&E track) LRA = +18 dB (-16 - (-34) = +18), the program loudness (PL) = -24 LKFS, and the dialog-to-program (or to residual or non-dialog or M&E) loudness (DPL or DRL) = +0 dB (since the IL of both Dx and M&E is -24 LKFS), as shown by the dashed box 320. The Dx and M&E audio tracks on lines 332A, 332B are provided to dialog enhancement (DE) logic 304A, which amplifies the Dx track and attenuates the M&E track while keeping the average or integrated loudness the same.
[0060] Referring to FIG. 7 , a FIG. 3A-FIG. 3Ea block diagram of an embodiment of dialog enhancement (DE) logic 304. In particular, the DE logic 304 receives the Dx audio track on line 332A and amplifies the Dx audio by multiplying the Dx by a gain on line 711 at a gain multiplication block 712 having a value greater than 1.0 (which is equal to the resulting gain indicated herein by -LKFS) to provide an amplified dialog (or amplified Dx) signal or track on line 334A. The DE logic 304 also receives the M&E audio track on line 332B and attenuates the M&E audio by multiplying the M&E by a gain on line 721 at a gain multiplication block 722 having a value less than 1.0 (which is equal to the resulting attenuation indicated herein by -LKFS) to provide an attenuated M&E (or non-dialog or residual) signal or track on line 334B.
[0061] The DE gains Gdx and Gmne provide the desired enhancement of dialog (Dx) over M&E while keeping the average or integrated loudness (IL) substantially constant. In some embodiments, this is accomplished by equalizing the amount of Dx amplification and M&E attenuation, as indicated by the Dx gain factor and the M&E gain factor on lines 702, 716, respectively, which in the example shown both have a value of 0.5, which provides an even split between the Dx amplification gain Gdx and the M&E attenuation gain Gmne. In particular, the Dx gain factor (e.g., 0.5) on line 702 is multiplied by the adjusted Dx amplification gain Gdx on line 709 at gain multiplication block 710 to provide the actual Dx gain on line 711, which is to be multiplied by the Dx on line 332A at gain block 712 to provide the amplified dialog on line 334A. Similarly, the M&E gain factor (e.g., 0.5) on line 716 is multiplied by the adjusted M&E attenuation gain Gmne on line 719 at gain block 720 to provide the actual M&E gain on line 721, which is to be multiplied by the M&E on line 332B at gain block 722 to provide the amplified M&E on line 334B.
[0062] In some embodiments, where the IL of the dialogue (Dx) is different from the IL of the M&E, e.g., the Dx IL is more responsive (less negative) than the M&E IL, there can be adjustments made to the gains of Dx and M&E as shown by summation blocks 708, 718. In particular, a DPL adjustment signal is provided on line 706 to summation blocks 708, 718 to equally adjust the gain values of both Dx and M&E. In that case, if the Dx IL is greater than the M&E IL, summation block 708 will decrease the Dx amplification gain Gdx on line 704 to provide an adjusted Dx gain on line 709, and summation block 718 will adjust the M&E attenuation gain Gmne on line 714 (i.e., make it less negative) to provide an adjusted M&E gain on line 719 that is equally adjusted as the Gdx adjustment. In some embodiments, the amount of DPL adjustment on line 706 is determined by DE logic 304 by calculating the difference between the Dx IL on line 732 and the M&E IL on line 734 as shown by summation block 730 (i.e., Dx IL - M&E IL). Further, in some embodiments, DCE logic 304A can calculate the IL of dialogue Dx and the IL of M&E (non-dialogue) from the input audio tracks Dx and M&E on lines 332A, 332B by using known methods or tools to perform a calculate Dx integrated loudness (IL) logic 736 and a calculate M&E IL logic 738, respectively, and by providing the Dx IL on line 732 to the positive (or add) input of summation block 730 and the M&E IL on line 734 to the negative (or subtract) input of summation block 730. The IL can be determined by known tools such as Insight2 or RX by iZotope, Inc. or open source software IL tools such as the loudness software described or by Youlean (see: https: / / youlean.co). Other techniques for determining IL can be used if desired. In some embodiments, the Dx IL and M&E IL on lines 732, 734 can be determined separately from DCE logic 304 and provided as inputs to DCE logic 304A. In that case, the calculate Dx IL logic 736 and the calculate M&E IL logic 738 will not be used. https: / / github.com / csteinmetz1 / pyloudnorm
[0063] Further, in some embodiments, instead of a single gain adjustment DPL adjustment, there can be separate adjustments (not shown) to the Dx and M&E gains based on the desired resulting audio output. Further, the Dx gain factor Gdx, the MNE gain factor Gmne can collectively be referred to as DE gains 342.
[0064] Reference is made to FIG. 4 and FIG. 7 , showing a pair of LRA bar graphs 402, 404 showing Dx IL offsets similar to those discussed above. Graph 402 (on the left) shows Dx and M&E with equal IL values (or DPL = 0 dB). In that case, Gdx and Gmne will be the same value. Graph 404 shows Dx and M&E where, given a non-zero DPL (e.g., 1.0 dB), the Dx IL is higher than the M&E IL. For example, if Dx IL = 1.0 dB, and M&E IL = 0 dB, then DPL will be 0.5 dB (midway between 0 (M&E IL) and 1 (Dx IL)). In that case, the DPL adjustment will reduce the dialogue gain Gdx( FIG. 7 ) to 3 dB (adjust Gdx), and change the M&E attenuation Gmne( FIG. 7 ) to -3 dB (adjust Gmne). After multiplication by the Dx gain factor and the M&E gain factor (each having a value of 0.5), this creates 1.5 dB of Dx amplification and -1.5 dB of M&E attenuation on lines 711, 721. The resulting value for Dx will be 2.5 dB, and the resulting value for M&E will be -1.5 dB, providing the desired separation of 4 LKFS (2.5 - (-1.5)). It also provides a resulting DPL of 0.5 (2.5 - 4 / 2), which in this example maintains the DPL value of 0.5. Thus, the DE logic 304A provides more than the desired dialogue (Dx) enhancement over non-dialogue (M&E) (e.g., 4 LKFS), while maintaining the dialogue-to-program (or to residual or M&E) loudness DPL unchanged (e.g., DPL = 0.5).
[0065] Referring to FIG. 3A , FIG. 8 and FIG. 9 , the dynamic range compression (DRC) logic 306A receives the attenuated M&E audio track on line 334B and the DCE gain profile on line 344, and provides a compressed M&E audio track on line 336. In particular, referring to FIG. 8 and FIG. 9 , the DRC logic 306A compresses the input M&E audio track on line 334B by multiplying the input M&E audio track by the value of the DRC gain curve (or profile) on line 344 through multiplication block 802( FIG. 8 ), such as shown in FIG. 9 . Referring to FIG. 9In particular, a known DRC gain curve 902 is shown plotted on a graph 900 of output level (vertical axis) versus input level (horizontal axis). The DRC gain curve 902 can have portions or ranges that provide desired M&E loudness range ALR compression, including: an enhancement range 908, a null band or unity gain region 910, an early cut range 912, and a cut range 914. Other types of DRC gain curves and ranges can be used if desired. The enhancement range 908 increases (or amplifies) low level (or quieter) audio M&E sounds with a gain greater than 1.0, the unity gain region preserves audio M&E sounds with a gain of 1.0, and the early cut range 912 provides a slight attenuation of audio M&E sounds, and the cut range 914 provides a large attenuation of audio M&E sounds with a gain less than 1.0. The early cut range 912 and the cut range 914 are set to provide a desired amount of M&E compression. In some embodiments, the amount of compression (and low-end enhancement) provided by the DRC logic 306A can be adjusted by the CP user / administrator by adjusting the DRC gain curve 902, which can be done via the CP user interface or portal UI, which is discussed more below.
[0066] Referring to FIG. 3A and FIG. 8 , the enhanced dialog insertion (EDI) logic 308A receives the attenuated and compressed M&E audio track on line 336 and the amplified dialog Dx on line 334A, and provides a recombined (or remix) composite audio track on line 338. In particular, referring to FIG. 8 , the EDI logic 308A adds the input amplified dialog Dx on line 334A to the compressed M&E on line 336 at a sum 804 to provide a recombined audio track on line 338, where the enhanced (or amplified) dialog Dx is inserted into the compressed M&E audio track.
[0067] Referring to FIG. 3A and FIG. 8 , the audio loudness normalization (ALN) logic 310A receives the recombined audio track on line 338 and an ALN gain on line 350, and provides a maximum accessibility mix (Max-Acc-Mix) audio track on line 92. In particular, referring to FIG. 8 , the ALN logic 310A multiplies the input audio signal on line 338 by the ALN gain on line 350 at a multiply gain block 806 to provide the Max-Acc-Mix on line 808 to maximum limiting logic 810, which ensures that the ALN gain does not cause the loudness range (LRA) of the Max-Acc-Mix to exceed an incorrect limit.
[0068] In particular, there are two equations that define limits on the values of the DCE gains (e.g., DE gain, DRC gain, ALN gain) of the dialog clarity engine 62. In particular, Equation 1 below defines a maximum value for the accessible mix Max-Acc-Mix: Max(TruePeak(MaxAccMix)) <-2.0 dB = MaxAudio Equation 1 where Max(TruePeak(MaxAccMix)) is the maximum value of the TruePeak value (in dB) of the Max-Acc-Mix track.
[0069] Further, if the noise floor (NF) is known, and Equation 1 is satisfied, then Equation 2 below must also be satisfied to provide the best user listening experience: Min(TruePeak(DX)) > NF Equation 2 where Min(TruePeak(DX)) is the minimum value of the TruePeak of the dialog Dx portion of the mix, and NF is the noise floor level. This can require adjustment of the DCE based on the results of Equation 2.
[0070] The maximum audio value for avoiding audio clipping is 0 dB True Peak. Some content providers require the maximum audio value to be -2 dB in order to provide some additional buffer before clipping occurs. Further, typically the maximum audio value is set to the dB True Peak value or dBTP, which is the maximum peak audio level of the content as known. The maximum limiting logic 810 within the ALN logic 310A ( FIG. 8 ) will not allow Max-Acc-Mix to exceed the requirements of Equation 1.
[0071] In particular, the maximum limiting logic 810 receives Max-Acc-Mix on line 808, and using the current noise floor NF setting from the CP user / administrator 21 or from the DCE parameter server 38 and the maximum audio from the DCE parameter server 38 on line 812, determines whether the value of Max-Acc-Mix satisfies Equation 1. If so, Max-Acc-Mix is passed unmodified to the output on line 92 as Max-Acc-Mix. Further, if NF is known and Equation 1 is satisfied, then the DCE parameters of the DCE can be adjusted as needed to make Equation 2 true.
[0072] If the value of Max-Acc-Mix does not satisfy Equation 1, then Max-Acc-Mix will be reduced to a value that satisfies Equation 1, and the modified version of Max-Acc-Mix is passed to the output on line 92 as Max-Acc-Mix.
[0073] Thus, if the CP user / administrator 21 attempts to set the ALN gain or DE gain or noise floor (NF) (or any other gain or factor in the DCE) such that it results in Max-Acc-Mix exceeding the requirement of Equation 1, then for a given content provider, the maximum limiting logic 810 will clamp the ALN gain to a value that allows Max-Acc-Mix to satisfy Equation 1. The maximum audio value can be an input to the maximum limiting logic 810 from the DCE parameter server 38 on line 812, and can be preset by the content provider to a value that satisfies the content provider's audio requirements, such as -2dbTP. Other values can be used if desired.
[0074] In some embodiments, the CP user / administrator 21 can set the ALN gain (discussed below) from the UI. In some embodiments, the system can automatically calculate the ALN gain using Equation 1 to maximize the LRA of Max-Acc-Mix based on the noise floor or DE gain that has been set by the CP user / administrator 21.
[0075] Referring to FIG. 3B , a second DCE model 300B (DCE Model 2) configuration is shown having a component order 306A, 302A, 304A, 308A, 310A from left to right of DRC-DI-DE-EDI-ALN, and corresponding loudness range (LRA) bar graphs 306B, 302B, 304B, 308B, 310B are shown, respectively, that illustrate the effect of each component on the input movie mix (Cin-Mix) LRA bar graph 301B, which is discussed more below. As discussed above, the LRA bar graphs 306B, 302B, 304B, 308B, 310B are plotted on a graph 311B having a vertical axis 312 of loudness measured in LKFS and a horizontal time axis (or timeline). The vertical axis 312 also shows a noise floor (NF) range 316, and a noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix, similar to the case of FIG. 3A .
[0076] FIG. 3A-FIG. 3E The DCE 62 processing of Applied to audio content Audio blocks of shorter size / durations , and can be performed on the blocks simultaneously (in parallel) or sequentially (one after the other). Furthermore, the DCE parameters can be different for each block, if necessary or desired.
[0077] In FIG. 3BIn DCE Model 2, the input Cin-Mix is first processed by dynamic range compression (DRC) logic 306A, which receives the Cin-Mix audio track on line 92 and the DCE gain profile on line 344 and provides a compressed Cin-Mix audio track on line 336. The DRC 306A functions similarly to the DRC logic described above. FIG. 3A In particular, refer to FIG. 8 and FIG. 9 , DRC logic 306A passes multiplication block 802 ( FIG. 8 ) compresses the input Cin-Mix audio track on line 96 by multiplying it by the value of the DRC gain curve (or profile) on line 344, such as FIG. 9 As shown in . FIG. 9 In particular, a known DRC gain curve 902 is shown plotted on a graph 900 of output level (vertical axis) versus input level (horizontal axis). FIG. 3A The model 1300A in the example differs in that it compresses the entire mix, not just the M&E.
[0078] The compressed Cin-Mix on line 336 is provided to the dialogue isolation (DI) logic 302A which separates the audio dialogue (Dx) from the non-dialogue (or M&E) portion of the audio (as described above using FIG. 3A discussed) and provides two audio tracks: a Dx track on line 332A and an M&E track on line 332B.
[0079] The Dx and M&E audio tracks on lines 332A, 332B are provided to dialogue enhancement (DE) logic 304A which receives the DE gain on line 342 and amplifies the Dx track and attenuates the M&E track while keeping the average or integrated loudness the same, similar to the above example using FIG. 3A discussed.
[0080] The Dx and M&E audio tracks on lines 334A, 334B from the dialogue enhancement (DE) logic 304A are provided to the enhanced dialogue insertion (EDI) logic 308A which receives the attenuated and compressed M&E audio track on line 334B and the amplified dialogue Dx on line 334A and provides the recombined (or remixed) composite audio track on line 338 to the audio loudness normalization (ALN) logic 310A, similar to the method described above utilizing FIG. 3A discussed.
[0081] Audio loudness normalization (ALN) logic 310A receives the recombined audio soundtrack on line 338 and the ALN gain on line 350 and provides a Max-Accessibility Mix (Max-Acc-Mix) soundtrack on line 92, similar to that discussed above with respect to FIG. 3A .
[0082] Referring FIG. 3C to FIG. 3C, a third DCE model 300C (DCE Model 3) configuration is shown having a component order 302A, 304A, 308A, 306A, 310A from left to right of DI-DE-EDI-DRC-ALN, and corresponding loudness range (LRA) bar graphs 302B, 304B, 308B, 306B, 310B are shown, respectively, that illustrate the effect of each component on an input cinema mix (Cin-Mix) LRA bar graph 301B, which is discussed more below. As discussed above, the LRA bar graphs 302B, 304B, 308B, 306B, 310B are plotted on a graph 311C having a vertical axis 312 of loudness measured in LKFS and a horizontal time axis (or chronology). The vertical axis 312 also shows a noise floor (NF) range 316, and a noise floor setting region 318 used by the DCE to determine a desired Max-Acc-Mix, similar to that discussed above with respect to FIG. 3A .
[0083] In the DCE Model 3 of FIG. 3C , the input Cin-Mix is first processed by dialogue isolation (DI) logic 302A that separates the audio dialogue (Dx) from the non-dialogue (or M&E) portion of the audio (as discussed above with respect to FIG. 3A ) and provides two audio soundtracks: a Dx soundtrack on line 332A and an M&E soundtrack on line 332B, similar to that discussed above with respect to FIG. 3A .
[0084] The Dx and M&E audio soundtracks on lines 332A, 332B are provided to dialogue enhancement (DE) logic 304A that receives a DE gain on line 342 and amplifies the Dx soundtrack and attenuates the M&E soundtrack while keeping the average or integrated loudness the same, similar to that discussed above with respect to FIG. 3A .
[0085] The Dx and M&E audio tracks on lines 334A, 334B from dialogue enhancement (DE) logic 304A are provided to enhanced dialogue insertion (EDI) logic 308A, which receives the attenuated and compressed M&E audio tracks on line 334B and the amplified dialogue Dx on line 334A and provides a recombined (or remix) composite audio track on line 338 to dynamic range compression (DRC) logic 306A, similar to that discussed above with respect to FIG. 3A .
[0086] Dynamic range compression (DRC) logic 306A receives the recombined (or remix) composite audio track on line 338 and DCE gain profile on line 344 and provides a compressed Cin-Mix audio track on line 336 to audio loudness normalization (ALN) logic 310A.
[0087] Audio loudness normalization (ALN) logic 310A receives the recombined audio track on line 338 and ALN gain on line 350 and provides a Max-Accessibility Mix (Max-Acc-Mix) track on line 92, similar to that discussed above with respect to FIG. 3A .
[0088] Referring to FIG. 3D , a fourth DCE model 300D (DCE Model 4) configuration is shown, which has a component order 302A, 306A, 304A, 308A, 310A from left to right of DI-DRC-DE-EDI-ALN, and corresponding loudness range (LRA) bar graphs 302B, 306B, 304B, 308B, 310B are shown, which illustrate the effect of each component on an input movie mix (Cin-Mix) LRA bar graph 301D, which is discussed more below. As discussed above, the LRA bar graphs 302B, 306B, 304B, 308B, 310B are plotted on a graph 311D, which has a vertical axis 312 of loudness measured in LKFS and a horizontal time axis (or timeline). The vertical axis 312 also shows a noise floor (NF) range 316, and a noise floor setting region 318 used by the DCE to determine a desired Max-Acc-Mix, similar to that discussed above with respect to FIG. 3A . In the DCE Model 4 of FIG. 3D , the configuration of components or logic is similar to that discussed above with respect to FIG. 3A , except that the positions of DE logic 304A and DRC logic 306A are swapped.
[0089] Referring to FIG. 3E, showing a fifth DCE model 300E (DCE Model 5) configuration, which has a component order 302A, 306A, 310A, 304A, 308A from left to right of DI-DRC-ALN-DE-EDI, and showing corresponding loudness range (LRA) bar graphs 302B, 306B, 310B, 304B, 308B, respectively, which show the effect of each component on the input cinematic mix (Cin-Mix) LRA bar graph 301E, which is discussed more below. As discussed above, the LRA bar graphs 302B, 306B, 310B, 304B, 308B are plotted on a graph 311E, which has a vertical axis 312 of loudness measured in LKFS and a horizontal time axis (or chronology). The vertical axis 312 also shows a noise floor (NF) range 316, and a noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix, similar to FIG. 3A the case of FIG. 3E In the DCE Model 5, the configuration of components or logic is similar to FIG. 3D the case of
[0090] Referring to FIG. 3A-FIG. 3E , in some embodiments, the DCE 62 can also provide a second output Mid-Acc-Mix on line 93, which is an audio track tapped from a predetermined location in the DCE 62. As shown in FIG. 3A-FIG. 3E , depending on the DCE model, the Mid-Acc-Mix is tapped after the dialogue enhancement logic (DE) 304A and before the dynamic range compression (DRC) logic 306A, or after the dynamic range compression (DRC) logic 306A and before the dialogue enhancement (DE) logic 304A, which allows the use of a 3-input X-Fade renderer as discussed herein and shown in FIG. 10C For some of the DCE models, such as DCE Model 1, Model 3, Model 4, and Model 5, the location of the tap requires recombining the audio tracks into a single track. In that case, there is additional EDI logic block 308A, shown as a dashed box, which performs the combining to create the Mid-Acc-Mix on line 102.
[0091] Referring back to FIG. 1 , the output of the DCE logic 62 is the Max-Acc-Mix track provided on line 92 to the audio encoder 26, which can also be stored on the non-encoded content server 36. The other input to the encoder 26 is the original Cin-Mix audio track on line 96. In some embodiments, the DCE can also provide a second output Mid-Acc-Mix on line 93, as discussed above withFIG. 1 As discussed, the second output Mid-Acc-Mix is an audio track tapped from a predetermined location in the DCE.
[0092] In general, the input to the audio encoder 26 is (one or more) audio source waveforms (e.g., PCM / WAV audio uncompressed format), where the resulting output is an audio bitstream (or bs), for example, a digitally compressed audio soundtrack encoded using MPEG-4 AAC (Advanced Audio Coding) or Dolby AC3 / EC3 developed by MPEG Audio. Other coding standards or codecs can be used for encoding if desired. The use of a bitstream (bs) allows for the transmission of complex surround sound formats such as Dolby Atmos or DTS:X. The audio encoder 26 uses known software running on hardware owned / provided by the streaming service (such as the CP computer 20). Certain AAC encoding and decoding techniques are available for licensing from the known AAC patent pool.
[0093] As discussed above, the output of encoder 26 is an encoded and interleaved audio stream, such as FIG. 5A 、 FIG. 5B 、 FIG. 5C The audio stream shown in , which is stored on the encoded content server 40 via the content API 30 shown by lines 94,95.
[0094] Reference again FIG. 1 and FIG. 12A The portal UI logic 60 also provides a UI to allow the CP user / administrator 21 to adjust the X-Fade via the display 22, and receives X-Fade adjustment selections from the CP user / administrator 21 via user interaction with the display 22 (or other input devices connected to the CP computer, such as a keyboard, mouse, or other devices). The UI for X-Fade control, monitoring, and selection on the display 22 can be as follows: FIG. 12A , which is discussed more below.
[0095] As discussed herein, the portal UI logic 60 may have X-Fade logic 100 (or cross-fade renderer logic) that provides a power-maintaining mix of the current values of Cin-Mix and Max-Acc-Mix. Specifically, the X-Fade logic 100 ( FIG. 1 ) provides an X-Fade slider 1210 ( FIG. 12A ) that allows the user to adjust the X-Fade slider that provides a cross-fade rendering or mix or mixing of the Cin-Mix and Max-Acc-Mix from the DCE logic 62 to adjust the X-Fade slider based on the X-Fade slider 1210 ( FIG. 12A) using a power-preserving gain curve to provide the resulting mix of the two input signals, the X-Fade slider 1210 providing an X-Fade adjustment or X-Fade Adj to the X-Fade logic 100 on line 80.
[0096] The power-preserving equation used by the X-Fade logic 100 is shown below as Equation 3: Sin 2 (θ) +Cos 2 (θ) = 1 Equation 3
[0097] where the value of θ (in degrees) ranges from 0 to 90 degrees or 0 to 180 degrees (depending on configuration) and corresponds to the position X of the X-Fade slider (0 degrees indicates slider lower limit X = 0 and 90 degrees indicates slider upper limit X = 1.0). The power-preserving nature of Equation 3 allows the system of the present disclosure to provide a change in loudness such that the stereo / space image is not affected, particularly for stereo, surround or immersive sound or audio applications. When using audio signal magnitudes to process audio signals, the sine and cosine functions can be used as coefficients (or multiplication attenuation gains of value < or = 1.0) to the crossfading input audio signals, respectively, to produce a power-preserving crossfading result.
[0098] Referring to FIG. 10AX-Fade logic 100 is illustrated by block diagram 1000 and corresponding graph 1010. X-Fade logic 100 has two inputs Cin-Mix and Max-Acc-Mix on lines 1001, 1005, which are each adjusted by gains g0 and g1, respectively, defined by cosine gain curve 1012 (or curve A) and sine gain curve 1014 (or curve B), respectively, and provides an output X-Fade mix audio signal on line 1009. In particular, X-Fade logic 100 receives a Cin-Mix signal on line 1001, which is fed to multiplication gain block 1002, which multiplies the Cin-Mix signal by gain g0, as shown in Equation 4 below. The result of gain block 1002 is an adjusted Cin-Mix value (Cin-Mix Adj) provided on line 1003 to summation block 1008. X-Fade logic 100 also receives a Max-Acc-Mix signal on line 1005, which is fed to multiplication gain block 1004, which multiplies the Max-Acc-Mix signal by gain g1, as shown in Equation 5 below. The result of gain block 1004 is an adjusted Max-Acc-Mix value (Max-Acc-Mix Adj) provided on line 1007 to summation block 1008. Summation block 1008 sums the values of Cin-Mix Adj and Max-Acc-Mix Adj on lines 1003, 1007, and provides the result of the summation as X-Fade MIX on line 1009. Further, graph 1010 illustrates cosine curve 1014 that determines the value of gain g0 and sine curve 1016 that determines the value of gain g1. Further, vertical line 1016 illustrates the value of X based on the position of the slider (X = 0 to 1.0), and where the value of X intersects curves 1012 and 1014 determines the values of gains g0 and g1. Equations 4 and 5 below determine the values of Cin-Mix Adj and Max-Acc-Mix Adj. Cin-Mix Adj = (Cin-Mix) * Cos(X*90) = (Cin-Mix) * g0 Equation 4 Max-Acc-Mix Adj = (Max-Acc-Mix) * Sin(X*90) = (Max- Acc-Mix) * g1 Equation 5 where the value of X ranges from 0 to 1.0 based on the percentage position of the X-Fade slider, for example, when the slider is at the low (minimum) end of the slider range X = 0, and when the slider is at the high (maximum) end of the slider range X = 1.0.
[0099] Referring toFIG. 10B X-Fade logic 100 for stereo audio content (with left / right channels indicated by L / R, respectively) is shown by block diagram 1020 and corresponding plot 1030 with sine function curve 1034 (curve B) and cosine function curve 1032 (curve A), similar to that shown in FIG. 10A Particularly, X-Fade logic 100 receives Cin-Mix stereo signals on lines 1021A, 1021B for left, right (L0 / R0) channels, respectively, which are fed to multiplication gain block 1022, where both Cin-MixL0and Cin-MixR0are multiplied by gain g0determined by sine function curve 1034 (B) such that when X-Fade position X is at the low end (X=0), gain g0= 1 (cos(0*90)=1) and Cin-MixL0and Cin-MixR0pass through multiplier (or gain) 1022 without any attenuation (gain=l), and when X-Fade position X is at the high end (X=l.0), gain g0= 0 (cos(l*90)=0) and Cin-MixL0= 0 and Cin-MixR0= 0 are completely attenuated (gain=0). Gain block 1022 provides adjusted Cin-Mix for left, right (L / R) channels on lines 1023A, 1023B, respectively.
[0100] Similarly, X-Fade logic 100 receives Max-Acc-Mix stereo signals on lines 1025A, 1025B for left, right (L1 / R1) channels, respectively, which are fed to multiplication gain block 1024, where each of Max-Acc-MixL1and Max-Acc-MixR1is multiplied by gain g1determined by sine function curve 1034 (B) such that when X-Fade position X is at the low end (X=0), gain g1= 0 (sin(0*90)=0) and Max-Acc-MixL1= 0 and Max-Acc-MixR1= 0, i.e., completely attenuated (gain=0), and when X-Fade position X is at the high end (X=l.0), gain g0= 1 (sin(l*90)=l) and Cin-MixL0and Cin-MixR0pass through multiplier (or gain) 1024 without any attenuation (gain=l). Gain block 1024 provides adjusted Max-Acc-Mix for left, right (L / R) channels on lines 1023A, 1023B, respectively.
[0101] Adjust Cin-Mix L0 on line 1023A and adjust Max-Cin-Mix L1 on line 1027B are provided to sum block 1028A and output as combined left channel adjusted signal or X-Fade mix L (left channel) on line 1009A. Similarly, adjust Cin-Mix R0 on line 1023B and adjust Max-Cin-Mix R1 on line 1027A are provided to sum block 1028B and the sum output is provided as combined right channel adjusted signal or X-Fade mix R (right channel) on line 1009B.
[0102] Reference is made to FIG. 10C the block diagram 1040 of a three-input cross-fade renderer (or X-Fade) and the corresponding plot 1050 with three gain curves 1051 (or curve A), 1052 (or curve B), 1054 (or curve C) illustrating the X-Fade logic 100 for input stereo audio content (with left / right channels indicated by L / R, respectively). In particular, in some embodiments, the DCE 62( FIG. 1 ) can provide an additional output Mid-Acc-Mix, illustrated by dashed line 93( FIG. 1 ). In that case, the X-Fade logic 100 has three inputs (or three pairs of inputs for stereo audio): Cin-Mix L0 / R0, Mid-Acc-Mix L1 / R1, and Max-Acc-Mix L2 / R2, which are adjusted or multiplied by three gains g0, g1, g2, respectively, with gain values defined by three gain curves 1051 (or curve A), 1052 (or curve B), 1054 (or curve C), respectively. In particular, the equations for the output left (L) channel and the output right (R) channel on lines 1064A, 1064B, respectively, are illustrated below: X-Fade (L) = L0g0+ L1g1+ L2g2 Equation 6 X-Fade (R) = R0g0+ R1g1+ R2g2 Equation 7 where the values of g0, g1, g2 are defined by three gain curves (sine or cosine curves) 1051, 1052, 1054, and g0, g1, g2 change or vary based on the X-Fade slider position X, as described below, where the gain curve 1051 is a cosine curve from X=0 to 0.5 (for g0 values, for the first half of the slider range between Cin-Mix and Mid-Acc-Mix), the gain curve 1052 (for gain g1 values) is a sine curve from X=0 to 1.0 (for g1 values, for the full range of the slider range from Cin-Mix to Mid-Acc-Mix to Max-Acc-Mix), and the gain curve 1054 (for gain g2 values) is a sine curve from X=0.5 to 1.0 (for g2 values, for the second half of the slider range between Mid-Acc-Mix to Max-Acc-Mix). Further, in some embodiments, the gain curve 1052 (for gain g1 values) can also be considered a sine curve from X=0 to 0.5 and a cosine curve from X=0.5 to 1.
[0103] In particular, the three-input X-Fade logic 100 receives Cin-Mix L0 / R0 (left and right channels, referred to herein as L0, R0, respectively), which is provided to a multiplication (or gain) block 1042, which also receives gain g0, and computes values of L0*g0 (or L0g0) and R0*g0 (or R0g0), which are the adjusted (or attenuated) values of the Cin-Mix L0 / R0 left and right channels. L0g0 is fed to a summing block 1060, which also receives a value of (L1g1+L2g2) from a summing block 1049, which is discussed below. The result of the summing 1060 provides the value of the X-Fade mix (L or left channel) from Equation 6 on line 1064A. Further, the value of R0g0 from the gain block 1042 is provided to a summing block 1048, which is discussed below.
[0104] In addition, the three-input X-Fade logic 100 receives Mid-Acc-Mix L1 / R1 (left and right channels, herein referred to as LI, R1, respectively), which is provided to gain multiplication block 1044, which also receives gain gl, which computes the values of LI*gl (or LIgl) and Rl*gl (or Rlgl), which are the adjustment (or attenuation) values for the Mid-Acc-Mix LI / R1 left and right channels. Rlgl is provided to sum block 1048 along with ROgO (from multiplication block 1042), and the result of the sum 1048 (Rlgo + Rlgl) is provided to sum block 1062, which adds the value ROgO + Rlgl (from sum block 1048) to R2g2 (from block 1046) to provide the value of the X-Fade mix (R or right channel) from equation 7 above on line 1064B. LIgl is provided to sum block 1049 along with L2g2 from gain block 1046, which is discussed more below, and the result of the sum 1049 (LIgl + L2g2) is provided to sum block 1060, as discussed above.
[0105] In addition, the three-input X-Fade logic 100 receives Max-Acc-Mix L2 / R2 (left and right channels, herein referred to as L2, R2, respectively), which is provided to gain multiplication block 1046, which also receives gain g2, which provides the values of L2*g2 (or L2g2) and R2*g2 (or R2g2), which are the adjustment (or attenuation) values for the Max-Acc-Mix L2 / R2 left and right channels. L2g2 is fed to sum block 1049 along with LIgl, and the result of the sum 1049 (LIgl + L2g2) is provided to sum block 1060 to provide the value of the X-Fade mix (L) from equation 6 on line 1064A.
[0106] With respect to the gain curves, the following equations 8, 9, and 10 can be used to determine the gain values for the gains go, gl, g2 for the left and right channels (for stereo applications), e.g., Cin-MixAdj (LOgO, ROgO), Mid-Acc-Mixadj (LIgl, Rlgl), Max-Acc-Mix adj (L2g2, R2g2). gO = Cos(X*180), for: (X = 0 to 0.5; if X > 0.5, set g2 = 0) Equation 8 gl = Sin(X*180), for: (X = 0 to 1.0) Equation 9 g2 = Sin(X*90) for: (X = 0.5 to 1.0; if X < 0.5, set g0 = 0) Equation 10 where the value of X ranges from 0 to 1.0 based on the position percentage of the X-Fade slider, e.g., as discussed herein, X = 0 when the slider is at the low (minimum) end of the slider range, and X = 1.0 when the slider is at the high (maximum) end of the slider range. Further, g0 is set to 0 in the second half of the slider range, and g2 is set to 0 in the first half of the slider range.
[0107] In particular, for the first half of the X-Fade slider range (X = 0 to 0.5), the X-Fade logic 100 performs cross-fade rendering or mixing of the Cin-Mix and Mid-Acc-Mix (e.g., from 0 to 90 degrees of the sine and cosine curves), providing values for gains g0 and g1, and the value for g2 is set to 0 as it is not relevant to the first half of the slider range. For the second half of the X-Fade slider range (X = 0.5 to 1.0), the X-Fade logic 100 performs cross-fade rendering or mixing of the Mid-Acc-Mix and Max-Acc-Mix (e.g., from 90 to 180 degrees of the sine and cosine curves), providing values for gains g1 and g2, and the value for g0 is set to 0 as it is not relevant to the first half of the slider range. The result is an X-Fade slider that ranges from X = 0 to 1.0 and provides progressive mixing of the Cin-Mix with the Mid-Acc-Mix and then the Mid-Acc-Mix with the Max-Acc-Mix. For each portion, the two input signals use a preserved gain curve based on the position of the X-Fade slider (or X-Fade adjustment or X-FadeAdj), as discussed herein.
[0108] Further, the plot 1050 shows a cosine curve 1051 that determines the value of g0, and a sine curve 1052 that determines the value of g1 and a sine curve 1054 that determines the value of g2. Further, the vertical line 1055 shows the value of X based on the slider position, and where it intersects the curves 1052 and 1054 determines the values of g1 and g2, respectively. In particular, the plot 1051 (cosine function) determines the value of g0, the plot 1052 (sine function) determines the value of g1 (from 0-90 and from 90-180), and the plot 1054 (cosine function) determines the value of g2 (from 90-180), as also shown by Equations 8, 9, and 10. Further, for the first half of the slider range (X = 0 to 0.5), the value of g2 can be set to 0, and for the second half of the slider range (X = 0.5 to 1.0), the value of g0 can be set to 0, as discussed herein.
[0109] The Portal UI logic 60 provides the X-Fade UI on line 84 and the content UI on line 82 to the display 22, which allows the CP user / administrator 21 to select content provided to the Portal UI logic 60 on line 78 and receive an X-Fade slider input (X-Fade Adj) from the CP user / administrator 21 on line 80, which indicates the current position of the X-Fade renderer. The output of the X-Fade logic 100 is the X-Fade mix audio signal or track provided to the speaker(s) 23 of the CP computer 20 on line 86, which allows the CP user / administrator 21 to listen to the resulting accessible mix (e.g., Max-Acc-Mix), as shown by line 76. The Portal UI logic 60 can also request and receive a content list (or index) on line 99, which lists the audio content or audio files available on the non-encoded content server 36 for viewing and selection by the CP user / administrator 21.
[0110] Referring to FIG. 2A , FIG. 2B , FIG. 2C and FIG. 2D , respectively, are several top-level block diagrams 200A, 200B, 200C, 200D showing different configurations or applications of the user / listener (or user / recipient) side of the system of the present disclosure. In particular, FIG. 2A is a block diagram 200A showing a smart TV as the smart playback device 12 interfacing with the user's main interface via the TV display 14 and speakers 11. In that case, the X-Fade feature and the personalized EQ feature are controlled by the user / listener 15 using the smart TV remote (or remote control) 220 or using a user device 12 such as a smart phone. In either case, the user 15 can control the X-Fade and various other Acc application features or parameters by viewing the smart TV display 14. Details of FIG. 2A are discussed more below. FIG. 2B is a block diagram 200B showing a single smart playback device (e.g., smart phone, tablet, laptop, etc.) being used by a single user, which can operate in two modes: a standalone mode (or single device mode) or in a hub receiver (or slave) mode, where the device is subordinate to a smart TV operating in hub mode. In standalone mode, the device 12 receives content and interacts with the user similar to the smart TV described above, except that there is no remote control. In hub receiver (or slave) mode, the device 12 receives and provides commands on line 231B and receives the C-Mix video on line 217, all from the smart TV operating in hub mode. The rest of the features are similar to FIG. 2AThe same as described in the
[0111] FIG. 2C is a block diagram 200C showing a smart TV 12 configured to work in hub mode, which allows multiple user devices 12, e.g., user device 1 through user device N. In this mode, each user / listener 15 can use their own device 12 to watch and listen to the content being played on the smart TV 12. In that case, each of the user devices 1-N receives the UI (including the X-Fade UI and the Pers. EQ UI) and is able to view and listen to selected content from the smart TV via lines 231A and 233A. This permits people in the same room to watch the same program on the smart TV to have a personalized audio experience. For example, each user / listener 1-N has their own X-Fade slider, which they can set to their own personal settings. This also applies to each user / listener’s Pers. EQ feature. FIG. 2D Similar to FIG. 2C which shows a block diagram 200D with a smart TV 12 configured to work in hub mode, except in this case there are multiple decoders 202 that allow multiple users / listeners to watch the same program as others in the room, but also can select the form of the video and language for the same program, e.g., language, subtitles, narration, etc. (for the UI, see FIG. 12C , discussed below). For example, in that case, user / listener 1 can be listening to a program in English, user / listener 2 can be listening to a program in French, user / listener 3 can be listening to Spanish subtitles, user / listener 4 can be listening to the same program with narration, etc. Thus, each user / listener 15 can use their own device 12 to watch and listen to the same program, but from different streams due to the type of content being played, as FIG. 5A-FIG. 5D illustrated in FIG. 5A-FIG. 5D shows how the data is stored and encoded and interleaved. In this case, each of the user devices 1-N receives the UI (including the X-Fade UI, the Pers. EQ UI, and the content UI) and is able to view and listen to selected content. The user / listener can also select whether to synchronize with other listeners based on a selectable synchronization (or smart TV synchronization) button on the Acc app UI 1270 FIG. 12C .
[0112] More specifically, referring to FIG. 2A, a top-level block diagram 200A of the user / listener (or user / recipient) side of a system for providing personalized audio streaming and rendering for a smart television with cross-fade audio mixing controlled by a remote control or user device, in accordance with embodiments of the present disclosure. In particular, a smart playback device 12, such as a smart television, is shown with a remote control (or remote) 220 or user device 12 for controlling cross-fade logic (X-Fade logic) 100 features and other AccApp UI features discussed herein via a smart television display 14. In some embodiments, a user / listener 15 can communicate with the smart television and AccApp logic 16 using a remote control 220 that communicates (wirelessly or wired) with a remote control interface or remote receiver 216 within the smart television 20. In that case, the user / listener 15 will press a button 218 on the remote 220, as shown by line 243, which will cause the AccApp logic 16 within the smart television to provide an X-Fade / EQ / content UI on line 223 to the smart television display 14, which is discussed below with reference to FIG. 2B. FIG. 12B More particularly, the UI can include an X-Fade slider and FIG. 12C other UI display features shown in FIG. 2B, including an audio content list, a personalized equalizer, and other AccApp UI features discussed herein.
[0113] More particularly, the AccApp logic 16 receives Max-Acc-Mix and Cin-Mix on lines 207, 209 from a known audio decoder 202, which is a known AAC-LC decoder such as currently provided and supported by the streaming application of Apple®TV®, Amazon®Fire TV®, Google®Chromecast®, and others. The audio decoder 202 receives an encoded audio signal or mix that can include the Cin-Mix and Max-Acc-Mix interleaved in a single bitstream, and decodes the bitstream (or bs) into separate streams using known decoding software, as provided on line 207 for the Cin-Mix and on line 209 for the Max-Acc-Mix. Again, it is well known that the audio decoder or decoder 202 ( FIG. 2A 、 FIG. 2B 、 FIG. 2C 、 FIG. 2D ) parses and decodes the input audio bitstream into the pre-encoded audio tracks, e.g., Cin-Mix and Max-Acc-Mix. The audio decoder 202 can use software running on the operating system (OS) of the device streaming the content (e.g., iOS, Android, Roku, etc.) and use the same codec as used by the encoder 26.
[0114] Further, the Acc application logic 16 receives a noise floor (NF) signal on line 213, which indicates the noise floor in the listening environment. The ambient noise floor NF can be estimated by receiving ambient sound from a microphone 206, which provides the ambient sound on line 211 to a known filter 204 that filters out (or blocks) the human voice frequency range and provides NF on line 213. The value of NF can be used to adjust the X-Fade slider position of the X-Fade logic 100. In particular, when NF is available, the system can first set the X-Fade slider position to a value that allows the minimum loudness LRA of the dialog (Dx) to be higher than the noise floor (NF) or the maximum value of the X-Fade slider position, whichever is lower. In some embodiments, the X-Fade logic 100 or the Acc application logic 16 can also check to ensure that the LRA of Dx and M&E do not exceed a maximum permissible audio level, such as -2 dB (true peak), which is discussed more below. Thus, the Acc application logic can also have known loudness detection software for determining LRA. In some embodiments, the LRA values of the content can be provided in the metadata of the video content or can be stored separately from the video on the encoded content server 40 or other server.
[0115] The Acc app logic 16 also receives content selection from the remote control 220 on line 229 and X-Fade / EQ adjustment signals from the remote control 220 on line 227. The Acc app logic 16 also provides X-Fade / EQ audio mix to the speakers 11 on line 225 and X-Fade / EQ and content UI to the display 14 on line 223. The X-Fade / EQ audio mix is a digital audio signal as it exits the X-Fade logic and is then processed through a known digital-to-analog converter (DAC) to convert the mix to an analog signal before it is provided to the speaker(s) 11. In addition, the Acc app logic 16 provides X-Fade / EQ and content UI to the user device 12 on a set of lines 230 and receives selected content requests and X-Fade / EQ adjustment signals on the set of lines 230 in communication with the user device 12. The Acc app logic 16 also communicates with the user attribute server 42 on line 219, which holds information and data associated with a given user device 12 or user / listener 15. The Acc app logic 16 also requests and receives an audio content list (or index or audio file list) on line 215 from the encoded content server 40 on line 201, which lists the audio content or audio files available on the encoded content server 40 for viewing and selection by the user 15. The Acc app logic 16 also provides content selection requests received from the user / listener 15 to the content API 30 on line 205, which obtains the requested content from the encoded content server 40 on line 255 and provides the selected content (e.g., Cin-Mix + Mac-Acc-Mix, which are interleaved in a single bitstream) to the audio decoder 202 on line 203.
[0116] In addition, the Acc app logic 16 has logic to perform various functions, such as the X-Fade logic 100, the Acc app UI logic 210, and the personalized equalizer (Pers. EQ) logic 208. In particular, the X-Fade logic 100 performs cross-fade rendering or mixing of the Cin-Mix and Max-Acc-Mix using the power-preserving X-Fade equation (Equation 3) based on the position of the X-Fade slider (or X-Fade adjustment) to provide the resulting mix of the two input signals using a power-preserving gain curve, as discussed above with FIG. 1 The personalized equalizer (Pers. EQ) logic 208 is described below with FIG. 11
[0117] The Acc app UI logic 210 provides a user interface (UI) to the display 14 of the smart TV including the X-Fade UI and content UI on line 223, such as discussed below FIG. 12B and FIG. 12C shown. The Acc app 16 receives X-Fade adjustment commands from the remote on line 227 or from the user device 12 on line set 231. The X-Fade logic 100 performs the X-Fade function based on the commands from the X-Fade slider, similar to the case discussed herein with FIG. 1 The output of the X-Fade logic 100 is the X-Fade mix (or Acc-Mix or personalized mix or personalized accessibility mix or Pers-Mix) provided on line 225 to the speakers 11, and the resulting output audio from the speakers is provided (through the air) to the user / listener 15, as shown by dashed line 239. In addition, the video portion of the selected content Cin-Mix (or C-Mix) is provided on line 217 from the content API 30 to the display 14 of the smart TV, which can be obtained on line 257 from the encoded content server 40. The video portion C-Mix video on line 223 and the X-Fade / EQ / content UI (described below) are provided by the display 14 to the user / listener 15, as shown by dashed line 241.
[0118] The Acc app logic 16 can also request and receive a content list (or index) on line 215, which lists the audio content or audio files available on the encoded content server 40 for viewing and selection by the user / listener 15.
[0119] Referring to FIG. 2B , a top-level block diagram 200B of the user / listener side of a system for providing personalized audio streaming and rendering for a single user device according to embodiments of the present disclosure as discussed above is shown. Referring to FIG. 2C , a top-level block diagram 200C of the user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices each having a personalized cross-fade audio mix according to embodiments of the present disclosure as discussed above is shown. Referring to FIG. 2D , a top-level block diagram 200D of the user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices each having a personalized cross-fade audio mix and content selection according to embodiments of the present disclosure as discussed above is shown.
[0120] Referring to FIG. 5A, a diagram 500 illustrating interleaved Cin-Mix and Max-Acc-Mix content audio streams of accessible audio provided by the audio encoder 26 according to an embodiment of the present disclosure. As discussed herein, the audio encoder 26 ( FIG. 1 ) analyzes and compresses the input audio source associated with the video source. FIG. 5A Shown on the left side of are two non-encoded files, a Cin-Mix English (original) audio file and a Max-Cin-Mix audio file (or audio track), which are the inputs to the encoder 26 for the selected content. FIG. 5A The right side of the figure shows the output of the encoder 26 as received by the content provider computer 20 ( FIG. 1 ) is an interleaved audio stream of accessible audio stored in an encoded content server, as discussed herein. In particular, the input audio files are each such as FIG. 5A are sampled and interleaved as shown in , such as sample 1, sample 2, etc. and each sample has a sample portion of each input file.
[0121] refer to FIG. 5B , diagram 515 showing interleaved Cin-Mix, Mid-Cin-Mix, and Max-Cin-Mix content audio streams according to an embodiment of the present disclosure.
[0122] refer to FIG. 5C , a diagram 530 of an interleaved Cin-Mix and Max-Acc-Mix content audio stream with accessible audio in two languages for voice-over audio according to an embodiment of the present disclosure is shown. In that case, for a stereo audio mix, two or three of the four mixes shown (if stereo is used) can be used together to provide a voice-over or adaptive voice-over feature. For example, one or both of the dashed boxes can be optional. Thus, the stream can include: if the audio mix is in mono (one channel per track), all four tracks shown can be used if desired. In that case, when four tracks are received, the Acc application can select three tracks to provide to the X-Fade renderer, or in some embodiments, a four-input X-Fade renderer can be used if desired.
[0123] In particular, the audio encoder 26 ( FIG. 1 ) provides interleaved data by placing samples from each audio channel (e.g., the left and right channels in a stereo mix) sequentially one after the other in a single stream, effectively "interleaving" the channels together so that the data for the first channel is followed by the data for the second channel, and so on, resulting in a single continuous data stream that can be easily decoded and played back by a compatible audio player.
[0124] like FIG. 5A 、FIG. 5B 、 FIG. 5C As shown in FIGS. 47A and 47B, the present disclosure also uses the encoder 26 to interleave two (or more) different audio tracks (e.g., Cin-Mix and Max-Acc-Mix), which provides a single stream with both audio tracks embedded in the stream, which allows the user / listener (after decoding) to cross-fade between the two (or more) audio tracks to obtain a personalized listening experience while only receiving a single stream. Thus, for a standard stereo mix, the present disclosure will use four channels, two for Cin-Mix (L / R) and two for Max-Acc-Mix (L / R), i.e., two additional channels. In the case of three stereo tracks, e.g., Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix, the present disclosure will use six channels, two for Cin-Mix (L / R), two for Mid-Acc-Mix, and two for Max-Acc-Mix (L / R), i.e., four additional channels. For a surround sound 5.1 mix with 5 channels and an immersive 5.1.4 mix with 10 channels (including subwoofers), only the center channel (or dialogue channel) of the mix will be DCE processed to create Max-Acc-Mix, so the encoder will only need to interleave one additional channel for Max-Acc-Mix, or two additional channels if Mid-Acc-Mix or Max-Acc-Mix is used.
[0125] Reference is made to FIG. 5D, a block diagram 550 showing non-encoded content server 36 having separately maintained non-encoded Cin-Mix and Max-Acc-Mix audio content streams as inputs to audio encoder 26, and encoded content server 40 having encoded and interleaved Cin-Mix / Max-Acc-Mix audio content files as outputs from audio encoder 26, in accordance with embodiments of the present disclosure. In particular, the Cin-Mix (original) files can come from a studio that created the movie mix, and the Max-Acc-Mix (from DCE) can come from DCE logic 62, as described herein. In some embodiments, the studio or content provider can provide Cin-Mix and Max-Acc-Mix in multiple languages (i.e., two files, Cin-Mix English (original) and Max-Acc-Mix English (from DCE)) to attract a large listening audience, as shown in table 550, for example, 40+ languages. Further, in some embodiments, for certain content, the studio or content provider can provide Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix in certain desired languages as needed (i.e., three files, Cin-Mix Japanese (original), Mid-Acc-Mix Japanese (from DCE), and Max-Acc-Mix Japanese (from DCE)), shown here as a Japanese language example.
[0126] Further, in some embodiments, the commentary options to be included in the encoded mix can include various desired combinations of desired audio files, such as: Cin-Mix English (original version OV), Max-Acc-Mix English (from DCE) + Max-Acc-Mix Spanish (from DCE), Max-Acc-Mix English (from DCE) + Max-Acc-Mix Spanish (from DCE) + separate AD Spanish (audio description (AD) in the same language as the commentary, as for this particular use case, this is the language that makes it easier for non-English native OV people to access the content). Other combinations can be used if desired. The “+” sign used above means the sum / addition / overlay of the 2 or 3 audio signals below. The audio description (AD) track here carries only the audio description, not the entire mix (unlike the method below).
[0127] In some embodiments, in addition to FIG. 5DIn addition to those shown in Table 550, the adaptive voiceover (VO) mix can include: Cin-Mix English / Cin-Mix Spanish, or Cin-Mix English / Cin-Mix Spanish VO-Beginner, or Cin-Mix English / Cin-Mix Spanish VO-Intermediate, or Cin-Mix English / Cin-Mix Spanish VO-Expert. FIG. 5D In addition to those shown in Table 550, the adaptive voiceover (VO) mix can include: Cin-Mix English / Cin-Mix Spanish, or Cin-Mix English / Cin-Mix Spanish VO-Beginner, or Cin-Mix English / Cin-Mix Spanish VO-Intermediate, or Cin-Mix English / Cin-Mix Spanish VO-Expert.
[0128] Reference is made to FIG. 5E Table 580 shows various encoding and decoding options with given stream types and user desired configurations, in accordance with embodiments of the present disclosure. In particular, the present disclosure provides embodiments that include a single multi-channel audio encoder / decoder. MPEGAAC-LC and Dolby E-AC3 support a “dual stereo” encoding mode, such as 2x 2.0 / stereo independent encoding. MPEGAAC-LC and Dolby E-AC3 support a “quad” encoding mode, such as 4x 1.0 / mono independent encoding. For example, it is well known that Dolby Digital Plus with Joint Object Coding (JOC), commonly referred to as Dolby Digital Plus with Dolby Atmos, allows for carrying main audio and associated audio, including Atmos object-based audio, in a single bitstream. In Table 580, 2.0 and 4.0 mean 2 channels and 4 channels, respectively. Furthermore, 100% of Dolby devices can decode AAC-LC. This is equally applicable to other audio compression technology codecs, such as DTS, Opus, ITU codecs, etc. It is well known that AC3 (or AC-3) is commonly referred to as Dolby Digittal, and EC3 (or Enhanced AC-3) is commonly referred to as Dolby Digital Plus, which has enhanced sound quality. In some embodiments, such as or Open source software tools such as FFmpeg®, Audacity®, and the like can also be used with the present disclosure for processing audio files. Further, in some embodiments, for Dolby Atmos 5.1.4 (movie immersive), the encoder can be a Dolby EC3 + JOC encoder with 5.1 core primary audio + JOC metadata per bsmod in ETSI TS 102 366 V1.4.1 and 1.0 associated audio. Further, in some embodiments, for Dolby Surround 5.1, the encoder can be a Dolby EC3, 7 channel encoder, such as 5.1 primary audio + 1.0 associated audio per bsmod in ETSI TS 102 366 V1.4.1. Further, in some embodiments, for stereo or audio, 2.0 dual stereo can have 2 enc / dec, or 1.0 quad can have 4 enc / dec, or 4.0 multi-channel can have 1 enc / dec. Further, in some embodiments, for surround sound, 6.1 = 5.1 + 1 (center channel) can have 1 enc / dec. Further, in some embodiments, for immersive sound, 6.1.4 = 5.1.4 + 1 (center channel) can have 1 enc / dec.
[0129] Further, as described herein, in some embodiments where Dolby / DTS (proprietary codecs) are used, there can be a 2.0 downmix before running DCE. Further, in some embodiments, the system can receive (or create) a Cin-Mix 2.0, and then create a Max-Acc-Mix 2.0, which is encoded and then decoded (at the user / listener side) and provided to an X-Fade renderer to provide an X-Fade mix in 2.0 format.
[0130] Referring to FIG. 6 , FIG. 1 and FIG. 2A-FIG. 2D The system of the present disclosure shown in FIGS. 1-3 can be implemented in a network environment 600. The system includes a content provider computer 20, one or more user devices 12, and various servers 35 that interact with the computer 20 or user devices 12 to perform the functions described herein.
[0131] In particular, various components of an embodiment of the system of the present disclosure include a plurality of computer-based user devices 12 (e.g., a smart TV and devices 2 through N) that can interact with respective users (user / listener 1 through user / listener N). A given user 15 can be associated with one or more of the devices 12. In some embodiments, the Acc application 19 can reside on the user device 12 or on a remote server and communicate with the user device(s) 12 via a network. In the case where the user device 12 is a smart TV 12, the user 15 can use a known remote control interface or receiver 216 ( FIG. 2A-FIG. 2D ) communicates with the smart TV via the remote control 220, as described above. In some embodiments, the user device 12 can communicate directly with the smart TV using a wired or wireless connection (e.g., Bluetooth, wifi, NFC or other wireless connection), as shown by line 71.
[0132] In some embodiments, one or more of the user devices 12 can be connected or communicate with each other via a communication network 70, such as a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a peer-to-peer network, or the Internet (wired or wireless), by sending and receiving digital data over the communication network 70, as indicated by line 72. If the user device 12 is connected via a local or private or secure network, the user device 12 can have a separate network connection to the Internet (or other network) for use by a web browser running on the device 12. The user devices 12 can also each have a web browser 17 to connect to or communicate with the Internet to obtain desired video / audio content in a standard client-server configuration to obtain the Acc application 19 or other files required to execute the logic of the present disclosure. The user device 12 can also have a local digital storage device located in the device itself (or directly connected to it, such as an external USB-connected hard drive, thumb drive, etc.) for storing data, images, audio / video, documents, etc. that can be accessed by the Acc application 16 running on the user device 12.
[0133] Further, the user device 12 can also communicate with separate computer servers 35 via the network 70, such as a content AP server 31 (which has a content selection API 30), a non-encoded content server 36, a DCE parameter server 38, an encoded content server 40, and a user attribute server 42. The servers 35 can be any type of computer server having the necessary software or hardware (including storage capabilities) for performing the functions described herein. Further, the servers 35 (or the functions performed thereby) can be located individually or collectively in one or more separate servers on the network 70, or can be located in whole or in part in one (or more) of the user devices 12 on the network 70.
[0134] As discussed herein, the content provider computer 20 can be a smartphone, computer, laptop, tablet, or other computer-based device that can receive instructions from a content provider user / administrator 21 and can have portal / DCE logic 28 running on the computer 20 that receives non-encoded (or unencoded) content from a non-encoded server 36 (as indicated by line 51) and creates encoded content and delivers the encoded content (e.g., streaming or VOD content of a movie or program or other content) to an encoded content server 40 via the communication network 70 using a web browser 24 (or web server or equivalent interface) (as indicated by line 52) for use by user devices 12, by users / listeners 15 (user / listener 1 through user / listener N). Similarly, as discussed herein, the user device 12 can be a smartphone, computer, laptop, tablet, smart television, or other computer-based device that can receive digital content (e.g., streaming or VOD programs) including encoded content from the content provider directly via the communication network 70, e.g., using a web browser 24 (or web server or equivalent interface) or from content stored on the encoded content server 40 to the device 12 for viewing and listening by the user / listeners 15.
[0135] The portal / DCE logic 28 running on the computer 20 can also provide audio / video to the CP user / administrator to obtain, update, and save DCE parameters to and from the DCE parameter server 38. The computer 20 can also have a known audio encoder 26 discussed herein that encodes audio content into a desired format (e.g., stereo, Dolby, Dolby ATMOS, etc.) for use by the users 15 as discussed herein.
[0136] The CP computer 20 has a portal UI that provides a model performance X-Fade control, a waveform window, DCE gain / model / noise floor selection data, audio output, and accessible audio approval selection capabilities to a user / administrator 21 of the content provider, as discussed herein. In addition, the portal / DCE logic 26 can obtain DCE parameters for the DCE and can display the parameters on a display, such as the display discussed herein FIG. 1 with the portal UI.
[0137] In addition, a content selection API 30 can reside on a content selection API server 31 that can communicate with the user devices 12 and with the CP computer 20 and other servers via the network 70 and with each other or any other network enabled device or logic as necessary to provide the functionality described herein.
[0138] Similarly, the user devices 12 can each also communicate with the logic or software applications described herein and any other network enabled device or logic necessary to perform the functionality described herein via the network 70.
[0139] By adding software or logic to the user devices 12, such as adding logic to the hub Acc application or Acc application software, or installing new / additional application software, firmware or hardware to perform some of the functionality described herein, or other functionality, logic or processes described herein, portions of the present disclosure shown herein as being implemented externally to the user devices 12 can be implemented within the user devices 12.
[0140] Referring to FIG. 11 , a block diagram showing components of the personalized equalizer (Pers. EQ) logic 208 FIG. 2A-FIG. 2D ) in accordance with embodiments of the present disclosure. In particular, the Pers. EQ logic 208 receives the X-Fade mix on line 1103 and performs digital frequency signal processing on the X-Fade mix through known digital frequency signal processing logic 1102 that amplifies certain frequency ranges based on frequency range factors from the user attribute server 42 on line 1105. The digital frequency signal processing logic 1102 can use similar logic as known hearing aid logic or other audio frequency range enhancement devices. The frequency range factors can be pre-stored in the user attribute server 42 based on hearing tests previously obtained via third party audiograms or user device hearing tests run by the user.
[0141] In some embodiments, the frequency range factor can be provided by the system of the present disclosure based on a user listening test. In that case, the content provider can provide a CP application listening test that provides a series of sounds to the user / listener 15 at various frequencies to test hearing and allows the user / listener 15 to selectively adjust the frequency enhancement for each frequency range being tested as shown by the frequency range adjustment block 1106, which can be a UI, and provides the frequency range factor for a given user / listener 15 to the digital frequency signal processing logic 1102 on line 1107. The Pers. EQ logic 208 can also save the frequency range factor for a given user / listener 15 on the user attributes server 42 for future use by the same user / listener 15. The UI for the listening test UI 1290 is shown in FIG. 12D FIG. 19, as discussed below, and provides a frequency range adjustment slider for the user / listener to adjust.
[0142] Referring to FIG. 12A, showing a screen illustration 1200 of a graphical user interface (UI) or portal UI that can be used by a content provider (CP) user / administrator with the ability to set / adjust DCE parameters and listen to results across cross-fade ranges for DCE parameter adjustment and Max-Acc-Mix monitoring, in accordance with embodiments of the present disclosure. In particular, the portal UI 1200 can contain a field 1206 for inputting an audio / video file of interest, which can be selected from a list of content audio files shown in window 1204, which lists available files and allows the user to select a file of interest for creating an accessible audio track. Once an audio file has been selected, the system displays that file in field 1206. The portal UI 1200 will also display an accessibility fade (or X-Fade and volume control) window 1202 (or screen portion). The X-Fade mix and volume can be controlled by X-Fade and volume sliders 1210, 1212, respectively. One end (upper or top) of the X-Fade slider 1210 provides only enhanced dialogue (or Max-Acc-Mix), and the other end (lower or bottom) of the X-Fade slider 1210 provides only original dialogue or original movie mix (Cin-Mix), and between the two ends of the slider, the X-Fade renderer provides a mix of Cin-Mix and Max-Acc-Mix, as discussed herein. The sliders 1210, 1212 can be manually moved by the user / listener by using a smart TV remote (as discussed herein) or by tapping (or clicking) on or touching the sliders 1210 using a smart phone or laptop or desktop computer or tablet computer, when in use. In some embodiments, voice commands can also be used to adjust the sliders, e.g., if the smart TV or user device is connected to a voice responsive application. This application supports the discussion in the case of most voice receiving products, such as Google etc. Additionally, there can be a waveform window 1208 displayed on the UI screen. The waveform window 1208 provides a display of the Cin-Mix and Max-Acc-Mix, as well as a view of the cross-fade and X-Fade gain curves discussed herein. In particular, the waveform window 1208 shows an example of audio signal monitoring on macOS using known pure data tools for the stereo 2.0 use case, which shows the original mix (Cin-Mix) and the dialogue-enhanced mix (Max-Acc-Mix). The window 1208 shows how a content provider (CP) user / administrator 21 can experience a real-time cross-fade between the Cin-Mix (original mix) (left / right) (on the left) and the Max-Acc-Mix (left / right) (dialogue-enhanced mix) (on the right), which allows the user / administrator 21 to see and hear the audio effects of the DCE used to create the accessible mix (or Max-Acc-Mix).
[0143] Further, the portal UI 1200 provides a window or screen portion 1220 that provides the ability for the CP user / administrator to select various DCE parameters (or DCE parameters), such as: DCE model type, DRC gain profile, Gmine, MNE gain factor, Gdx, Dx gain factor, noise floor, ALN gain value input or automatic calculation. The DCE parameter window 1220 can also include the ability to set or adjust the dialogue to program loudness or dialogue to residual loudness (DPL or DRL), e.g., the gap in dB between dialogue and residual / MNE (not shown). Other parameters can be included in the DCE parameter window 1220 if desired. Further, the portal UI has a button 1222 that can be selected when the Max-Acc-Mix is at an acceptable level (Max-Acc-Mix OK). The DCE parameter window 1220 allows the ALN gain to be set to a desired value, or if the user selects (or clicks) the automatic box, the system will automatically select an ALN gain that will maximize the accessible audio mix (Max-Acc-Mix) dynamic range (or loudness range) above the set noise floor (NF) while not exceeding the requirements for maximum audio levels described herein, as described with equation 1. The content owner can set a given noise floor, e.g., based on a preset audio standard noise floor or based on target audience or user preferences.
[0144] Further, the DCE parameter window 1220 allows the DRC gain profile (or DRC gain curve) to be set based on a desired gain profile provided in a file that the user can select. As discussed herein, different profiles can provide different levels or amounts of compression to the M&E mix, as represented by the equation 2, which is discussed herein. FIG. 9DRC gain curve 900 shown in FIG. 9, including a boost range, a null range (or unity gain region), an early cut range, and a cut range. Other ranges and labels can be used to describe the DRC gain curve if desired.
[0145] Referring to FIG. 12B , a screen shot 1260 of a smart TV display 14 showing a remote control 220 graphical user interface (UI) of a cross-fade (X-Fade) renderer (or accessibility cross-fader) window 1202 for manually adjusting FIG. 2A , FIG. 2C and FIG. 2D in accordance with embodiments of the present disclosure. In particular, when a user holds an up or down arrow for a "long press", for example, greater than 2 seconds, the smart TV knows this is a special command and the Acc application 210 FIG. 2A within the smart TV 12 will display the accessibility cross-fader (X-Fade and volume control) window 1202 on the TV display screen 14. Further, if the plus (+) volume arrow 1264 on the remote control 220 is pressed for a long time (e.g., more than 2 seconds), a signal will be sent to the smart TV, as shown by the dashed line 1262, which will be interpreted by the Acc application 210 as a command to increase the X-Fade slider 1210 position. Conversely, if the plus (+) volume arrow 1264 on the remote control 220 is pressed for a short time (e.g., less than 2 seconds), a signal will be sent to the smart TV and it will be interpreted as a command to increase the volume 1212. Similar results occur if the minus (-) arrow 1266 on the remote control 220 is pressed for a long time, a signal will be sent to the smart TV, as shown by the dashed line 1262, which will be interpreted by the Acc application 210 as a command to increase the X-Fade slider 1210 position. Conversely, if the minus (-) volume arrow 1266 on the remote control 220 is pressed for a short time (e.g., less than 2 seconds), a signal will be sent to the smart TV and it will be interpreted as a command to decrease the volume 1212. As discussed herein, the X-Fade slider 1210 transitions between the original dialogue or original cinema mix (Cin-Mix) and the enhanced dialogue or maximum accessibility mix (Max-Acc-Mix).
[0146] Referring to FIG. 12C , a screen shot 1270 of a graphical user interface (UI) of an Acc application used by a user having the ability to set / adjust cross-fade and enable certain features in accordance with embodiments of the present disclosure. In particular, the UI 1270 can contain a field 1206 for inputting an audio / video file of interest that can be selected from a list of content audio files shown in window 1204 that lists available files and allows the user to select a file of interest, similar toFIG. 12A the case shown in FIG. 12. Once an audio file has been selected, the system displays the file in field 1206. The UI 1270 will also display the accessibility X-Fade control (or X-Fade and volume control) window 1202 (or screen portion). The X-Fade mix and volume can be controlled by sliders 1210, 1212, respectively, similar to the case shown in FIG. 12. FIG. 12A the case shown in FIG. 12. One end of slider 1210 provides full enhanced dialogue (or Max-Acc-Mix), and the other end of slider 1210 provides the original movie mix (Cin-Mix). When slider 1210 is between the two ends, the output audio is a hybrid combination of the Cin-Mix and Max-Acc-Mix, using the power hold curve shown in FIG. 13, and discussed above for the X-Fade logic in FIG. 10A-FIG. 10C and FIG. 1 and FIG. 2A-FIG. 2D is discussed above.
[0147] In addition, the user / listener UI 1270 can also display several other X-Fade displays, such as: video voice-over 1280, live sports X-Fade 1 1282, live sports X-Fade 2 1284. In addition, the user / listener UI 1270 can also display an Acc features window 1288 that allows the user 15 to select various features of the Acc application, such as: remote, hub mode, single device mode, sync smart TV, microphone (presence or absence), live sports, personalized EQ, adaptive voice-over (VO), and audio description (or narration), as discussed herein. The user / listener UI 1270 can also display an X-Fade mix OK button 1294 that allows the user to select or click when the settings for a given selected content are acceptable to the user / listener 15.
[0148] In particular, when the Mic button is selected, the X-Fade logic will automatically adjust the X-Fade slider 1210 to a level above the measured noise floor (NF). If desired, the user can then further adjust the slider 1210 to provide the best listening experience for the user 15. In addition, when the live sports button is selected, the live sports X-Fade 1 and X-Fade 2 windows 1282, 1284 will appear on the UI. In addition, when the audio description button is selected, the accessibility X-Fade control window 1202 can display the range under the bottom X-Fade slider range limit for the mix according to the original version (OV or English OV) and the range under the top X-Fade slider range limit for the mix according to the audio description (AD or English AD), as shown in FIG. 14EFurthermore, when hub mode is selected, device 12 knows to communicate with a hub device (e.g., a smart TV) that will provide content directly to the user device. Furthermore, when single device (or standalone) mode is selected, device 12 knows that it must obtain audio content directly from the video / audio source, rather than through the smart TV. Furthermore, when remote mode is selected, device 12 knows to communicate that it can control the UI on the smart TV using the user device UI.
[0149] When the Sync with Smart TV mode is selected, the device 12 will synchronize a single user device with the Smart TV while playing content, so that the user device and the Smart TV play the same content in sync with each other. Even if there are multiple different streams of the same content provided from the Smart TV to different devices (e.g., original, alternative language, voice-over, adaptive voice-over, audio description, etc.), the content is played in sync with each other. FIG. 2D (with multiple decoders)), this can also be done. The synchronization request can be part of the communication with the hub Acc application 16 on the smart TV. The status of each Acc feature in Acc features 1288 can also be included in the communication with the hub Acc application as needed (for example, FIG. 2C 、 FIG. 2D on lines 231B, 233B, 231A, 231B).
[0150] Additionally, when the Personal EQ button is selected, the Pers.EQ window 1286 appears and allows the user to select a frequency range factor or weighting factor ( FIG. 11 ) comes from, such as using FIG. 11 Further described, including allowing the user to run a hearing test to determine the frequency range factor. Thus, the present disclosure can be customized to the hearing ability of the user / listener 15 to provide a personalized audio experience. In addition, when the adaptive voice-over button is selected, a language level selection window 1289 appears and allows the user to select the level of the voice-over language proficiency, for example, beginner, intermediate, expert. Other levels can be provided if desired. As discussed herein, content providers can provide additional mixes for certain levels of translation based on the level desired by the user. In that case, the system will filter the content list based on the level desired by the user. If desired, the language window 1290 can also be used as a language filter at any time to provide only content in a given language in the content list 1204.
[0151] The user / listener UI 1270 can also display a personalized EQ window 1286 that allows the user to select between types of personalized equalizer sources, including third party audiograms, device applications or OS (e.g., IOS or Android), or content provider (CP) application listening test modes, as discussed herein. The user / listener UI 1270 can also display a language selection window 1290 that allows the user to select content in a particular language, as discussed herein. The user / listener UI 1270 can also display a language understanding window 1290 that allows the user to select a level of understanding of voice-over (VO) language for selecting VO content in a particular language, as discussed herein.
[0152] Referring to FIG. 12D , a screen illustration 1270 of a graphical user interface (UI) for a listening test to determine parameters of a personal equalizer, in accordance with an embodiment of the present disclosure, is shown. In particular, the UI 1270 can display an input frequency range control window 1292 that allows the user to adjust the volume (or gain) of different frequency ranges based on the position of the slider associated with each range. The user / listener UI 1270 can also display a pre-adjusted hearing result window 1294 that shows the output results before frequency adjustment has been made for a given ear of the user / listener. In addition, the user / listener UI 1270 can also display an adjusted hearing result window 1296 that shows the output results after frequency adjustment has been made for a given ear of the user / listener. In addition, the user / listener UI 1270 can also display a button 1291 that can be selected or clicked to start the test. The user / listener UI 1270 can be launched when the user selects a CP application listening test mode option in the Acc application UI 1270, as discussed herein.
[0153] Referring to FIG. 13A , a flowchart 1300 illustrates one embodiment of a process or logic for executing the portal / DCE logic 28( FIG. 1 ) in accordance with an embodiment of the present disclosure. The process 1300 begins at block 1302, which launches the portal UI 28( FIG. 12A) from the CP user / administrator. Next, block 1304 retrieves the Cin-Mix audio for the selected content from the non-encoded content server. Next, block 1305 retrieves the default (or saved) DCE model, DCE gain, and noise floor (NF) from the DCE parameter server. Next, block 1306 executes the DCE model with the DCE gain and NF and creates or updates Max-Acc-Mix. This block 1305 can also receive adjustments to the NF, DCE gain, and DCE model, as shown in block 1314, depending on the determination of whether the Max-Acc-Mix is acceptable that occurs in block 1312. After block 1306, block 1308 provides the Max-Acc-Mix to the CP computer display UI (portal UI logic) of the CP user / administrator. Next, block 1310 executes X-Fade (crossfade) logic using X-Fade position gain on the Cin-Mix and Max-Acc-Mix to create an X-Fade mix and play the X-Fade mix on the speaker(s) for the CP user or administrator. Next, block 1312 determines whether the Max-Acc-Mix is acceptable. If not, block 1314 receives adjustments to the NF, DCE gain, and DCE model, and this is passed to block 1306, which executes the DCE model with the DCE gain and NF and creates or updates Max-Acc-Mix. Next, if yes, the Max-Acc-Min is acceptable, block 1316 saves or updates the DCE model, DCE gain, and NF for the selected content in the DCE parameter server and provides the Max-Acc-Mix to the audio encoder. After block 1316 is complete, the process exits.
[0154] Reference FIG. 13B , flowchart 1320 illustrates the portal UI logic 60 FIG. 1One embodiment of the process or logic of the Cin-Mix and Max-Acc-Mix (per channel) is illustrated in flowchart 1320 of FIG. 13B. Process 1320 begins at block 1322, which retrieves the latest Cin-Mix audio and Max-Acc-Mix (per channel) and provides these to the waveform display tool (real data). Next, block 1324 displays the portal UI with the waveform window (with Cin-Mix Cin-Mix and Max-Acc-Mix waveforms) and user selected parameters on the content provider computer display, including: X-Fade and volume controls, audio file list, and DCE parameters such as: DCE model type, DRC gain profile, Gmne, MNE gain factor, Gdx, Dx gain factor, noise floor, ALN gain value input or automatic calculation. Next, block 1326 determines if an updated value for the X-Fade slider has been received from the user. If so, block 1328 updates the X-Fade gain value in the X-Fade logic and updates on the user interface (UI). Next, or if the result of block 1326 is no, block 1330 determines if an updated value for the DCE model type has been received. If so, block 1332 updates the DCE model to the selected DCE model and updates on the UI. Next, or if the result of block 1330 is no, block 1334 determines if an updated value for any of the DCE parameters has been received. If so, block 1336 updates the user selected values for the DCE parameters in the appropriate equations and models and updates on the UI. Next, or if the result of block 1334 is no, block 1338 determines if an updated value for the volume slider has been received. If so, block 1340 updates the total volume for the speakers and updates on the UI. Next, or if the result of block 1338 is no, process 1320 exits.
[0155] Reference is made to FIG. 13C , flowchart 1300 illustrates one embodiment of a process for executing accessibility (Acc) application logic 16 FIG. 2A 、 FIG. 2C 、 FIG. 2D) of FIG. 13. Process 1350 begins at block 1352, which receives content selection from a user / listener, which can be from a remote or user device. Next, block 1353 sends a request to the content API and receives Cin-Mix and Max-Acc-Mix from the decoder and subtitles if needed. Next, block 1354 determines whether hub mode should be performed. If so, block 1356 performs hub mode. Next, or if the result of block 1354 is no, block 1357 determines whether single mode should be performed. If so, block 1358 performs single mode. Next, or if the result of block 1357 is no, block 1359 performs remote control mode. Next, block 1360 determines whether a noise floor signal is available. If so, block 1361 receives a noise floor (NF) audio signal from a microphone 206 FIG. 2A-FIG. 2D ) connected to the user device 12 (or smart television). Next, or if the result of block 1360 is no, block 1362 performs X-Fade logic 100 on the Cin-Mix and Max-Acc-Mix to create X-Fade mix values based on measured NF or based on the position of the X-Fade slider as set by the user. In particular, if NF is measured or sensed, the X-Fade logic 100 can set the value of the slider such that the integrated loudness IL or, if needed, the entire dynamic range (or audio loudness range) of the output audio (X-Fade mix) is higher than the noise floor NF or at the maximum value of the X-Fade slider. In some embodiments, the logic can also check to ensure that the dynamic range (or audio loudness range) of the content does not exceed the maximum loudness permitted by the content provider or any required maximum loudness standard (maximum audio), such as -2dB (true peak).
[0156] Next, block 1364 determines whether synchronization is selected. If so, block 1366 performs synchronization with the current audio and video. Next, or if the result of block 1364 is no, block 1368 determines whether personalized EQ is enabled. If so, block 1370 performs personalized EQ on the X-Fade Mic based on the personalized EQ parameters on the user attribute server. Next, or if the result of block 1370 is no, block 1372 plays the X-Fade / EQ mix for the user or listener on the speaker(s). Next, block 1374 determines whether adaptive voice over (VO) is enabled. If so, block 1376 displays adaptive voice over content. Next, or if the result of block 1374 is no, process 1350 exits.
[0157] Referring to FIG. 13D , flowchart 1300 illustrates an embodiment of the Acc application UI logic 210FIG. 2A 、 FIG. 2C 、 FIG. 2D ) of the process or logic. Process 1380 begins at block 1382, which displays the Acc app UI on a user device or smart TV display, and user selected parameters, including: X-Fade and volume control of the following: Acc Fader, VO Fader, Live Sports Fader 1 & 2, Language, Audio File List, Personalized EQ, Level, and Acc Features. Next, block 1384 determines if an updated value of the X-Fade slider has been received from the user. If so, block 1386 updates the X-Fade gain value in the X-Fade logic and updates on the UI. Next, or if the result of block 1384 is no, block 1388 determines if an updated Acc app feature selection has been received from the user. If so, block 1389 updates the state of the ACC feature and updates on the UI. Next, or if the result of block 1389 is no, block 1390 determines if an updated Personalized EQ, Language, or Level selection has been received. If so, block 1391 updates the state of the Personalized EQ, Language, or Level and updates on the UI. Next, or if the result of block 1390 is no, block 1392 determines if an updated value of the volume slider has been received. If so, block 1393 updates the total volume of the speakers and updates on the UI. Next, or if the result of block 1392 is no, block 1394 determines if the selection of the X-Fade mix is OK. If so, block 1395 saves the value of the X-Fade on the server. Next, or if the result of block 1394 is no, process 1380 exits.
[0158] Referring to FIG. 14A , elimination 14B, FIG. 14C 、 FIG. 14D and FIG. 14E , a block diagram illustrating various embodiments of different audio inputs to be streamed and adjustably rendered by a cross-fade renderer, in accordance with embodiments of the present disclosure.
[0159] In particular, referring to FIG. 14A , a block diagram 1402 of a video on demand (VOD) application illustrates embodiments of the present disclosure, similar to utilizing FIG. 2A-FIG. 2DThe case discussed, where Cin-Mix from the studio is provided to the content provider (CP) system on line 96 as discussed herein. The Cin-Mix soundtrack is provided to the DCE 62, which can be tuned (using DCE gains or parameters, as discussed herein) to provide dialogue (Dx) enhancement and M&E compression Max-Acc-Mix on line 92, as discussed herein. The Max-Acc-Mix is provided on line 92 to the encoder 26, which also receives Cin-Mix on line 96. The encoder 26 provides on line 94 an interleaved and encoded soundtrack with Cin-Mix and Max-Acc-Mix, as discussed herein.
[0160] On the receiver or user / listener (right) side of FIG. 1402, the Cin-Mix / Max-Acc-Mix interleaved encoded soundtrack is received on line 203 by the decoder 202, which separates the interleaved soundtrack into Cin-Mix and Max-Acc-Mix soundtracks on lines 209, 207, respectively, which are provided to the X-Fade renderer logic 100. The output of the X-Fade renderer logic 100 is a power-preserving X-Fade mix of the Cin-Mix and Max-Acc-Mix soundtracks, as described herein, where the fade amount is based on the position (or location) of the user-selected X-Fade slider 1282 (as discussed herein) provided on line 225. FIG. 12C ) as discussed herein.
[0161] In addition, in some embodiments, the DCE 62 can provide a second output mix to the encoder 26, e.g., Mid-Acc-Mix (on line 93). In that case, the decoder 202 would provide three audio soundtracks Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix to the three input X-Fade logic 100 within the Acc application logic 16, as described above. FIG. 1 FIG. 2A-FIG. 2D
[0162] Reference is made to FIG. 14B , a block diagram 1420 of live sports illustrating embodiments of the present disclosure, similar to utilizing FIG. 2A-FIG. 2D The case in point, where the input signal is a live (or pre-recorded) video feed from a sporting event or other live event (e.g., a sports mix or live mix) provided to the content provider (CP) system from an outside broadcast (OB) truck on line 96. The sports mix is provided to the DCE 62, which can be tuned (using DCE gains or parameters, as discussed herein) to provide on line 92 a dialog-enhanced commentator mix (or ComMix) with commentator dialog enhanced over background stadium audio sounds or M&E (e.g., crowd and other general stadium noise or audio), or just commentator dialog audio without stadium audio or M&E. In some embodiments, the DCE 62 can be tuned to provide on line 92 a stadium mix (or Stadium-Mix) with background stadium audio sounds or M&E (e.g., crowd and other general stadium noise or audio) enhanced over commentator dialog, or just stadium audio without commentator dialog audio.
[0163] The Com-Mix or Stadium-Mix is provided on line 92 to the audio encoder 26, which also receives on line 96 the sports mix including the Com-Mix and M&E (Stadium-Mix). The encoder 26 provides on line 94 an interleaved and encoded soundtrack (Sports-Mix / Com-Mix or Sports-Mix / Stadium-Mix) with the sports mix along with the Com-Mix or Stadium-Mix.
[0164] In this embodiment, on the receiver or user / listener (right) side of FIG. 1420, the Sports-Mix / Com-Mix or Sports-Mix / Stadium-Mix interleaved encoded soundtrack is received on line 203 by the decoder 202, which separates the interleaved soundtrack into the Sports-Mix soundtrack on line 209 and the Com-Mix or Stadium-Mix soundtrack on line 207 (depending on which the content provider selected when tuning the DCE 62, as discussed herein), which are provided to the X-Fade renderer logic 100. The output of the X-Fade renderer logic 100 is a power-preserving X-Fade mix of the Sports-Mix and Com-Mix or Sports-Mix and Stadium-Mix (depending on which pair is provided in the received stream), with the fade amount based on the user-selected X-Fade slider 1282 FIG. 12C ), as discussed herein, provided on line 225.
[0165] Referring againFIG. 14B In some embodiments, a second (lower) DCE 62A can be provided that can provide a mix (Com-Mix or Stadium-Mix) that is not provided by the upper DCE 62. For example, the upper DCE 62 can provide the Com-Mix on line 92, and the lower DCE 62A can provide the Stadium-Mix on line 92A. In that case, the encoder 26 will receive the Sports-Mix, the Com-Mix, and the Stadium-Mix, and provide on line 94 an interleaved and encoded track (Sports-Mix / Com-Mix / Stadium-Mix) with the Sports-Mix, Com-Mix, and Stadium-Mix tracks, similar to FIG. 5B as shown in FIG. 5B FIG. 14 illustrates an example of three track interleaving by the audio encoder 26.
[0166] At the receiver or user / listener (right) side of FIG. 14, the Sports-Mix / Com-Mix / Stadium-Mix interleaved encoded track is received on line 203 by the decoder 202, which separates the interleaved track into the Com-Mix on line 207, the Sports-Mix track on line 209, and the Stadium-Mix on line 250A, which are provided to the three input X-Fade renderer logic 100 FIG. 10C ). The output of the three input X-Fade renderer logic 100 is a power hold X-Fade mix of the Com-Mix, Sports-Mix, and Stadium-Mix tracks, as described in FIG. 10C , based on the position of the user selected X-Fade slider 1282 FIG. 12C ) provided on line 225 as discussed herein. In that case, one end of the X-Fade slider 1282 provides only the Stadium-Mix (for those who do not want to hear the commentator dialog audio), a middle position of the X-Fade slider 1282 provides the Sports-Mix (with both the stadium and commentator for those who want to hear it as originally delivered from the broadcast truck), and the other end of the X-Fade slider 1282 provides only the Com-Mix (for those who do not want to hear the stadium background noise / non-commentary audio).
[0167] The three-input X-Fade renderer can be used with any audio input where the content provider desires to provide a controlled, selectable transition between two audio components in a given audio mix (e.g., Cin-Mix / Stadium-Mix, Cin-Mix / Cin-Mix-AD, etc.). In that case, one end of the X-Fade slider provides only the first audio component, the opposite end of the X-Fade slider provides only the second audio component, and the middle of the X-Fade slider provides a portion of both audio components. Other crossfade renderer configurations and equations described herein for the three-input and two-input X-Fade renderers can be used if desired, as long as they provide the same functionality and performance as described herein. Additionally, where possible, it is generally desirable to use the same M&E track for each input to the X-Fade renderer to minimize the risk of holes in the M&E caused by studio or other processing of a given mix or track. In particular, when creating a translation from an English original video to an alternative language (AL), blank spots (or holes or gaps) may be left in the audio due to differences in time or word count from English to the alternative language (AL). This is also true for non-English original content.
[0168] refer to FIG. 14C , a block diagram 1440 for a live sports application illustrates an embodiment of the present disclosure, similar to utilizing FIG. 2A-FIG. 2D The case in question, where the first input signal on line 96A has the FIG. 14B A live (or pre-recorded) sports mix is described where the commentator dialogue audio portion is that of the home sports team (Sports-Home-Mix) and the second input signal on line 96B has the same characteristics as described above using FIG. 14B A live (or pre-recorded) sports mix is described, where the commentator dialogue audio portion is for the away team (Sports-Away-Mix). The Sports-Home-Mix and Sports-Away-Mix inputs on lines 96A, 96B, respectively, are provided to the audio encoder 26, which provides an interleaved and encoded audio track (Sports-Home-Mix / Sports-Away-Mix) having Sports-Home-Mix and Sports-Away-Mix on line 94.
[0169] At the receiver or user / listener (right) side of Figure 1440, the Sports-Home-Mix / Sports-Away-Mix interleaved encoded audio track is received by decoder 202 on line 203, which separates the interleaved audio track into Sports-Home-Mix and Sports-Away-Mix on lines 207, 209, respectively, which are provided to the X-Fade renderer logic 100.
[0170] The output of the X-Fade renderer logic 100 is a power- preserving X-Fade mix of the Sports-Home-Mix and Sports-Away-Mix audio tracks as described herein, where the fade amount is based on the position of the user-selected X-Fade slider 1284 FIG. 12C ) provided on line 225, as discussed herein. In that case, one end of the X-Fade slider 1284 provides only the Sports-Home-Mix (for those who want to hear only the home team commentator dialog audio), the middle position of the X-Fade slider 1282 provides a mix of the Sports-Home-Mix and Sports-Away-Mix (with both home and away commentator dialog audio), and the other end of the X-Fade slider 1284 provides only the Sports-Away-Mix (for those who want to hear only the away or visiting team commentator dialog audio).
[0171] In some embodiments, the DCE 62 can be applied to the Sports-Home-Mix and Sports-Away-Mix input signals on lines 96A, 96B to enhance the audio dialog relative to the background noise or M&E, as indicated by the dashed box 62. In that case, the DCE parameters of the DCE 62 can be set based on the desired dialog audio, e.g., enhanced dialog relative to M&E or dialog only without M&E, as discussed herein.
[0172] Referring to FIG. 14D , a block diagram 1460 of the commentary application shows embodiments of the present disclosure, similar to the utilization of the X-Fade slider 1282 to provide a mix of the Sports-Home-Mix and Sports-Away-Mix audio tracks as described herein. FIG. 2A-FIG. 2DThe case in point, where the first input signal on line 96A has the original version (OV) cinematic mix (Cin-Mix-English or OV) in English from the studio, as discussed above, and the second input signal on line 96B has the alternate language (AL) cinematic mix (Cin-Mix-Spanish or AL) in English from the studio. The Cin-Mix-English (OV) and Cin-Mix-Spanish (AL) inputs on lines 96A, 96B are provided to the audio encoder 26, which provides an interleaved and encoded soundtrack (Cin-Mix-English / Cin-Mix-Spanish) with Cin-Mix-English and Cin-Mix-Spanish on line 94.
[0173] At the receiver or user / listener (right) side of Fig. 1460, the Cin-Mix-English / Cin-Mix-Spanish interleaved encoded soundtrack is received on line 203 by the decoder 202, which separates the interleaved soundtrack into Cin-Mix-English and Cin-Mix-Spanish on lines 207, 209, respectively, which are provided to the X-Fade renderer logic 100.
[0174] As described herein, the output of the X-Fade renderer logic 100 is a power- preserving X-Fade mix of the Cin-Mix-English and Cin-Mix-Spanish soundtracks, where the fade amount is based on the position of the user-selected X-Fade slider 1280( FIG. 12C ) provided on line 225, as discussed herein. In that case, one end of the X-Fade slider 1280 provides only Cin-Mix-English (OV) (for those who want to hear only the original version language dialogue audio), a middle position of the X-Fade slider 1280 provides a mix of Cin-Mix-English and Cin-Mix-Spanish (with both original version OV language dialogue and alternate language AL dialogue audio), and the other end of the X-Fade slider 1280 provides only Cin-Mix-Spanish (AL) (for those who want to hear only the AL dialogue audio).
[0175] In some embodiments, DCE 62 may be used for Cin-Mix-English and Cin-Mix-Spanish input signals on lines 96A, 96B to enhance audio dialogue relative to M&E, as indicated by dashed box 62. In that case, DCE parameters of DCE 62 may be set based on the desired dialogue audio, e.g., enhanced dialogue relative to M&E or dialogue only without M&E, as discussed herein.
[0176] refer to FIG. 14E , a block diagram 1480 for audio description (AD) or narration illustrates an embodiment of the present disclosure, similar to using FIG. 2A-FIG. 2D The case in question is one in which the first input signal on line 96A has the original version (OV) film mix in English from the studio (Cin-Mix-English or OV or Cin-Mix), as discussed above, and the second input signal on line 96B has the audio description from the studio (AD or Cin-Mix-AD). The Cin-Mix (OV) and Cin-Mix-AD (AD) inputs on lines 96A, 96B, respectively, are provided to the audio encoder 26, which provides an interleaved and encoded audio track (Cin-Mix / Cin-Mix-AD) with Cin-Mix and Cin-Mix-AD on line 94.
[0177] On the receiver or user / listener (right) side of Figure 1480, the Cin-Mix / Cin-Mix-AD interleaved encoded audio track is received on line 203 by decoder 202, which separates the interleaved audio track into Cin-Mix (OV) and Cin-Mix-AD on lines 207 and 209, respectively, which are provided to the X-Fade renderer logic 100 described herein.
[0178] The output of the X-Fade renderer logic 100 is a power-preserving X-Fade mix of the Cin-Mix (OV) and Cin-Mix-AD tracks provided on line 225, as described herein, where the amount of fade is based on the position of a user-selected X-Fade slider similar to slider 1280 ( FIG. 12C), except that the upper end of the slider can be said to provide only audio description (or only AD). In that case, one end of the X-Fade slider 1280 provides only Cin-Mix (OV) (for those who want to hear only the original version language dialogue audio), the middle position of the X-Fade slider 1280 provides a mix of Cin-Mix (OV) and Cin-Mix-AD (with both original version OV dialogue and audio description or voice-over (AD) dialogue audio), and the other end of the X-Fade slider 1280 provides only Cin-Mix-AD (for those who want to hear only the audio description or voice-over dialogue (AD) audio).
[0179] In some embodiments, the DCE 62 can be used on the Cin-Mix (OV) and Cin-Mix-AD input signals on lines 96A, 96B to enhance the audio dialogue relative to the M&E, as shown by the dashed box 62. In that case, the DCE parameters of the DCE 62 can be set based on the desired dialogue audio, e.g., enhanced dialogue relative to the M&E or dialogue only without the M&E, as discussed herein.
[0180] Referring to FIG. 15A and FIG. 15B , block diagrams are shown of various embodiments for streaming and different types of audio inputs to be adjustably rendered by a crossfade renderer for video on demand (VOD) and live sports / event (live) applications, respectively, in accordance with embodiments of the present disclosure.
[0181] Referring to FIG. 15A , block diagram 1502 for a video on demand (VOD) application shows embodiments of the present disclosure, similar to the cases discussed with respect to FIG. 14A , FIG. 14D and FIG. 14E , in which a film mix (Cin-Mix) from a studio is provided to a content provider (CP) system 10 on line 96 as discussed herein. FIG. 1). The Cin-Mix soundtrack on line 96 is provided to the DCE 62, which can be tuned (using DCE gains or parameters, as discussed herein) to provide an enhanced dialogue (Dx) mix, e.g., Max-Acc-Mix on line 92, as discussed herein, which can also be referred to herein generally as an accessible mix. The Max-Acc-Mix is provided on line 92 to the encoder 26, which also receives the Cin-Mix on line 96. The encoder 26 provides on line 94 an interleaved and encoded soundtrack with the Cin-Mix and the Max-Acc-Mix, as discussed herein. The accessible mix output of the DCE 62 (or Mac-Acc-Mix or Acc-Mix) can be an enhanced dialogue mix (EDx-Mix) = enhanced dialogue (Dx or EDx) + M&E, or an audio description mix (AD-Mix or Voice-Over Mix) = audio description dialogue (Dx(AD) or EADx or ENDx) + M&E, or an alternate language mix (AL-Mix or VO-Mix or Voice-Over-Mix) = alternate language dialogue (Dx(AL) or ALDx or VODx). The film mix is the film mix (Cin-Mix) = dialogue (original version or OV) + M&E, as discussed herein.
[0182] More specifically, for the audio description mix (AD-Mix) output on line 92 from the DCE 62, the DCE 62 can be configured or tuned to provide only the audio description dialogue (Dx(AD)) and M&E, without providing the main or talent dialogue. Further, for the alternate language mix output on line 92 from the DCE 62, the DCE 62 can be configured or tuned to provide only the alternate language dialogue (Dx(AL) or Dx(ALT)) and M&E, without providing the main original version (OV) language dialogue.
[0183] In some embodiments, as shown by dashed line 1506, the CP user / administrator 21 can be able to adjust the DCE parameters (as discussed herein) to optimize the Max-Acc-Mix. In that case, the Max-Acc-Mix and the Cin-Mix are provided to the X-Fade logic 100, and the user / administrator 21 controls the slider (as described herein), as shown by dashed line 1508, and the X-Fade logic 100 provides the X-Fade Mix on dashed line 1510 to the user / administrator to determine whether the DCE parameters are providing the desired audio experience.
[0184] The accessible mix or Acc-Mix is provided on line 92 to the audio encoder 26 which also receives the Cin-Mix on line 96. The encoder 26 provides on line 94 an interleaved and encoded soundtrack (Acc-Mix / Cin-Mix) with the Acc-Mix and Cin-Mix.
[0185] At the receiver or user / listener (right) side of Fig. 1502, the Acc-Mix / Cin-Mix interleaved encoded soundtrack is received on line 203 by the decoder 202 which separates the interleaved soundtrack into the Acc-Mix (EDx+M&E or Dx(AD)+M&E or Dx(ALT)+M&E) (e.g. enhanced dialogue, audio dialogue or alternative language) on line 207 and the Cin-Mix (Dx(OV)+M&E) on line 209 depending on the application, which are provided to the X-Fade renderer logic 100 described herein.
[0186] The output of the X-Fade renderer logic 100 is a power preserving X-Fade mix of the Acc-Mix and Cin-Mix soundtracks provided on line 225 as described herein, where the fade amount is based on the position of the user selected X-Fade slider 1504, which is similar to the slider described herein with FIG. 12C In that case, one end of the X-Fade slider provides only the Cin-Mix (OV) (for those who want to hear only the original version dialogue audio) and depending on the application, the opposite end of the X-Fade provides only the Acc-Mix (EDx+M&E or Dx(AD)+M&E or Dx(ALT)+M&E) (for those who want to hear only the accessible mix (Acc-Mix) for the given application), and depending on the application, the middle position of the X-Fade slider provides a mix of the Cin-Mix (OV) and the Acc-Mix (EDx+M&E or Dx(AD)+M&E or Dx(ALT)+M&E).
[0187] Referring to FIG. 15B , a block diagram 1520 for a live sports / event (live) application shows embodiments of the present disclosure, similar to the cases discussed with FIG. 14B and 14C , where the Sports-Home-Mix from the OB truck is provided on line 96 to the content provider (CP) system 10 ( FIG. 1 ) as discussed herein.
[0188] The Sports-Home-Mix track on line 96 is provided to the DCE 62, which can be tuned (using DCE gains or parameters, as discussed herein) to provide a home commentary mix (Com-Home-Mix) on line 92 or a stadium mix (Stadium-Mix) on line 92A, as discussed above.
[0189] The Comm-Home-Mix or Stadium-Mix is provided on lines 92, 92A, respectively, to the audio encoder 26 (depending on which the content provider selected when tuning the DCE 62, as discussed herein), which also receives the Sports-Mix on line 96. The encoder 26 provides an interleaved and encoded track on line 94 with the Comm-Home-Mix or Stadium-Mix and the Sports-Home-Mix (Sports-Home-Mix / Com-Home-Mix or Sports-Home-Mix / Stadium-Mix).
[0190] Alternatively, the away sports mix track on line 96A is provided to the DCE 62, which can be tuned (using DCE gains or parameters, as discussed herein) to provide an away commentary mix (Com-Away-Mix) on line 92B or a stadium mix (Stadium-Mix) on line 92C, as discussed above.
[0191] In that case, the Comm-Away-Mix or Stadium-Mix is provided on lines 92B, 92C, respectively, to the audio encoder 26 (depending on which the content provider selected when tuning the DCE 62, as discussed herein), which also receives the Sports-Mix on line 96. The encoder 26 provides an interleaved and encoded track on line 94 with the Comm-Away-Mix or Stadium-Mix and the Sports-Away-Mix (Sports-Away-Mix / Com-Away-Mix or Sports-Away-Mix / Stadium-Mix).
[0192] Alternatively, Sports-Home-Mix and Sports-Away-Mix inputs are provided on lines 96, 96A, respectively, to audio encoder 26, which provides on line 94 an interleaved and encoded audio track with Sports-Home-Mix and Sports-Away-Mix (Sports-Home-Mix / Sports-Away-Mix).
[0193] In some embodiments, as shown by dashed line 1526, CP user / administrator 21 can be able to adjust DCE parameters (as discussed herein) to provide the desired live sports / event mix described above. In that case, the mix on lines 92, 92A, 92B, 92C, 96, 96A can be provided to X-Fade logic 100, and user / administrator 21 controls X-Fade slider 1529 (as described herein), as shown by dashed line 1528, and X-Fade logic 100 provides the X-Fade mix on dashed line 1527 to user / administrator 21 to determine whether the DCE parameters are providing the desired audio experience.
[0194] On the receiver or user / listener (right) side of FIG. 1520, in upper portion 1530, Sports-Home-Mix / Com-Home-Mix or Sports-Home-Mix / Stadium-Mix interleaved encoded audio tracks are received on line 203 by decoder 202, which separates the interleaved audio tracks into Com-Home-Mix or Stadium-Mix on line 207 (depending on which is selected) and Sports-Home-Mix on line 209, which are provided to X-Fade renderer logic 100 described herein, which has X-Fade slider 1522, which crossfades between Sports-Home-Mix and Com-Home-Mix or Stadium-Mix (depending on which is selected) based on the position of X-Fade slider 1522, similar to the use of FIG. 14B Sports-Home-Mix / Sports-Away-Mix described above. Similar X-Fades can be made for Sports-Away-Mix / Com-Away-Mix or Sports-Away-Mix / Stadium-Mix.
[0195] In the receiver or user / listener (right) side of Figure 1520, in lower portion 1532, the Sports-Home-Mix / Sports-Away-Mix interleaved encoded audio track is received by decoder 202 on line 203, which separates the interleaved audio track into Sports-Home-Mix on line 207 and Sports-Away-Mix on line 209, which are provided to the X-Fade renderer logic 100 described herein, which has an X-Fade slider 1524 that crossfades between Sports-Home-Mix and Sports-Away-Mix based on the position of the X-Fade slider 1524, similar to utilizing FIG. 14C the situations described herein.
[0196] As discussed herein, the present disclosure uses a crossfade renderer or X-Fade 100 FIG. 1 , FIG. 2A-FIG. 2D and FIG. 14A-FIG. 14E as well as FIG. 15A-FIG. 15B to mix two (or more) input audio signals, which can generally be referred to as audio Mix A and audio Mix B. The X-Fade renderer has an X-Fade slider that can be moved (automatically or manually) from an initial position of the X-Fade slider that provides an audio output of only "Mix-A" (initial or first end) to a final position of the slider that provides a crossfade of only "Mix-B" (opposite or second end), and that provides an infinite number of positions between the initial (first end) and final (second end) positions, each of which corresponds to a power- held mix of audio Mix A and audio Mix B.
[0197] As discussed herein, audio MixA and audio MixB may be any two audio signals that are desired to be mixed or faded from a first signal (MixA) to a second signal (MixB), for example, MixA and MixB may be as follows: MixA = English original version (or Cin-Mix) and MixB = Enhanced Dialogue English (Max-Acc-Mix); MixA = English original version; MixB = English audio description; MixA = English original version and MixB = Spanish audio voice-over; MixA = English original version and MixB = Spanish audio description; MixA = Home team commentary mix and MixB = Away team commentary mix; MixA = Home team commentary mix and MixB = Stadium sound (no commentary) mix. Other audio mixes may be used as inputs to the X-Fade logic 100 if desired, however, it may be desirable that the M&E portions of both the audio of MixA and the audio of MixB be similar or substantially identical to optimize the listening experience.
[0198] When the input is a Cin-Mix (a studio VOD content application), the label Max-Acc-Mix may be used herein to refer to the maximum accessible mix for the dialog-enhanced output of the DCE. However, it may also be used herein for the endpoints of the X-Fade slider for any of the other embodiments or applications discussed herein. Additionally, the term Maximum Personalized Mix (or Max-Pers-Mix) may also be used. However, other labels for the maximum or upper endpoint of the X-Fade slider may be used for other applications, such as FIG. 14B-FIG. 14E and those shown in 15B, such as: Max-Com-Mix for maximum commentary mix for live sports content, Alt-Com-Mix for alternative commentary mix for live sports content, Max-AD-Mix for maximum audio description mix for audio description of any live / VOD content, and Max-VO-Mix for maximum voice-over mix for any live or VOD content. Depending on the application, other labels may be used if desired. In some embodiments, the terms MixA and MixB may be used to generally describe two audio mixes or tracks that define the lower boundary (or first end) and relative upper boundary (or second end) of the X-Fade slider for a dual-input X-Fade renderer, as described herein. Furthermore, for the above voice-over example of MixA and MixB, English may be replaced by any primary (or primary) language desired by the user, and Spanish may be replaced by any alternative language (other than the primary language) desired by the user, as long as the content owner or content provider provides the mixes.
[0199] Further, while X-Fade has been described as using sine and cosine functions to mix two (or three) input signals to produce a power-preserving output mix or X-Fade mix, those skilled in the art will appreciate that other cross-fade methods can be used if desired, such as the methods described below.
[0200] In some embodiments, the X-Fade logic 100 can include step switching and smoothing. This method uses step switching, where the transition from mix A to mix B is smoothed in different steps, rather than a continuous crossfade. It applies a short smoothing ramp (e.g., 200ms - 500ms) to avoid sudden jumps in volume, making the transition feel more natural. More specifically, with this method, when the transition is triggered, mix A fades out quickly, while mix B fades in over a short period (e.g., 200ms - 500ms). The fade curve can be linear, logarithmic, or s-shaped to optimize the perceived smoothness. Further, the transition can be applied to natural switching points, such as sentence boundaries in dialog, scene cuts, or pauses in background music. In some embodiments, a delay or hold function can be used, which can prevent fast, repetitive switching, ensuring a smoother experience. When compared to traditional cross-fading, such as that utilized herein with reference to FIG. 1, there are certain advantages and disadvantages associated with this method. For the advantages, this method reduces the "mixed audio" period, minimizing the moment where both mixes overlap unnaturally. Further, it works well with different audio streams (e.g., original audio vs. dubbed audio, commentary vs. stadium sound), and prevents the speech intelligibility issues that arise when both sounds are heard simultaneously during a slow crossfade. For the disadvantages, when both mixes contain similar elements (e.g., different language versions of the same dialog), this method is not as smooth as gradual cross-fading, and requires careful selection of the transition timing to avoid jarring switches. FIG. 10A-FIG. 10C
[0201] In some embodiments, the X-Fade logic 100 can include intelligent gain-based mixing (not shown). This method applies frequency-dependent gain interpolation, rather than a uniform crossfade across the entire frequency spectrum. It prioritizes certain frequency bands based on the audio content, making the transition feel more natural. More specifically, with this method, instead of fading all frequencies equally, the method applies different gain curves to different frequency bands. In particular, for multi-dialog content, frequencies between 1 kHz and 4 kHz (the speech intelligibility range) are faded more smoothly, and low-frequency elements (background music, effects) are faded faster or held temporarily to avoid gaps. For music or environmental sounds, low and mid frequencies are transitioned first, while high frequencies are faded in gradually to maintain clarity. When compared to traditional cross-fading, such as that utilized herein with reference to FIG. 1, there are certain advantages and disadvantages associated with this method. For the advantages, this method is more flexible and can be tailored to different audio streams (e.g., original audio vs. dubbed audio, commentary vs. stadium sound), and can be used to prioritize certain frequencies to maintain speech intelligibility. For the disadvantages, this method is more complex to implement than traditional cross-fading, and can require more processing power. FIG. 10A-FIG. 10C As shown, there are certain advantages and disadvantages associated with this approach. For advantages, this approach prevents the "fuzzy" middle transition caused by overlapping full-spectrum audio. In addition, it improves speech intelligibility by prioritizing dialog clarity over background noise. Furthermore, it can create more natural transitions between different types of content. For disadvantages, this approach requires real-time frequency analysis, making it computationally more expensive. In addition, if not carefully tuned, it can introduce unintended artifacts.
[0202] In some embodiments, the X-Fade logic 100 can include adaptive cross-fade with content awareness (not shown). This approach uses real-time audio analysis to determine the optimal fade shape and duration based on the similarity of mix A and mix B. More specifically, with this approach, the system analyzes audio characteristics (e.g., spectral content, loudness, speech presence). If mix A and mix B contain similar elements (e.g., both are multi-dialog), a longer, more gradual fade is used. If mix B introduces a completely different element (e.g., moving from spoken commentary to crowd noise), a faster fade is applied to avoid confusion. In some embodiments, a machine learning model can be used to classify the audio types and dynamically adjust the fade parameters based on past data. When compared to traditional cross-fades, such as utilized herein FIG. 10A-FIG. 10C As shown, there are certain advantages and disadvantages associated with this approach. For advantages, this approach reduces perceptual disruption by adapting the transition to the content. In addition, it avoids awkward overlaps between similar content types. Furthermore, it can be optimized for accessibility use cases (e.g., ensuring clear speech during the transition). For disadvantages, this approach requires real-time analysis and adaptive decision-making, making it computationally more burdensome. In addition, it is more difficult to predict or manually control compared to fixed cross-fades.
[0203] In some embodiments, the X-Fade logic 100 can include multi-channel matrix mixing with dynamic weighting (not shown). In this approach, instead of a cross-fade between two stereo streams, this approach keeps all potential audio mixes available in a multi-channel format (e.g., 4.0, 5.1) and dynamically adjusts their levels based on user preferences. More specifically, with this approach, all audio options (e.g., original mix, enhanced dialog mix, Spanish mix) are saved in a multi-channel format. In addition, a matrix mixer dynamically adjusts the gain of each mix, blending between them without actually fading them in and out. Furthermore, it can be controlled using a crossfader UI, automation, or an AI-driven recommendation system. When compared to traditional cross-fades, such as utilized herein FIG. 10A-FIG. 10CAs shown, there are certain advantages and disadvantages associated with this approach. For advantages, this approach allows for instantaneous and seamless transitions without noticeable crossfade artifacts. In addition, it reduces the need for destructive crossfading, thus preserving audio quality and being more flexible for users who want to fine-tune their experience (e.g., 75% of original mix + 25% of enhanced dialogue). For disadvantages, this approach requires a playback system capable of handling multi-channel audio and can be more complex to implement in a traditional stereo playback environment.
[0204] In some embodiments, the X-Fade logic 100 can include a spectral cross adaptive mixing with dynamic weighting (not shown). This is an approach that analyzes the spectral components of mix A and mix B and cross-adapts them to create a smooth transition with phase coherence. More specifically, with this approach, a real-time frequency analysis of both mixes is performed. Overlapping spectral components are intelligently mixed instead of applying a simple volume-based crossfade. In addition, a phase coherence morphing technique is used to prevent phase cancellation artifacts. Furthermore, it can use machine learning to predict the optimal spectral adjustments for a seamless mix. When compared to traditional crossfading, such as utilized herein FIG. 10A-FIG. 10C As shown, there are certain advantages and disadvantages associated with this approach. For advantages, this approach prevents phase issues that can occur when mixing similar content, creates a seamless and transparent transition without noticeable overlap, and works well for complex audio environments that require different layers to be preserved. For disadvantages, this approach is computationally expensive and requires high-performance processing. In addition, it is more difficult to implement in real-time streaming applications.
[0205] Other techniques or approaches can be used for the X-Fade logic 100 if desired, as long as they provide the functionality and performance of the X-Fade logic described herein.
[0206] The present disclosure describes an accessible audio streaming system with an automatic personalized crossfade (per language) renderer on a user device. In some embodiments, as described herein, if no measurement (or sensing or recording) of the ambient noise floor (NF) is available (e.g., no microphone), the accessibility X-Fade renderer (or accessibility crossfader or crossfade slider) as described herein can become available to the user through the UI user device. In some embodiments, the X-Fade renderer can be available to the user independent of the sensing of the noise floor NF. In that case, the system can provide an initial X-Fade setting based on the measured NF (or NF capture), and if desired, the user can further adjust that setting for the user’s optimal listening experience.
[0207] Further, the present disclosure provides for personalized audio streaming. In some embodiments, there can be cross-fading between the Cin-Mix and Max-Acc-Mix audio tracks, which equalizes the non-dialog (or M&E), and in some embodiments, can also provide high-fidelity voice-over (VO), e.g., rendering two languages or a narration (AD), e.g., audio description dialog added to the main dialog, simultaneously, or automatic use of NF capture, or manual use of X-Fade slider by user / listener, as discussed herein.
[0208] In some embodiments, dialog isolation (DI) is performed with the rest of the audio mix, and the perceived loudness of the mix is preserved. Further, various known AI machine learning (ML) models (open source and commercial) can be used to extract dialog from the video stream, e.g.,, (Deezer Research), demucs (Meta), (AudioShake), etc., as discussed herein. In that case, the known ML models will be trained using existing datasets for extracting dialog from videos across a wide range of content.
[0209] For the ALN logic, a target loudness normalization gain (ALN gain) is applied to the accessible mix to ensure that all elements of the accessible mix are above the noise floor, e.g., -16 LKFS is the industry standard for mobile noisy environments. Further, the ALN gain can be a function of the content type, e.g., soft / quiet scenes, action / noisy scenes, etc. Further, the same processing will apply to Dolby ATMOS (with more channels, which do not need to be modified).
[0210] In some embodiments, the input audio signal can be in 5.1 or 5.1.4 immersive sound, including Dolby Atmos. In that case, a dialog intelligibility engine (DCE) including dialog enhancement (DE) and dynamic range control (DRC) can also handle these formats. In some embodiments, the dialog intelligibility engine (DCE) includes a dialog aware remix engine (or enhanced dialog insertion (EDI)) for better dialog intelligibility surround audio 5.1 delivery, along with a combination of dialog enhancement (DE), dynamic range control or compression (DRC), or audio loudness normalization (ALN), as discussed herein.
[0211] In some embodiments, for example, when the input audio is 5.1 surround sound or immersive 5.1.4 and all loudspeaker layouts, the story-telling dialog (or Dx) (from actors or talent reading a script) can already be separated from the video on its own channel (e.g., center channel), in which case the dialog isolation (DI) logic is optional, and the DCE gain (or DCE parameter or compensation gain) can be applied directly to the center channel (C) of the audio input signal. In addition, during the final normalization (ALN) step, the DCE gain is also applied to all remaining channels (in addition to the C channel) to maintain overall loudness (surround normalization). Furthermore, the DCE gain (or DCE parameter) can be a function of the type of content, e.g., soft / quiet scenes, action / noisy scenes, etc. The same processing as described above for surround 5.1 and the like would apply to immersive sound such as Dolby ATMOS (i.e., sound setups with more channels and self-contained independent dialog objects).
[0212] The techniques of the present disclosure are device agnostic. Furthermore, for embodiments that require it, both dual or multi-audio decoding and audio capture are supported natively by and the operating system (OS), which significantly facilitates integration and testing of the present disclosure, and can also be provided in streaming service application software or APIs, which can include or be part of the application logic 16 discussed herein. Furthermore, the present disclosure supports HLS / DASH streaming protocols, which support multimedia streaming and associated bitrates for 2 stereo streams, which are equivalent to surround audio bitrates, and have been vetted from a CDN delivery perspective for over a decade, which have shown no negative impact on streaming quality of experience (QoE). Surround and immersive payload delivery would require 5.1+1 (surround + enhanced center channel (Max-Acc-Mix)) or 5.1.4+1 (immersive / Atmos + enhanced center channel (Max-Acc-Mix)). An alternative personalized Dolby audio delivery can simply apply to stereo creative mix or auto-stereo downmix from immersive / surround creative (near-field) mix, e.g., Dolby dual stereo encoding of (2.0 Cin-Mix + 2.0 Max-Acc-Mix (or Max-Pers-Mix)).
[0213] As discussed herein, in some embodiments, the system of the present disclosure uses measured (or captured) noise floor (NF or Accessibility Fader) levels to tune X-Fade rendering to an environment (e.g., airplane, car, bedroom, etc.) in real-time using both the original movie and the accessible audio track (or Max-Acc-Mix) (also referred to as Max-Pers-Mix).
[0214] In some embodiments, as discussed herein, the present disclosure provides that the environmental noise floor (NF) can be estimated while filtering out the human voice frequency range, the crossfader (left and right) can be power-preserving so that loudness and stereo / space image are not affected, DRC can be applied pre-encode and affect the M&E loudest and softest parts (non-dialog).
[0215] Further, as discussed herein, PL stands for Program Loudness as specified in ITU-R BS.1770-1. LRA or Loudness Range is a measure of the range between the loudest and softest parts in all channels of an audio track. DPL or Dialog to Program Loudness; LKFS = Loudness K-Weighted Full Scale. M&E stands for Music and Effects (i.e., sound effects or special effects, etc.; everything other than dialog / storytelling). In some embodiments, once DRC is applied, a target loudness normalization gain is applied to the accessible mix to ensure that all elements of the accessible mix are above the noise floor, e.g., -16 LKFS is the industry standard for mobile noisy environments, as discussed herein.
[0216] In some embodiments, the accessible stereo mix (or Mac-Acc-Mix) can be produced directly by the video recording studio and / or generated by the content provider or streaming service performing the media processing DCE, which can be scaled securely to the entire video on demand (VOD) catalog at the content holder’s site. Accessibility rendering (movie crossfader to DE + DRC or Max-Acc-Mix) can be scaled to all streaming media environments and devices.
[0217] As discussed herein, the X-Fade renderer can be driven by the measured (or captured) noise floor (NF) or manually controlled by the user / listener within the streaming service application with an accessibility slider or X-Fade renderer or X-Fade logic as discussed herein. For example, in a noisy environment, the X-Fade slider can be set such that the DRC effect in the DCE can be high and the effect of the DE is present or on. In another example, a user can be watching through a streaming device in their home and select an X-Fade slider position where the DRC can be low and the DE is on. Further, when the user is watching in their home in an optimal environment, the user can set the X-Fade slider position where the effect of the DRC is low (can be off or not active) and the user can select the slider position where the effect of the DE is only the amount needed for the user to understand the dialogue (i.e., accessible audio) based on preference.
[0218] The present disclosure can also be used with accessible stereo / downmix. In that case, the crossfade (or X-Fade) can be applied to each L / R channel of the 2.0 downmix. The DE is applied to the active dialogue of the 2.0 downmix. All channels of the mix should have the same DRC applied. A single binaural EC3 encoder / decoder can be used, such as a 2x2.0 independent encoded mix.
[0219] The present disclosure provides embodiments including accessible audio delivery for all Dolby capable devices. A smart TV or device manufacturer (OEM) that downmixes from Atmos 5.1.4 (with five speakers at ear level (left front, center, right front, left surround, right surround), a subwoofer, and four speakers for overhead audio (two front and two back overhead channels) to 3.0 (with three channels, left, right, and center) or 2.0 (with two channels, left and right) will work with the present disclosure but can reduce dialogue intelligibility. Further, an OEM that virtualizes Atmos 5.1.4 on a limited number of speakers will also work with the present disclosure but can reduce intelligibility (e.g., Dolby certification does measure intelligibility). Streaming Dolby encoding accessible streams that can be used with the present disclosure for “consumer grade” Dolby systems in the home can include the following options: Atmos / 5.1 <> 2.0 downmix, such as movie <> accessible; stereo 2.0 native mix, such as movie <> accessible; surround 5.1 native mix, such as center channel movie <> accessible; Atmos 5.1.4 native mix, such as center channel movie <> accessible.
[0220] The present disclosure also provides embodiments that include a single multi-channel audio encoder / decoder. MPEG AAC-LC and Dolby E-AC3 support "Dual Stereo" encoding modes, such as 2x 2.0 / stereo independent encoding. MPEG AAC-LC and Dolby E-AC3 support "Four Channel" encoding modes, such as 4x 1.0 / mono independent encoding. Dolby (E-AC3+ Joint Object Coding) = Atmos encoder supports primary audio and associated audio in a single bitstream. In some embodiments, the present invention can use known sources or Software tools for audio signal processing.
[0221] As discussed herein, the present disclosure includes embodiments that provide personalized accessibility that accommodates hearing ability. In that case, the accessible audio (or Max-Acc-Mix) is decoded and further equalized to compensate for each user's hearing loss frequency response. For example, the hearing loss function (e.g., audiogram) can be provided by a third-party audiogram API, device operating system (e.g., iOS® and the like), or streaming media service application listen test mode.
[0222] In some embodiments, as discussed herein, the present disclosure includes embodiments that include a system for providing personalized dubbing that accommodates a user's language skills (beginner, intermediate, expert), where the user can select the loudness level of the dubbing dialog. For example, an inexperienced English speaker can choose to listen to two languages with their native language in the background for assistance, such as a dubbing mode. Similarly, an intermediate English speaker can choose to listen to their primary language (or native language) in the background only for detected difficult sentences. The content based on the user's language ability can be provided by a content provider that has specific mixes for each desired dubbing level, thereby providing an adaptive experience for the user / listener.
[0223] As discussed herein, in some embodiments, the present disclosure includes a dialog clarity engine (DCE) 62 for stereo 2.0 audio delivery FIG. 3A ). In this embodiment, dialogue isolation (DI) logic extracts active story-telling dialogue from the rest of the mix (e.g., M&E) in both the left and right channels of the mix using a dual word-separation (ML) model. Dialogue enhancement (DE) logic applies amplification to the separated story-telling dialogue L / R channels and attenuation to the residual L / R channels such that the original program loudness is maintained. In that case, dynamic range compression (DRC) logic can be applied only to the residual channels. Stereo remix (or enhanced dialogue insertion - EDI) logic combines the left dialogue (or vocal) and M&E (or residual) channels together to generate dialogue enhanced left / right channels. Finally, in some embodiments, accessibility target loudness (ATL) logic can drive attenuated residual DRC and final audio loudness normalization (ALN) steps.
[0224] As discussed herein, the present disclosure also includes embodiments that provide a dialogue intelligibility engine for receiving audio in surround sound 5.1 audio delivery. In that case, dialogue isolation (DI) can extract active story-telling dialogue from the rest of the mix (e.g., M&E) in the center channel of the mix using a dual word-separation (ML) model. Dialogue in the other audio channels can be considered to be sound effects, rather than story-telling. Dialogue enhancement (DE) logic can apply amplification to the separated story-telling dialogue channel (vocal) and attenuation to the residual channel such that the original program loudness is maintained. Dynamic range compression (DRC) logic can be applied only to the attenuated residual channel. Mono remix (or enhanced dialogue insertion - EDI) logic steps combine the vocal and residual channels into a dialogue enhanced audio channel.
[0225] The present disclosure will also work with input audio having immersive 5.1.4 audio delivery and DOLBY Atmos 5.1.4 and DOLBY Surround 5.1. In that case, a cross-fade (X-Fade) is applied to the center channel of the mix where there is active story-telling dialogue (other channel dialogue is assumed to be sound M&E or FX specific). In that case, DE logic is applied to the active dialogue of the movie center channel and, in some embodiments, an additional single mono EC3 encoder / decoder can be used. The associated audio encoding bs mode of the EC3 encoder can be used to deliver a "accessible center channel".
[0226] As discussed herein, the present disclosure provides embodiments of a dialogue intelligibility engine (DCE) 62 that can have 2.0 stereo audio as input audio signals and can provide dynamic range control of "dialogue enhanced" movie stereo as FIG. 3EAs shown in FIG. 3, the DCE model #3). In such an embodiment, dialogue isolation (DI) logic extracts active story-telling dialogue from the rest of the mix (e.g., M&E) in the L&R channels. Dialogue enhancement (DE) logic applies amplification to the isolated story-telling dialogue L&R channels and attenuation to the residual L&R channels. The amplified dialogue and attenuated residual can be remixed after the DE logic. Next, dynamic range compression (DRC) logic can be applied to the dialogue-enhanced stereo audio, e.g., to the M&E to compress the M&E loudness range. Audio loudness normalization (ALN) logic can be applied to normalize the output stereo to a desired target loudness, e.g., -16 LKFS or the same LKFS as the integrated loudness of the source measurement. The system can use a dual stem separation model, where the story-telling dialogue is extracted from the L&R channels and the residual is extracted from the M&E channels. Manufactured Products may be used to provide audio separation.
[0227] As discussed herein, the present disclosure provides an alternative embodiment of a dialogue intelligibility engine for 2.0 stereo audio: dialogue enhancement of a “DRC’d” movie stereo. In this embodiment, dynamic range compression (DRC) is applied to the source movie mix. Then, dialogue isolation (DI) extracts active story-telling dialogue from the rest of the mix (e.g., M&E). Next, dialogue enhancement (DE) applies amplification to the isolated story-telling dialogue L&R channels and attenuation to the residual L&R channels. The amplified dialogue and attenuated residual are mixed after DE but before applying ALN. Audio loudness normalization (ALN) is applied to normalize the output stereo to a desired target loudness, e.g., 16 LKFS or the same LKFS as the integrated loudness of the source measurement.
[0228] As discussed herein, the present disclosure provides an alternative embodiment of a dialogue intelligibility engine for surround and immersive audio dynamic range control of a “dialogue-enhanced” center channel. In this embodiment, dialogue isolation (DI) uses a dual stem separation (ML) model to extract active story-telling dialogue from the rest of the mix (e.g., M&E) in the center channel. Dialogue in the other audio channels can be considered to be sound effects, rather than story-telling. Dialogue enhancement (DE) applies amplification to the isolated story-telling dialogue channel and attenuation to the residual channels, such that the original program loudness is maintained. Dynamic range compression (DRC) is applied to the dialogue-enhanced center channel resulting from the sum of the amplified dialogue and attenuated residual. Audio loudness normalization (ALN) is applied to normalize the “DE+DRC enhanced” center channel to a desired target loudness, e.g., -16 LKFS or the same LKFS as the integrated loudness of the center channel measurement.
[0229] As discussed herein, the present disclosure provides an alternative embodiment of a dialog intelligibility engine for surround sound and immersive audio dialog enhancement of a centrally panned channel that is "DRCed." In this embodiment, dynamic range compression (DRC) is applied to the source central channel. Dialog isolation (DI) uses a dual word-separation (ML) model to extract active story-telling dialog from the rest of the mix of the central channel (e.g., M&E). Dialog in the other audio channels can be considered sound effects, rather than story-telling spoken each time (or words that are crucial to the story). Dialog enhancement (DE) applies amplification to the separated story-telling dialog channel and attenuation to the residual channel, such that the original program loudness is maintained. Audio loudness normalization (ALN) can be applied to normalize the "DRC+DE enhanced" central channel to a desired target loudness, such as -16 LKFS or the same LKFS as the integrated loudness measured for the central channel.
[0230] As discussed herein, the present invention allows for providing accessible audio, meaning that a listener can hear and understand the dialog being spoken in the content. In some embodiments, one option to activate accessible audio includes automatically activating and a streaming software application running on, for example, a smartphone, laptop, smart television, or other smart playback device, at which point the device can access a device microphone that is capable of measuring the ambient noise and adjusting the cross-fade renderer (X-Fade) to provide a desired listening experience.
[0231] In some embodiments, where there is no access to a microphone or similar device, as discussed herein, another option to activate accessible audio is for the user to use the remote control of the smart television to activate the accessibility renderer implemented within the streaming service application on the device, as discussed herein. This indicates that the user is capable of using the remote (or similar device) paired with the streaming service application to adjust the accessibility features. For example, variations of a short or long press of different buttons on the remote (e.g., volume up / down buttons) can be used to control different functions, such as the accessibility renderer. This embodiment can be beneficial when the user's device does not have a microphone capable of measuring the ambient noise, etc. The customization of the functions of the remote can be tailored to accommodate the user's ability to control the accessibility renderer (or Acc application and decoder(s)), which can be part of the smart television OS streaming service application.
[0232] As discussed herein, the present disclosure provides an accessibility cross-fader (or X-Fade), which can be displayed in the UI (or UX, user experience) (see FIG. 12A-FIG. 12CThe UI can be part of a streaming service application (e.g., Apple TV+, Hulu, Amazon Prime Video, etc.) to activate the accessibility renderer. Further, there are many ways to implement an accessibility renderer or X-Fade renderer. This disclosure shows a power-preserving cross-fade renderer that can be used to implement an accessibility renderer. Other equations or configurations or techniques can be used for X-Faders if desired, some of which are discussed above.
[0233] In some embodiments, as discussed herein, an accessibility toggle can be activated on the UI / UX (user interface / user experience) of any streaming application that delivers audio (premium entertainment, podcasts, music, etc.) to allow the end user to control the degree to which they want the content to become accessible with the help of an accessibility cross-fader (or X-Fade) that mixes both the movie audio and the accessible audio (Mac-Acc-Mix or DRC+DE processed) in real-time according to a limited set of predefined accessibility levels that are pre-rendered and available at the origin, on launch, or as a default condition. In one example, 5 stereo streams are pre-rendered and available for streaming to the end user according to the accessibility level selected on the X-Fade or UI / UX cross-fader. This post-decoding toggle solution can be applied to every language delivered to the end user and does not require all languages to have an accessibility payload - this solution is fully backward compatible with legacy ABR (adaptive bitrate) toggle technology and the payload remains unchanged. In some embodiments, the “accessibility toggle” can be a UI / UX element that allows the user to select a pre-rendered accessible audio stream. The “legacy ABR” (adaptive bitrate) toggle is an audio toggle technology known in the art that is used by streaming services to select the best available stream by the device based on the device capabilities, available network speed, customer membership plan (ad or extra fee, etc.).
[0234] As discussed herein, in some embodiments, the present disclosure can have only 2 pre-rendered streams: the movie and Max-Acc-Mix (or “drc+de” mix) accessible audio streams that utilize only 2 streams / stereo pairs to implement the overall accessibility rendering. However, in alternative embodiments, the renderer can select 2 out of 3 pre-rendered streams: the movie (Cin-Mix), Mid-Acc-Mix (or “drc only” mix) accessible, and Max-Acc-Mix (or “drc+de” mix) accessible audio streams that rely on a three-input progressive accessibility cross-fader with a 3 gain structure, as shown in FIG. 10CAs shown in FIG. 1, in some embodiments, the accessibility renderer can be a crossfade renderer that can access 3 pre-rendered accessible audio streams and can render multiple accessibility layers in real-time more progressively, as discussed herein.
[0235] As discussed herein, the present disclosure provides an option to activate accessible audio through user control of the accessibility renderer of a streaming service application by using a device (e.g., smartphone, tablet, etc.) paired with the streaming service application. This embodiment can be beneficial when the user’s device does not have a microphone capable of measuring the floor noise, etc.
[0236] According to embodiments of the present disclosure, the accessibility renderer (or X-Fade) can be activated on the UI / UX of any streaming application that delivers audio (premium entertainment, podcasts, music, etc.) to the end user, who can very precisely and continuously control the degree to which they want the content to become accessible by means of the accessibility crossfader (or X-Fade), which mixes both the movie audio and the accessible audio (Mac-Acc-Mix or DRC+DE processed) in real-time. Such an audio post-decoding rendering solution is applicable to every language delivered to the end user and does not require all languages to have an accessibility payload - this solution is fully backward compatible with traditional ABR switching / decoding techniques. However, the accessibility payload is twice the traditional payload because a minimum of 2 stereo pairs should be delivered to the renderer. In some embodiments, the “total” accessibility crossfader can be a crossfade renderer that can only access 2 pre-rendered accessible audio streams (1 stream is accessible and the other is not accessible) and can render all accessibility layers in real-time, e.g., all at once.
[0237] As discussed herein, the present disclosure provides for NF adaptation and full audio accessibility, Dialogue Enhancement (DE) and Dynamic Range Control (DRC) are combined. In that case, the ambient noise floor (NF) can be estimated while filtering out the human voice frequency range. The crossfade (left and right) should be power preserving so that loudness and stereo / space image are fully preserved. DRC is applied to the pre-encoded and affects the M&E loudest and softest parts (not the dialogue). Once DRC is applied, a target loudness normalization gain is applied to the accessible mix to ensure that all sound elements of the accessible mix are above the noise floor, e.g. -16 LKFS is the industry standard for noisy environments. PL stands for Program Loudness in LKFS as specified in ITU-R BS.1770-1. LRA or Loudness Range represents a measure of the range between the loudest and softest parts of an audio track in all of its channels. DPL or Dialogue to Program Loudness in LKFS is specified in ITU-R BS.1770-1.
[0238] In some embodiments, Dynamic Range Control (DRC) and Dialogue Enhancement (DE) can be combined. In that case, DRC can be applied first with a target loudness normalization gain to ensure that all sound elements of the accessible mix are above the "worst case" noise floor (e.g. -16 LKFS is the industry standard for noisy environments). After DRC, DE is applied as a final enhancement to further boost the dialogue relative to the M&E.
[0239] In some embodiments, the present disclosure includes providing personalized accessibility of adaptive hearing capabilities. Here, the accessible audio (e.g., Max-Acc-Mix) is decoded and rendered with the movie audio (1) and then further equalized (2) to compensate for the hearing loss frequency response of each user of a given streaming service. The hearing loss function can be provided (for compensation by personalized EQ) by: a third-party audiogram API (such as ); a device operating system (e.g., iOS, etc.); and / or a streaming service application listening test mode.
[0240] In some embodiments, as described herein, the present disclosure includes providing accessible and personalized audio delivery for voice on demand (VOD). The accessible stereo mix can be produced by the studio and / or generated by a media processing DCE, which can be scaled safely to the entire VOD catalog at the site. The accessibility rendering of Max-Acc-Mix (or DE + DRC) (movie (or Cin-Mix) crossfade scaling to all streaming environments and devices. The X-Fade renderer is driven by the captured NF, or by the user in the streaming service application (e.g., Apple TV, etc.). FIG. 2A-FIG. 2DThe Acc app or the central Acc app 16) with an accessibility fader or X-Fader to control. Personalized EQ (PEQ) can be applied in the playback experience after (or post) the X-Fade renderer on a per user basis. Hearing loss functionality can be measured during streaming service distribution setup or provided by the device OS. PEQ compensates for the hearing loss functionality of the active member distribution by boosting certain frequency ranges in the X-Fade mix to enable personalized hearing adjustment.
[0241] In some embodiments, the present disclosure includes a method for providing a multi-channel accessible stereo encoder (MASE) according to embodiments of the present disclosure. MPEGAAC-LC and Dolby E-AC3 support a “dual stereo” encoding mode, such as 2x2.0 / stereo independent encoding. Dolby (E-AC3+ joint object encoding) = Atmos encoder supports primary audio and associated audio in a single bitstream. As discussed herein, known open source software audio processing tools, such as or , can be utilized to provide native support for implementing the present disclosure.
[0242] The following defined terms and concepts can be referenced throughout this disclosure and can be a component of the systems and methods and / or various alternative embodiments of the systems and methods of the present disclosure as well as the aforementioned co-owned provisional patent applications.
[0243] A cross-fade renderer (which can also be referred to herein as an X-Fade, X-Fader, renderer, cross-fade, or CFR) combines or mixes multiple audio files or experiences with a power- preserved amplitude panning curve, rather than abruptly switching between audio streams with different functionality or focus. In one embodiment, this involves taking input of multiple audio assets or files or content with multiple audio channels in WAV / PCM format, where the resulting output is 1 audio file with multiple audio channels in WAV / PCM format. The cross-fade renderer can use software running on a streaming service (or content provider) app, as discussed herein, running on and the like.
[0244] A dialog clarity engine (DCE) is a digital audio signal processing engine that performs one or more of the following: (1) dialog isolation (DI) from other elements of the mix (typically implying a source separation technique with a neural network); (2) dialog enhancement (DE) in terms of dialog lift, for increasing overall dialog intelligibility; and (3) dynamic range control (DRC) of the mix to make dialog intelligible on any size of speaker and any type of noisy environment. The input to the engine is typically an audio asset or content with active dialog (for story telling), where the resulting output is an audio asset or content file with elevated active dialog, such that the overall loudness of the mix remains unchanged compared to the input audio source. As discussed herein, the DCE can be performed using software running on hardware owned / provided by the streaming service.
[0245] Dialog isolation (DI) includes isolation of active story-telling dialog from an audio source mix. In particular, dialog isolation (DI) involves using software (typically a machine learning (ML) model, such as an open source artificial intelligence (AI) or machine learning (ML) model like Meta’s Demucs) running on the streaming service (or content provider) hardware to isolate active story-telling dialog from an asset audio source mix. DI involves input of an asset audio source in uncompressed PCM format (e.g., N = 2ch, 6ch, 10ch, etc.), where the resulting output is audio with only dialog content (e.g., K = 2ch, 1ch, 1ch, etc.) and audio with only non-story-telling dialog (or non-dialog or music and effects or M&E) content (e.g., residual, L = 2ch, 5ch, 9ch, etc.).
[0246] Dialog enhancement (DE) applies amplification to the isolated story-telling dialog channel and attenuation to the residual channel, such that the original program loudness is maintained. DE involves digital audio signal analysis and processing for boosting active dialog (for story telling) and attenuating other components of the mix to make the dialog more prominent without modifying the overall loudness of the resulting mix compared to the source or input audio mix. The input to DE is an audio asset with active dialog (for story telling), where the resulting output is an audio asset with elevated active dialog (for story telling), such that the overall loudness of the mix remains unchanged. As discussed herein, DE can be performed by software running on hardware owned / provided by the streaming service.
[0247] Dynamic Range Compression / Control (DRC) involves the analysis and processing of digital audio signals to limit the span / range of the amplitude of the audio samples to avoid the listener experiencing drastic loudness jumps from quiet to loud scenes. The terms compression and control are used interchangeably for the purposes of DRC. DRC involves the input of a full dynamic range audio asset, where the resulting output is a compressed dynamic range audio asset. DRC is implemented through the use of software running on hardware owned / provided by the streaming service. Examples of applied DRC profiles include: “Film Standard”, “Film Light”, or “Noisy Environment”, etc.
[0248] Enhanced Dialogue Insertion (EDI) (or stereo EDI or stereo remix) uses known audio signal combining software to separately combine the left channel of the dialogue (or vocal) mix and the residual (or M&E) mix and the right channel of the dialogue (or vocal) mix and the residual (or M&E) mix to generate dialogue enhanced left / right (L / R) channels. The input to the stereo EDI or remix can be either (1) the output of the DRC+DE processing, or (2) the output of the DE+DRC processing, where the resulting output of the EDI logic is a fully accessible audio stereo pair L / R.
[0249] Loudness Range (LRA) is a measure of the range between the loudest and softest parts of all channels of an audio track. Integrated Loudness (IL) is defined by the LoudLAB company at https: / / www.loudlab-app.com / sonicatom / en / 2.html, which defines IL and various other audio terms described herein.
[0250] Accessibility Target Loudness (ATL) involves the use of software to normalize audio channels to a target audio loudness in units of LKFS. The input is a stereo pair L / R channels and a target LKFS value (typically a negative number such as -20 or -16 LKFS), where the resulting output is normalized stereo pair L / R with a measured audio loudness at the target LKFS value. The term gain can also represent the accessibility target loudness (a value in units of LKFS).
[0251] Audio Loudness Normalization (ALN) is used to set expectations about how loud the ALN audio output can become with respect to the level settings when played on consumer audio electronics / devices. By applying ALN, the integrated audio loudness of the processed audio will average to X LKFS over the entire duration of the asset, where X is set by the content provider in conjunction with best practices and industry standards. The input to ALN is the accessible audio track (or Max-Acc-Mix) consisting of audio processed by DRC+DE or DE+DRC from DCE 62, or 2 channels for stereo sources or 1 channel (center) for surround / immersive audio sources. The output of ALN is the normalized audio output of ALN and DCE, which is 2 channels for stereo sources or 1 channel (center) for surround / immersive audio sources. ALN involves the use of software running on hardware owned / provided by the content provider / streaming service. Audio Loudness Normalization (ALN) is applied to normalize the output stereo to a desired target loudness, such as -16 LKFS or the same LKFS as the integrated loudness (IL) measured for the source.
[0252] In some embodiments, the present disclosure includes at least a method of increasing the audio dialogue level of digital cinema content while preserving the creative intent of the movie content, comprising: receiving a movie audio signal; receiving an accessible audio signal (e.g., Max-Acc-Mix); and combining the movie audio (Cin-Mix) signal and the accessible audio (Max-Acc-Mix) signal to provide an improved dialogue audio signal that allows a listener to hear dialogue over the background noise.
[0253] The systems and methods of the present disclosure will work with any codec that provides the functionality and performance described herein. Examples of codecs (or encoding and decoding formats) that can be used with the present invention are shown at: https: / / en.wikipedia.org / wiki / Comparison_of_audio_coding_formats or from the known IAMF standard codecs of AOM at https: / / aomedia.org / specifications / iamf / As is well known, a "codec" is a hardware- or software-based process that compresses and decompresses large amounts of data. Codecs are used in applications such as in the present disclosure to efficiently send media files over a network, as well as to receive and play media files for users on the receiving end.
[0254] The term "accessible" audio mix or audio track as used herein means an audio track with dialog intelligibility in all noise environments, devices, and capabilities so that the dialog can be understood as if the listener is in the movie experience (without subtitles). Further, in this disclosure, the audio or sound units of LKFS, dB, and TruePeak are used in various examples herein for illustrative purposes, and one of skill in the art should understand how the units relate to each other.
[0255] Further, while some embodiments of the system of the present disclosure are shown as storing encoded (and interleaved) digital audio tracks in an encoding server at the content provider side, and then retrieving the digital audio tracks for use at the user / listener side, it should be understood that the encoded digital audio tracks can alternatively or additionally be sent directly (e.g., by the encoder or other hardware or software) over a communication network to the user / listener device(s), which decodes the digital audio tracks for use by the user playback device, as described herein.
[0256] Embodiments of the present disclosure can also include the methods described herein, where creating an accessible audio signal includes using a dialog intelligibility engine. Embodiments of the present disclosure can also include the methods described above, where the dialog intelligibility engine includes: receiving a movie audio signal; and performing at least one of: dialog isolation (DI), dialog enhancement (DE), dynamic range compression (DRC), audio loudness normalization, dialog to program (or non-dialog or M&E or residual) loudness (DPL), mono-enhanced dialog insertion, or EDI (or remixing), and stereo EDI (or remixing). Embodiments of the present disclosure can also include the methods described above, where combining includes using a cross-fade renderer.
[0257] In some embodiments, the cross-fade experience can also rely on an accessible audio track mix (or accessible mix) generated by a mixer, rather than by a dialog intelligibility engine (DCE), which can be AI-based. Further, in some embodiments, to allow for broad compatibility with current production means from studios, the mixer (which generates the accessible mix or Max-Acc-Mix) can be customized or preset to worst-case scenarios, such as an elderly customer with a hearing impairment, poor TV speakers, poor sound connectivity, and / or high ambient background noise.
[0258] The systems and methods of the present disclosure include the necessary computers, servers, devices, etc., and the necessary electronics, computer processing power, interfaces, memory, hardware, software, firmware, logic / state machines, databases, microprocessors, communication links, displays or other visual or audible user interfaces, printing devices, and any other input / output interfaces to provide functionality or achieve the results described herein. Unless otherwise explicitly or implicitly indicated herein, the processes or method steps described herein can be implemented in software modules (or computer programs) that are executed on one or more general purpose computers. Special purpose hardware can alternatively be used to perform certain operations. Thus, any of the methods described herein can be performed by hardware, software, or any combination of the methods. Moreover, a computer readable storage medium can have stored thereon instructions that, when executed by a machine (such as a computer), cause performance of the steps in accordance with any of the embodiments described herein.
[0259] Further, the computers or computer-based devices described herein can include any number of computing devices capable of performing the functions described herein, including but not limited to: tablet computers, laptop computers, desktop computers, smartphones, smart televisions, set-top boxes, e-readers / players, etc.
[0260] Although the present disclosure has been described herein using exemplary techniques, algorithms or processes for implementing the present disclosure, those skilled in the art will appreciate that other techniques, algorithms and processes or other combinations and orders of techniques, algorithms and processes described herein can be used or performed to achieve the same function(s) and result(s) described herein and these are included within the scope of the present disclosure.
[0261] Any process descriptions, steps, or blocks in the process or logic flow diagrams provided herein represent ease of one possible implementation, and do not imply a fixed order or sequence of steps, and alternative implementations include those in which steps from two or more of the diagrams are performed in the same order as shown or discussed, in a different order from that shown or discussed, or in an order that is substantially different from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art.
[0262] It should be understood that any of the features, characteristics, alternatives or modifications described herein in relation to a particular embodiment can also be applied, used or combined with any of the other embodiments described herein, unless the context clearly indicates otherwise. Further, the drawings herein are not drawn to scale, unless otherwise indicated.
[0263] Conditional language, such as "can," "could," "might," or "may," unless specifically stated otherwise, or otherwise understood within the context as used, generally contemplated by the inventors that a certain embodiment can include, while other embodiments do not include, certain features, elements, or steps. Thus, such conditional language is not generally intended to imply that certain features, elements, or steps are in some way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, or steps are included or are to be performed in any particular embodiment.
[0264] While the present application has been described and illustrated herein with reference to exemplary embodiments thereof, it will be apparent to those of ordinary skill in the art that various alterations and modifications can be made thereto without departing from the spirit and scope of the present disclosure.
Claims
1. A computer-based method for generating a personalized accessible audio mix (AccMix) of audio content having dialog audio and non-dialog audio, the personalized accessible audio mix allowing at least one user to hear and understand the dialog audio when played on a user playback device in a listening environment having an actual noise floor (NF) level, the method comprising: receiving at the user playback device an original cinematic audio mix (CinMix) having the dialog audio and the non-dialog audio as set from a recording studio; receiving at the user playback device a maximum accessible mix (MaxAccMix) of the dialog audio and the non-dialog audio, the Max-Acc-Mix being an audio modification of the CinMix such that a loudness range of the dialog audio and a loudness range of the non-dialog audio is greater than a predetermined acceptable noise floor and less than a predetermined maximum audio loudness limit; and combining the CinMix and MaxAccMix to create the personalized AccMix audio mix, the personalized AccMix audio mix being power-conserving and having a loudness of the dialog audio that allows the user to hear and understand the dialog audio when played on the user playback device in the listening environment having the actual noise floor (NF) level, wherein the values of the personalized AccMix have a range from the CinMix values to the MaxAccMix values, and are power-conserving combinations of CinMix and MaxAccMix therebetween, and the personalized AccMix values are set based on at least one of user input and the actual noise floor level.
2. The computer-based method of claim 1, wherein, AccMix is automatically set based on the actual noise floor in the listening environment, the actual noise floor being measured by a sound sensor that filters out human voice from the audio signal to provide the actual noise floor level.
3. The computer-based method of claim 1, wherein, The CinMix and the Max-Acc-Mix are received in a single stream or digital file, and further comprising decoding the single stream to provide separate audio tracks of CinMix and Max-Acc-Mix.
4. The computer-based method of claim 1, wherein, AccMix is manually adjusted by a user interface using an X-Fade slider.
5. The computer-based method of claim 1, wherein, The at least one user includes a plurality of users, each user being able to set their own personalized AccMix.
6. The computer-based method of claim 1, wherein, The power-conserving combination of CinMix and MaxAccMix is performed using a cross-fade renderer that uses sine and cosine curves as factors based on a position value of an X-Fade slider.
7. The computer-based method of claim 1, further comprising providing a personalized equalization of the AccMix based on a frequency range weighting factor corresponding to the user.
8. The computer-based method of claim 1, further comprising providing a UI that provides at least one of the following: a content list for creating an AccMix, an accessibility fader, a voice-over fader, a live sports fader, an audio description fader, adaptive voice-overs, language selection, language level selection, personalized EQ selection.
9. The computer-based method of claim 1, further comprising receiving a command from a user to adjust an X-Fade.
10. A computer-based method for creating a maximum accessible audio mix (Max-Acc-Mix), comprising: receiving a raw cinematic mix (Cin-Mix) audio content; and adjusting the Cin-Mix using a dialog clarity engine (DCE) with a predetermined DCE gain to create the Max-Acc-Mix.
11. The method of claim 10, wherein, the dialog clarity engine comprises: receiving a cinematic audio signal; and performing at least one of the following: dialog isolation (DI), dialog enhancement (DE), dynamic range compression (DRC), audio loudness normalization (ALN), enhanced dialog insertion (EDI).
12. The method of claim 10, wherein, the dialog clarity engine (DCE) comprises at least one of the following: dialog isolation (DI) that separates dialog audio and non-dialog audio; dialog enhancement (DE) that amplifies the dialog audio relative to the non-dialog audio by a predetermined DE gain; dynamic range compression (DRC) that compresses the non-dialog audio or combined dialog and non-dialog audio; and audio loudness normalization (ALN) that amplifies the dialog audio or the combined dialog and non-dialog audio using a predetermined ALN gain.
13. The method of claim 11, wherein, the dynamic range compression (DRC) comprises attenuating a predetermined upper or lower segment of a range of the non-dialog audio or the combined dialog and non-dialog audio by a predetermined DRC gain profile.
14. The method of claim 10, wherein, the creating the MaxAccMix comprises adjusting the DCE amplification gain such that both the dialog audio and the non-dialog audio are greater than a predetermined acceptable floor noise.
15. The method of claim 10, further comprising providing a UI that allows a content provider user / administrator to adjust DCE parameters, the DCE parameters comprising at least one of the following: floor noise, Dx gain factor, MNE gain factor, Gdx, Gmne, DCE mode, DRC gain profile, and ALN gain.
16. The method of claim 10, wherein, the Cin-Mix is in a format of at least one of the following: mono, stereo, surround sound, and immersive sound.
17. The method of claim 10, wherein, the dialog enhancement (DE) comprises at least one of the following: amplifying the dialog and attenuating the non-dialog while maintaining the same overall loudness; and amplifying only the dialog.
18. The method of claim 10, wherein, the dynamic range compression (DRC) comprises attenuating a predetermined upper or lower segment of a range of the non-dialog audio or the combined dialog and non-dialog audio by a predetermined DRC gain profile.
19. A computer-based method for generating an accessible audio mix (AccMix) of dialog audio and non-dialog audio (residual music / effects) that allows a listener user to understand the dialog audio played on a user playback device (or listening device) in a listening environment having a plurality of different actual ambient noise levels, the method comprising: receiving at a user device from the user device a cinematic mix (CinMix) and a maximum accessible mix (MaxAccMix) associated with video / audio content to be viewed and listened to by a user; creating the AccMix audio mix by mixing the CinMix and the MaxAccMix such that the resulting AccMix audio mix is power-conserved and greater than an ambient noise (NF) in the listening environment of the user device, the AccMix having a range from the CinMix value to the MaxAccMix value, the value of AccMix based on at least one of a user input and the actual ambient noise level.
20. The computer-based method of claim 19, wherein, the AccMix is automatically set based on an actual ambient noise in the listening environment, the actual ambient noise measured by a sound sensor that filters out human speech from the audio signal to provide the actual ambient noise level.
21. A computer-based method of improving an audio dialog level of digital cinema content while maintaining creative intent of the cinema content, the method comprising: receiving a cinematic audio mix; receiving an accessible audio mix; and combining the cinematic audio mix and the accessible audio mix to provide an improved dialog audio mix that allows a listener to hear the dialog over an ambient noise.
22. The method of claim 21, wherein, the combining includes using a cross-fade renderer.
23. A computer-based method of fading from a first audio mix (MixA) to a second audio mix (MixB), comprising: receiving the MixA audio mix; receiving the MixB audio mix; combining MixA and MixB using a cross-fade renderer having an output that transitions from MixA to MixB based on a position of an X-Fade slider.
24. The computer-based method of claim 23, wherein, the cross-fade renderer is power-conserved.
25. The computer-based method of claim 23, wherein, MixA and MixB include at least one of: MixA = primary language original version (or Cin-Mix) and MixB = enhanced dialog primary language (Max-Acc-Mix); MixA = primary language original version and MixB = primary language audio description; MixA = primary language original version and MixB = alternative language audio commentary; MixA = primary language original version and MixB = alternative language audio description; MixA = home team commentary mix and MixB = away team commentary mix; and MixA = home team commentary mix and MixB = stadium sound (no commentary) mix.