System and method to provide personalized audio streaming and rendering
The system generates personalized audio mixes with dialogue enhancement and dynamic range compression, addressing the limitations of current streaming solutions by ensuring optimal dialogue clarity and adaptability to noise levels, enhancing user experience.
Patent Information
- Application Number
- JP2025069131
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-12
- Filing Date
- 2025-04-18
- Publication Date
- 2025-11-05
AI Technical Summary
Current audio streaming solutions lack the ability to personalize the audio experience dynamically based on environmental noise levels and require users to manually switch streams, leading to potential delays and suboptimal dialogue intelligibility.
A system and method that generates personalized audio mixes by applying dialogue enhancement and dynamic range compression, allowing simultaneous playback of multiple audio streams with cross-fade capabilities, ensuring optimal dialogue clarity across varying noise environments.
Provides an optimized, synchronized, and personalized audio experience across devices, enhancing dialogue intelligibility and adaptability to noise levels without requiring stream switching, thus improving user experience.
Smart Images

Figure 2025165903000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 636,400, filed April 19, 2024, and U.S. Provisional Patent Application No. 18 / 999,628, filed December 23, 2024, both of which are incorporated by reference herein to the fullest extent permitted by applicable law. [Background technology]
[0002] Current audio streaming solutions typically rely on a single audio bitstream and decoder per streaming application, limiting the ability of content owners and streaming service providers to personalize the audio experience they provide to end users.
[0003] In particular, existing streaming services offer separate streams for different versions of augmented dialogue, e.g., English Dialogue Amplified High, English Dialogue Amplified Medium, etc. In this case, the user must select the version they want to hear. To change to a different audio version, for example because the background environmental noise has changed, the user must request it, and the new version is retrieved from the appropriate content server. Furthermore, most television devices apply post-processing, such as AI sound enhancement, to the decoded audio bitstream to reduce noise or amplify dialogue, but these do not preserve the integrity of the original sound mix or the content owner's creative intent. Summary of the Invention [Problem to be solved by the invention]
[0004] Switching audio streams can be tedious and time-consuming, potentially resulting in buffering / loading delays for digital streaming files during audio delivery from the content delivery network (CDN) to the end user. Furthermore, conventional systems only offer a predefined number of dialogue-audio-audio-audio versions (e.g., 3-5 versions) without guaranteeing that the selected version is optimal for the listener's / user's noise environment. This approach can make it difficult for users to hear the dialogue in audio content if the background noise becomes louder than intended for the selected stream. For example, if a user selects an "English - Low Dialogue Amplification" track but the ambient noise is higher than expected, the selected audio stream may not provide full dialogue intelligibility to the end user, and "English - High Dialogue Amplification" may be more appropriate.
[0005] Additionally, existing streaming services require each user in a room or listening environment to use the exact same audio mix. [Means for solving the problem]
[0006] Therefore, what is desired is a system and method that overcomes the shortcomings of current approaches and improves the user's audio experience without switching streams and going back to the content provider's original server or CDN where all predefined streams are available. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a top-level block diagram of content provider (CP)-side components of a system for providing personalized audio streaming and rendering to generate a maximal version of an accessible audio mix, according to an embodiment of the present disclosure. [Figure 2]Figure 2A is a top-level block diagram of a user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with a remote- or user-device-controlled cross-fade audio mix according to an embodiment of the present disclosure. Figure 2B is a top-level block diagram of a user / listener side of a system for providing personalized audio streaming and rendering for a single user device according to an embodiment of the present disclosure. Figure 2C is a top-level block diagram of a user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices, each with a personalized cross-fade audio mix, according to an embodiment of the present disclosure. Figure 2D is a top-level block diagram of a user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices, each with a personalized cross-fade audio mix and content selection, according to an embodiment of the present disclosure. [Figure 3] 3A, 3B, 3C, 3D, and 3E are block diagrams of alternative embodiments of the Dialogue Clarity Engine (DCE) of FIG. 1, including combinations of Dialogue Isolation (DI), Dialogue Enhancement (DE), Dynamic Range Compression (DRC), Enhanced Dialogue Insertion (EDI), and Audio Loudness Normalization (ALN) at various signal processing locations and the corresponding effects on dynamic range, according to embodiments of the present disclosure. [Figure 4] 1A-1C are two diagrams of dynamic range adjustment for dialogue enhancement (DE) according to an embodiment of the present disclosure, the first diagram showing the integrated loudness (IL) of the dialogue (Dx) audio being the same as the non-dialogue (i.e., music and effects, or M&E or MNE) audio, and the second diagram showing the integrated loudness (IL) of the dialogue (Dx) being different from the non-dialogue audio (i.e., music and effects, or M&E or MNE). [Figure 5]Figure 5A is a diagram of an audio stream of interleaved original cinematic mix (Cin-Mix) and maximum accessibility mix (Max-Acc-Mix) content to provide accessible audio according to an embodiment of the present disclosure. Figure 5B is a diagram of an audio stream of interleaved Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix content to provide accessible audio according to an embodiment of the present disclosure. Figure 5C is a diagram of an audio stream of interleaved Cin-Mix and Max-Acc-Mix content to provide accessible audio in two languages for voice-over audio according to an embodiment of the present disclosure. 5D is a block diagram illustrating a non-encoding content server with separately stored unencoded Cin-Mix and Max-Acc-Mix and Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix audio content streams as input to an audio encoder, and an encoding content server with interleaved Cin-Mix / Max-Acc-Mix and Cin-Mix / Mid-Acc-Mix / Max-Acc-Mix audio content streams as output from the audio encoder, according to an embodiment of the present disclosure. FIG. 5E is a table illustrating various encoding and decoding options for provided stream types and user-desired configurations, according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a block diagram of various components of a system for providing personalized, networked audio streaming and rendering, according to an embodiment of the present disclosure. [Figure 7] FIG. 4 is a block diagram of the dialogue enhancement (DE) logic of FIGS. 3A-3E according to an embodiment of the present disclosure. [Figure 8] 3A-3E are block diagrams of the Dynamic Range Compression (DRC), Enhanced Dialogue Insertion (EDI), and Audio Loudness Normalization (ALN) of FIGS. 3A-3E according to an embodiment of the present disclosure. [Figure 9] 3A-3E to vary the DRC gain to provide range compression according to an embodiment of the present disclosure. [Figure 10] Figure 10A is a block diagram of a crossfade renderer (X-Fade) according to an embodiment of the present disclosure, and a graph showing an X-Fade power retention gain curve. Figure 10B is a block diagram of a crossfade renderer (X-Fade) for stereo (L / R) input audio according to an embodiment of the present disclosure, and a corresponding graph showing an X-Fade power retention gain curve versus a position range of an X-Fade slider adjustment. Figure 10C is a block diagram of a crossfade renderer (X-Fade) for stereo (L / R) input audio with three inputs according to an embodiment of the present disclosure, and a corresponding graph showing an X-Fade power retention gain curve versus a position range of an X-Fade slider adjustment. [Figure 11] FIG. 11 is a block diagram of components of personalized equalizer (Pers.EQ) logic according to an embodiment of the present disclosure. [Figure 12] 12A is a screenshot of a graphical user interface (UI) for DCE gain adjustment and Max-Acc mix monitoring used by a content provider's CP user / administrator with the ability to set / adjust DCE gain and noise floor and listen to the results across the crossfade range according to an embodiment of the present disclosure. FIG. 12B is a diagram of a smart TV remote control for manually adjusting the crossfade (X-fade) renderer of FIGS. 2A, 2C, and 2D according to an embodiment of the present disclosure. FIG. 12C is a screenshot of an Acc app graphical user interface (UI) used by a user with the ability to set / adjust crossfades and enable certain features according to an embodiment of the present disclosure. FIG. 12D is a screenshot of a graphical user interface (UI) for a listening test to determine personal equalizer parameters according to an embodiment of the present disclosure. [Figure 13]13A is a flow diagram of the Portal / DCE logic of FIG. 1 according to an embodiment of the present disclosure. FIG. 13B is a flow diagram of the Portal UI logic of FIG. 1 according to an embodiment of the present disclosure. FIG. 13C is a flow diagram of the Hub Accessibility (Acc) app logic of FIGS. 2A, 2C, and 2D according to an embodiment of the present disclosure. FIG. 13D is a flow diagram of the Acc app UI logic of FIGS. 2A, 2C, and 2D according to an embodiment of the present disclosure. [Figure 14] 14A, 14B, 14C, 14D, and 14E are block diagrams of various embodiments for different types of audio inputs to be streamed and adjustably rendered by the crossfade renderer of FIG. 10A, 10B, or 10C according to an embodiment of the present disclosure. [Figure 15] 15A and 15B are block diagrams of various embodiments for different types of audio inputs to be streamed and adjustably rendered by the crossfade renderer of FIG. 10A, FIG. 10B, or FIG. 10C for video-on-demand (VOD) and live sports / events (Live) applications, respectively, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0008] As described in further detail below, in some embodiments, the present disclosure is directed to systems and methods that provide personalized audio streaming and rendering, including cross-fade content delivery in accordance with content owner rules, requirements, and needs.
[0009] The present disclosure provides end users with a personalized audio streaming experience that enables users to understand storytelling dialogue (Dx) across a wide range of user devices and environmental background noise levels by providing accessible audio mixes or audio tracks generated from an original studio cinematic mix. The accessible mix is generated by providing dialogue-intelligible and / or dynamic range-compressed (to adapt to the environmental noise floor), enhanced audio (applied to language, sports commentary, etc.), and multiple dialogue audio streams simultaneously to one or more devices.
[0010] The personalized audio streaming and rendering systems and methods of the present disclosure provide for a variety of applications and configurations, such as in video-on-demand (VOD) applications, providing accessible streaming audio with enhanced dialogue, audio description (or narration), and voice-over (VO), and in live sports / event audio streaming applications, providing options such as home and away sports team commentary and sports team commentary focus, or stadium sound focus, as described herein.
[0011] The present disclosure provides for rendering of decoded streams based on what the end user wants to hear to provide an optimized individual listening experience or personalized audio experience. The present disclosure also provides audio in multiple streams and can perform software decoding of one or more audio streams within the streaming app itself to enhance the audio experience (with accessibility, localization, and synchronization). No existing streaming service can simultaneously play and render at least two different audio mixes as described in this disclosure.
[0012] Current commercial solutions for delivering enhanced listening experiences are standardized, as illustrated by the ATSC-3.0 standard, which is designed to provide personalized listening and immersive audio experiences with enhanced dialogue clarity. The ATSC-3.0 standard supports audio decoders for both the Dolby AC-4 and MPEG-H codecs. However, these existing commercial solutions have not yet been adopted by the entertainment industry and rely on audio production formats that require complex multidimensional audio renderers that are difficult for smart TV OEM partners and streaming services to implement.
[0013] One advantage of the disclosed system and method is its ability to provide an accessible, personalized, localized, and synchronized audiovisual (AV) experience across one or more devices. The solution is "accessible" because of the native and adaptive dialogue enhancement solution provided by the present disclosure. The solution is also "localized" due to its nature of supporting not only simultaneous dominant and ducted (non-dominant) dialogue tracks, but also multi-dialogue audio with alternative language rendering of the native / enhanced (or accessible) dialogue tracks. This applies to alternative language tracks, narration (audio description or AD) tracks, voice-over (VO), live sports commentary (home / away team commentary), and more. The solution is "synchronized" because only one streaming application performs the audio decoding, allowing multiple users watching different streams of the same program, for example in different languages, to be synchronized. Conventional solutions for watching together using multiple streaming applications on different devices do not provide perfect synchronization.
[0014] While current AC-4 and MPEG-H compatible products offer some dialogue enhancement features, they do not use the noise floor of the listening environment to automatically adjust the ratio, i.e., blend ratio, of the original audio cinematic mix (dialog and non-dialogue) to the dialogue-enhanced (i.e., accessible) audio track, as done by the crossfade renderer of this disclosure. Also, they do not provide the ability for users to manually adjust the ratio of the original audio cinematic mix (dialog and non-dialogue) to the dialogue-enhanced (i.e., accessible) audio track in real time, providing fine tuning capabilities without changing (or switching) streams. Furthermore, conventional products do not support the ability to reproduce multiple self-contained audio tracks, such as cinematic / accessible mixes with ED (enhanced dialogue), AD (audio description), and VO (voice-over), as well as sports / personalized mixes with home / away team commentary and stadium mixes, from a single streaming application, as described herein.
[0015] Referring to FIG. 1 , a top-level block diagram of content provider (CP)-side components of a system for providing personalized audio streaming and rendering that generates a maximum version of an accessible audio mix (Max-Acc-Mix or Max-Acc-Mix) according to an embodiment of the present disclosure is shown. Specifically, system 10 includes a content provider computer 20. Content provider computer 20 receives instructions (e.g., content selection) from a content provider (CP) user / administrator 21 and includes portal / DCE logic 28 running on CP computer 20. Portal / DCE logic 28 receives the non-encoded (or unencoded) audio portion of the selected content, e.g., the original theatrical near-field cinematic audio mix (i.e., Cin-Mix or CinMix), from a content selection API (i.e., content API) 30, as shown by line 96. This audio portion can be obtained from a non-encoded content server 36 via line 87. In theatrical audio, a "near-field audio mix" or "near-field mix" refers to an audio mix created for home entertainment, such as streaming or DVD, using a smaller near-field speaker setup to simulate a home theater environment. Such a mix is typically created after the standard theatrical mix has been approved and finalized by the studio. A near-field mix is designed to be listened to in a more intimate environment, with speakers closer to the listener, as opposed to the wide, reverberant space of a movie theater. This disclosure may also apply to original cinematic mixes, which are standard theatrical mixes designed for movie theaters and the like. The content provider computer 20 may be a smartphone, computer, laptop, tablet, or other computer-based device.
[0016] The portal / DCE logic 28 provides a maximum accessibility mix (Max-Acc-Mix) audio track on line 92 (described below) based on the original Cin-Mix content selected by the CP user / administrator 21 and the audio signal processing performed by the dialogue clarity engine 62 with DCE parameters, i.e., DCE Params (e.g., DCE Gain, Noise Floor, etc.), either defaulted from previously stored values or set (or adjusted) by the CP user / administrator 21. The portal / DCE logic 28 provides the Max-Acc-Mix on line 92 to a known audio encoder 26 as described herein with reference to FIG. 5E (described below). The audio encoder 26 also receives the original Cin-Mix audio signal (or track) from the DCE 62 on line 96, encodes the audio content of both tracks (stereo / surround / immersive) into the desired digital audio format (e.g., encoding dual stereo, enhanced surround, enhanced immersive into Dolby E-AC3, Dolby ATMOS, MPEG AAC-LC, Opus, etc.) as described later in this specification using Figures 5A, 5B, and 5C, and interleaves them into a single bitstream, as shown by line 94. The audio encoder 26 stores the output of the interleaved encoded content (encoded Cin-Mix + Max-Acc-Mix) on the encoding content server 40 using the content API 30, as shown by line 95. Once the encoded content is stored on the encoding content server 40, it is available for streaming by the content provider to authorized users (or listeners) of the CP content. In some embodiments, as described in more detail below, portal / DCE logic 28 can also provide a Mid-Acc-Mix audio track on line 93 to encoder 26 having a loudness level between Cin-Mix and Max-Acc-Mix or an enhanced level (partial DRC + DE or DRC only or DE only).
[0017] The Portal / DCE logic 28 may include components such as Portal UI logic 60 and a Dialog Clarity Engine 62, which together provide a Max-Acc-Mix and an optional Mid-Acc-Mix.
[0018] In particular, the portal UI logic 60 provides a portal user interface (i.e., a portal UI) for the CP user / administrator 21, which may have a content UI section provided at line 82 to the CP user / administrator 21 via the display 22, as indicated by line 74. The content UI section may display a list of audio content (or audio files) available for generating an accessible audio mix (see FIG. 12A), which may be obtained at line 96 from a content API. The portal UI logic 60 also receives a content selection at line 78 from the CP user / administrator 21 via user interaction with the portal UI display 22 (or other input device connected to the CP computer 20, such as a keyboard, mouse, or other device). The content selection at line 78 may be provided to the DCE 62, either directly or from the portal UI logic 60, so that the DCE can select content as needed. As indicated by line 74 in FIG. 12A, the UI for content selection may be displayed on the display 22 as a "content audio file," as described below.
[0019] The portal UI logic 60 also requests and receives the original Cin-Mix audio from the non-encoding content server 36 via the content API 30 on line 96. The portal UI logic 60 also provides DCE Params, such as DCE gain / model type and noise floor (NF), to the dialogue clarity engine 62 on line 90, which uses the DCE Params to perform audio signal processing on the selected Cin-Mix audio track, as described below. The portal UI logic 60 can also communicate with the DCS Params server 38 on line 85 to retrieve or store the DCE Params. The portal UI logic 60 may receive the DCE Params, for example, from a user via the portal UI (see FIG. 12A ), from a DCE Params server associated with the selected content, or from metadata embedded within the digital audio content (if predetermined for the content). The Portal UI logic 60 provides the maximum accessible audio mix (Max-Acc-Mix) on line 92, which is also provided (or fed back) to the Portal UI logic 60 as described herein, allowing the CP user / administrator 21 to adjust the accessible mix (Max-Acc-Mix) via the Portal UI as needed.
[0020] The dialogue clarity engine 62 also receives DCE parameters (i.e., DCE Params), such as NF, DCE Gain, and DCE Model Type, from the portal UI logic 60 on line 90, receives a Max-Acc-Mix=OK signal from the portal UI logic 60 on line 88, and further receives Cin-Mix audio on line 96 from the encoded content server 36 via the content API 30.
[0021] In some embodiments, the DCE 62 or portal UI logic 60 may obtain default or initial values for the DCE Params (e.g., gain, model type, NF, and any other required parameters) associated with the selected content from the DCE parameter server 38 at line 83. Additionally, the DCE 62 or portal UI logic 60 stores the values of the DCE Params provided by the CP user / administrator 21 for the selected content in the DCE parameter server 38.
[0022] The output of the DCE logic 62 is a Max-Acc-Mix track provided on line 92 to the audio encoder 26, as described below. The DCE 62 may store the Max-Acc-Mix track on the non-encoding content server 36. The other input to the encoder 26 is the original Cin-Mix audio track on line 96. In some embodiments, the DCE may provide a second output, Mid-Acc-Mix, on line 93, which is an audio track branched from a predetermined intermediate (i.e., mid) position within the DCE, as described below. The output of the encoder 26 is an encoded interleaved audio stream having both the Cin-Mix and Max-Acc-Mix in a single bitstream, as shown in Figures 5A, 5B, and 5C, which is stored on the encoding content server 40 via the content API 30 on lines 94 and 95, as described below.
[0023] 2A, 2B, 2C, and 2D, there are shown several top-level block diagrams 200A, 200B, 200C, and 200D, respectively, for different configurations or applications of the user / listener side of the disclosed systems and methods, which, as described herein, access Cin-Mix / Max-Acc-Mix encoded interleaved content from an encoding content server 40, decode the digital stream into individual tracks of Cin-Mix and Max-Acc-Mix (or Cin-Mix and Max-Pers-Mix) for use by a crossfade renderer (X-Fade) 100 (described later herein) to provide a personalized, accessible crossfade mix for one or more users / listeners 15. The user / listener side of the disclosed systems and methods is described further below.
[0024] Returning to FIG. 1 and referring to FIGS. 3A-3E, block diagrams of alternative embodiments of the Dialogue Clarity Engine (DCE) of FIG. 1 and the corresponding effect on audio dynamic range or audio loudness range are shown, according to embodiments of the present disclosure. In particular, the DCE includes several components or logic, including Dialogue Isolation (DI) 302A, Dialogue Enhancement (DE) 304A, Dynamic Range Compression (DRC) 306A, Enhanced Dialogue Insertion (EDI) 308A, and Audio Loudness Normalization (ALN) 310A, which may be configured in several different ways to provide the maximum accessible mix audio (Max-Acc-Mix) on line 92. The configuration of the DCE model depends on the selected content and the desired accessible audio mix result from the content provider's CP user / administrator 21. The DCE may also receive DCE parameters on line 83 from the DCE Params server 38 (as described below).
[0025] Referring to FIG. 3A, a first DCE model 300A (DCE Model 1) configuration is shown having, from left to right, DI-DE-DRC-EDI-ALN components 302A, 304A, 306A, 308A, and 310A, and corresponding loudness range (LRA) bar graphs 302B, 304B, 306B, 308B, and 310B, respectively, illustrating the effect of each component on an input cinematic mix (Cin-Mix) LRA bar graph 301, as described below.
[0026] LRA bar graphs 301B, 302B, 304B, 306B, 308B, and 310B are plotted on graph 311 having a vertical axis 312 of loudness measured in LKFS (Loudness K-Weighted Full Scale) and a horizontal time axis (i.e., time series). Vertical axis 312 also shows the noise floor (NF) range 316 and noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix.
[0027] The LRA bar graphs 301B, 302B, 304B, 306B, 308B, 310B have the audio loudness range (LRA) of the entire program (i.e., program loudness PL) along with the loudness range (LRA) of the dialogue (Dx) portion of the audio (i.e., dialogue LRA or Dx LRA) 301C, 302C, 304C, 306C, 308C, 310C and the loudness range (LRA) of the non-dialogue (i.e., music and effects, M&E, or MNE) portion of the audio (i.e., non-dialogue LRA or M&E LRA) 301D, 302D, 304D, 306D, 308D, 310D.
[0028] The LRA bar graphs 301B, 302B, 304B, 306B, 308B, 310B may also include the average loudness or integrated loudness (IL) of the dialogue (Dx) portions of the audio (i.e., dialogue LRAs or Dx LRAs) 301E, 302E, 304E, 306E, 308E, 310E, the average loudness or integrated loudness (IL) of the non-dialogue (i.e., music and effects (M&E)) portions of the audio (i.e., non-dialogue LRAs or M&E LRAs) 301D, 302D, 304D, 306D, 308D, 310D, and the average loudness or integrated loudness (IL) of the overall program loudness 310G (i.e., program loudness (PL)).
[0029] In DCE Model 1 of FIG. 3A , the input Cin-Mix is processed by dialogue separation (DI) logic 302A, which uses known tools to separate the dialogue (Dx) portion of the audio from the non-dialogue (i.e., M&E) portion of the audio and provide two audio tracks: a Dx track on line 332A and an M&E track on line 332B. Tools that may be used by DI logic 302A include AI machine learning (ML) models (open source and commercial) for extracting dialogue from video streams, such as Spleeter™ (Deezer Research), Demucs™ (Meta™), and AudioShake™ (AudioShake). The ML models used for DI are trained using existing datasets for extracting dialogue from videos across a wide range of content. Other methods, tools, or models for isolating (or extracting or separating) Dx from M&E may be used as needed, provided they provide the functionality and performance described herein.
[0030] In this example, the loudness range LRA for the two tracks (in this case dominated by the M&E track) is +18 dB (-16 - (-34) = +18), the program loudness (PL) is -24 LKFS, and the dialog for the program (or reverberant, non-dialogue, or M&E) loudness (DPL or DRL) is +0 dB (because the IL for both Dx and M&E is -24 LKFS), as shown by dashed box 320. The Dx and M&E audio tracks are provided on lines 332A, 332B to the dialogue enhancement (DE) logic 304A, which amplifies the Dx track and attenuates the M&E track while keeping the average or integrated loudness the same.
[0031] 7, a block diagram of an embodiment of the dialogue enhancement (DE) logic of FIGS. 3A-3E is shown, in accordance with an embodiment of the present disclosure. In particular, the DE logic 304 receives a Dx audio track on line 332A and amplifies the Dx audio in gain multiplication block 712 by multiplying the Dx by a gain having a value greater than 1.0 (equal to the resulting gain indicated herein by -LKFS) on line 711 to provide an amplified dialogue (i.e., amplified Dx) signal or track on line 334A. The DE logic 304 also receives an M&E audio track on line 332B and attenuates the M&E audio in gain multiplication block 722 by multiplying the M&E by a gain having a value less than 1.0 (equal to the resulting attenuation indicated herein by -LKFS) on line 721 to provide an attenuated M&E (i.e., non-dialogue or reverberant) signal or track on line 334B.
[0032] The DE gains Gdx and Gmne provide the desired amplification of the dialogue (Dx) relative to the M&E without substantially changing the average loudness or integrated loudness (IL), which in some embodiments is done by evenly splitting the Dx amplification and M&E attenuation, as shown by the Dx gain factor and M&E gain factor on lines 702 and 716, respectively, which in the illustrated example both have a value of 0.5, providing an even split between the Dx amplification gain Gdx and the M&E attenuation gain Gmne. In particular, the Dx gain factor (e.g., 0.5) on line 702 is multiplied by the adjusted Dx amplification gain Gdx on line 709 in a gain multiplication block 710, which provides the actual Dx gain on line 711, which is multiplied by Dx on line 332A in a gain block 712, which provides the amplified dialogue on line 334A. Similarly, the M&E gain factor (e.g., 0.5) on line 716 is multiplied by the adjusted M&E attenuation gain Gmne on line 719 in a gain block 720, which provides the actual M&E gain on line 721, which is multiplied by the M&E on line 332B in a gain block 722, which provides the amplified M&E on line 334B.
[0033] In some embodiments, if the dialogue (Dx) IL is not the same as the M&E IL, for example, if the Dx IL is more (less negative) than the M&E IL, adjustments may be made to the Dx gain and the M&E gain, as indicated by summing blocks 708, 718. In particular, a DPL adjustment signal is provided on line 706 to summing blocks 708, 718, which adjust both the Dx and M&E gain values equally. In that case, if the Dx IL is greater than the M&E IL, summing block 708 decreases the Dx amplification gain Gdx on line 704 to provide an adjusted Dx gain on line 709, and summing block 718 adjusts the M&E attenuation gain Gmne on line 714 (i.e., make it less negative) to provide an adjusted M&E gain on line 719 equal to the Gdx adjustment. In some embodiments, the DPL adjustment amount on line 706 is determined by the DE logic 304 by calculating the difference between the Dx IL on line 732 and the M&E IL on line 734 (i.e., Dx IL-M&E IL), as indicated by summing block 730. Also, in some embodiments, the DCE logic 304A may use known methods or tools to perform calculate Dx IL logic 736 and calculate M&E IL logic 738, respectively, to calculate the integrated loudness (IL) of the dialogue Dx and the IL of the M&E (non-dialogue) from the input audio tracks Dx and M&E on lines 332A, 332B by providing the Dx IL on line 732 to the positive (or plus) input of the summing block 730 and the M&E IL on line 734 to the negative (or minus) input of the summing block 730. The IL may be determined by known tools such as Insight2 or RX by iZotope, or by open source software IL tools such as those described in https: / / github.com / csteinmetz1 / pyloudnorm or in the loudness software by Youlean (see: https: / / youlean.co). Other techniques for determining the IL may be used if desired.In some embodiments, the Dx IL and M&E IL of lines 732, 734, respectively, may be determined separately outside of DCE logic 304 and provided as inputs to DCE logic 304A, in which case Calculate Dx IL logic 736 and Calculate M&E IL are not used.
[0034] Also, in some embodiments, instead of a single gain adjustment, the DPL adjustment, there may be separate adjustments (not shown) to the Dx gain and the M&E gain based on the desired resulting audio output. Also, the Dx gain factor, Gdx, and the MNE gain factor, Gmne, may be collectively referred to as the DE gain 342.
[0035] 4 and 7, a pair of LRA bar graphs 402, 404 are shown illustrating Dx IL offsets similar to those described above. Graph 402 (left) shows Dx and M&E with equal IL values (or DPL = 0 dB). In this case, Gdx and Gmne are the same value. Graph 404 shows Dx and M&E where the Dx IL is higher than the M&E IL, giving a non-zero DPL, e.g., 1.0 dB. For example, if Dx IL = 1.0 dB and M&E IL = 0 dB, the DPL will be 0.5 dB (halfway between 0 (M&E IL) and 1 (Dx IL)). In this case, adjusting the DPL will reduce the dialogue gain Gdx (FIG. 7) to 3 dB (adjusted Gdx) and change the M&E attenuation Gmne (FIG. 7) to -3 dB (adjusted Gmne). Multiplication by the Dx gain factor and the M&E gain factor (each 0.5) results in a Dx amplification of 1.5 dB and an M&E attenuation of −1.5 dB, as shown in lines 711 and 721. The resulting Dx value is 2.5 dB and the resulting M&E value is −1.5 dB, providing the desired separation of 4 LKFS (2.5 − (−1.5)). A resulting DPL of 0.5 (2.5 − 4 / 2) is also provided, preserving the 0.5 DPL value in this example. Thus, DE logic 304A provides the desired dialogue (Dx) enhancement (e.g., 4 LKFS) for non-dialogue (M&E) while leaving the dialogue unchanged (e.g., DPL = 0.5) for the program (i.e., reverberation or M&E) loudness DPL.
[0036] 3A, 8, and 9, dynamic range compression (DRC) logic 306A receives the attenuated M&E audio track on line 334B and the DCE gain profile on line 344 and provides a compressed M&E audio track on line 336. With particular reference to FIGS. 8 and 9, DRC logic 306A compresses the input M&E audio track by multiplying the input M&E audio track on line 334B by the value of a DRC gain curve (or profile) on line 344, as shown in FIG. 9, via multiplication block 802 (FIG. 8). With particular reference to FIG. 9, a known DRC gain curve 902 is shown plotted on a graph 900 of output level (vertical axis) versus input level (horizontal axis). The DRC gain curve 902 may have sections or ranges that provide ALR compression of a desired M&E loudness range, including an amplification range 908, a null band or unity gain region 910, an initial cut range 912, and a cut range 914. Other types of DRC gain curves and ranges may be used if desired. The amplification range 908 increases (or amplifies) low-level (i.e., quiet) audio M&E sounds using a gain greater than 1.0, the unity gain region preserves audio M&E sounds using a gain of 1.0, the initial cut range 912 provides slight attenuation of audio M&E sounds, and the cut range 914 provides greater attenuation of audio M&E sounds using a gain less than 1.0. The initial cut range 912 and cut range 914 are set to provide the desired amount of M&E compression. In some embodiments, the amount of compression (and low-frequency amplification) provided by the DRC logic 306A may be adjusted by a CP user / administrator by adjusting the DRC gain curve 902, which may be done via the CP user interface or portal UI, as further described below.
[0037] 3A and 8, the enhanced dialogue insertion (EDI) logic 308A receives the attenuated and compressed M&E audio track on line 336 and the amplified dialogue Dx on line 334A and provides a resynthesized (or remixed) composite audio track on line 338. In particular, with reference to FIG. 8, the EDI logic 308A adds the input amplified Dx on line 334A to the compressed M&E on line 336 at summer 804 and provides a resynthesized audio track on line 338 with the enhanced (or amplified) dialogue Dx inserted into the compressed M&E audio track.
[0038] 3A and 8, audio loudness normalization (ALN) logic 310A receives a resynthesized audio track on line 338, an ALN gain on line 350, and provides a maximum accessibility mix (Max-Acc-Mix) track on line 92. In particular, referring to FIG. 8, ALN logic 310A multiplies the input audio signal on line 338 by the ALN gain on line 350 in multiplication gain block 806, which provides the Max-Acc-Mix on line 808 to maximum limiting logic 810 to ensure that the ALN gain does not inappropriately exceed the loudness range (LRA) of the Max-Acc-Mix limit.
[0039] In particular, there are two equations that define the limits on the values of the DCE gains (e.g., DE gain, DRC gain, ALN gain) of the dialogue clarity engine 62. In particular, Equation 1 below defines the maximum value of the Max-Acc mix, which is an accessible mix. Equation 1: Max(TruePeak(MaxAccMix)) < -2.0dB = Max Audio Here, Max(TruePeak(MaxAccMix)) is the maximum true peak value (in dB) of the Max-Acc-Mix track.
[0040] Also, if the noise floor (NF) is known and Equation 1 is satisfied, then Equation 2 below must also be satisfied to provide the best user listening experience. Equation 2: Min(TruePeak(DX)) > NF where Min(TruePeak(DX)) is the minimum true peak value of the dialogue DX portion of the mix, and NF is the noise floor level. Therefore, it may be necessary to adjust the DCE based on the result of Equation 2.
[0041] The maximum audio value to avoid audio clipping is 0 dB true peak. Some content providers require the maximum audio value to be -2 dB before clipping occurs to add some buffering. As is well known, the maximum audio value is typically set as dB true peak, or dBTP, which is the maximum peak audio level of the content. Maximum limit logic 810 (FIG. 8) within ALN logic 310A will not allow Max-Acc-Mix to exceed the requirements of Equation 1.
[0042] In particular, max limiting logic 810 receives Max-Acc-Mix on line 808 and uses the current noise floor NF setting from the CP user / administrator 21 or from the DCE Params server 38 and the maximum audio on line 812 from the DCE Params server 38 to determine if the value of Max-Acc-Mix satisfies Equation 1. If so, Max-Acc-Mix is passed unchanged to the output on line 92 as Max-Acc-Mix. Alternatively, if the NF is known and Equation 1 is satisfied, the DCE parameters of the DCE may be adjusted as necessary so that Equation 2 is true.
[0043] If the value of Max-Acc-Mix does not satisfy Equation 1, then Max-Acc-Mix is reduced to a value that does satisfy Equation 1, and the modified version of Max-Acc-Mix is passed to the output on Max-Acc-Mix line 92 .
[0044] Thus, if the CP user / administrator 21 attempts to set the ALN gain, DE gain, or noise floor (NF) (or any other gain or coefficient within the DCE) such that Max-Acc-Mix exceeds the requirements of Equation 1 for a given content provider, max limiting logic 810 will clamp the ALN gain so that Max-Acc-Mix satisfies Equation 1. The maximum audio value may be input to max limiting logic 810 from DCE Params server 38 on line 812 and may be pre-set by the content provider to a value that meets the content provider's audio requirements, for example, -2dbTP. Other values may be used if desired.
[0045] In some embodiments, the CP user / administrator 21 may set the ALN gain from the UI (described later in this specification). In some embodiments, the system may automatically calculate the ALN gain using Equation 1 to maximize the LRA of Max-Acc-Mix based on the DE gain or noise floor set by the CP user / administrator 21.
[0046] 3B, there is shown a configuration of a second DCE model 300B (DCE Model 2) having, from left to right, the components DRC-DI-DE-EDI-ALN 306A, 302A, 304A, 308A, and 310A, and corresponding loudness range (LRA) bar graphs 306B, 302B, 304B, 308B, and 310B, respectively, illustrating the effect of each component on an input cinematic mix (Cin-Mix) LRA bar graph 301B, as described hereinabove. As previously described herein, the LRA bar graphs 306B, 302B, 304B, 308B, and 310B are plotted on a graph 311B having a vertical axis 312 of loudness measured in LKFS and a horizontal axis of time (i.e., era). Similar to FIG. 3A, vertical axis 312 also indicates a noise floor (NF) range 316 and noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix.
[0047] 3A-3E may be performed on the entire audio file, or may be applied to audio chunks of shorter size / duration of the audio content, and may be performed on the chunks simultaneously (in parallel) or sequentially (in order), and the DCE parameters may be set to different values for each chunk, if necessary or desired.
[0048] In DCE Model 2 of FIG. 3B , the input Cin-Mix is first processed by dynamic range compression (DRC) logic 306A, which receives the Cin-Mix audio track on line 92 and the DCE gain profile on line 344 and provides a compressed Cin-Mix audio track on line 336. The DRC 306A functions similarly to that previously described herein with reference to FIG. 3A . In particular, with reference to FIGS. 8 and 9 , the DRC logic 306A compresses the input Cin-Mix audio track by multiplying the input Cin-Mix audio track on line 96 by the value of a DRC gain curve (or profile) as shown in FIG. 9 via multiplication block 802 ( FIG. 8 ). With reference to FIG. 9 , a known DRC gain curve 902 is shown plotted on a graph 900 of output level (vertical axis) versus input level (horizontal axis). The difference between Model 1 in Figure 3A and 300A is that the entire mix is compressed, not just M&E.
[0049] As previously described with reference to FIG. 3A, the compressed Cin-Mix is provided on line 336 to the dialogue separation (DI) logic 302A, which separates the dialogue (Dx) portion of the audio from the non-dialogue (i.e., M&E) portion of the audio and provides two audio tracks: a Dx track on line 332A and an M&E track on line 332B.
[0050] As previously described with reference to FIG. 3A, the Dx and M&E audio tracks on lines 332A, 332B are provided to the dialogue enhancement (DE) logic 304A, which receives a DE gain on line 342 and amplifies the Dx track and attenuates the M&E track while keeping the average or integrated loudness the same.
[0051] The Dx audio track and M&E audio track on lines 334A, 334B from the dialogue enhancement (DE) logic 304A are provided to the enhanced dialogue insertion (EDI) logic 308A, which receives the attenuated and compressed M&E audio track on line 334B and the amplified dialogue Dx on line 334A, and provides a resynthesized (or remixed) composite audio track (similar to that described above with reference to FIG. 3A) on line 338 to the audio loudness normalization (ALN) logic 310A.
[0052] Similar to that described above with reference to FIG. 3A, audio loudness normalization (ALN) logic 310A receives the resynthesized audio track on line 338 and the ALN gain on line 350, and provides a maximum accessibility mix (Max-Acc-Mix) track on line 92.
[0053] 3C, a third DCE model 300C (DCE Model 3) configuration is shown having, from left to right, the components DI-DE-EDI-DRC-ALN 302A, 304A, 308A, 306A, 310A, and corresponding loudness range (LRA) bar graphs 302B, 304B, 308B, 306B, 310B, respectively, illustrating the effect of each component on an input cinematic mix (Cin-Mix) LRA bar graph 301B, as further described below. As previously described herein, the LRA bar graphs 302B, 304B, 308B, 306B, 310B are plotted on a graph 311C having a vertical axis 312 of loudness measured in LKFS and a horizontal axis of time (i.e., time series). Similar to FIG. 3A, vertical axis 312 also indicates a noise floor (NF) range 316 and noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix.
[0054] In DCE Model 3 of FIG. 3C, the incoming Cin-Mix is first processed by dialogue separation (DI) logic 302A, which, as described above with reference to FIG. 3A, separates the dialogue (Dx) from the non-dialogue (i.e., M&E) portions of the audio, providing two audio tracks, a Dx track on line 332A and an M&E track on line 332B, similar to FIG. 3A.
[0055] As previously described with reference to FIG. 3A, the Dx and M&E audio tracks on lines 332A, 332B are provided to dialogue enhancement (DE) logic 304A, which receives a DE gain on line 342 and amplifies the Dx track and attenuates the M&E track while keeping the average or integrated loudness the same.
[0056] The Dx and M&E audio tracks from the dialogue enhancement (DE) logic 304A are provided on lines 334A, 334B to the enhanced dialogue insertion (EDI) logic 308A, which receives the attenuated and compressed M&E audio track on line 334B and the amplified dialogue Dx on line 334A, and provides a resynthesized (or remixed) composite audio track (similar to that described above with reference to FIG. 3A) on line 338 to the dynamic range compression (DRC) logic 306A.
[0057] The dynamic range compression (DRC) logic 306A receives the resynthesized (or remixed) composite audio track on line 338 and the DCE gain profile on line 344, and provides the compressed Cin-Mix audio track on line 336 to the audio loudness normalization (ALN) logic 310A.
[0058] Similar to that described above with reference to FIG. 3A, audio loudness normalization (ALN) logic 310A receives the resynthesized audio track on line 338 and the ALN gain on line 350, and provides a maximum accessibility mix (Max-Acc-Mix) track on line 92.
[0059] 3D, a fourth DCE model 300D (DCE Model 4) configuration is shown having, from left to right, DI-DRC-DE-EDI-ALN components 302A, 306A, 304A, 308A, and 310A, respectively, and corresponding loudness range (LRA) bar graphs 302B, 306B, 304B, 308B, and 310B illustrating the effect of each component on an input cinematic mix (Cin-Mix) LRA bar graph 301D, as described above. As previously described herein, the LRA bar graphs 302B, 306B, 304B, 308B, and 310B are plotted on a graph 311D having a vertical axis 312 of loudness measured in LKFS and a horizontal axis of time (i.e., era). Similar to the graph of Figure 3A, vertical axis 312 also shows noise floor (NF) range 316 and noise floor setting region 318 used by the DCE to determine the desired Max-Acc mix. In DCE Model 4 of Figure 3D, the components or logic configuration is similar to that of Figure 3A, except that DE logic 304A and DRC logic 306A have swapped positions.
[0060] 3E, a fifth DCE model 300E (DCE Model 5) configuration is shown having, from left to right, the following components: DI-DRC-ALN-DE-EDI 302A, 306A, 310A, 304A, and 308A, along with corresponding loudness range (LRA) bar graphs 302B, 306B, 310B, 304B, and 308B, respectively, illustrating the effect of each component on an input cinematic mix (Cin-Mix) LRA bar graph 301E, as described above. As previously described herein, the LRA bar graphs 302B, 306B, 310B, 304B, and 308B are plotted on a graph 311E having a vertical axis 312 of loudness measured in LKFS and a horizontal axis of time (i.e., chronology). Similar to Figure 3A, vertical axis 312 also indicates noise floor (NF) range 316 and noise floor setting region 318 used by the DCE to determine the desired Max-Acc-Mix. In DCE Model 5 of Figure 3E, the component or logic configuration is similar to that of Figure 3D, except that ALN logic 310A is located after DRC logic 306A.
[0061] 3A-3E, in some embodiments, the DCE 62 may also provide a second output, Mid-Acc-Mix, on line 93, which is an audio track tapped from a predetermined location within the DCE 62. The Mid-Acc-Mix may be tapped after the dialogue enhancement logic (DE) 304A and before the dynamic range compression (DRC) logic 306A, or after the dynamic range compression (DRC) logic 306A and before the dialogue enhancement (DE) logic 304A, depending on the DCE model, as shown in FIGS. 3A-3E, thereby enabling the use of a three-input X-fade renderer, as described herein and shown in FIG. 10C. In some DCE models, such as DCE Models 1, 3, 4, and 5, the location of the tapping requires the audio tracks to be recombined into a single track. In that case, there is an additional EDI logic block 308A, indicated by the dashed box, that performs the tapping to generate the Mid-Acc-Mix on line 102.
[0062] Returning to Figure 1, the output of DCE logic 62 is a Max-Acc-Mix track, which may be provided to audio encoder 26 on line 92 and stored on non-encoded content server 36. The other input to encoder 26 is the original Cin-Mix audio track on line 96. In some embodiments, as described herein above with reference to Figure 1, the DCE may also provide a second output Mid-Acc-Mix on line 93, which is an audio track that has been dropped at a predetermined location within the DCE.
[0063] Typically, the input to the audio encoder 26 is an audio source waveform (e.g., PCM / WAV audio uncompressed format), and the resulting output is an audio bitstream (i.e., bs), a digitally compressed audio track encoded using, for example, MPEG-4 AAC (Advanced Audio Coding) developed by MPEG Audio or Dolby AC3 / EC3. Other coding standards or codecs may be used for encoding if desired. The use of the bitstream (bs) enables the transmission of complex surround sound formats such as Dolby Atmos or DTS:X. The audio encoder 26 uses known software running on hardware owned / provided by a streaming service such as the CP computer 20. Certain AAC encoding and decoding techniques may be licensed from known AAC patent pools.
[0064] As previously described herein, the output of the encoder 26 is an encoded interleaved audio stream, as shown in Figures 5A, 5B, and 5C, which is stored in the encoding content server 40 via the content API 30, as shown by lines 94 and 95.
[0065] 1 and referring again to FIG. 12A, the portal UI logic 60 also provides a UI that allows the CP user / administrator 21 to adjust X-fade via the display 22, and receives X-fade adjustment selections from the CP user / administrator 21 via user interaction with the display 22 (or other input device connected to the CP computer, such as a keyboard, mouse, or other device). The UI for X-fade control, monitoring, and selection on the display 22 may be as shown in FIG. 12A, which is described further below.
[0066] As described herein, the portal UI logic 60 may include X-fade logic 100 (i.e., cross-fade renderer logic) that preserves power and provides a blend between the current values of Cin-Mix and Max-Acc-Mix. In particular, the X-fade logic 100 (FIG. 1) provides a portal UI having an X-fade slider 1210 (FIG. 12A) that allows a user to adjust the X-fade slider to provide a cross-fade rendering, blending, or mixing of the Cin-Mix and Max-Acc-Mix from the DCE logic 62, whereby the X-fade logic 100 provides a resulting blend of the two input signals on line 80 using a power-preserving gain curve based on the position of the X-fade slider 1210 (FIG. 12A), which provides the X-fade adjustment, i.e., X-fade Adj.
[0067] The power conservation equation used by the X-fade logic 100 is shown in Equation 3. Equation 3: Sin 2 (θ) + Cos 2 (θ) = 1
[0068] where the value of θ (degrees) ranges from 0 to 90 degrees or 0 to 180 degrees (depending on the configuration) and corresponds to the position X of the X-fade slider (0 degrees represents the lower limit of the slider, X=0, and 90 degrees represents the upper limit of the slider, X=1.0). The power-preserving nature of Equation 3 enables the system of the present disclosure to provide loudness changes such that the stereo / spatial image is not affected, particularly in stereo, surround, or immersive sound or audio applications. As the audio signals are processed using the amplitude of the audio signals, sin and cos functions may be applied as coefficients (i.e., multiplicative attenuation gains with values less than or equal to 1.0) to the crossfade input audio signals, respectively, to produce a power-preserving crossfade result.
[0069] 10A, the X-fade logic 100 is shown by a block diagram 1000 and a corresponding graph 1010. The X-fade logic 100 has two inputs, Cin-Mix and Max-Acc-Mix, on lines 1001 and 1005, which are scaled by gains g0 and g1, respectively, defined by a cosine gain curve 1012 (i.e., curve A) and a sine gain curve 1014 (i.e., curve B), and provides an output X-fade mix audio signal on line 1009. In particular, the X-fade logic 100 receives the Cin-Mix signal on line 1001, which is provided to a multiplication gain block 1002 where it is multiplied by a gain g0, shown below as Equation 4. The result of the gain block 1002 is an adjusted Cin-Mix value (Cin-Mix Adjust), which is provided to a summation block 1008 on line 1003. X-fade logic 100 also receives a Max-Acc-Mix signal on line 1005, which is provided to multiplication gain block 1004 where it is multiplied by a gain g1 as shown in Equation 5 below. The result of gain block 1004 is an adjusted Max-Acc-Mix value (Max-Acc-Mix Adjust) provided on line 1007 to summation block 1008. Summation block 1008 adds the Cin-Mix Adj and Max-Ac-Mix Adj values on lines 1003 and 1007 and provides the result of the addition as the X-fade mix on line 1009. Graph 1010 also shows a cosine curve 1014 that determines the value of gain g0 and a sine curve 1016 that determines the value of gain g1. Furthermore, vertical line 1016 indicates the X value (X=0 to 1.0) based on the slider position, and the values of gains g0 and g1 are determined by the positions where the X value intersects with curve 1012 and curve 1014. Below are equations 4 and 5 that determine the values of Cin-Mix Adj and Max-Acc-Mix Adj. Equation 4: Ci-Mix Adj = (Cin-Mix)*Cos(X*90) = (Ci-Mix)*g0 Equation 5: Max-Acc-Mix Adj = (Max-Ac-Mix)*Sin(X*90)=(Max-Acc-Mix)*g1 Here, the value of X ranges from 0 to 1.0 based on the position percentage of the X Fade slider, e.g., X=0 when the slider is at the low end (minimum) end of the slider range, and X=1.0 when the slider is at the high end (maximum) of the slider range.
[0070] Referring to FIG. 10B, the X-fade logic 100 is illustrated for stereo audio content (with left / right channels designated L / R, respectively) by a block diagram 1020 and corresponding graph 1030 having a sine function curve 1034 (curve B) and a cosine function curve 1032 (curve A), similar to that shown in FIG. 10A. In particular, the X-fade logic 100 receives Cin-Mix stereo signals for the left and right (L0 / R0) channels on lines 1021A and 1021B, respectively. The Cin-Mix stereo signals are provided to a multiplication gain block 1022, where both Cin-Mix L0 and Cin-Mix R0 are multiplied by a gain g0 determined by a cosine function curve 1034(B). When the X-fade position X is at the low end (X=0), the gain g0=1 (cos(0*90)=1), and Cin-Mix L0 and Cin-Mix R0 pass through the multiplier (i.e., gain) 1022 without attenuation (gain=1). When the X-fade position X is at the high end (X=1.0), the gain g0=0 (cos(1*900)=0), and Cin-Mix L0=0 and Cin-Mix R0 Let R0 = 0 for full attenuation (gain = 0). Gain block 1022 provides adjusted Cin-Mix on lines 1023A and 1023B for the left and right (L / R) channels, respectively.
[0071] Similarly, the X-fade logic 100 receives Max-Acc-Mix stereo signals for the left and right (L1 / R1) channels on lines 1025A and 1025B, respectively. The Max-Acc-Mix stereo signals are provided to a multiplication gain block 1024, where Max-Acc-MixL1 and Max-Acc-MixR1 are each multiplied by a gain g1 determined by a sinusoidal curve 1034(B) to produce an X-fade position X that is the low end ( When X fade position X is at the high end (X=1.0), gain g1=0 (sin(0*90)=0), Max-Acc-MixL1=0, and Max-Acc-MixR1=0, i.e., full attenuation (gain=0), and when X fade position X is at the high end (X=1.0), gain g0=1 (sin(1*90)=1), causing Cin-MixL0 and Cin-MixR0 to pass through multiplier (i.e., gain) 1024 unattenuated (gain=1). Gain block 1024 provides adjusted Max-Acc-Mix for the left and right (L / R) channels on lines 1023A and 1023B, respectively.
[0072] The adjusted Cin-MixL0 on line 1023A and the adjusted Max-Cin-MixL1 on line 1027B are provided to summing block 1028A, and the output is provided on line 1009A as a combined left channel adjusted signal, i.e., X-fade mix L (left channel). Similarly, the adjusted Cin-MixR0 on line 1023B and the adjusted Max-Cin-MixR1 on line 1027A are provided to summing block 1028B, and the summed output is provided on line 1009B as a combined right channel adjusted signal, i.e., X-fade mix R (right channel).
[0073] 10C, the X-fade logic 100 is illustrated for input stereo audio content (with left / right channels, respectively, indicated by L / R) by a block diagram 1040 for a three-input cross-fade renderer (i.e., X-fade) and a corresponding graph 1050 having three gain curves 1051 (i.e., curve A), 1052 (i.e., curve B), and 1054 (i.e., curve C). In particular, in some embodiments, the DCE 62 (FIG. 1) may provide an additional output, Mid-Acc-Mix, indicated by dashed line 93 (FIG. 1). In that case, the X-fade logic 100 has three inputs (or three sets of inputs for stereo audio) that are Cin-MixL0 / R0, Mid-Ac-MixL1 / R1, and Max-Acc-MixL2 / R2, which are adjusted or multiplied by three gains g0, g1, and g2, respectively, with gain values defined by three gain curves 1051 (i.e., curve A), 1052 (i.e., curve B), and 1054 (i.e., curve C). In particular, the equations for the output left (L) and right (R) channels at lines 1064A and 1064B, respectively, are shown below: Equation 6: X Fade(L)=L0g0+L1g1+L2g2 Equation 7: X Fade(R)=R0g0+R1g1+R2g2 Here, as will be described later, the values of g0, g1, and g2 are defined by three gain curves (sine or cosine) 1051, 1052, and 1054, respectively, where g0, g1, and g2 are varied or changed based on the X-fade slider position X. Gain curve 1051 is a cosine curve from X=0 to 0.5 (used for the first half of the slider range between Cin-Mix and Mid-Acc-Mix at the g0 value), and X=0 to 0.5 (used for the first half of the slider range between Cin-Mix and Mid-Acc-Mix at the g1 value). Gain curve 1052 (for gain g1 value) is a sine curve from X=0 to 1.0 (used for the entire range of the slider from Cin-Mix to Mid-Acc-Mix to Max-Acc-Mix at the g1 value), and gain curve 1054 (for gain g2 value) is a sine curve from X=0.5 to 1.0 (used for the latter half of the slider range from Mid-Acc-Mix to Max-Acc-Mix at the g2 value). Also, in some embodiments, gain curve 1052 (for gain g1 value) may be considered as a sine curve from X=0 to 0.5 and a cosine curve from X=0.5 to 1.
[0074] Specifically, the three-input X-fade logic 100 receives Cin-MixL0 / R0 (left and right channels, referred to herein as L0 and R0, respectively), which is provided to a multiplication (or gain) block 1042, which also receives a gain g0 and calculates values for L0*g0 (i.e., L0g0) and R0*g0 (i.e., R0g0), which are adjustment (or attenuation) values for the Cin-MixL0 / R0 left and right channels. L0g0 is provided to a summation block 1060, which also receives a value of (L1g1+L2g2) from a summation block 1049, described later in this specification. The result of the summation 1060 provides the X-fade mix (L, i.e., left channel) value from Equation 6 above on line 1064A. The R0g0 value from gain block 1042 is also provided to a summation block 1048, described later.
[0075] Additionally, the 3-input X-fade logic 100 receives Mid-Acc-MixL1 / R1 (left and right channels, referred to herein as L1 and R1, respectively), which is provided to a gain multiplication block 1044. The gain multiplication block 1044 also receives a gain g1 and calculates values for L1*g1 (i.e., L1g1) and R1*g1 (i.e., R1g1), which are adjustment (or attenuation) values for the Mid-Acc-MixL1 / R1 left and right channels. R1g1 is provided along with R0g0 (from multiplication block 1042) to summation block 1048, and the result of summation 1048 (R1g0+R1g1) is provided to summation block 1062, which adds the value R0g0+R1g1 (from summation block 1048) to R2g2 (from block 1046) to provide the value of the X fade mix (R, i.e., right channel) from equation 7 above on line 1064B. L1g1 is provided along with L2g2 from gain block 1046, as described further below herein, to summation block 1049, and the result of summation 1049 (L1g1+L2g2) is provided to summation block 1060, as described above.
[0076] Additionally, the three-input X-fade logic 100 receives Max-Acc-MixL2 / R2 (left and right channels, herein referred to as L2 and R2, respectively), which is provided to a gain multiplication block 1046, which also receives a gain g2 and provides the adjustment (or attenuation) values for the Max-Acc-MixL2 / R2 left and right channels, L2*g2 (i.e., L2g2) and R2*g2 (i.e., R2g2). L2g2 is provided along with L1g1 to a summation block 1049, and the result of summation 1049 (L1g1 + L2g2) is provided to a summation block 1060, which provides the value of X-fade mix (L) from Equation 6 on line 1064A.
[0077] Regarding the gain curves, the following equations 8, 9, and 10 may be used to determine the gain values of the gains g0, g1, and g2 for the left and right channels (in a stereo application), e.g., the Cin-Mix adjustment (L0g0, R0g0), the Mid-Acc-Mix adjustment (L1g1, R1g1), and the Max-Acc-Mix adjustment (L2g2, R2g2). Equation 8: g0=Cos(X*180) (X=0~0.5, if X>0.5, set g2=0) Equation 9: g1=Sin(X*180) (X=0 to 1.0) Equation 10: g2=Sin(X*90) (X=0.5~1.0. If X<0.5, set g0=0) where the value of X ranges from 0 to 1.0 based on the position percentage of the X-fade slider, e.g., X=0 when the slider is at the low (minimum) end of the slider range and X=1.0 when the slider is at the high (maximum) end of the slider range, as described herein, and g0 is set to 0 in the latter half of the slider range and g2 is set to 0 in the first half of the slider range.
[0078] In particular, in the first half of the X-fade slider range (X=0 to 0.5), the X-fade logic 100 performs cross-fade rendering or blending on the Cin-Mix and Mid-Acc-Mix (e.g., from 0 degrees to 90 degrees for sine and cosine curves) and provides gain g0 and g1 values, with the value of g2 being irrelevant for the first half of the slider range and set to 0. In the second half of the X-fade slider range (X=0.5 to 1.0), the X-fade logic 100 performs cross-fade rendering or blending on the Mid-Acc-Mix and Max-Acc-Mix (e.g., from 90 degrees to 180 degrees for sine and cosine curves) and provides gain g1 and g2 values, with the value of g0 being irrelevant for the first half of the slider range and set to 0. The resulting X-Fade slider ranges from X=0 to 1.0 and is blended progressing from Cin-Mix to Mid-Acc-Mix and from Mid-Acc-Mix to Max-Acc-Mix. As described herein, in each section, the two input signals use a maintained gain curve based on the position of the X-Fade slider (or X-Fade Adjustment, or X-Fade Adj).
[0079] Graph 1050 also shows cosine curve 1051, which determines the value of g0, sine curve 1052, which determines the value of g1, and sine curve 1054, which determines the value of g2. Vertical line 1055 indicates the X value based on the slider position, and the positions where this X value intersects with curves 1052 and 1054 determine the values of g1 and g2, respectively. In particular, as also shown by Equations 8, 9, and 10, graph 1051 (cosine function) determines the value of g0, graph 1052 (sine function) determines the value of g1 from 0 to 90 and from 90 to 180, and graph 1054 (cosine function) determines the value of g2 from 90 to 180. Also, as described herein, the value of g2 may be set to 0 in the first half of the slider range (X=0 to 0.5), and the value of g0 may be set to 0 in the second half of the slider range (X=0.5 to 1.0).
[0080] The portal UI logic 60 provides an X-fade UI on line 84 and a content UI on line 82 to the display 22, which allows the CP user / administrator 21 to select content, which is provided to the portal UI logic 60 on line 78. The portal UI logic 60 receives an X-fade slider (X-fade Adj) input from the CP user / administrator 21 on line 80, which indicates the current position of the X-fade renderer. The output of the X-fade logic 100 is an X-fade mix audio signal or track that is provided to the speakers 23 of the CP computer 20 on line 86, allowing the CP user / administrator 21 to hear the resulting accessible mix (e.g., Max-Acc-Mix), as shown on line 76. The portal UI logic 60 can also request and receive a content list (or index) on line 99 that lists audio content or audio files available on the non-encoded content server 36 for review and selection by the CP user / administrator 21.
[0081] 2A, 2B, 2C, and 2D, several top-level block diagrams 200A, 200B, 200C, and 200D are shown, respectively, of different configurations or applications on the user / listener side of a system according to the present disclosure. In particular, FIG. 2A shows block diagram 200A illustrating a smart TV as smart playback device 12 as the primary interface with the user via television display 14 and speaker 11. In that case, the X-Fade and personalized EQ functions are controlled by user / listener 15 using a smart TV remote control (or remote) 220 or a user device 12, such as a smartphone. In either case, user 15 can control X-Fade and various other Acc app functions or parameters by viewing the smart TV display 14. Details of FIG. 2A are provided further below in this specification. FIG. 2B is a block diagram 200B illustrating a single smart playback device, such as a smartphone, table, laptop, etc., used by a single user. The smart playback device can operate in two modes: standalone mode (i.e., single device mode) or hub receiver (i.e., slave) mode, in which the device is a slave to a smart TV operating in hub mode. In standalone mode, device 12 receives content and interacts with the user similarly to the smart TV described earlier in this specification, except that there is no remote control. In hub receiver (i.e., slave) mode, device 12 receives all C-Mix video on line 217 from the smart TV operating in hub mode, and receives and provides instructions on line 231B. The remaining functionality is the same as that described in FIG. 2A. In some embodiments, when configured in a hub mode configuration or operating mode, the Acc app may be referred to as a Hub Acc app.
[0082] FIG. 2C is a block diagram 200C illustrating a smart TV 12 configured to operate in a hub mode that allows multiple user devices 12, e.g., user device 1 through user device N. In this mode, each user / listener 15 may use their own device 12 to view content being played on the smart TV 12. Each user device 1 through N then receives a UI (including the X-fade UI and Pers. EQ UI) and can view selected content from the smart TV via lines 231A and 233A. This allows people in the same room watching the same program on a smart TV to have a personalized audio experience. For example, each user / listener 1 through N has their own X-fade slider that can be set to their own personal setting. This also applies to each user / listener's Pers. EQ function. FIG. 2D is similar to FIG. 2C and shows a block diagram 200D with a smart TV 12 configured to operate in hub mode, where there are multiple decoders 202, allowing multiple users / listeners not only to watch the same program as other users / listeners in the room, but also to select video and language formats for the same program, e.g., language, voiceover, narration, etc. (See FIG. 12C for a UI described later in this specification). For example, user / listener 1 may be listening to a program in English, user / listener 2 may be listening to a program in French, user / listener 3 may be listening to a voiceover in Spanish, user / listener 4 may be listening to the same program with narration, etc. Thus, each user / listener 15 may use their device 12 to watch the same program, albeit from a different stream due to the type of content being played, as shown in FIGS. 5A-5D. FIGS. 5A-5D show how data is stored, encoded, and interleaved. In this case, each of the user devices 1 to N receives the UI (including the X-fade UI, Pers. EQ UI, and content UI) and can view the selected content.The user / listener can also choose to sync with other listeners based on a selectable Sync (i.e., Smart TV Sync) button on the Acc app UI 1270 (FIG. 12C) (discussed further later herein).
[0083] 2A , a top-level block diagram 200A of the user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with a remote-controlled or user device-controlled crossfade audio mix, according to an embodiment of the present disclosure, is shown. In particular, a smart playback device 12, such as a smart TV, is shown having a remote control (i.e., remote) 220 or user device 12 that is used to control the crossfade logic (X-Fade Logic) 100 functionality and other Acc App UI functions described herein via the smart TV display 14. In some embodiments, a user / listener 15 may communicate with the smart TV and Acc App Logic 16 using the remote control 220, which communicates (wirelessly or wired) with a remote control interface or remote receiver 216 within the smart TV 20. In that case, the user / listener 15 presses button 218 on the remote 220, as indicated by line 243, which causes the Acc App logic 16 in the Smart TV to provide an X-Fade / EQ / Content UI on line 223 to the Smart TV display 14, as further described below with reference to FIG. 12B, which may include the X-Fade slider shown in FIG. 12C and other UI display features (including an audio content list, a personalized equalizer, and other Acc App UI features described herein).
[0084] More specifically, the Acc app logic 16 receives the Max-Acc-Mix and Cin-Mix on lines 207, 209 from a known audio decoder 202, such as a known AAC-LC decoder currently provided and supported by streaming apps from Apple®, Amazon®, Roku®, etc. The audio decoder 202 receives the encoded audio signal or mix, which may include the Cin-Mix and Max-Acc-Mix interleaved in a single bitstream, and uses known decoding software to decode the bitstream (i.e., bs) into individual streams as the Cin-Mix provided on line 207 and the Max-Acc-Mix provided on line 209. The audio decoder or decoder 202 (FIGS. 2A, 2B, 2C, 2D) also parses and decodes the input audio bitstream into pre-encoded tracks, such as the Cin-Mix and Max-Acc-Mix, as is well known. The audio decoder 202 uses software that runs on the operating system (OS) of the device streaming the content (iOS, Android, Roku, etc.) and can use the same codecs as those used by the encoder 26.
[0085] Additionally, the Acc applilogic 16 receives a noise floor (NF) signal on line 213, which indicates the noise floor of the listening environment. The environmental noise floor NF may be estimated by receiving environmental sound from a microphone 206, which provides the environmental sound to a known filter 204 on line 211, which filters (or blocks) the frequency range of the human voice and provides the NF on line 213. The NF value may be used to adjust the position of the X-fade slider of the X-fade logic 100. In particular, if the NF is available, the system may first set the X-fade slider position to a value that allows the minimum loudness of the dialogue (Dx) LRA (loudness range) to exceed the noise floor (NF) or the maximum value of the X-fade slider position, whichever is lower. In some embodiments, as described further herein, the X-fade logic 100, i.e., the Acc App Logic 16, may ensure that the LRA of Dx and M&E does not exceed the maximum allowed audio level, e.g., -2 dB (true peak). Accordingly, the Acc App Logic may also include known loudness detection software for determining the LRA. In some embodiments, the content LRA values may be provided in the video content metadata or stored separately from the video on the Encoded Content Server 40 or other server.
[0086] The Acc App Logic 16 also receives a content selection from the remote control 220 on line 229 and an X-Fade / EQ adjustment signal from the remote control 220 on line 227. The Acc App Logic 16 also provides an X-Fade / EQ audio mix to the speaker 11 on line 225 and provides an X-Fade / EQ and content UI to the display 14 on line 223. The X-Fade / EQ audio mix is a digital audio signal when output from the X-Fade Logic and is then processed through a known digital-to-analog converter (DAC) to convert the mix to an analog signal before being provided to the speaker 11. The Acc App Logic 16 also provides the X-Fade / EQ and content UI to the user device 12 on a group of lines 230 and receives a selected content request signal and an X-Fade / EQ adjustment signal on a group of lines 230 for communication with the user device 12. The Acc app logic 16 also communicates with a user attribute server 42 (line 219) that contains information and data related to a given user device 12 or user / listener 15. The Acc app logic 16 also requests and receives an audio content list (or index or audio file list) (line 215) from the encoding content server 40 (line 201), the audio content list listing audio content or audio files available on the encoding content server 40 for review and selection by the user 15. The Acc app logic 16 also provides content selection requests received from the user / listener 15 to the content API 30 (line 205), which retrieves the requested content from the encoding content server 40 (line 255) and provides the selected content (e.g., Cin-Mix+Mac-Acc-Mix interleaved into a single bitstream) to the audio decoder 202 (line 203).
[0087] Additionally, Acc App Logic 16 includes logic for performing various functions, such as X-Fade Logic 100, Acc App UI Logic 210, and Personalized Equalizer (Pers. EQ) Logic 208. In particular, X-Fade Logic 100 performs crossfade rendering or blending of the Cin-Mix and Max-Acc-Mix, providing the resulting blending of the two input signals based on the position of the X-Fade slider (or X-Fade adjustment) using a power-preserving gain curve using the power-preserving X-Fade equation (Equation 3) described above with reference to Figure 1. Personalized Equalizer (Pers. EQ) Logic 208 is described below with reference to Figure 11.
[0088] The Acc App UI logic 210 provides a user interface (UI) including an X-fade UI and a Content UI on the smart TV display 14 on line 223, as shown in Figures 12B and 12C described later in this specification. The Acc App 16 receives X-fade adjustment commands from a remote on line 227 or from the user device 12 on group on line 231. The X-fade logic 100 performs the X-fade function based on the commands from the X-fade slider, similar to that described with reference to Figure 1. The output of the X-fade logic 100 is an X-fade mix (i.e., Acc-Mix or personalized mix or personalized accessible mix or Pers-Mix) on line 225 that is provided to the speakers 11, and the resulting output audio from the speakers is provided (over the air) to the user / listener 15, as indicated by dashed line 239. Additionally, the video portion of the selected content Cin-Mix (i.e., C-Mix) may be provided from the content API 30 to the smart TV's display 14 on line 217, and the video portion may be retrieved from the encoding content server 40 on line 257. The video portion, the C-Mix video, and the X-fade / EQ / content UI (described herein below) on line 223 are provided to the user / listener 15 by the display 14, as indicated by dashed line 241.
[0089] Acc App Logic 16 can also request and receive a content list (or index) at line 215. The content list lists audio content or audio files available on Encoded Content Server 40 for review and selection by User / Listener 15.
[0090] Referring to Figure 2B, a top-level block diagram 200B of a user / listener side of a system for providing personalized audio streaming and rendering for a single user device according to an embodiment of the present disclosure is shown, as previously described. Referring to Figure 2C, a top-level block diagram 200D of a user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices, each with a personalized cross-fade audio mix, according to an embodiment of the present disclosure. Referring to Figure 2D, a top-level block diagram 200 of a user / listener side of a system for providing personalized audio streaming and rendering for a smart TV with multiple user devices, each with a personalized cross-fade audio mix and content selection, according to an embodiment of the present disclosure, as previously described.
[0091] Referring to FIG. 5A, diagram 500 illustrates interleaved Cin-Mix and Max-Acc-Mix content audio streams for accessible audio provided by audio encoder 26 in accordance with an embodiment of the present disclosure. As described herein, audio encoder 26 (FIG. 1) analyzes and compresses input audio sources associated with a video source. The left side of FIG. 5A illustrates two unencoded files, the (original) Cin-Mix English audio file and the Max-Cin-Mix audio file (or track), that are input to encoder 26 for selected content. The right side of FIG. 5A illustrates the output of encoder 26 as an interleaved audio stream for accessible audio, which is stored by content provider computer 20 (FIG. 1) on an encoded content server as described herein. In particular, the input audio files are sampled and interleaved as shown in FIG. 5A, respectively, as sample 1, sample 2, etc., with each sample comprising a sampled portion of each input file.
[0092] Referring to FIG. 5B, a diagram 515 of interleaved Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix content audio streams is shown, according to an embodiment of the present disclosure.
[0093] Referring to FIG. 5C , a diagram 530 of interleaved Ci-Mix and Max-Acc-Mix content audio streams for accessible audio with two languages for voice-over audio is shown, according to an embodiment of the present disclosure. In this case, for a stereo audio mix, two or three of the four displayed mixes (if stereo) can be combined to provide voice-over or adaptive voice-over functionality. For example, one or both of the dashed boxes can be optional. Thus, if the audio mix is mono (one channel per track), the stream can include all four displayed tracks available for use if desired. In this case, upon receiving four tracks, the Acc App can select three to provide to the X-Fade Renderer, or, in some embodiments, can use a four-input X-Fade Renderer if desired.
[0094] In particular, audio encoder 26 (FIG. 1) provides interleaved data by placing samples from each audio channel (such as the left and right channels of a stereo mix) one after the other into a single stream, effectively "interweaving" the channels together, so that data from the first channel is followed by data from the second channel, and so on, resulting in a single continuous data stream that can be easily decoded and played back by a compatible audio player, as is known.
[0095] As shown in Figures 5A, 5B, and 5C, the present disclosure uses encoder 26 to interleave two (or more) different tracks (e.g., Cin-Mix and Max-Acc-Mix) and provide a single stream with both tracks embedded. This allows a user / listener (after decoding) to crossfade the two (or more) tracks to create a personalized listening experience while receiving only a single stream. Thus, for a standard stereo mix, the present disclosure uses four channels: two for the Cin-Mix (L / R) and two for the Max-Acc-Mix (L / R), i.e., two additional channels. For three stereo tracks (e.g., Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix), the present disclosure uses six channels: two for the Cin-Mix (L / R), and two for the Mid-Acc-Mix and Max-Acc-Mix (L / R), for a total of four additional channels. For 5-channel surround 5.1 mixes and 10-channel (including subwoofer) immersive 5.1.4 mixes, the DCE processes only the center channel (or dialogue channel) of the mix to create the Max-Acc-Mix, so the encoder only interleaves one additional channel for the Max-Acc-Mix. Two additional channels are required if both the Mid-Acc-Mix and Max-Acc-Mix are used.
[0096] 5D , a block diagram 550 is shown of a non-encoding content server 36 having unencoded Cin-Mix and Max-Acc-Mix audio content streams stored separately as input to audio encoder 26, and an encoding content server 40 having encoded and interleaved Cin-Mix / Max-Acc-Mix audio content files as output from audio encoder 26, according to an embodiment of the present disclosure. In particular, the Cin-Mix (Original) file may be provided by a studio that created the cinematic mix, as described herein, and the Max-Acc-Mix (From DCE) may be provided by DCE logic 62. In some embodiments, a studio or content provider may provide the Cin-Mix and Max-Acc-Mix (i.e., two files, Cin-Mix English (Original) and Max-Acc-Mix English (From DCE)) in multiple languages to appeal to a wide audience, e.g., in 40 or more languages, as shown in table 550. Additionally, in some embodiments, for particular content, a studio or content provider may provide a Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix (i.e., three files: Cin-Mix Japanese (Original), Mid-Acc-Mix Japanese (From DCE), and Max-Acc-Mix Japanese (From DCE)) in a particular desired language, as needed, as shown here for example for Japanese.
[0097] Additionally, in some embodiments, voiceover options for inclusion in the encoded mix can include various combinations of desired audio files, such as Cin-Mix English (Original Version OV), Max-Acc-Mix English (from DCE) + Max-Acc-Mix Spanish (from DCE), or Max-Acc-Mix English (from DCE) + Max-Acc-Mix Spanish (from DCE) + Separate-AD Spanish (audio description (AD) in the same language as the voiceover, in this particular use case, making the content more accessible to non-native speakers of English OV). Other combinations can also be used as needed. The "+" sign above refers to the addition / addition / overlay of two or three underlying audio signals. The audio description (AD) track here carries only the audio description, not the entire mix (unlike the approach below).
[0098] In some embodiments, in addition to what is shown in Figure 5D, the narration or audio description (AD) mix can include Cin-Mix English / Cin-Mix English AD / Narration, Cin-Mix Spanish / Cin-Mix Spanish AD / Narration, etc. In this case, the bottom end of the X-Fade slider displays only the selected language without narration, and raising the slider increases the volume of the selected narration language. Also, in some embodiments, in addition to what is shown in Figure 5D, the adaptive voice-over (VO) mix can include Cin-Mix English / Cin-Mix Spanish, or Cin-Mix English / Cin-Mix Spanish VO Beginner, or Cin-Mix English / Cin-Mix Spanish VO Intermediate, or Cin-Mix English / Cin-Mix Spanish VO Expert.
[0099] Referring to FIG. 5E, a table 580 is shown with various encoding and decoding options for a given stream type and user-desired configuration, according to an embodiment of the present disclosure. In particular, the present disclosure provides an embodiment including a single multi-channel audio encoder / decoder. MPEG AAC-LC and Dolby E-AC3 support a "dual stereo" encoding mode, for example, where two 2.0 / stereo streams are independently encoded. MPEG AAC-LC and Dolby E-AC3 support a "quad mono" mode, for example, where four 1.0 / mono streams are independently encoded. For example, Dolby Digital Plus with Joint Object Coding (JOC), often referred to as Dolby Digital Plus with Dolby Atmos, can transmit primary and associated audio, including object-based audio for Atmos, in a single bitstream. In table 580, 2.0 and 4.0 represent two and four channels, respectively. Also, 100% of Dolby devices can decode AAC-LC. The same is true for other audio compression technology codecs such as DTS, Opus, and ITU codecs. AC3 (or AC-3) is commonly referred to as Dolby Digital, and EC3 (or enhanced AC-3) is commonly referred to as Dolby Digital Plus, and as such, provides enhanced audio quality. In some embodiments, open-source software tools such as FFmpeg® or Gstreamer® may also be used with the present disclosure to process audio files. Furthermore, in some embodiments, for Dolby Atmos 5.1.4 (cinematic immersion), the encoder may be a Dolby EC3 + JOC encoder with 5.1 core main + JOC metadata and associated bsmod 1.0 per ETSI TS 102 366 V1.4.1.Also, in some embodiments, for Dolby Surround 5.1, the encoder may be a Dolby EC3, 7-channel encoder, e.g., 5.1 main + 1.0 associated per bsmod of ETSI TS 102 366 V1.4.1. Also, in some embodiments, for stereo sound or audio, the encoder may be 2 x enc / dec for 2.0 dual stereo sound, 4 x enc / dec for 1.0 quad mono sound, or 1 x enc / dec for 4.0 multi-channel sound. Also, in some embodiments, for surround sound, the encoder may be 1 x enc / dec = 5.1 + 1 (center channel) for 6.1. And in some embodiments, for immersive sound, the encoder may be 1 x enc / dec = 5.1.4 + 1 (center channel).
[0100] Also, as described herein, in some embodiments where Dolby / DTS (proprietary codecs) are used, a 2.0 downmix can be performed before running DCE. Also, in some embodiments, the system can receive (or create) a Cin-Mix 2.0 and then create a Max-Acc-Mix 2.0. The Max-Acc-Mix 2.0 can be encoded (at the user / listener) and then decoded and provided to the X-Fade Renderer to provide an X-Fade mix in 2.0 format.
[0101] 6, the system of the present disclosure shown in Figures 1 and 2A-2D may be implemented in a network environment 600. The system includes a content provider computer 20, one or more user devices 12, and various servers 35 that interact with the computer 20 or user devices 12 to perform the functions described herein.
[0102] In particular, the various components of one embodiment of the disclosed system include multiple computer-based user devices 12 (e.g., smart televisions and device 2 through device N) that may interact with respective users (user / listener 1 through user / listener N). A given user 15 may be associated with one or more devices 12. In some embodiments, Acc app 19 may reside on user device 12 or on a remote server and communicate with user device 12 over a network. If user device 12 is a smart television 12, user 15 may communicate with the smart television via remote control 220 using known remote control interfaces (FIGS. 2A-2D) within smart television 20, or receiver 216, as previously described herein. In some embodiments, user device 12 may communicate directly with the smart television using a wired or wireless connection (e.g., Bluetooth, Wi-Fi, NFC, or other wireless connection), as indicated by line 71.
[0103] In some embodiments, one or more user devices 12 may be connected to or communicate with each other via a wired or wireless communications network 70, such as a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a peer-to-peer network, or the Internet, as indicated by line 72, by transmitting and receiving digital data via the communications network 70. If the user devices 12 are connected via a local, private, or secure network, the user devices 12 may have a separate network connection to the Internet (or other network) used by a web browser operating on the devices 12. The user devices 12 may also have a web browser 17 for connecting to or communicating with the Internet to obtain desired video / audio content and to obtain Acc App 19 or other necessary files necessary to execute the logic of the present disclosure in a standard client-server based configuration. The user device 12 may also have local digital storage located on the device itself (or connected directly to the device, such as an external USB-connected hard drive, thumb drive, etc.) for storing data, images, audio / video, documents, etc. that can be accessed by the Acc app 16 running on the user device 12.
[0104] The user device 12 may also communicate over the network 70 with separate computer servers 35, such as a content AP server 31 (having a content selection API 30), a non-encoding content server 36, a DCE parameter server 38, an encoding content server 40, and a user attribute server 42. The servers 35 may be any type of computer server having the necessary software or hardware (including storage capabilities) to perform the functions described herein. The servers 35 (or the functions performed by them) may also be located, individually or collectively, in one or more separate servers on the network 70, or may be located in whole or in part on one or more user devices 12 on the network 70.
[0105] The content provider computer 20, as described herein, may be a smartphone, computer, laptop, tablet, or other computer-based device, and may receive instructions from the content provider user / administrator 21, have portal / DCE logic 28 running on the computer 20, receive non-encoded (or unencoded) content from the non-encoding server 36, as shown by line 51, generate encoded content (e.g., streaming content, cinematic or programmatic VOD content, or other content) for use by the user / listeners 15 (user / listener 1 through user / listener N) via their user devices 12, and distribute it to the encoded content server 40 via a communications network 70, as shown by line 52, using a web browser 24 (or web server or equivalent interface). Similarly, as described herein, user device 12 may be a smartphone, computer, laptop, tablet, smart TV, or other computer-based device, and may receive digital content (e.g., streaming programs or VOD programs), including encoded content from a content provider, over a communications network 70, either directly or from content stored on an encoded content server 40, for viewing by a user / listener 15, for example, using a web browser 24 (or web server or equivalent interface).
[0106] Portal / DCE logic 28 running on computer 20 may also provide audio / video for CP users / administrators to obtain, update, and save DCE parameters to DCE Params server 38. Computer 20 may also include a known audio encoder 26, as described herein, that encodes audio content into a desired format (e.g., stereo, Dolby, Dolby ATMOS, etc.) for use by user 15, as described herein.
[0107] The CP computer 20 has a Portal UI that provides model performance X-fade control, waveform window, DCE gain / model / noise floor selection data, audio output, and accessible audio approval selection capabilities to the content provider user / administrator 21, as described herein. Also, as described herein with reference to FIG. 1, the Portal / DCE logic 26 may obtain the DCE parameters of the DCE and display them in the Portal UI on the display.
[0108] Further, the content selection API 30 may reside in a content selection API server 31, which may communicate with the user device 12, the CP computer 20, other servers, each other, or other network-enabled devices or logic as needed via a network 70 to provide the functionality described herein.
[0109] Similarly, each user device 12 may also communicate via network 70 with the logic or software applications described herein and other network-enabled devices or logic necessary to perform the functions described herein.
[0110] Portions of the present disclosure shown in this specification as being implemented external to the user device 12 may also be implemented within the user device 12 by adding software or logic to the user device 12 to perform some of the functions described herein or other functions, logic, or processes described herein, for example, by adding logic to the Hub Acc App or Acc App software, or by installing new / additional application software, firmware, or hardware.
[0111] Referring to FIG. 11, a block diagram of the components of personalized equalizer (Pers. EQ) logic 208 (FIGS. 2A-2D) is shown, according to an embodiment of the present disclosure. In particular, Pers. EQ logic 208 receives an X-fade mix on line 1103 and performs digital frequency signal processing on the X-fade mix via known digital frequency signal processing logic 1102, which amplifies specific frequency ranges based on frequency range coefficients on line 1105 from user attribute server 42. Digital frequency signal processing logic 1102 may use logic similar to known hearing aid logic or other audio frequency range amplification devices. The frequency range coefficients may be pre-stored in user attribute server 42 based on third-party audiograms or hearing tests previously obtained via a user-performed hearing test on the user device.
[0112] In some embodiments, frequency range coefficients may be provided by the system of the present disclosure based on a user's listening test. In that case, a content provider may provide a CP app listening test that provides a series of sounds at various frequencies to the user / listener 15 to test their hearing, allowing the user / listener 15 to selectively adjust the frequency amplification at each tested frequency range, as indicated by frequency range adjustment boxes 1106 (which may be UI), and provide the frequency range coefficients for the given user / listener 15 to the digital frequency signal processing logic 1102 on line 1107. The Pers. EQ logic 208 may also store the frequency range coefficients for the given user / listener 15 in the user attribute server 42 for future use by the user / listener 15. The UI for the listening test UI 1290 is shown in FIG. 12D , described later in this specification, and provides frequency range adjustment sliders for the user / listener to adjust.
[0113] Referring to FIG. 12A, a screen view 1200 of a graphical user interface (UI) or portal UI for adjusting DCE parameters and monitoring Max-Acc-Mix is shown, according to an embodiment of the present disclosure, which may be used by a content provider (CP) user / administrator with the ability to set / adjust DCE parameters across a crossfade range and hear the results. In particular, the portal UI 1200 may include a field 1206 for inputting an audio / video file of interest, which may be selected from a content audio file list shown in window 1204, which lists available files and allows the user to select the file of interest to generate an accessible track. Once an audio file is selected, the system displays the file in field 1206. The portal UI 1200 also displays an accessibility fader (or X-fade and volume control) window 1202 (or portion of the screen). The X-fade mix and volume may be controlled by X-fade and volume sliders 1210 and 1212, respectively. One end (upper or top side) of the X-Fade slider 1210 provides the enhanced dialogue alone (i.e., Max-Acc-Mix), while the other end (lower or bottom side) of the X-Fade slider 1210 provides the original cinematic mix (Cin-Mix). Between the two ends of the slider, the X-Fade renderer provides a blend of the Cin-Mix and Max-Acc-Mix, as described herein. The sliders 1210, 1212 can be moved manually by the user / listener using a smart TV remote control (as described herein) or by tapping (or clicking) or touching the slider 1210 when used with a smartphone, laptop, desktop, or tablet computer. In some embodiments, the sliders can be adjusted using voice commands, for example, when the smart TV or user device is connected to a voice-enabled app.The present application supports interaction with most voice-enabled products, such as Siri®, Google Assistant®, and Alexa®. Additionally, a waveform window 1808 may be displayed on the UI screen. The waveform window 1208 provides a display of the Cin-Mix and Max-Acc-Mix while showcasing the crossfader and X-fade gain curves described herein. In particular, the waveform window 1208 shows an example of monitoring an audio signal using known PureData utilities on macOS for a Stereo 2.0 use case showing the original mix (Cin-Mix) and the dialogue-enhanced mix (Max-Acc-Mix). Window 1208 shows how a content provider (CP) user / administrator 21 can experience a real-time crossfade between (on the left) Cin-Mix (original mix) (left / right) and (on the right) Max-Acc-Mix (left / right) (dialogue enhanced mix), allowing the user / administrator 21 to see and hear the DCE's audio effects to generate the accessible mix (i.e., Max-Acc-Mix).
[0114] The portal UI 1200 also provides a window or screen section 1220 that provides the CP user / administrator with the ability to select various DCE parameters (or DCE Params), such as DCE model type, DRC gain profile, Gmne, MNE gain factor, Gdx, Dx gain factor, noise floor, and input or automatic calculation of ALN gain values. The DCE parameter window 1220 may include the ability to set or adjust the dialogue-to-program loudness or dialogue-to-reverberation loudness (DPL or DRL), e.g., the clearance in dB between the dialogue and the reverberation / MNE (not shown). Other parameters may also be included in the DCE parameter window 1220, as desired. Additionally, the Portal UI also has a button 1222 that may be selected when the Max-Acc-Mix is at an acceptable level (Max-Acc-Mix OK). The DCE parameter window 1220 allows the user to set the ALN gain to a desired value, or if the user selects (or clicks) the Auto box, the system will automatically select an ALN gain that maximizes the dynamic range (i.e., loudness range) of the accessible audio mix (Max-Acc-Mix) above the set noise floor (NF), as described using Equation 1, without exceeding the requirements described herein for maximum audio level. The DCE parameters window 1220 also allows a DRC gain profile (or DRC gain curve) to be set based on a desired gain profile provided in a user-selectable file. As described herein, different profiles can provide different levels or amounts of compression to the M&E mix, including an amplification range, a null band range (or unity gain region), an initial cut range, and a cut range, as shown by the DRC gain curve 900 in FIG. 9. Other ranges or labels can be used to describe the DRC gain curve if desired. Content owners can set a specific noise floor, for example, based on a pre-defined audio standard noise floor or based on target audience or user preferences.
[0115] 12B, a screen view 1260 of a smart TV display 14 is shown that is a graphical user interface (UI) for the remote control 220 for manually adjusting the crossfade (X-fade) renderer (or accessibility fader) window 1202 of FIGS. 2A, 2C, and 2D in accordance with an embodiment of the present disclosure. In particular, when a user performs a "long press," e.g., press and hold the up or down arrow for more than two seconds, the smart TV recognizes this as a special command, and the Acc app 210 (FIG. 2A) in the smart TV 12 displays the accessibility fader (X-fade and volume control) window 1202 on the TV display screen 14. Furthermore, when the plus (+) volume arrow 1264 is pressed and held (e.g., more than two seconds) on the remote 220, a signal is sent to the smart TV, as indicated by dashed line 1262, which is interpreted by the Acc app 210 as a command to increase the X-fade slider 1210 position. Conversely, if the plus (+) volume arrow 1264 is pressed for a short time (e.g., less than 2 seconds) on the remote 220, a signal is sent to the smart TV that is interpreted as a command to increase the volume 1212. A similar result occurs when the minus (-) arrow 1266 is pressed and held on the remote control 220, as indicated by dashed line 1262, a signal is sent to the smart TV that is interpreted by the Acc app 210 as a command to increase the X-fade slider 1210 position. Conversely, if the minus (-) volume arrow 1266 is pressed for a short time (e.g., less than 2 seconds) on the remote 220, a signal is sent to the smart TV that is interpreted as a command to decrease the volume 1212. As described herein, the X-Fade slider 1210 transitions between the original dialogue or original cinematic mix (Cin-Mix) and the enhanced dialogue or maximum accessible mix (Max-Acc-Mix).
[0116] Referring to FIG. 12C, a screen view 1270 of the Acc app's graphical user interface (UI) used by a user with the ability to set / adjust crossfades and enable certain features is shown, according to an embodiment of the present disclosure. In particular, the UI 1270 may include a field 1206 for inputting an audio / video file of interest, which may be selected from a content audio file list shown in window 1204, similar to that shown in FIG. 12A, which lists available files and allows the user to select a file of interest. Once an audio file is selected, the system displays the file in field 1206. The UI 1270 also displays an accessibility fader (or X-fade and volume control) window 1202 (or portion of the screen). The X-fade mix and volume may be controlled by sliders 1210 and 1212, respectively, similar to that shown in FIG. 12A. One end of slider 1210 provides fully enhanced dialogue (i.e., Max-Acc-Mix), and the other end of slider 1210 provides the original cinematic mix (Cin-Mix). When slider 1210 is between the two ends, the output audio is a combination of Cin-Mix and Max-Acc-Mix blended using the power conservation curves shown in Figures 10A-10C and described earlier in this specification for the X-fade logic of Figures 1 and 2A-2D.
[0117] Additionally, the user / listener UI 1270 may display several other fader displays, such as video voice over 1280, live sports fader 1 1282, and live sports fader 2 1284. Additionally, the user / listener UI 1270 may display an Acc function window 1288, as described herein, which allows the user 15 to select various features of the Acc app, such as remote, hub mode, single device mode, smart TV sync, microphone (presence or absence), live sports, personalized EQ, adaptive voice over (VO), and audio description (or narration). The user / listener UI 1270 may also display an X-Fade Mix OK button 1294, which allows the user to select or click if the settings for a given selected content are acceptable to the user / listener 15.
[0118] In particular, when the microphone button is selected, the X-fade logic automatically adjusts the X-fade slider 1210 to a level above the measured noise floor (NF). The user may then further adjust the slider 1210 as needed to provide an optimal listening experience for the user 15. Additionally, when the live sports button is selected, a live sports fader 1 window 1282 and a live sports fader 2 window 1282 appear on the UI. Additionally, when the audio description button is selected, the accessibility fader window 1202 may display a range from an original version (OV or English OV) mix at the lower end of the X-fade slider range to an audio description (OD or English AD) mix at the upper end of the X-fade slider range, similar to that described in FIG. 14E. When the accessibility mode is selected, the device 12 recognizes that it is communicating with a hub device, e.g., a smart TV, which provides content directly to the user device. Additionally, when single device (or standalone) mode is selected, device 12 recognizes that it must obtain audio content directly from the video / audio source, rather than through the smart TV, and when remote mode is selected, device 12 recognizes that it will communicate with the user device UI to control the UI on the smart TV.
[0119] When the smart TV synchronization mode is selected, device 12 synchronizes each user device with the smart TV while playing content, so that the user device and smart TV synchronize with each other and play the same content. This may also be done (using multiple decoders) when there are multiple different streams of the same content being provided from the smart TV to different devices, such as original content, alternate languages, voice-over, adaptive voice-over, and audio descriptions, as shown in FIG. 2D. The synchronization request may be part of communication with the hub Acc app 16 on the smart TV. If desired, the status of each Acc function 1288 may also be included in the communication with the hub Acc app (e.g., lines 231B, 233B, 231A, and 231B in FIGS. 2C and 2D).
[0120] Additionally, when the Personalize EQ button is selected, a Pers. EQ window 1286 appears, allowing the user to select where the frequency range coefficients or weighting coefficients (FIG. 11) will come from, as further described with reference to FIG. 11. This includes allowing the user to perform a hearing test to determine the frequency range coefficients. Thus, the present disclosure can provide a personalized audio experience tailored to the hearing ability of the user / listener 15. Additionally, when the Adaptive Voice-Over button is selected, a Language Level Selection window 1289 appears, allowing the user to select a level of narration language level, for example, beginner, intermediate, or expert. Other levels may be provided as needed. As described herein, content providers may offer additional mixes for specific translation levels based on the user's desired level. In that case, the system filters the content list based on the user's desired level. Additionally, the Language window 1290 may be used at any time as a language filter to provide only content of a given language in the content list 1204, if desired.
[0121] The user / listener UI 1270 may also display a personalized EQ window 1286 that allows the user to select between types of personalized equalizer sources, including third-party audiograms, device apps, or OS (e.g., iOS or Android) or content provider (CP) app listening test modes, as described herein. The user / listener UI 1270 may also display a language selection window 1290 that allows the user to select content in a particular language, as described herein. The user / listener UI 1270 may also display a language comprehension window 1290 that allows the user to select a voice-over (VO) language comprehension level to select VO content in a particular language, as described herein.
[0122] 12D , a screenshot 1290 of a graphical user interface (UI) for a listening test to determine parameters of a personal equalizer is shown, according to an embodiment of the present disclosure. In particular, the UI 1290 may display an input frequency range control window 1292, which allows a user to adjust the volume (or gain) of different frequency ranges based on the position of a slider associated with each range. The user / listener UI 1270 may also display a before-adjustment hearing results window 1294, which shows the output results before frequency adjustments are made to a given ear of the user / listener. The user / listener UI 1270 may also display a after-adjustment hearing results window 1296, which shows the output results after frequency adjustments are made to a given ear of the user / listener. The user / listener UI 1270 may also display a button 1291, which may be selected or clicked to initiate the test. As described herein, the User / Listener UI 1270 may be launched when the user selects the CP App Listening Test Mode option on the Acc App UI 1270.
[0123] Referring to FIG. 13A, flow diagram 1300 illustrates one embodiment of a process or logic for executing portal / DCE logic 28 (FIG. 1) in accordance with an embodiment of the present disclosure. Process 1300 begins at block 1302, which receives content from a CP user / administrator via the portal UI (FIG. 12A). Next, block 1304 retrieves the Cin-Mix audio for the selected content from a non-encoding content server. Next, block 1305 retrieves a default (or saved) DCE model, DCE gain, and noise floor (NF) from a DCE parameter server. Next, block 1306 executes the DCE model using the DCE gain and NF to generate or update the Max-Acc-Mix. This block 1305 may also receive an adjusted NF, DCE gain, and DCE model in response to a determination of whether the Max-Acc-Mix is acceptable, as shown in block 1314. Following block 1306, block 1308 provides the Max-Acc-Mix to the CP computer display UI (Portal UI logic) for the CP user / administrator. Next, block 1310 executes X-fade (cross-fade) logic using the X-fade position gains of the Cin-Mix and Max-Acc-Mix to generate an X-fade mix and play the X-fade mix on the CP user's or administrator's speaker. Next, block 1312 determines whether the Max-Acc-Mix is acceptable. If not, then block 1314 receives the adjusted NF, DCE gain, and DCE model, which are delivered to block 1306, which executes the DCE model using the DCE gain and NF to generate or update the Max-Acc-Mix. Next, if yes, and Max-Acc-Min is acceptable, then block 1316 stores or updates the DCE model, DCE gain, and NF for the selected content in the DCE Params server and provides the Max-Acc-Mix to the audio encoder. After block 1316 is completed, the process ends.
[0124] Referring to FIG. 13B, a flow diagram 1320 illustrates one embodiment of a process or logic for executing the portal UI logic 60 (FIG. 1) in accordance with an embodiment of the present disclosure. Process 1320 begins at block 1322, which obtains the latest Cin-Mix audio and Max-Acc-Mix (by channel) and provides them to a waveform display tool (actual data). Next, block 1324 displays a portal UI on the content provider computer display, including a waveform window with the Cin-Mix and Max-Acc-Mix waveforms and parameters selected by the user. These parameters include X-fade and volume control, an audio file list, and DCE parameters, such as the DCE model type, DRC gain profile, Gmne, MNE gain factor, Gdx, Dx gain factor, noise floor, and input or automatic calculation of ALN gain values. Next, block 1326 determines whether an updated value for the X-fade slider has been received from the user. If yes, then block 1328 updates the X-fade logic and X-fade gain values of the user interface (UI) on the user interface (UI). Otherwise, if the result of block 1326 is NO, then block 1330 determines whether a DCE model type has been received. If yes, then block 1332 updates the DCE model to the selected DCE model on the UI. Otherwise, if the result of block 1330 is NO, then block 1334 determines whether updated values of any DCE parameters have been received. If yes, then block 1336 updates the user-selected values of the DCE parameters in the appropriate equations and models on the UI. Otherwise, if the result of block 1334 is NO, then block 1338 determines whether updated values of the volume sliders have been received. If yes, then block 1340 updates the overall volume for the speaker on the US. Otherwise, if the result of block 1338 is NO, then the process 1320 ends.
[0125] Referring to FIG. 13C , a flow diagram 1300 illustrates one embodiment of a process or logic for executing the accessibility (Acc) app logic 16 (FIGS. 2A, 2C, and 2D) according to an embodiment of the present disclosure. Process 1350 begins at block 1352, which receives a content selection from a user / listener, which may be from a remote or user device. Next, block 1353 sends a request to the content API to receive the Cin-Mix and Max-Acc-Mix from the decoder and subtitles (if needed). Next, block 1354 determines whether hub mode should be implemented. If yes, then block 1356 implements hub mode. Otherwise, if the result of block 1354 is no, then block 1357 determines whether single mode should be implemented. If yes, then block 1358 implements single mode. Otherwise, if the result of block 1357 is no, then block 1359 implements remote control mode. Next, block 1360 determines whether a noise floor signal is available. If yes, block 1361 receives a noise floor (NF) audio signal from the microphone 206 (FIGS. 2A-2D) connected to the user device 12 (or smart TV). Otherwise, if the result of block 1360 is no, then block 1362 executes the X-fade logic 100 on the Cin-Mix and Max-Acc-Mix to generate an X-fade mix value based on the measured NF or based on the position of the X-fade slider set by the user. In particular, if the NF is measured or sensed, the X-fade logic 100 can set the slider value so that the integrated loudness IL, or optionally the total dynamic range (or audio loudness range) of the output audio (X-fade mix), is above the noise floor NF or the maximum value of the X-fade slider.In some embodiments, the logic may ensure that the dynamic range (or audio loudness range) of the content does not exceed the maximum loudness allowed by the content provider or any requested maximum loudness standard (Max Audio), for example -2dB (True Peak).
[0126] Next, block 1364 determines if synchronization is selected. If yes, block 1366 performs synchronization with the current audio and video. Otherwise, if the result of block 1364 is no, then block 1368 determines if personalized EQ is enabled. If yes, then block 1370 performs personalized EQ on the X-fade mix based on the personalized EQ parameters in the user attribute server. Otherwise, if the result of block 1370 is no, then block 1372 plays the X-fade / EQ mix on the user's or listener's speaker. Next, block 1374 determines if adaptive voice-over (VO) is enabled. If yes, then block 1376 displays the adaptive voice-over content. Otherwise, if the result of block 1374 is no, then process 1350 ends.
[0127] Referring to FIG. 13D , a flow diagram 1300 illustrates one embodiment of a process or logic for executing the Acc app UI logic 210 (FIGS. 2A, 2C, and 2D) in accordance with an embodiment of the present disclosure. Process 1380 begins at block 1382, which displays the Acc app UI on a user device or smart TV display, displaying user-selected parameters, including the Acc fader, VO fader, live sports faders 1 and 2, language, audio file list, Pers. EQ, level, and X-fade and volume for the Acc function. Next, block 1384 determines whether an updated value for the X-fade slider has been received from the user. If yes, block 1386 updates the X-fade gain value of the X-fade logic on the UI. Otherwise, if the result of block 1384 is no, then block 1388 determines whether an updated Acc app function selection has been received from the user. If yes, block 1389 updates the status of the ACC function on the UI. Otherwise, if the result of block 1389 is NO, then block 1390 determines if an updated personalized EQ, language, or level selection has been received. If yes, then block 1391 updates the personalized EQ, language, or level status on the UI. Otherwise, if the result of block 1390 is NO, then block 1392 determines if an updated volume slider value has been received. If yes, then block 1393 updates the overall volume to the speakers on the UI. Otherwise, if the result of block 1392 is NO, then block 1394 determines if the X-fade mix selection is OK. If yes, then block 1395 saves the X-fader value to the server. Otherwise, if the result of block 1394 is NO, then process 1380 ends.
[0128] Referring to Figures 14A, 14B, 14C, 14D, and 14E, block diagrams of various embodiments for different audio inputs to be streamed and adjustably rendered by a crossfade renderer are shown, in accordance with embodiments of the present disclosure.
[0129] 14A , block diagram 1402 for a video-on-demand (VOD) application illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 2A-2D , where a Cin-Mix from a studio is provided on line 96 to a content provider (CP) system, as described herein. The Cin-Mix track is provided to DCE 62, which may adjust (using DCE gains or parameters, as described herein) and provide a dialogue (Dx) enhanced and M&E compressed Max-Acc-Mix on line 92, as described herein. The Max-Acc-Mix is provided on line 92 to encoder 26, which also receives the Cin-Mix on line 96. As described herein, encoder 26 provides an interleaved encoded track on line 94 having the Cin-Mix and Max-Acc-Mix.
[0130] On the receiving or user / listener (right) side of diagram 1402, the Cin-Mix / Max-Acc-Mix interleaved encoded tracks are received on line 203 by decoder 202, which separates the interleaved tracks into Cin-Mix and Max-Acc-Mix tracks on lines 209 and 207, respectively, and the Cin-Mix and Max-Acc-Mix tracks are provided to X-fade renderer logic 100. As described herein, the output of X-fade renderer logic 100 is an X-fade mix of the Cin-Mix and Max-Acc-Mix tracks, retaining power, where the amount of fade is based on the position (or location) of user-selected X-fade slider 1282 (FIG. 12C), provided on line 225, as discussed herein.
[0131] Also, in some embodiments, DCE 62 may provide a second output mix, e.g., Mid-Acc-Mix (FIG. 1, line 93), to encoder 26. In that case, decoder 202 provides three audio tracks, Cin-Mix, Mid-Acc-Mix, and Max-Acc-Mix, to three-input X-fade logic 100 in Acc appli logic 16 (FIGS. 2A-2D), as previously described herein.
[0132] 14B, a block diagram 1420 for live sports illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 2A-2D, where the input signal is a live (or pre-recorded) video feed (e.g., Sports Mix or Live Mix) from a sporting or other live event provided on line 96 from an external broadcast (OB) truck to a content provider's (CP) system. The Sports Mix is provided to the DCE 62, which may be adjusted (using DCE gains or parameters as described herein) to provide a dialogue-enhanced commentary mix (i.e., Com-Mix) on line 92, with enhanced commentary dialogue over background stadium audio sounds or M&E (e.g., crowd and other general stadium noises or audio), or with only the commentary dialogue audio without stadium audio or M&E. In some embodiments, DCE 62 may be configured to provide a Stadium-Mix (Sports Mix) on line 92, which has enhanced background stadium audio sounds or M&E (e.g., crowd and other general stadium noises or audio) over commentary dialogue, or just stadium audio without commentary dialogue audio.
[0133] The Com-Mix or Stadium-Mix is provided to audio encoder 26 on line 92, which also receives a Sports Mix containing the Com-Mix and M&E (Stadium-Mix) on line 96. Encoder 26 provides an interleaved encoded track (Sports-Mix / Com-Mix or Sports-Mix / Stadium-Mix) on line 94 having the Sports-Mix along with either the Com-Mix or Stadium-Mix.
[0134] In this embodiment, on the receiving or user / listener (right) side of block diagram 1420, the Sports-Mix / Com-Mix or Sports-Mix / Stadium-Mix interleaved and encoded tracks are received on line 203 by decoder 202, which separates the interleaved tracks (depending on what is selected by the content provider when DCE 62 is tuned, as described herein) into a Sports-Mix track on line 209 and a Com-Mix or Stadium-Mix track on line 207, and the Sports-Mix track and the Mix or Stadium-Mix track are provided to X-Fade renderer logic 100. The output of the X-Fade Renderer Logic 100 is an X-Fade mix, as described herein, retaining the powers of the Sports-Mix and Com-Mix, or the Sports-Mix and Stadium-Mix (depending on which set is provided in the received stream), with the fade amount based on the user-selected X-Fade slider 1282 (FIG. 12C), as provided on line 225, as described herein.
[0135] 14B, in some embodiments, there may be a second (lower) DCE 62A that may provide a mix (Com-Mix or Stadium-Mix) not provided by the upper DCE 62. For example, the upper DCE 62 may provide a Com-Mix on line 92, and the lower DCE 62A may provide a Stadium-Mix on line 92A. In that case, encoder 26 receives the Sports-Mix, Com-Mix, and Stadium-Mix and provides an interleaved encoded track (Sports-Mix / Com-Mix / Stadium-Mix) on line 94 having a Sports-Mix track, a Com-Mix track, and a Stadium-Mix track, similar to that shown in FIG. 5B, which illustrates an example of three-track interleaving by audio encoder 26.
[0136] On the receiving or user / listener (right) side of diagram 1420, the interleaved encoded Sports-Mix / Com-Mix / Stadium-Mix tracks are received on line 203 by decoder 202, which separates the interleaved tracks into a Com-Mix on line 207, a Sports-Mix track on line 209, and a Stadium-Mix on line 250A, which are provided to three-input X-fade renderer logic 100 (FIG. 10C). The output of the three-input X-fade renderer logic 100 is a power-preserving X-fade mix of the Com-Mix, Sports-Mix, and Stadium-Mix tracks described in FIG. 10C, which X-fade mix is provided on line 225 based on the position of X-fade slider 1282 (FIG. 12C) selected by the user as described herein. In that case, one end of the X-fade slider 1282 provides only the Stadium-Mix (for those who do not want to hear the commentary dialogue audio), the middle position of the X-fade slider 1282 provides the Sports-Mix (which has both the stadium and commentary as originally delivered from the broadcast track), and the other end of the X-fade slider 1282 provides only the Com-Mix (for those who do not want to hear the stadium background noise / non-commentary audio).
[0137] The three-input X-fade renderer may be used with any audio input for which a content provider desires to provide a controlled, selectable transition between two audio components in a given audio mix (e.g., Com-Mix / Stadium-Mix, Cin-Mix-Dx / Cin-Mix-AD, etc.). In this case, one end of the X-fade slider provides only the first audio component, the other end of the X-fade slider provides only the second audio component, and the middle of the X-fade slider provides portions of both audio components. Other cross-fade renderer configurations and formulas for the three-input and two-input X-fade renderers described herein may be used as desired, provided they provide the same functionality and performance as those described herein. It is also generally desirable to use the same M&E track for each input to the X-fade renderer, if possible, to minimize the risk of holes in the M&E of a given mix or track caused by studio or other processing. In particular, when an Alternate Language (AL) translation is generated from an original English video, differences in duration and word count from English to AL can leave blank spots (or holes or gaps) in the audio. This also applies to original non-English content.
[0138] Referring to FIG. 14C, a block diagram 1440 for a live sports application illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 2A-2D, where a first input signal on line 96A comprises a live (or pre-recorded) Sports Mix, as described above with reference to FIG. 14B, with the commentator dialogue audio portion being the commentator dialogue audio of the home sports team (Sports-Home-Mix), and a second input signal on line 96B comprises a live (or pre-recorded) Sports Mix, as described above with reference to FIG. 14B, with the commentator dialogue audio portion being the commentator dialogue audio of the away sports team (Sports-Away-Mix). The Sports-Home-Mix and Sports-Away-Mix inputs on lines 96A and 96B, respectively, are provided to the audio encoder 26, which provides an interleaved encoded track on line 94 comprising the Sports-Home-Mix and the Sports-Away-Mix (Sports-Home-Mix / Sports-Away-Mix).
[0139] On the receiving or user / listener (right) side of block diagram 1440, the Sports-Home-Mix / Sports-Away-Mix interleaved and encoded track is received on line 203 by decoder 202, which separates the interleaved track into the Sports-Home-Mix and Sports-Away-Mix, which are provided to X-Fade renderer logic 100 on lines 207 and 209, respectively.
[0140] The output of the X-fade renderer logic 100 is a power-preserving X-fade mix of the Sports-Home-Mix and Sports-Away-Mix tracks, provided on line 225, as described herein, with the amount of fade based on the position of the X-fade slider 1284 (FIG. 12C) selected by the user, also provided on line 225. In this case, one end of the X-fade slider 1284 provides only the Sports-Home-Mix (for those who want to hear only the home team commentator's dialogue audio), an intermediate position of the X-fade slider 1282 provides a blend of the Sports-Home-Mix and the Sports-Away-Mix (with both home and away commentator dialogue audio), and the other end of the X-fade slider 1284 provides only the Sports-Away-Mix (for those who want to hear only the away or visiting team commentator's dialogue audio).
[0141] In some embodiments, DCE 62 may be used with the Sports-Home-Mix and Sports-Away-Mix input signals on lines 96A, 96B to enhance audio dialogue relative to background noise or M&E, as indicated by dashed box 62. In that case, the DCE parameters of DCE 62 may be set based on the desired dialogue audio, e.g., dialogue enhanced relative to M&E, or dialogue only without M&E, as described herein.
[0142] 14D, block diagram 1460 for a voice-over application illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 2A-2D, where a first input signal on line 96A comprises an original version (OV) cinematic mix in English from the studio (Cin-Mix English or OV), as described herein, and a second input signal on line 96B comprises an alternative language (AL) cinematic mix in English from the studio (Cin-Mix Spanish or AL). The Cin-Mix English (OV) and Cin-Mix Spanish (AL) inputs on lines 96A and 96B, respectively, are provided to audio encoder 26, which provides an interleaved and encoded track on line 94 comprising Cin-Mix English and Cin-Mix Spanish (Cin-Mix English / Cin-Mix Spanish).
[0143] On the receiving or user / listener (right) side of block diagram 1460, the Cin-Mix English / Cin-Mix Spanish interleaved and encoded track is received on line 203 by decoder 202, which separates the interleaved track into Cin-Mix English and Cin-Mix Spanish on lines 207, 209, respectively, which are provided to X-fade render logic 100.
[0144] The output of the X-fade renderer logic 100 is an X-fade mix, provided on line 225, retaining the power of the Cin-Mix English track and the Cin-Mix Spanish track, where the amount of fading is based on the position of the X-fade slider 1280 (FIG. 12C) selected by the user, where one end of the X-fade slider 1280 provides only Cin-Mix English (OV) (for those who want to hear only the dialogue audio in the original version language), a middle position of the X-fade slider 1280 provides a blend of Cin-Mix English and Cin-Mix Spanish (having both dialogue audio in the original version OV language and dialogue audio in the alternative language AL), and the other end of the X-fade slider 1280 provides only Cin-Mix Spanish (AL) (for those who want to hear only the dialogue audio in AL).
[0145] In some embodiments, DCE 62 may be used on the Cin-Mix English and Cin-Mix Spanish input signals on lines 96A and 96B to enhance the audio dialogue for M&E, as indicated by dashed box 62. In that case, the DCE parameters of DCE 62 may be set based on the desired dialogue audio, e.g., dialogue enhanced for M&E or dialogue only without M&E, as described herein.
[0146] 14E, a block diagram 1480 for audio description (AD) or narration illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 2A-2D, in which a first input signal on line 96A comprises an original version (OV) cinematic mix in English from the studio (Cin-Mix English, OV, or Cin-Mix), as previously described herein, and a second input signal on line 96B comprises an audio description (AD or Cin-Mix-AD) from the studio. The Cin-Mix (OV) and Cin-Mix-AD (AD) inputs on lines 96A and 96B, respectively, are provided to audio encoder 26, which provides interleaved encoded tracks on line 94 comprising Cin-Mix and Cin-Mix-AD (Cin-Mix / Cin-Mix-AD).
[0147] On the receiving or user / listener (right) side of block diagram 1480, the Cin-Mix / Cin-Mix-AD interleaved encoded track is received on line 203 by decoder 202, which separates the interleaved track into Cin-Mix(OV) and Cin-Mix-AD on lines 207, 209, respectively, and the Cin-Mix(OV) and Cin-Mix-AD are provided to the X-fade renderer logic 100 described herein.
[0148] The output of the X-fade render logic 100 is a power-preserving X-fade mix of the Cin-Mix (OV) track and the Cin-Mix-AD track provided on line 225, as described herein, with the amount of fade based on the position of the X-fade slider selected by the user, which is similar to slider 1280 (FIG. 12C), except that the top end of the slider can say Aud. Desc. only (i.e., AD only). In that case, one end of the X-fade slider 1280 provides only Cin-Mix (OV) (for those who want to hear only the original version language dialogue audio), a middle position of the X-fade slider 1280 provides a blend of Cin-Mix (OV) and Cin-Mix-AD (having both original version OV dialogue audio and audio description or narration (AD) dialogue audio), and the other end of the X-fade slider 1280 provides only Cin-Mix-AD (for those who want to hear only audio description or narration dialogue (AD) audio).
[0149] In some embodiments, DCE 62 is used on the Cin-Mix (OV) and Cin-Mix-AD input signals on lines 96A, 96B to enhance the audio dialogue for M&E, as indicated by dashed box 62. In that case, the DCE parameters of DCE 62 can be set based on the desired dialogue audio, e.g., dialogue enhanced for M&E or dialogue only without M&E, as described herein.
[0150] Referring to Figures 15A and 15B, block diagrams of various embodiments for different types of audio inputs streamed and adjustably rendered by a crossfade renderer for video-on-demand (VOD) and live sports / events (Live) applications, respectively, are shown in accordance with embodiments of the present disclosure.
[0151] 15A, block diagram 1502 for a video-on-demand (VOD) application illustrates an embodiment of the present disclosure similar to that described with reference to FIGS. 14A, 14D, and 14E, in which a cinematic mix (Cin-Mix) from a studio is provided to a content provider's (CP) system 10 (FIG. 1) on line 96, as described herein. The Cin-Mix track on line 96 is provided to a DCE 62, which may be adjusted (using DCE gains or parameters, as described herein) to provide a dialogue (Dx) enhanced mix, e.g., a Max-Acc-Mix, sometimes generally referred to herein as an accessible mix, on line 92. The Max-Acc-Mix is provided on line 92 to an encoder 26, which also receives the Cin-Mix on line 96. The encoder 26 provides an interleaved encoded track on line 94 having a Cin-Mix and a Max-Acc-Mix as described herein. The accessible mix output (i.e., Mac-Acc-Mix or Acc-Mix) of the DCE 62 can be an enhanced dialogue mix (EDx-Mix) = enhanced dialogue (Dx or EDx) + M&E, an audio description mix (AD-Mix or narration mix) = audio description dialogue (Dx(AD), EADDx, or ENDx) + M&E, or an alternative language mix (AL-Mix, VO-Mix, or Voice-Over-Mix) = alternative language dialogue (Dx(AL), ALDx, or VODx). A cinematic mix as described herein is a cinematic mix (Cin-Mix) = dialogue (original version or OV) + M&E.
[0152] More specifically, for an audio description mix (AD-Mix) output from the DCE 62 on line 92, the DCE 62 may be configured or adjusted to provide only audio description dialogue (Dx(AD)) and M&E, but not primary or talent dialogue. Also, for an alternate language mix output from the DCE 62 on line 92, the DCE 62 may be configured or adjusted to provide only alternate language dialogue (Dx(AL) or Dx(ALT)) and M&E, but not the primary language dialogue of the original version (OV).
[0153] In some embodiments, the CP user / administrator 21 may be able to adjust the DCE Params (as described herein) to optimize the Max-Acc-Mix, as shown by dashed line 1506. In that case, the Max-Acc-Mix and Ci-Mix are provided to the X-fade logic 100, the user / administrator 21 controls the sliders (as described herein), as shown by dashed line 1508, and the X-fade logic 100 provides the X-fade mix to the user / administrator at dashed line 1510 to determine if the DCE Params are providing the desired audio experience.
[0154] The accessible mix, or Acc-Mix, is provided on line 92 to the audio encoder 26, which also receives the Cin-Mix on line 96. The encoder 26 provides an interleaved encoded track on line 94 comprising the Acc-Mix and the Cin-Mix (Acc-Mix / Cin-Mix).
[0155] On the receiving or user / listener (right) side of block diagram 1502, the Acc-Mix / Cin-Mix interleaved encoded track is received on line 203 by decoder 202, which separates the interleaved track into an Acc-Mix (EDx+M&E, Dx(AD)+M&E, or Dx(ALT)+M&E) on line 207, e.g., enhanced dialogue, audio dialogue, or alternative language, depending on the application, and a Cin-Mix (Dx(OV)+M&E) on line 209, which is provided to the X-fade renderer logic 100 described herein.
[0156] The output of the X-Fade Renderer Logic 100 is an X-Fade mix holding the power of the Acc-Mix and Cin-Mix tracks provided on line 225 as described herein, with the amount of fade based on the position of the X-Fade slider 1504 selected by the user, similar to the slider described herein with reference to FIG. 12C. In that case, one end of the X-fade slider provides only the Cin-Mix (OV) (for those who only want to hear the original version of the dialogue audio), the other end of the X-fade provides only the Acc-Mix (EDx+M&E, Dx(AD)+M&E, or Dx(ALT)+M&E) depending on the application (for those who only want to hear the Accessible Mix (Acc-Mix) for a given application), and an intermediate position of the X-fade slider provides a blend of the Cin-Mix (OV) and the Acc-Mix (EDx+M&E, Dx(AD)+M&E, or Dx(ALT)+M&E) depending on the application.
[0157] Referring to FIG. 15B, a block diagram 1520 for a live sports / events (live) application shows an embodiment of the present disclosure similar to that described with reference to FIGS. 14B and 14C, in which a sports home mix from an OB truck is provided on line 96 to a content provider's (CP) system 10 (FIG. 1) as described herein.
[0158] The Sports Home Mix track is provided to DCE 62 on line 96, which, as previously described herein, may be adjusted (using DCE gains or parameters as described herein) to provide a Home Commentary Mix (Com-Home-Mix) on line 92 or a Stadium Mix (Stadium-Mix) on line 92A.
[0159] The Comm-Home-Mix or Stadium-Mix (depending on which is selected by the content provider when tuning DCE 62, as described herein) is provided to audio encoder 26 on lines 92, 92A, respectively, which also receives the Sports Mix on line 96. Encoder 26 provides an interleaved encoded track on line 94 having the Comm-Home-Mix or Stadium-Mix and the Sports-Home-Mix (Sports-Home-Mix / Com-Home-Mix or Sports-Home-Mix / Stadium-Mix).
[0160] Alternatively, the away sports mix track may be provided to DCE 62 on line 96A, which may be adjusted (using DCE gains or parameters as described herein) to provide an away commentary mix (Com-Away-Mix) on line 92B or a stadium mix (Stadium-Mix) on line 92C, as previously described herein.
[0161] In that case, the Comm-Away-Mix or Stadium-Mix (depending on which was selected by the content provider when tuning DCE 62, as described herein) is provided to audio encoder 26 on lines 92B, 92C, respectively, which also receives the Sports-Mix on line 96. Encoder 26 provides an interleaved encoded track on line 94 having the Comm-Away-Mix or Stadium-Mix and the Sports-Away-Mix (Sports-Away-Mix / Com-Away-Mix or Sports-Away-Mix / Stadium-Mix).
[0162] Alternatively, the sports home mix input and sports away mix input on lines 96, 96A, respectively, are provided to audio encoder 26, which provides an interleaved encoded track on line 94 having the sports home mix and sports away mix (sports home mix / sports away mix).
[0163] In some embodiments, the CP user / administrator 21 can adjust the DCE Params (as described herein) to provide the desired live sports / event mix described above, as indicated by dashed line 1526. In that case, the mix on lines 92, 92A, 92B, 92C, 96, 96A may be provided to X-fade logic 100, the user / administrator 21 controls the X-fade slider 1529 (as described herein), as indicated by dashed line 1528, and the X-fade logic 100 provides the X-fade mix at dashed line 1527 to the user / administrator 21 to determine if the DCE parameters are providing the desired audio experience.
[0164] On the receiving or user / listener (right) side of block diagram 1520, at the top 1530, the Sports-Home-Mix / Com-Home-Mix or Sports-Home-Mix / Stadium-Mix interleaved and encoded tracks are received by decoder 202 on line 203, which then encodes the interleaved tracks into Com-Home-Mix or Stadium-Mix (depending on which is selected) on line 207 and Sports-Home-Mix on line 209. The Com-Home-Mix or Stadium-Mix, which separates the Sports-Home-Mix and the Com-Home-Mix, are provided to the X-fade renderer logic 100 described herein, which has an X-fade slider 1522 similar to that described in Figure 14B that fades between the Sports-Home-Mix and the Com-Home-Mix or Stadium-Mix (as selected) based on the position of the X-fade slider 1522. A similar X-fade may be performed for the Sports-Away-Mix / Com-Away-Mix or the Sports-Away-Mix / Stadium-Mix.
[0165] On the receiving or user / listener (right) side of block diagram 1520, at the bottom 1532, the sports home mix / sports away mix interleaved and encoded tracks are received on line 203 by decoder 202, which separates the interleaved tracks into the sports home mix on line 207 and the sports away mix on line 209. The sports home mix and sports away mix are provided to X-fade renderer logic 100, described herein, which has an X-fade slider 1524 that fades between the sports home mix and the sports away mix based on the position of X-fade slider 1524, similar to that described in FIG. 14C.
[0166] As described herein, the present disclosure uses a crossfade renderer or X-fade 100 (FIGS. 1, 2A-2D, 14A-14E, and 15A-15B) to blend two (or more) input audio signals, sometimes commonly referred to as Audio Mix A and Audio Mix B. The X-fade renderer has an X-fade slider that can be moved (automatically or manually) from an initial position (initial or first extreme) of the X-fade slider, which provides an audio output that is exclusively "Mix A," to a final position (other or second extreme) of the crossfade slider, which is exclusively audio "Mix B," with the X-fade slider providing an unlimited number of positions between the initial (first extreme) and final (second extreme) positions, each corresponding to a blend that preserves the power of Audio Mix A and Audio Mix B.
[0167] As described herein, audio mix A and audio mix B may be any two audio signals desired to be blended or faded from a first signal (Mix A) to a second signal (Mix B), for example, Mix A and Mix B may be as follows: Mix A = English original version (i.e., Cin-Mix) and Mix B = Enhanced Dialogue English (Max-Acc-Mix), Mix A = English original version and Mix B = English audio description, Mix A = English original version and Mix B = Spanish audio voice-over, Mix A = English original version and Mix B = Spanish audio description, Mix A = Home team commentary mix and Mix B = Away team commentary mix, Mix A = Home team commentary mix and Mix B = Stadium sounds (no commentary) mix. Other audio mixes may be used as desired when input is provided to the X-fade slider. However, it may be desirable for the M&E portions of the audio in both MixA and MixB to be similar or substantially the same to optimize the listening experience.
[0168] The label Max-Acc-Mix may be used herein to refer to the maximum accessible mix of the dialogue-enhanced output of a DCE when the input is a Cin-Mix (studio VOD content application). However, it may also be used to represent the endpoint of the X-Fade slider in other embodiments or applications described herein. Also, the term Max Personalized Mix (or Max-Pers-Mix) may also be used. The term Maximum Personalized Mix (or Max-Pers-Mix) may also be used. However, other labels representing the maximum or upper limit endpoints of the X-Fade slider may be used in other applications, such as Max-Com-Mix, representing a maximum commentary mix for live sports content; Alt-Com-Mix, representing an alternative commentary mix for live sports content; Max-AD-Mix, representing a maximum audio description mix for audio descriptions of live / VOD content; and Max-VO-Mix, representing a maximum voice-over mix for live or VOD content, as shown in FIGS. 14B-14E and 15B. Other labels may be used as needed depending on the application. In some embodiments, the terms MixA and MixB may be used collectively to represent two audio mixes or tracks that define the lower limit (or first end) and the opposite upper limit (or second end) of the X-Fade slider of a two-input X-Fade renderer described herein. Also, in the MixA and MixB voice-over examples above, English can be replaced with any main (or primary) language desired by the user, and Spanish can be replaced with any alternate language (a language different from the main language) desired by the user, as long as the mix is provided by the content owner or content provider.
[0169] Also, although X-fade has been described as using sine and cosine functions to blend two (or three) input signals to produce a power-conserving output mix, i.e., an X-fade mix, those skilled in the art should understand that other cross-fade techniques, such as those described below, may be used if desired.
[0170] In some embodiments, the X-fade logic 100 can provide gradual switching and smoothing. This technique uses gradual switching, with a smooth transition from Mix A to Mix B in distinct steps rather than a continuous fade. It applies a short, smooth ramp (e.g., 200 ms to 500 ms) to avoid abrupt jumps in volume and make the transition feel more natural. More specifically, with this technique, when the transition is triggered, Mix A fades out quickly and Mix B fades in over a short period (e.g., 200 ms to 500 ms). The fade curve can be linear, logarithmic, or sigmoidal to optimize perceived smoothness. Transitions can also be applied at natural switching points, such as dialogue sentence boundaries, scene cuts, pauses in background music, etc. In some embodiments, a delay or hold function can be used to prevent abrupt and repetitive switching, ensuring a smoother experience. Compared to traditional crossfades, such as those shown in Figures 10A-10C herein, this technique has certain advantages and disadvantages. On the one hand, this technique shortens the duration of "blended audio," minimizing moments when both mixes unnaturally overlap. This technique is also effective for separate audio streams (e.g., original audio vs. dubbed audio, commentary audio vs. stadium audio) and can prevent speech intelligibility issues that occur when two voices are heard simultaneously during a slow fade. On the other hand, this technique is less smooth than a gradual crossfade when the two mixes contain similar elements (e.g., different language versions of the same dialogue), and the timing of the transition must be carefully selected to avoid a jarring transition.
[0171] In some embodiments, the X-fade logic 100 may include smart gain-based blending (not shown). This technique applies frequency-dependent gain interpolation rather than a uniform fade across the entire spectrum. It prioritizes certain frequency bands based on the audio content, making the transition perceptually more natural. More specifically, rather than fading all frequencies equally, this technique applies different gain curves to different frequency bands. Specifically, for dialogue-heavy content, frequencies between 1 kHz and 4 kHz (the speech intelligibility range) fade more smoothly, while low-frequency elements (background music, effects) fade more quickly or remain transient to avoid gaps. For music or ambient sounds, low and mid-frequencies transition first, while high frequencies gradually fade in to maintain clarity. Compared to traditional crossfades, such as those illustrated herein using Figures 10A-10C, this method has certain advantages and disadvantages. Advantageously, this technique prevents the "muddiness" of intermediate transitions caused by overlapping full-spectrum audio. This technique also improves speech intelligibility by prioritizing dialogue intelligibility over background noise. Furthermore, this technique can produce more natural transitions between different types of content. On the downside, this technique requires real-time frequency analysis, which makes it computationally expensive. Also, this technique can introduce unintended artifacts if not carefully tuned.
[0172] In some embodiments, the X-fade logic 100 may include a content-aware adaptive crossfade (not shown). This approach uses real-time audio analysis to determine optimal fade shape and duration based on the similarity between Mix A and Mix B. More specifically, with this approach, the system analyzes audio characteristics (e.g., spectral content, loudness, and speech presence). If Mix A and Mix B contain similar elements (e.g., both are dialogue-heavy), a longer, more gradual fade is used. If Mix B introduces dramatically different elements (e.g., moving from spoken commentary to crowd noise), a faster fade is applied to avoid confusion. In some embodiments, machine learning models may be used to classify audio types and dynamically adjust fade parameters based on historical data. Compared to traditional crossfades, such as those shown in Figures 10A-10C herein, this approach has certain advantages and disadvantages. On the advantage, this approach reduces perceptual distraction by adapting the transition to the content. Also, this approach avoids unwieldy overlaps between similar content types. Additionally, this technique may be optimized for accessibility use cases (e.g., ensuring speech clarity during transitions). On the downside, this technique requires real-time analysis and adaptive decision-making, making it computationally intensive, and is more difficult to predict or manually control compared to a fixed crossfade.
[0173] In some embodiments, the X-fade logic 100 may include multi-channel matrix mixing with dynamic weighting (not shown). Instead of crossfading between two stereo streams, this approach keeps all potential audio mixes available in multi-channel format (e.g., 4.0, 5.1) and dynamically adjusts their levels based on user preference. More specifically, with this approach, all audio options (e.g., original mix, enhanced dialogue mix, Spanish mix) are kept in multi-channel format. The matrix mixer also dynamically adjusts the gain of each mix, blending them smoothly without actual fading in and out. Furthermore, the mixes may be controlled using a fader UI, automation, or an AI-driven recommendation system. Compared to traditional crossfades, such as those shown in Figures 10A-10C herein, this approach has certain advantages and disadvantages. On the plus side, this approach allows for instant, seamless transitions without noticeable fade artifacts. This technique also reduces the need for destructive crossfades, preserves audio quality, and is more flexible for users who want to fine-tune their experience (e.g., 75% original mix + 25% enhanced dialogue).On the downside, this technique requires a playback system capable of handling multi-channel audio, which can be complex to implement in traditional stereo playback environments.
[0174] In some embodiments, the X-fade logic 100 may include spectral cross-adaptive mixing with dynamic weighting (not shown). This is a technique that analyzes the spectral content of Mix A and Mix B and cross-adapts them to create a smooth transition with phase coherence. More specifically, in this technique, frequency analysis of both mixes is performed in real time. Instead of applying a simple volume-based fade, overlapping spectral components are intelligently blended. It also uses phase-coherent morphing techniques to prevent phase cancellation artifacts. Furthermore, this technique may use machine learning to predict optimal spectral adjustments for seamless blending. Compared to traditional crossfades, such as those illustrated in FIGS. 10A-10C herein, this technique has certain advantages and disadvantages. On the advantages, this technique prevents phase issues that can occur when blending similar content, creates seamless and transparent transitions without noticeable overlaps, and works well in complex audio environments where different layers need to be preserved. On the disadvantages, this technique is computationally expensive and requires high-performance processing. It is also more difficult to implement in real-time streaming applications.
[0175] Other techniques or methods may be used in the X-fade logic 100, if desired, provided that they provide the functionality and performance described herein.
[0176] This disclosure describes an accessible audio streaming system with an automatically personalized (per language) cross-fade renderer on a user device. In some embodiments, as described herein, if measurement (i.e., sensing or recording) of the environmental noise floor (NF) is not available (e.g., no microphone is present), an accessibility X-fade renderer (i.e., an accessibility fader, cross-fader, or cross-fade slider) may be made available to the user via a UI user device. In some embodiments, the X-fade renderer may be available to the user independently of sensing the noise floor NF. In that case, the system may provide an initial X-fade setting based on the measured NF (or NF capture), and the user may further adjust the setting, if necessary, for an optimal listening experience for the user.
[0177] Additionally, the present disclosure provides for personalized audio streaming. In some embodiments, there may be a crossfade between the Cin-Mix audio track and the Max-Acc-Mix audio track, which evens out non-dialogue (or M&E), and in some embodiments, as described herein, high fidelity voice-over (VO, e.g., two languages rendered simultaneously) or narration (AD, e.g., audio-described dialogue added to the main dialogue) may be provided automatically using NF capture or manually using an X-fade slider by the user / listener.
[0178] In some embodiments, separation (DI) of dialogue from the rest of the audio mix is performed to preserve perceived loudness in the mix. Also, as described herein, various known AI machine learning (ML) models (open source and commercial) (e.g., Spleeter® (Deezer Research), demucs (Meta), AudioShake® (AudioShake), etc.) may be used to extract dialogue from video streams. In this case, the known ML models are trained using existing datasets for extracting dialogue from videos across a wide range of content.
[0179] In the ALN logic, a target loudness normalization gain (ALN gain) is applied to the accessible mix to ensure that all elements of the accessible mix are above the noise floor, e.g., -16 LKFS, which is the industry standard for mobile noise environments. The ALN gain may also be a function of content type, e.g., soft / quiet scenes, action / loud scenes, etc. The same process is also applied to Dolby ATMOS (which does not need to be changed due to its high channel count).
[0180] In some embodiments, the input audio signal may be 5.1 or 5.1.4 immersive sound, including Dolby Atmos. In that case, the dialog clarity engine (DCE), which includes dialogue enhancement (DE) and dynamic range control (DRC), can also handle these formats. In some embodiments, the dialog clarity engine (DCE) includes a dialog-aware remixing engine (or enhanced dialogue insertion (EDI)) for better dialogue clarity surround audio 5.1 delivery, along with a combination of dialogue enhancement (DE), dynamic range control or compression (DRC), or audio loudness normalization (ALN).
[0181] In some embodiments, for example, when the input audio is 5.1 surround sound or immersive 5.1.4 and all speaker layouts, storytelling dialogue (or Dx) (from a script read by actors or talent) may already be separated from the video on its own channel (e.g., the center channel), in which case the dialogue separation (DI) logic is optional and DCE gain (or DCE Params or correction gain) may be applied directly to the center channel (C) of the audio input signal. DCE gain is also applied to all remaining channels (other than the C channel) during a final normalization step (ALN) (surround normalization) to preserve overall loudness. Furthermore, DCE gain (or DCE Params) may be a function of the type of content, such as soft / quiet scenes, action / loud scenes, etc. The same process applies to Dolby ATMOS (which does not need to be changed due to the large number of channels). The same process described above for surround sound 5.1 etc. also applies to immersive sound such as Dolby ATMOS (i.e. a sound setup with more channels which does not need to be changed) and separate individual dialogue objects.
[0182] The technology of this disclosure is device-agnostic. Furthermore, in embodiments requiring it, dual or multiple audio decoding and audio capture are all natively supported by the Apple®, Amazon®, and Roku® operating systems (OS), greatly facilitating integration and testing of this disclosure. They may also be provided in streaming service app software or APIs, which may include or be part of the app logic 16 described herein. This disclosure also supports the HLS / DASH streaming protocol, which supports multimedia streaming, supports associated bitrates for two stereo streams equivalent to the surround audio bitrate, and has been proven for over a decade from a CDN delivery perspective, showing that CDN delivery does not adversely affect streaming quality of experience (QoE). Surround and immersive payload delivery requires 5.1+1 (Surround + Enhanced Center Channel (Max-Acc-Mix)) or 5.1.4+1 (Immersive / Atoms + Enhanced Center Channel (Max-Acc-Mix)). An alternative personalized Dolby audio delivery can be applied to an automatic stereo downmix from a stereo creative mix or an immersive / surround creative (near-field) mix, for example Dolby dual stereo encoding (2.0 Cin-Mix + 2.0 Max-Acc-Mix (or Max-Pers-Mix)).
[0183] As described herein, in some embodiments, the system of the present disclosure uses measured (or captured) noise floor (NF or accessibility fader) levels to adjust X-fade rendering in real time in the environment (e.g., airplane, car, bedroom, etc.) using both the original cinematic track and the accessible audio track (or Max-Acc-Mix), also referred to as the maximum personalized audio track (Max-Pers-Mix).
[0184] In some embodiments, as described herein, the present disclosure provides that the ambient noise floor (NF) may be estimated while filtering the frequency range of the human voice, crossfading (left / right) preserves power so that loudness and stereo / spatial image are unaffected, and DRC is applied to pre-encoding to affect the loudest and quietest (non-dialogue) parts of the M&E.
[0185] Also, as discussed herein, PL refers to program loudness as specified in ITU-R BS.1770-1. LRA, or loudness range, is a measure of the range between the loudest and quietest parts of an audio track across all channels. DPL refers to dialogue to program loudness, and LKFS refers to loudness K-weighted full scale. M&E refers to music and effects (i.e., sound effects, special effects, etc., everything other than dialogue / storytelling). In some embodiments, as discussed herein, when DRC is applied, a target loudness normalization gain is applied to the accessible mix to ensure all elements of the accessible mix are above the noise floor, e.g., -16 LKFS, which is the industry standard for mobile noise environments.
[0186] In some embodiments, an accessible stereo mix (or Mac-Acc-Mix) may be produced directly by a video recording studio and / or generated by a content provider or streaming service running a media processing DCE and scaled across the entire video-on-demand (VOD) catalog securely on the content owner's premises. The accessibility rendering (cinematic crossfade to DE+DRC or Max-Acc-Mix) may be scaled to all streaming environments and devices.
[0187] As described herein, the X-fade renderer may be driven by a measured (or captured) noise floor (NF) or may be manually controlled by a user / listener within a streaming service app with an accessibility fader, X-fade renderer, or X-fade logic as described herein. For example, in a noisy environment, the X-fade slider may be set so that the effect of DRC in the DCE is high and the effect of DE is present or on. In another example, when a user is listening at home through a streaming device, the user may select a position for the X-fade slider so that the DRC is low and the DE is on. Furthermore, when a user is listening at home in an optimal environment, the user may set the X-fade slider position so that the effect of DRC is low (may be off or not as effective), depending on their preference, and the user can select a slider position where the effect of DE is only as much as needed to understand the dialogue (i.e., accessible audio).
[0188] This disclosure may also be used with accessible DOLBY® stereo / downmixes. In that case, a crossfade (i.e., X-fade) may be applied to each L / R channel of the 2.0 downmix. DE is applied to active dialogue in the 2.0 downmix. The same DRC is applied to all channels of the mix. A single dual stereo EC3 encoder / decoder, e.g., 2x 2.0, independently encoded mixes, may be used.
[0189] This disclosure provides embodiments that include Dolby® audio distribution accessible to all Dolby-enabled devices. Smart TV or device manufacturers (OEMs) that downmix Atmos 5.1.4 (with five speakers at ear height (left front, middle, right front, left ambient, and right ambient) plus a subwoofer and four speakers for overhead audio (with two front and two rear height channels)) from 3.0 (with three channels: left, right, and middle) or 2.0 (with two channels: left and right) will work with this disclosure but may reduce dialogue intelligibility. OEMs that virtualize Atmos 5.1.4 with a limited number of speakers will also work with this disclosure but may reduce intelligibility (e.g., Dolby certification measures intelligibility). Accessible Dolby-encoded streaming streams that may be used with this disclosure for in-home "consumer-grade" Dolby systems include the following options: For example, an Atmos / 5.1 <> 2.0 downmix that is Cinematic <> Accessible, for example, a stereo 2.0 native mix that is Cinematic <> Accessible, for example, a surround 5.1 native mix that is center channel Cinematic <> Accessible, for example, an Atmos 5.1.4 native mix that is center channel Cinematic <> Accessible.
[0190] This disclosure also provides embodiments including a single multi-channel audio encoder / decoder. MPEG AAC-LC and Dolby E-AC3 support a "dual stereo" encoding mode, for example, where 2x 2.0 / stereo are encoded independently. MPEG AAC-LC and Dolby E-AC3 support a "quad mono" encoding mode, for example, where 4x 1.0 / mono are encoded independently. The Dolby (E-AC3 + Joint-Object Coding) = Atmos encoder supports primary and associated audio in a single bitstream. In some embodiments, this disclosure may use known source FFmpeg® or Gstreamer® software tools for audio signal processing.
[0191] As described herein, the present disclosure includes several embodiments that provide personalized accessibility that adapts to hearing ability, where accessible audio (or Max-Acc-Mix) is decoded and further equalized to compensate for hearing loss frequency response on a per-user basis. For example, hearing loss functionality (e.g., audiograms) may be provided by a third-party audiogram API such as Knisper®, the device's operating system (e.g., iOS®), or a listening test mode in a streaming service app.
[0192] In some embodiments, as described herein, the present disclosure includes some embodiments including a system for providing personalized voice-over that adapts to a user's language skill (beginner, intermediate, expert), allowing the user to select the loudness level of the voice-over dialogue. For example, an inexperienced English speaker may choose to hear both languages in the background with their native language in an assisted, e.g., voice-over, mode. Similarly, an intermediate English speaker may choose to hear their primary language (or native language) in the background only when a difficult sentence is detected. Content based on the user's language ability may be provided by a content provider, with a specific mix for each desired voice-over level, providing an adaptive experience for the user / listener.
[0193] As described herein, in some embodiments, the present disclosure includes a dialog clarity engine (DCE) 62 for use in Stereo 2.0 audio distribution (FIG. 3A). In this embodiment, the dialogue separation (DI) logic uses a two-way separation (ML) model to extract active storytelling dialogue from the remainder of the mix (e.g., M&E) in both the left and right channels of the mix. The dialogue enhancement (DE) logic applies amplification to the separated storytelling dialogue L / R channels and attenuation to the remaining L / R channels so that the loudness of the original program is maintained. In that case, dynamic range compression (DRC) logic may be applied only to the remaining channels. The stereo remix (or enhanced dialogue insertion—EDI) logic combines the left dialogue (i.e., vocal) channel and the M&E (i.e., reverberation) channel to generate a dialogue-enhanced L / R channel. Finally, Accessible Target Loudness (ATL) logic may, in some embodiments, drive the DRC of the attenuated reverberation and the final Audio Loudness Normalization (ALN) step.
[0194] As described herein, the present disclosure also includes some embodiments that provide a dialogue clarity engine for receiving audio in a surround 5.1 audio distribution. In this case, dialogue separation (DI) may use a two-way separation (ML) model to extract active storytelling dialogue in the center channel of the mix from the rest of the mix (e.g., M&E). Dialogue in other audio channels may be considered sound effects, not storytelling. Dialogue enhancement (DE) logic applies amplification to the separated storytelling dialogue channel (vocals) and attenuation to the reverberant channel, such that the loudness of the original program is maintained. Dynamic range compression (DRC) logic may be applied only to the attenuated reverberant channel. A mono remix (or enhanced dialogue insertion - EDI) logic step combines the vocal and reverberant channels into a dialogue-enhanced audio channel.
[0195] This disclosure also works with input audio having immersive 5.1.4 audio distribution, DOLBY Atmos 5.1.4, and DOLBY Surround 5.1. In that case, a crossfade (X-fade) is applied to the center channel of the mix where active storytelling dialogue is present (dialog in other channels is assumed to be dedicated to sound M&E or FX). In that case, DE logic is applied to the active dialogue in the cinematic center channel, and in some embodiments, an additional single mono channel EC3 encoder / decoder may be used. The associated audio encoding bs mode of the EC3 encoder may be used to provide an "accessible center channel."
[0196] As described herein, this disclosure provides an embodiment of a dialogue clarity engine (DCE) 62 that may have an input audio signal 2.0 stereo audio and may provide dynamic range control for "dialogue-enhanced" cinematic stereo (as shown in FIG. 3E, DCE Model #3). In such an embodiment, dialogue separation (DI) logic extracts active storytelling dialogue from the rest of the mix (e.g., M&E). Dialogue enhancement (DE) logic applies amplification to the L&R channels of the separated storytelling dialogue and attenuation to the L&R channels of the reverberation. The amplified dialogue and attenuated reverberation may be remixed after the DE logic. Dynamic range compression (DRC) logic may then be applied to the dialogue-enhanced stereo audio, for example, to compress the loudness range of the M&E. Audio loudness normalization (ALN) logic may be applied to normalize the output stereo to a desired target loudness, e.g., -16 LKFS, or the same LKFS as the measured, integrated loudness of the sound source. The system may use a two-channel separation model, and Meta's Demucs product may be used to provide the audio separation.
[0197] As described herein, this disclosure provides an alternative embodiment of a "DRCed" cinematic stereo dialogue enhancement engine, a dialogue intelligibility engine for 2.0 stereo audio. In this embodiment, dynamic range compression (DRC) is applied to the cinematic mix source. Dialogue separation (DI) then extracts the active storytelling dialogue from the rest of the mix (e.g., M&E). Dialogue enhancement (DE) then applies amplification to the L&R channels of the separated storytelling dialogue and attenuation to the L&R channels of the reverberation. The amplified dialogue and attenuated reverberation are mixed after DE is applied but before ALN is applied. Audio loudness normalization (ALN) is applied to normalize the output stereo to a desired target loudness, such as 16 LKFS, or the same LKFS as the measured integrated sound source.
[0198] As described herein, this disclosure provides an alternative embodiment of a dynamic range controlled dialogue clarity engine for surround and immersive audio in a "dialogue enhanced" center channel. In this embodiment, dialogue separation (DI) uses a two-way separation (ML) model to extract active storytelling dialogue from the rest of the mix (e.g., M&E) in the center channel. Dialogue in the other audio channels is essentially sound effects, not storytelling. Dialogue enhancement (DE) applies amplification to the separated storytelling dialogue channel and attenuation to the reverberant channel so that the loudness of the original program is maintained. Dynamic range compression (DRC) is applied to the dialogue enhanced center channel, resulting from the addition of the amplified dialogue and the attenuated reverberation. Audio loudness normalization (ALN) is applied to normalize the "DE+DRC enhanced" center channel to a desired target loudness, e.g., -16 LKFS, or the same LKFS as the measured and integrated center channel.
[0199] As described herein, this disclosure provides an alternative embodiment of a dialogue clarity engine for dialogue enhancement of surround and immersive audio in a "DRCed" center channel. In this embodiment, dynamic range compression (DRC) is applied to the source center channel. Dialogue separation (DI) uses a two-way separation (ML) model to extract active storytelling dialogue from the rest of the mix (e.g., M&E) in the center channel. Dialogue in other audio channels may be considered sound effects and not storytelling (or speech essential to the story). Dialogue enhancement (DE) applies amplification to the separated storytelling dialogue channel and attenuation to the reverberant channel so that the original program loudness is maintained. Audio loudness normalization (ALN) may be applied to normalize the "DRC+DE-enhanced" center channel to a desired target loudness, e.g., -16 LKFS, or the same LKFS as the measured and integrated center channel.
[0200] As described herein, the present invention enables the provision of accessible audio, meaning that a listener can hear and understand the dialogue spoken within the content. In some embodiments, one option for activating accessible audio includes automatic activation and a streaming software application running on a device, such as a smartphone, laptop, smart TV, or other smart playback device, when the streaming software application has access to the device's microphone, where the device can measure the noise floor and adjust a crossfade renderer (X-fade) to provide the desired listening experience.
[0201] In some embodiments, as another option for activating accessible audio in situations where a microphone or similar device is not accessible, a user uses a smart TV remote control, as described herein, to activate an accessibility renderer implemented within the device's streaming service app, as described herein. This indicates the user's ability to adjust accessibility features using a remote (or similar device) paired with the streaming service app. For example, variations of short presses or long presses of different buttons on the remote (e.g., volume up / down buttons) may be used to control different functions (e.g., the accessibility renderer). This embodiment may be beneficial if the user's device does not have a microphone capable of measuring the noise floor, etc. Customization of the remote's functionality can be tailored to accommodate the user's ability to control the accessibility renderer (or Acc app and decoder), which may be part of the streaming service application in the smart TV's OS.
[0202] As described herein, the present disclosure provides an accessibility fader (or X-fade) that may be displayed in a UI (or UX, user experience) (see FIGS. 12A-12C) that may be part of a streaming service application (e.g., Apple Plus, Hulu, Amazon Prime Video, etc.) that invokes an accessibility renderer. Furthermore, there are many ways to implement an accessibility renderer or X-fade renderer. The present disclosure indicates that a power-saving cross-fade renderer may be used to implement the accessibility renderer. If desired, other formulas, configurations, and techniques may be used for the X-fader, some of which have been previously described herein.
[0203] In some embodiments, as described herein, an accessibility switcher may be activated on the UI / UX (user interface / user experience) of any streaming app that delivers audio (premium entertainment, podcasts, music, etc.), allowing end users to control the degree to which they want their content to be accessible via an accessibility fader (or X-fade). The accessibility fader (or X-fade) blends both cinematic audio and accessible audio (Mac-Acc-Mix or processed DRC+DE) in real time according to a limited set of pre-rendered, pre-defined levels of accessibility, and is natively available at launch as a default state. In one example, five stereo streams may be pre-rendered and streamed to the end user depending on which accessibility level is selected via the X-fade or UI / UX fader. This post-decode switching solution may be applied for each language delivered to the end user; not all languages need have an accessibility payload. This solution is fully backward compatible with traditional ABR (adaptive bitrate) switching techniques, and the payload remains unchanged. In some embodiments, the "accessibility switcher" may be a UI / UX element that allows a user to select a pre-rendered accessible audio stream. A "traditional ABR" (adaptive bitrate) switcher is an audio switching technique known in the art and is used by streaming services to select the best stream available to a device based on the device's capabilities, available network speed, and the customer's membership plan (e.g., ad or premium).
[0204] As described herein, in some embodiments, the present disclosure may have only two pre-rendered streams. The two pre-rendered streams are the cinematic mix and the Max-Acc-Mix (or "drc+de" mix) accessible audio stream, allowing for full accessibility rendering with only two streams / stereo pairs. However, in alternative embodiments, the renderer may select two of three pre-rendered streams: the cinematic (Cin-Mix) audio stream, the accessible Mid-Acc-Mix (or "drc-only" mix) audio stream, and the accessible Max-Acc-Mix (or "drc+de" mix) audio stream, relying on a three-input progressive accessibility fader with a three-gain structure, as shown in FIG. 10C . In some embodiments, the "progressive" accessibility fader may be a cross-fade renderer that can access the three pre-rendered accessible audio streams and more gradually render multiple layers of accessibility in real time, as described herein.
[0205] As described herein, the present disclosure provides an option for a device (e.g., a smartphone, tablet, etc.) paired with a streaming service app to activate accessible audio through user control of the streaming service application's accessibility renderer. This embodiment may be beneficial if the user's device does not have a microphone capable of measuring the noise floor, etc.
[0206] According to embodiments of the present disclosure, an accessibility renderer (or X-Fade) may be invoked on the UI / UX of any streaming app delivering audio (premium entertainment, podcasts, music, etc.) to end users, allowing them to have very precise and continuous control over how accessible they want their content to be, via an accessibility fader (or X-Fade) that blends both cinematic audio and accessible audio (Mac-Acc-Mix or processed DRC+DE) in real time. Such a post-audio decode rendering solution is applied per language delivered to end users, and not all languages need to have an accessibility payload. This solution is fully backward compatible with traditional ABR (adaptive bitrate) switching / decoding techniques. However, the accessibility payload is twice as large as the traditional payload, since a minimum of two stereo pairs must be delivered to the renderer. In some embodiments, a "total" accessibility fader may be a cross-fade renderer that has access to only two pre-rendered accessible audio streams (one stream accessible and one inaccessible) and can fully render all layers of accessibility in real-time, e.g., all at once.
[0207] As described herein, the present disclosure provides a combination of NF-adaptive and full audio accessibility, dialogue enhancement (DE), and dynamic range control (DRC). In this case, the ambient noise floor (NF) may be estimated while filtering the frequency range of the human voice. Crossfades (left and right) maintain power so that loudness and stereo / spatial image are fully preserved. DRC is applied before encoding and affects the loudest and quietest parts of the M&E (but not the dialogue). Once DRC is applied, a target loudness normalization gain is applied to the accessible mix to ensure that all sonic elements in the accessible mix are above the noise floor, e.g., -16 LKFS, the industry standard for noisy environments. PL represents the program loudness in LKFS units as specified in ITU-R BS.1770-1. LRA, or loudness range, represents a measure of the range between the loudest and quietest parts of an audio track across all channels. DPL, or Dialogue to Program Loudness, represents the program loudness in LKFS as specified in ITU-R BS.1770-1.
[0208] In some embodiments, dynamic range control (DRC) and dialogue enhancement (DE) may be combined, where DRC is applied first with a target loudness normalization gain to ensure that all sonic elements of the accessible mix are above the "worst case" noise floor (e.g., -16 LKFS, an industry standard for noisy environments), and DE is applied after DRC as a final enhancement to further amplify the dialogue for M&E.
[0209] In some embodiments, the present disclosure includes providing personalized accessibility to adaptive hearing capabilities, where accessible audio (e.g., Max-Acc-Mix) is decoded and rendered along with cinematic audio (1) and further equalized to compensate for each user's hearing loss frequency response for a given streaming service (2). Hearing loss functionality may be provided by a third-party audiogram API such as Knisper® (for personalized EQ compensation), the device's operating system (e.g., iOS), and / or a listening test mode in the streaming service app.
[0210] In some embodiments, as described herein, the present disclosure includes providing accessible, personalized audio delivery for voice-on-demand (VOD). The accessible stereo mix may be produced by a studio and / or generated by a media processing DCE that securely scales internally to the entire VOD catalog. Accessibility rendering (cinematic (i.e., Cin-Mix) crossfade to Max-Acc-Mix (i.e., DE+DRC)) accommodates all streaming environments and devices. The X-Fade renderer is driven by the captured NF or controlled by the user within the streaming service app (e.g., the Acc app or Hub Acc app 16 in Figures 2A-2D) using an accessibility fader or X-fader. Personalized EQ (PEQ) may be applied after (or post) the X-Fade renderer in the playback experience, per user. Hearing loss capabilities may be measured during streaming service profile setup or provided by the device's OS. PEQ compensates for hearing loss features in the active member profile by amplifying specific frequency ranges within the Xfade mix, allowing for personalized hearing adjustments.
[0211] In some embodiments, the present disclosure includes a method for providing a multi-channel accessible stereo encoder (MASE) in accordance with an embodiment of the present disclosure. MPEG AAC-LC and Dolby E-AC3 support a "dual stereo" encoding mode, for example, where 2x 2.0 / stereo are encoded independently. The Dolby (E-AC3 + Joint-Object Coding) = Atmos encoder supports primary and associated audio in a single bitstream. As described herein, native support for implementing the present disclosure may be provided in known open-source software audio processing tools, such as FFMPEG® or Gstreamer®.
[0212] The following defined terms and concepts may be referenced throughout this disclosure, may be components of the systems and methods of the present disclosure, and / or may be referenced in various alternative embodiments of the systems and methods of the present disclosure, and may be referenced in the aforementioned commonly owned provisional patent applications.
[0213] A crossfade renderer, sometimes referred to herein as an X-fade, X-fader, renderer, crossfade, or CFR, combines or blends multiple audio files or audio experiences with a power-preserving amplitude panning curve instead of abruptly switching between audio streams with different functionality or focus. In one embodiment, this involves inputting multiple audio assets, audio files, or audio content with multiple audio channels in WAV / PCM format, resulting in an output of a single audio file with multiple audio channels in WAV / PCM format. The crossfade renderer, as described herein, may use software running on a streaming service (or content provider) application running on iOS®, Android®, Roku®, etc.
[0214] A dialogue clarity engine (DCE) is a digital audio signal processing engine that performs one or more of the following: (1) separation of dialogue (DI) from other elements of a mix (usually referring to neural network-based source separation techniques); (2) dialogue enhancement (DE) that amplifies dialogue to improve overall dialogue intelligibility; and (3) dynamic range control (DRC) of a mix to improve dialogue intelligibility over loudspeakers of any size and in any type of noisy environment. The input to this engine is typically an audio asset or audio content file containing active dialogue (used for storytelling), and the resulting output is an audio asset or audio content file in which the active dialogue has been amplified so that the overall loudness of the mix remains unchanged compared to the input source. The DCE may be implemented using software running on hardware owned / provided by the streaming service, as described herein.
[0215] Dialogue separation (DI) involves separating active storytelling dialogue from an audio source mix. Specifically, dialogue separation (DI) involves separating active storytelling dialogue from an asset audio source mix using software, typically a machine learning (ML) model, such as an open-source artificial intelligence (AI) or machine learning (ML) model like Meta's Demucs, running on the streaming service's (or content provider's) hardware. DI involves inputting an asset audio source in uncompressed PCM format (e.g., N=2ch, 6ch, 10ch, etc.) and resulting in audio outputs containing only dialogue content (e.g., K=2ch, 1ch, 1ch, etc.) and audio containing only non-storytelling dialogue (or non-dialogue, music and effects, or M&E) content (e.g., reverberation, L=2ch, 5ch, 9ch, etc.).
[0216] Dialogue enhancement (DE) applies amplification to isolated storytelling dialogue channels and attenuation to reverberant channels so that the loudness of the original program is maintained. DE involves analyzing and processing digital audio signals to amplify active dialogue (used for storytelling) and attenuate other components of the mix to make the dialogue more prominent without changing the overall loudness of the resulting mix compared to the source or input audio mix. The input to DE is an audio asset with active dialogue (used for storytelling), and the resulting output is an audio asset with active dialogue (used for storytelling) amplified so that the overall loudness of the mix is unchanged. DE may be performed by software running on hardware owned / provided by the streaming service, as described herein.
[0217] Dynamic Range Compression / Control (DRC) involves the analysis and processing of digital audio signals to limit the amplitude length / range of audio samples so that the listener avoids experiencing a sudden loudness spike when going from a quiet scene to a noisy scene. For the purposes of DRC, the terms compression and control are used interchangeably. It involves inputting a full dynamic range audio asset and outputting a compressed dynamic range audio asset. DRC is achieved using software running on hardware owned / provided by the streaming service. Examples of DRC profiles that can be applied include "Film Standard," "Film Light," "Noisy Environment," etc.
[0218] Enhanced Dialogue Insertion (EDI) (or Stereo EDI or Stereo Remix) uses known audio signal synthesis software to separately combine the left channel of a dialogue (or vocal) mix and a reverberation (or M&E) mix with the right channel of a dialogue (or vocal) mix and a reverberation (or M&E) mix to produce a dialogue-enhanced left / right (L / R) channel. The input to Stereo EDI or Remix can be (1) the output of DRC+DE processing, or (2) the output of the EDI logic, which is the output of DE+DRC processing, resulting in a fully accessible stereo pair L / R.
[0219] Loudness Range (LRA) is a measure of the range between the loudest and quietest parts across all channels of an audio track. Integrated Loudness (IL) is defined by LoudLAB (https: / / www.loudlab-app.com / sonicatom / en / 2.html), which defines IL and various other audio terms described herein.
[0220] Accessible Target Loudness (ATL) involves using software to normalize audio channels to a target audio loudness in LKFS units. The input is a stereo pair of L / R channels and a target LKFS value (typically a negative number such as -20 or -16 LKFS), and the resulting output is a normalized stereo pair of L / R with the audio loudness measured at the target LKFS value. The term gain can refer to the accessible target loudness (value in LKFS units).
[0221] Audio Loudness Normalization (ALN) is used to set expectations for how loud the ALN audio output can be when played on commercial audio electronics equipment / devices. By applying ALN, the integrated loudness of the processed audio reaches X LKFS, averaged over the entire duration of the asset. X is set by the content provider, along with best practices and industry standards. The input for ALN is the accessible audio track (or Max-Acc-Mix) from the DCE62, consisting of DRC+DE or DE+DRC processed audio, with two channels for stereo sources and one channel (center) for surround / immersive audio sources. The output of ALN is the normalized audio output of the ALN and DCE, which may be two channels for stereo sources or one channel (center) for surround / immersive audio sources. ALN involves the use of software running on hardware owned / provided by the content provider / streaming service. Audio loudness normalization (ALN) is applied to normalize the output stereo to a desired target loudness, for example -16 LKFS or the same LKFS as the integrated loudness (IL) measured at the sound source.
[0222] In some embodiments, the present disclosure includes at least one method for increasing audio dialogue levels in digital cinematic content while preserving the creative intent of the cinematic content, the method comprising receiving a cinematic audio signal, receiving an accessible audio signal (e.g., Max-Acc-Mix), and combining the cinematic audio (Cin-Mix) signal and the accessible audio (Max-Acc-Mix) signal to provide an enhanced dialogue audio signal that allows a listener to hear the dialogue above the noise floor.
[0223] The systems and methods of the present disclosure work with any codec that provides the functionality and capabilities described herein. Examples of codecs (or encoding and decoding formats) that can be used in the present disclosure are listed at https: / / en.wikipedia.org / wiki / Comparison_of_audio_coding_formats. Also, known IAMF standard codecs from AOM are listed at https: / / aomedia.org / specifications / iamf / . As known, a "codec" is a hardware- or software-based process that compresses and decompresses large amounts of data. Codecs are used in applications such as the present disclosure to efficiently transmit media files over a network and to receive and play media files for a receiving user.
[0224] As used herein, the term "accessible" audio mix or track means an audio track in which the dialogue is intelligible in all noise environments, devices, and capabilities, allowing a listener to understand the dialogue as if they were in a movie theater (without subtitles). Also, in this disclosure, LKFS, dB, and TruePeak audio or sound units are used in various examples herein for illustrative purposes, and one of ordinary skill in the art will understand how these units relate to each other.
[0225] Also, although some embodiments of the disclosed system are shown storing the encoded (and interleaved) digital audio tracks on an encoding server at the content provider's end and then retrieving them for use at the user / listener's end, it will be understood that the encoded digital audio tracks may alternatively or additionally be transmitted (e.g., by an encoder or other hardware or software) over a communications network directly to the user / listener's device, which can decode the digital audio tracks for use on the user's playback device as described herein.
[0226]
[0010] Embodiments of the present disclosure also include a method described herein, wherein generating an accessible audio signal may include using a dialog clarity engine. Embodiments of the present disclosure also include the method described above, wherein the dialog clarity engine may include receiving a cinematic audio signal and performing at least one of dialog isolation (DI), dialog enhancement (DE), dynamic range compression (DRC), audio loudness normalization, dialog loudness relative to the program (or non-dialogue, M&E, or reverberation) (DPL), mono enhanced dialog insertion or EDI (or remix), and stereo EDI (or remix). Embodiments of the present disclosure also include the method described above, wherein combining may include using a crossfade renderer.
[0227] Also, in some embodiments, the crossfade experience may rely on an accessible audio track mix (or accessible mix) generated by a mixer rather than a (possibly AI-based) dialogue clarity engine (DCE). Also, in some embodiments, to allow for broad compatibility with current production methods from studios, the mixer (which generates the accessible mix or Max-Acc-Mix) may be tuned or pre-configured for worst-case scenarios such as elderly customers with hearing impairments, poor TV speakers, poor sound connectivity, and / or a high ambient noise floor.
[0228] The systems and methods of the present disclosure include necessary computers, servers, devices, etc., having the necessary electronics, computing power, interfaces, memory, hardware, software, firmware, logic / state machines, databases, microprocessors, communications links, displays or other visual or audio user interfaces, printing devices, and any other input / output interfaces to provide the functionality or achieve the results described herein. Except as expressly or implicitly indicated herein, the processes or method steps described herein may be implemented in software modules (or computer programs) running on one or more general-purpose computers. Specially designed hardware may alternatively be used to perform certain operations. Thus, any of the methods described herein may be performed by hardware, software, or any combination of these approaches. Additionally, a computer-readable storage medium may have stored thereon instructions that, when executed by a machine (such as a computer), result in the performance according to any of the embodiments described herein.
[0229] Additionally, the computers or computer-based devices described herein may include any number of computing devices capable of performing the functions described herein, including, but not limited to, tablets, laptop computers, desktop computers, smartphones, smart televisions, set-top boxes, e-readers / players, etc.
[0230] Although the present disclosure describes the use of example techniques, algorithms, or processes for implementing the present disclosure, it should be understood by those skilled in the art that other techniques, algorithms, and processes, or other combinations and sequences of the techniques, algorithms, and processes described herein, may be used or implemented to achieve the same functions and results described herein and are included in the present disclosure.
[0231] As will be reasonably understood by those skilled in the art, the process descriptions, steps, or blocks in the process flow diagrams or logic flow diagrams provided herein are indicative of one potential implementation and do not imply a fixed order; alternative implementations are included within the preferred implementations of the systems and methods described herein, and functions or steps may be omitted or performed in an order different from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved.
[0232] Unless otherwise indicated herein, either expressly or implicitly, it is to be understood that any feature, characteristic, alternative, or modification described with respect to a particular embodiment herein may also be applied to, used in, or incorporated into other embodiments described herein. Additionally, the drawings herein are not drawn to scale unless otherwise indicated.
[0233] In particular, conditional language such as "can," "potential," "might," or "may," unless otherwise specified or understood within the context in which it is used, is intended to generally convey that a particular embodiment can include, but does not require, a particular feature, element, or step. Thus, such conditional language is not generally intended to imply that a feature, element, or step is somehow required by one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, or steps are included in or performed in any particular embodiment, with or without user input or prompting.
[0234] While the present invention has been described and illustrated with respect to exemplary embodiments thereof, the foregoing and other various additions and omissions can be made therein and therein without departing from the spirit and scope of the disclosure.
Claims
1. 1. A computer-based method for generating a personalized accessible audio mix (Acc-Mix) of audio content having dialog audio and non-dialogue audio that enables at least one user to understand the dialog audio when played back on a user playback device in a listening environment having a realistic noise floor (NF) level, comprising: receiving, at the user playback device, an original cinematic audio mix (Cin-Mix) having dialog audio and non-dialogue audio set in a recording studio; receiving, at the user playback device, a maximum accessible mix (Max-Acc-Mix) of the dialog audio and the non-dialogue audio, the Max-Acc-Mix being an audio update of the Cin-Mix such that the loudness range of the dialog audio and the loudness range of the non-dialogue audio are greater than a predetermined allowable noise floor and less than a predetermined maximum audio loudness limit; and generating a personalized Acc-Mix audio by combining the Cin-Mix and MaxAcc-Mix to preserve an output and having the loudness of the dialogue audio, so that the user hears and understands the dialogue audio when played on the user playback device in a listening environment having the actual noise floor (NF), wherein the volume of the personalized Acc-Mix ranges from the value of the Cin-Mix to the value of the MaxAcc-Mix and is an output that preserves a combination of CinMix and MaxAccMix therebetween, and the volume of the personalized Acc-Mix is set based on at least one of a user input and an actual noise floor level.
2. 2. The computer-based method of claim 1, wherein Acc-Mix is automatically set based on an actual noise floor in the listening environment, the actual noise floor being measured by a sound sensor that filters human voices from the audio signal to provide a level of the actual noise floor.
3. 2. The computer-based method of claim 1, wherein the Cin-Mix and the Max-Acc-Mix are received in a single stream or digital file, and further comprising decoding the single stream to provide separate tracks for the Cin-Mix and the Max-Acc-Mix.
4. 10. The computer-based method of claim 1, wherein the Acc-Mix is manually adjusted using an X-fade slider via a user interface.
5. 10. The computer-based method of claim 1, wherein the at least one user comprises a plurality of users, each user being able to set their own personalized AccMix.
6. The power retention of the combination of CinMix and MaxAccMix is 10. The computer-based method of claim 1, implemented using a crossfade renderer that uses sine and cosine curves as coefficients based on the position value of an X-fade slider.
7. The computer-based method of claim 1 , further comprising providing personalized equalization to the Acc-Mix based on frequency range weighting factors corresponding to a user.
8. 10. The computer-based method of claim 1, further comprising providing a UI that provides at least one of a content list, an accessibility fader, a voice-over fader, a live sports fader, an audio description fader, an adaptive voice-over, a language selection, a language level selection, and a personalized EQ selection for generating an Acc-Mix.
9. 10. The computer-based method of claim 1, further comprising receiving an instruction from a user to adjust the X-fade.
10. 1. A computer-based method for generating a maximum accessible audio mix (Max-Acc-Mix), comprising: Receiving original cinematic mix (Cin-Mix) audio content; and adjusting the Cin-Mix using a dialogue clarity engine (DCE) having a predetermined DCE gain to generate the Max-Acc-Mix; A computer-based method comprising:
11. The dialogue clarity engine includes: receiving the cinematic audio signal; performing at least one of dialogue isolation (DI), dialogue enhancement (DE), dynamic range compression (DRC), audio loudness normalization (ALN), and enhanced dialogue insertion (EDI); The method of claim 10, comprising:
12. 11. The method of claim 10, wherein the dialog clarity engine (DCE) includes at least one of a dialog separation (DI) that separates the dialog audio from the non-dialog audio, a dialog enhancement (DE) that amplifies the dialog audio relative to the non-dialog audio with a predetermined DE gain, a dynamic range compression (DRC) that compresses the non-dialog audio or synthesized dialog and non-dialog audio, and an audio loudness normalization (ALN) that amplifies the dialog audio or the synthesized dialog and non-dialog audio with a predetermined ALN gain.
13. 12. The method of claim 11, wherein the dynamic range compression (DRC) comprises attenuating a predetermined upper or lower range in the non-dialogue audio or the combined dialogue and non-dialogue audio with a predetermined DRC gain profile.
14. 11. The method of claim 10, wherein generating the MaxAcc-Mix comprises adjusting the DCE amplification gain so that both the dialogue audio and non-dialogue audio are above a predetermined acceptable noise floor.
15. 11. The method of claim 10, further comprising providing a UI that allows a content provider user / administrator to adjust DCE parameters, the DCE parameters including at least one of a noise floor, a Dx gain factor, an MNE gain factor, Gdx, Gmne, a DCE mode, a DRC gain profile, and an ALN gain.
16. The method of claim 10, wherein the Cin-Mix is in at least one of the following formats: mono, stereo, surround sound, and immersive sound.
17. 11. The method of claim 10, wherein the dialogue enhancement (DE) comprises at least one of amplifying the dialogue and attenuating the non-dialogue while maintaining the same overall loudness, and amplifying only the dialogue.
18. 11. The method of claim 10, wherein the dynamic range compression (DRC) comprises attenuating a predetermined upper or lower range of the non-dialogue audio or synthesized dialogue and non-dialogue audio by a predetermined DRC gain profile.
19. 1. A computer-based method for generating an accessible audio mix (Acc-Mix) of dialogue audio and non-dialogue audio (reverberant music / effects) sounds that enables a listener-user to understand dialogue audio played on a user playback device (or listening device) in listening environments having a plurality of different actual noise floor levels, comprising: receiving, at a user device, a cinematic mix (Cin-Mix) and a maximum accessible mix (MaxAcc-Mix) associated with video / audio content viewed by the user from the user device; a computer-based method for blending the Cin-Mix and the MaxAcc-Mix to generate the Acc-Mix audio mix, wherein the resulting Acc-Mix audio mix conserves power and is above a noise floor (NF) in the listening environment of the user device, the Acc-Mix having a range from a value of the Cin-Mix to a value of the MaxAcc-Mix, the value of the Acc-Mix being based on at least one of a user input and an actual noise floor level.
20. 20. The computer-based method of claim 19, wherein the Acc-Mix is automatically set based on an actual noise floor in the listening environment, the actual noise floor being measured by a sound sensor that filters human voices from the audio signal to provide the level of the actual noise floor.
21. 1. A computer-based method for enhancing the audio dialogue level of digital cinematic content and preserving the creative intent of the cinematic content, comprising: receiving a cinematic audio mix; receiving an accessible audio mix; combining the cinematic audio mix and the accessible audio mix to provide an enhanced dialogue audio mix that enables a listener to hear the dialogue above the noise floor; A computer-based method comprising:
22. 22. The method of claim 21, wherein the combining includes using a crossfade renderer.
23. 1. A computer-based method for performing a fade from a first audio mix (MixA) to a second audio mix (MixB), comprising: receiving the MixA audio mix, receiving the MixB audio mix, Compositing MixA and MixB using a crossfade renderer having an output that transitions from MixA to MixB based on the position of a crossfade slider; A computer-based method comprising:
24. 24. The method of claim 23, wherein the combining includes using a crossfade renderer.
25. MixA and MixB are MixA = the original version of the main language (or Cin-Mix) and MixB = the main language of the enhanced dialogue (Max-Acc-Mix) MixA = original version of the main language and MixB = audio description of the main language, Mix A = original version in the main language and Mix B = audio voiceover in the alternative language. MixA = original version in the main language and MixB = audio description in the alternative language, Mix A = home team commentary mix and Mix B = away team commentary mix, and 24. The method of claim 23, wherein Mix A = home team commentary mix and Mix B = stadium sounds (no commentary).
Citation Information
Patent Citations
Speech communication apparatus
JP2004187165A
Data processing device and method
JP2009531926A
Electronic device
JP2016082459A
MIXING CONTROL DEVICE, AUDIO SIGNAL GENERATOR, AUDIO SIGNAL SUPPLY METHOD AND COMPUTER PROGRAM
JP2016522640A
Sound signal processing device
JP2020014138A