Audio signal processing system and method
The adaptive audio system addresses the limitations of current cinema audio systems by combining channel-based and object-based processing with metadata-driven rendering, enabling precise sound positioning and enhanced audio-visual coherence across diverse playback environments.
Patent Information
- Application Number
- JP2025081687
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-04-20
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-20
AI Technical Summary
Current cinema audio systems struggle to accurately reproduce sound sources in various playback environments, lacking flexibility and precision in positioning sounds relative to the listener, and often require significant processing and knowledge of the playback environment, limiting the ability to create immersive and realistic audio experiences.
An adaptive audio system that combines channel-based and object-based audio processing, utilizing metadata to describe the intended position of audio streams, allowing for flexible rendering based on the playback environment, and incorporating novel speaker layouts and spatial description formats to enhance audio-visual coherence and sound quality.
The system enables precise sound positioning, improved audio-visual coherence, and increased flexibility in audio reproduction across different environments, simplifying distribution and installation while maintaining artistic intent, and providing a more immersive and realistic audio experience.
Smart Images

Figure 2025122045000001_ABST
Abstract
Description
[Technical Field]
[0001] One or more embodiments relate generally to audio signal processing, and more particularly to hybrid object and channel-based audio processing for use in cinema, home, and other environments. [Background technology]
[0002] The subject matter described in the Background Art section should not be considered prior art merely because it is mentioned in the Background Art section. Similarly, problems mentioned in or related to the subject matter of the Background Art section should not be considered to have been previously recognized in the prior art. The subject matter of the Background Art section may simply represent multiple different approaches and may itself be inventions.
[0003] Since the advent of film-accompanied sound, there has been a steady evolution in the technology used to capture the creator's artistic intent for a film soundtrack and accurately reproduce it in a cinematic environment. The fundamental role of cinema sound is to support the story shown on screen. A typical cinema soundtrack contains many different sound elements that correspond to the on-screen images: dialogue, noise, and sound effects emanating from the different on-screen elements and combining with background music and ambient effects to form the overall audience experience. The creator's or producer's artistic intent dictates a desire to reproduce these sounds so that they correspond as closely as possible to what is shown on screen, in terms of source location, intensity, movement, and other similar parameters.
[0004] Current cinema authoring, distribution, and playback methods have limitations that restrict the creation of truly immersive, lifelike audio. Traditional channel-based audio systems send audio content in the form of speaker feeds to individual speakers in a playback environment, such as a stereo or 5.1 system. The advent of digital cinema has created new standards for on-film sound, such as the incorporation of up to 16 channels of audio, giving content creators greater creativity and audiences a more enveloping and realistic acoustic experience. The emergence of 7.1 surround systems has provided a new format that increases the number of surround channels by splitting the existing left and right surround channels into four zones, increasing the scope and control of sound designers and mixers for positioning audio elements in the theater.
[0005] To further enhance the listener's experience, playing sound in virtual three-dimensional environments has become an area of research and development. Spatial representation of sound makes use of audio objects, which are parametric source descriptions of an audio signal and their associated apparent source location (e.g., 3D coordinates), apparent source width, and other parameters. Object-based audio is increasingly being used in many current multimedia applications, including digital movies, video games, simulations, and 3D video.
[0006] As a means of delivering spatial audio, it is essential to extend beyond traditional speaker feeds and channel-based audio, and there has been considerable interest in model-based audio descriptions. Model-based audio descriptions offer the freedom for listeners / exhibitors to choose a playback configuration that suits their individual needs and budget, and to have the audio rendered in the configuration of their choice. At a high level, there are currently four main spatial audio description formats: speaker feeds, where the audio is described as signals intended for speakers at nominal speaker positions; microphone feeds, where audio is described as signals captured by virtual or real microphones in a given array; a model-based description in which the audio is described in terms of a sequence of audio events at described locations; and Binaural describes audio by the signal that reaches the listener's ears. These four description formats are often associated with one or more rendering techniques that convert the audio signal into speaker feeds. Current rendering techniques include panning, where the audio stream is converted into speaker feeds using a set of panning laws and known or assumed speaker positions (typically rendered before distribution); Ambisonics, where microphone signals are converted to feed a scalable speaker array (typically rendered after distribution); WFS (wave field synthesis), where sound events are converted into appropriate speaker signals to synthesize a sound field (typically rendered after distribution); and binaural, where L / R (left / right) binaural signals are sent to the left and right ears (rendered before or after distribution), typically using headphones but also using speakers and crosstalk cancellation. Of these formats, the speaker feed format is the most common due to its simplicity and effectiveness. The best sonic result (most accurate and most reliable) can be achieved by mixing / monitoring and delivering directly to the speaker feeds, since there is no processing between the content creator and the listener. Speaker feed descriptions generally provide the highest fidelity when the playback system is known in advance; however, in many real-world applications, the playback system is unknown. Model-based descriptions are likely to be the most adaptive, since they make no assumptions about the rendering technique and are therefore easily adaptable to any rendering technique. Model-based descriptions capture spatial information efficiently, but become inefficient as the number of sound sources increases.
[0007] For many years, cinema systems have featured discrete screen channels in the form of left, center, right, and sometimes "inner left" and "inner right" channels. These discrete sources generally have sufficient frequency response and power handling to accurately place sounds in different areas of the screen and tonal matching as the sounds move or pan between locations. Recent developments to enhance the listener experience have attempted to accurately reproduce the location of sounds relative to the listener. In 5.1 systems, surround "zones" comprise an array of multiple speakers, all of which contain the same audio information in each of the left and right surround zones. While such arrays may be useful for "ambient" or diffuse surround effects, in everyday life, sound effects emanate from randomly positioned point sources. For example, in a restaurant, while ambient music is blaring throughout, subtle but distinct sounds emanate from multiple points—for example, people speaking from one point and knives clinking on plates from another. Placing these sounds directly around the auditorium enhances realism, even if they are not noticeable. Overhead sound is also an important component of surround definition. In the real world, sounds emanate from all directions, not necessarily from a single horizontal plane. Realism is enhanced when sounds come from above, or from the "upper hemisphere." However, current systems cannot provide truly accurate reproduction of different audio types in a variety of different playback environments. Attempting to accurately represent sound location using existing systems requires significant processing, knowledge of the actual playback environment, and configuration, making current rendering systems impractical for most applications.
[0008] What's needed is a system that supports multiple screen channels to increase on-screen sound and dialogue definition and improve audio-visual coherence, as well as the ability to precisely position sound sources anywhere in the surround zone to improve audio-visual transitions from screen to room. For example, if an on-screen character is looking toward a sound source in the room, the sound engineer ("mixer") should have the ability to precisely position that sound to match the character's line of sight and ensure the effect is consistent throughout the audience. However, traditional 5.1 or 7.1 surround sound mixes are at a disadvantage in large listening environments because their effect is highly dependent on the listener's seating position. The increased resolution of surround creates new opportunities for room-centric sound, as opposed to the traditional approach of creating content by assuming a single listener is in a "sweet spot."
[0009] Aside from spatial issues, current state-of-the-art multi-channel systems also present challenges with regard to sound quality. For example, timbral qualities, such as the hiss of steam escaping from a broken pipe, must be reproduced using an array of multiple speakers. The ability to direct sound to a single speaker offers mixers the opportunity to eliminate array reproduction artifacts and deliver a more realistic experience to the audience. Traditionally, multiple surround speakers do not support the same full range of audio frequencies and levels that large screen channels support. Historically, this has presented challenges for mixers, reducing their ability to freely move full-range sound from the screen into the room. As a result, theater owners have felt reluctant to upgrade their surround channel configurations, hindering the widespread adoption of higher-quality sound systems. [CROSS-REFERENCE TO RELATED APPLICATIONS]
[0010] This application claims priority to U.S. Provisional Application No. 61 / 504,005, filed July 1, 2011, and U.S. Provisional Application No. 61 / 636,429, filed April 20, 2012, both of which are incorporated by reference in their entirety for all purposes. Summary of the Invention
[0011] This paper describes a cinema sound format and a processing system that includes a new speaker layout (channel configuration) and associated spatial description format. It defines an adaptive audio system and format that supports multiple rendering techniques. Audio streams are transmitted along with metadata that describes the "mixer's intent," including the desired position of the audio stream. The position is expressed as a specified channel (within a predefined channel configuration) or as 3D position information. This channel and object format optimally combines channel-based and model-based audio scene description methods. Audio data for the adaptive audio system includes multiple independent monophonic audio streams. Each stream has associated metadata that indicates whether the stream is a channel-based stream or an object-based stream. Channel-based streams have rendering information encoded with the channel name.
[0012] Object-based streams have location information encoded using mathematical formulas that are encoded in other associated metadata. The original, independent audio streams are packaged as a single serial bitstream containing all of the audio data. This configuration allows the sound to be rendered in an allocentric frame of reference. The rendering location of the sound is based on the characteristics of the playback environment (e.g., room size, shape, etc.) to correspond to the mixer's intent. The object position metadata contains the appropriate allocentric frame of reference information needed to correctly play the sound using the available speaker positions in the room configured to play the adaptive audio content. This allows the sound to be optimally mixed for the playback environment, which may differ from the mixing environment experienced by the sound engineer.
[0013] The adaptive audio system improves the quality of audio in different rooms, with benefits such as improved room equalization and surround bass management, allowing mixers to freely address speakers (whether on-screen or off-screen) without having to worry about timbre matching. The adaptive audio system adds the flexibility and power of dynamic audio objects to traditional channel-based workflows. These audio objects allow creators to control individual sound elements regardless of the playback speaker configuration, including overhead speakers. The system also introduces new efficiencies to the post-production process, allowing sound engineers to efficiently capture all of their intent and monitor or automatically generate surround sound 7.1 and 5.1 versions in real time.
[0014] This adaptive audio system simplifies distribution by encapsulating the essence and artistic intent of audio within a single track file within the digital cinema processor. This single track file can be faithfully reproduced in a wide range of theater configurations. The system provides optimal reproduction of artistic intent when downmixing, i.e., when mix and render using a single inventory that is downward-adapted to the same channel and rendering configuration.
[0015] These and other advantages are provided through embodiments relating to a cinema sound platform, which overcomes current system limitations and delivers an audio experience that goes beyond currently available systems. [Brief explanation of the drawings]
[0016] In the following drawings, the same reference numerals are used to refer to the same elements. The following drawings show various examples, but the implementation(s) are not limited to the examples shown in the drawings. [Figure 1] FIG. 1 illustrates a top-level overview of an audio production and playback environment utilizing an adaptive audio system, according to one embodiment. [Figure 2] FIG. 1 illustrates combining channel and object-based data to create an adaptive audio mix, according to one embodiment. [Figure 3] FIG. 1 is a block diagram illustrating a workflow for generating, packaging, and rendering adaptive audio content according to one embodiment. [Figure 4] FIG. 2 is a block diagram illustrating the rendering stages of an adaptive audio system, according to one embodiment. [Figure 5] 1 is a table listing metadata types and associated metadata elements for an adaptive audio system, according to one embodiment. [Figure 6] FIG. 1 illustrates post-production and mastering of an adaptive audio system, according to one embodiment. [Figure 7]FIG. 1 illustrates an example workflow for a digital cinema packaging process using adaptive audio files, according to one embodiment. [Figure 8] FIG. 1 is a top view showing an example layout of suggested speaker locations for use with an adaptive audio system in a typical auditorium. [Figure 9] FIG. 1 is a front view showing an example of a suggested screen speaker location arrangement for use in a typical auditorium. [Figure 10] FIG. 1 is a side view showing an example layout of suggested speaker locations for use in an adaptive audio system in a typical auditorium. [Figure 11] 1A and 1B are diagrams illustrating example placements of top surround speakers and side surround speakers relative to a reference point according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] Adaptive audio systems and methods, associated audio signals, and data formats that support multiple rendering techniques are described. Aspects of one or more embodiments described herein can be implemented in audio or audiovisual systems that process source audio information in mixing, rendering, and playback systems that include one or more computers or processing devices executing software instructions. Any of the described embodiments can be used alone or in combination with each other. Various deficiencies in the prior art motivated various embodiments, which are described in one or more places herein, but embodiments may not necessarily overcome these deficiencies. In other words, different embodiments overcome different deficiencies described herein. Some embodiments may only partially overcome some deficiencies described herein, some may overcome only one deficiency, and some may overcome none of these deficiencies.
[0018] For purposes of this description, the following terms have the associated meanings: Channel or Audio Channel: A monophonic audio signal or audio stream and metadata whose location is coded as a channel ID, e.g., left front or right top surround. A channel object drives multiple speakers, e.g., the left surround channel (Ls) feeds all speakers in the Ls array.
[0019] Channel Configuration: A predetermined set of speaker zones with associated nominal locations, e.g., 5.1, 7.1, etc.; 5.1 refers to a six-channel surround sound audio system with left and right channels, a center channel, two surround channels, and a subwoofer channel; 7.1 refers to an eight-channel surround system that adds two additional surround channels to a 5.1 system. Examples of 5.1 and 7.1 configurations include the Dolby® Surround system.
[0020] Speaker: An audio transducer or set of transducers that renders an audio signal.
[0021] Speaker Zone: An arrangement of one or more speakers that can be uniquely referenced and that receives a single audio signal, such as left surround, as commonly seen in cinema, and that is specifically excluded or included for object rendering.
[0022] Speaker Channel or Speaker Feed Channel: An audio channel associated with a specified speaker or speaker zone within a defined speaker configuration. A speaker channel is nominally rendered with the associated speaker zone.
[0023] Speaker Channel Group: A set of one or more speaker channels corresponding to a channel configuration (e.g., stereo track, mono track, etc.).
[0024] Object or Object Channel: One or more audio channels with a numerical source description, such as an apparent source position (e.g., 3D coordinates), an apparent source width, etc. An audio stream with metadata where the position is coded as a 3D position in space. Audio Program: The complete set of speaker channels and / or object channels and associated metadata that describes the desired spatial audio presentation.
[0025] Allocentric reference: A spatial reference in which audio objects are defined relative to features in the rendering environment, such as the walls and corners of a room, the standard speaker location, or the screen location (e.g., the front-left corner of a room).
[0026] Egocentric reference: A spatial reference in which an audio object is defined relative to the (audience) listener's point of view, often specified in terms of an angle relative to the listener (e.g., 30° to the right of the listener).
[0027] Frame: A frame is an independently decodable segment into which the entire audio program is divided. Audio frame rates and boundaries are generally aligned with video frames.
[0028] Adaptive Audio: Channel-based and / or object-based audio signals and metadata that renders the audio signal based on the playback environment.
[0029] The cinema sound format and processing system described herein, also known as an "adaptive audio system," utilizes novel spatial audio description and rendering techniques to enhance audience immersion, provide greater artistic control, increase system flexibility and scalability, and simplify installation and maintenance. An embodiment of the cinema audio platform includes multiple individual components, including a mixing tool, packer / encoder, unpacker / decoder, in-theater final mix rendering components, novel loudspeaker designs, and networked amplifiers. The system includes suggestions for new channel configurations for use by content creators and exhibitors. The system utilizes model-based descriptions that support several features: a single inventory that allows optimal use of available loudspeakers, adapting downward and upward to the rendering configuration; Improved sound envelopment, including optimized downmixing to avoid inter-channel correlation; Increased spatial resolution through steer-thru arrays (e.g., audio objects dynamically assigned to one or more loudspeakers in a surround array); Support for alternative rendering methods.
[0030] Figure 1 illustrates a top-level overview of an audio production and playback environment utilizing an adaptive audio system, according to one embodiment. As shown in Figure 1, the comprehensive end-to-end environment 100 includes content creation, packaging, distribution, and playback / rendering components across a wide range of endpoint devices and use cases. The overall system 100 originates with content captured from a number of different use cases, including different user experiences 112. The content capture element 102 includes audio / visual or pure audio content, including, for example, cinema, TV, live broadcast, user-generated content, recorded content, games, music, etc. As content progresses through the system 100 from the capture stage 102 to the final user experience 112, it passes through individual system components and undergoes several key processing steps. These process steps include audio preprocessing 104, authoring tools and processing 106, and audio codec encoding 108, which captures audio data, additional metadata and playback information, and object channels, for example. Various processing effects, such as compression (lossy or lossless), encryption, etc., are applied to the object channels for efficient and secure delivery over various media. Appropriate endpoint-specific decryption and rendering processes 110 are applied to reproduce and deliver an adaptive audio user experience 112. The audio experience 112 represents the playback of audio or audio / visual content through appropriate speakers and playback devices, and can represent any environment in which a listener experiences playback of the captured content, such as a cinema, concert hall, outdoor theater, home or room, listening booth, automobile, game console, headphone or headset system, public address (PA) system, or other playback environment.
[0031] An embodiment of system 100 includes an audio codec 108 that allows for efficient delivery and storage of multi-channel audio programs and is therefore sometimes referred to as a "hybrid" codec. Codec 108 combines traditional channel-based audio data with associated metadata to create audio objects that facilitate the generation and delivery of audio that is adapted and optimized for rendering and playback in environments that may differ from the mixing environment. This allows sound engineers to encode their intentions for how the final audio will sound to a listener based on the listener's actual listening environment.
[0032] Traditional channel-based audio codecs operate under the assumption that an audio program is played back through a speaker array at a predetermined location relative to the listener. To create a complete multi-channel audio program, a sound engineer typically mixes multiple audio streams (e.g., dialogue, music, and sound effects) to create a desired overall impression. Audio mixing decisions are typically made by listening to the audio program played back through a speaker array at a predetermined location, such as a 5.1 or 7.1 system in a particular theater. The final mixed signal is the input to the audio codec. During playback, a spatially accurate sound field is achieved only when the speakers are positioned at their predetermined locations.
[0033] A new audio coding format, called audio object coding, provides distinct audio sources (audio objects) as input to an encoder in the form of separate audio streams. Examples of audio objects include dialogue tracks, single instruments, individual sound effects, and other point sources. Each audio object is associated with spatial parameters, including, but not limited to, sound position, sound width, and velocity information. The audio objects and associated parameters are coded for delivery and storage. Final audio object mixing and rendering occurs at the receiving end of the audio delivery chain as part of audio program playback. This step is based on knowledge of actual speaker positions, allowing the audio delivery system to customize the results to a user's specific listening conditions. Two coding formats, channel-based and object-based, operate optimally under different input signal conditions. Channel-based audio coders are generally more efficient at coding input signals containing a dense mixture of different audio sources. Conversely, audio object coders are more efficient at coding a small number of highly directional audio sources.
[0034] In one embodiment, the method and system 100 components comprise an audio encoding, distribution, and decoding system configured to generate one or more bitstreams that include both traditional channel-based audio elements and audio object-coded elements. Such a combined approach provides greater coding efficiency and rendering flexibility compared to either channel-based or object-based approaches.
[0035] Another aspect of the described embodiments includes backward-compatible extension of a given channel-based audio codec to include audio object coding elements. A new "enhancement layer" containing the audio object coding elements is defined and added to the "base" or "backward-compatible" layer of the channel-based audio codec bitstream. This approach allows one or more bitstreams containing the enhancement layer to be processed by legacy codecs, while providing an enhanced listener experience for users with newer decoders. One example of an enhanced listener experience includes control over audio object rendering. An additional advantage of this approach is that audio objects may be added or modified anywhere in the distribution chain without the need to decode / mix / re-encode multi-channel audio encoded with the channel-based audio codec.
[0036] With respect to frames of reference, the spatial effect of an audio signal is important for providing an immersive experience to the listener. Sounds intended to emanate from a viewing screen or a certain area of a room should be emitted by speakers positioned in the same relative location. Thus, the primary audio metadata of a sound event in a model-based description is its position, but other parameters, such as size, direction, velocity, and acoustic dispersion, may also be described. To convey position, a model-based 3D audio spatial description requires a 3D coordinate system. The coordinate system used for transmission (e.g., Euclidean, spherical) is generally chosen for its convenience and compactness, but other coordinate systems may be used for the rendering process. In addition to the coordinate system, a frame of reference is required to represent the location of objects in space. Choosing an appropriate frame of reference is a key factor for a system to accurately reproduce position-based sounds in a variety of different environments. In allocentric frames of reference, sound source positions are defined relative to features in the rendering environment, such as room walls and corners, standard speaker locations, and screen locations. In egocentric frames of reference, location is expressed from the listener's perspective, e.g., "in front of me, slightly to the left." Scientific research into spatial perception (audio and otherwise) has led to the almost universal use of egocentric perspective. However, in cinema, allocentricity is generally appropriate for several reasons. For example, the precise location of audio objects is paramount when there are related objects on the screen. Using an allocentric criterion, for every listening position and for any screen size, sounds are localized to the same relative position on the screen, e.g., one-third left of the middle of the screen. Another reason is that mixers think and mix allocentrically, with panning tolls configured in allocentric frames (room walls), and mixers expect this sound to be rendered onscreen, this sound offscreen, or from the left wall, etc.
[0037] Despite the use of an allocentric frame of reference in cinema environments, there are cases where an egocentric frame of reference is useful and more appropriate. These include non-narrative sounds, i.e., sounds that are not "story-world," e.g., mood music, for which a uniform egocentric presentation is desirable. Other cases are near-field effects (e.g., a mosquito buzzing in the listener's left ear) that require egocentric representation. Currently, there is no way to render such sound fields without headphones or very near-field speakers. Also, infinitely distant sound sources (and the resulting plane waves) appear to come from a fixed egocentric position (e.g., 30 degrees to the left), and such sounds are easier to describe in egocentric rather than allocentric terms.
[0038] In some cases, the use of an allocentric frame of reference is possible as long as a nominal listening position is defined, while other instances require egocentric representations that cannot yet be rendered. While allocentric referencing is more useful and appropriate, the audio representation should be scalable, as many new features, including egocentric representations, are desirable in certain applications and listening environments. Embodiments of the adaptive audio system include a hybrid spatial description approach that uses egocentric referencing to render diffuse or complex multi-point sources (e.g., stadium crowds, ambiance, etc.) and allocentric model-based sound descriptions for optimal fidelity, including recommended channel configurations for increased spatial resolution and scalability.
[0039] System Components Referring to FIG. 1 , original sound content data 102 is first processed in a pre-processing block 104. The pre-processing block 104 of the system 100 includes an object channel filtering component. Often, audio objects contain individual sound sources, allowing for independent sound panning. In some cases, such as when creating an audio program using natural or "production" sound, it is necessary to extract individual sound objects from a recording containing multiple sound sources. Embodiments include a method for isolating an independent source signal from a more complex signal. Unwanted elements to be separated from the independent source signal include, but are not limited to, other independent sound sources and background noise. Additionally, reverb is removed to restore the "dry" sound source.
[0040] The pre-processor 104 also includes source separation and content type detection functionality. The system provides for automatic generation of metadata through analysis of input audio. Positional metadata is derived from multi-channel recordings by analyzing the relative levels of corresponding inputs between channel pairs. Detection of content types such as "speech" or "music" can be achieved, for example, through feature extraction and classification.
[0041] Authoring Tools The authoring tools block 106 includes features that improve the authoring of audio programs by optimizing the input and encoding of a sound engineer's creative intent and allowing the engineer to generate a final audio mix that is optimized for playback in realistically any playback environment. This is achieved through the use of audio objects and positional data that was associated with and encoded in the original audio content. To properly position sounds in the auditorium, the sound engineer needs to control how the sounds are ultimately rendered based on the actual constraints and characteristics of the playback environment. An adaptive audio system provides this control by allowing the sound engineer to alter how audio content is designed and mixed through the use of audio objects and positional data.
[0042] An audio object can be thought of as a group of sound elements perceived to emanate from a physical location in the auditorium. Such objects can be static or moving. In the adaptive audio system 100, audio objects are controlled by metadata, which details, among other things, the location of the sound at a given time. When monitored or played back in a theater, the object is not necessarily output to a physical channel, but is rendered using the speakers present there, according to the positional metadata. Tracks in a session are audio objects, and standard panning data is analogous to positional metadata. In this way, content placed on the screen is panned in a manner similar to channel-based content, while content placed in the surrounds can be rendered to individual speakers if desired. While the use of audio objects allows for the desired control of individual effects, other aspects of a movie soundtrack work well in a channel-based environment. For example, many ambient effects and reverberations actually benefit from being input to an array of speakers. These can be treated as objects with sufficient width to fill the array, while retaining some channel-based functionality.
[0043] In one embodiment, the adaptive audio system supports "beds" in addition to audio objects, where beds are effectively channel-based submixes or stems. These can be delivered individually or combined into a single bed for final rendering, depending on the content creator's intent. These beds can be generated in different channel-based configurations, such as 5.1, 7.1, etc., and are scalable to larger formats, such as 9.1, and arrangements that include overhead speakers.
[0044] 2 is a diagram illustrating the combination of channel- and object-based data to create an adaptive audio mix, according to one embodiment. As shown in process 200, channel-based data 202, e.g., 5.1 or 7.1 surround sound data provided in the form of pulse-code modulation (PCM) data, is combined with audio object data 204 to form an adaptive audio mix 208. The audio object data 204 is generated by combining elements of the original channel-based data with associated metadata that defines parameters related to the location of the audio objects.
[0045] As conceptually illustrated in FIG. 2, the authoring tool provides the ability to create audio programs that simultaneously contain a combination of speaker channel groups and object channels. For example, an audio program may include one or more speaker channels, optionally organized into groups (or tracks, such as stereo or 5.1 tracks), descriptive metadata for one or more speaker channels, one or more object channels, and descriptive metadata for one or more object channels. Within an audio program, each speaker channel group and each object channel is represented using one or more different sample rates. For example, digital cinema (D-Cinema) applications support sample rates of 48 kHz and 96 kHz, but may also support other sample rates. Furthermore, capture, storage, and editing of channels with different sample rates may also be supported.
[0046] Creating an audio program requires a sound design step, which involves combining sound elements as a sum of level-adjusted constituent sound elements to create a new desired sound effect. The adaptive audio system's authoring tool enables the creation of sound effects as a collection of sound objects with relative positions using a spatial-visual sound design graphical user interface. For example, a visual representation of a sound-producing object (e.g., a car) can be used as a template to assemble audio elements (exhaust, tire noise, engine noise) into object channels containing sounds and appropriate spatial locations (tailpipe, tire, hood). Individual object channels can be linked and manipulated as groups. The authoring tool 106 includes multiple user interface elements that allow sound engineers to input control information and view mix parameters to improve system functionality. The sound design and authoring process is also improved by allowing object channels and speaker channels to be linked and manipulated as groups. One example is combining object channels with individual dry sound sources with a set of speaker channels containing associated reverberation signals.
[0047] The audio authoring tool 106 supports the ability to combine multiple audio channels, commonly referred to as mixing. Multiple methods of mixing are supported, including traditional level-based mixing and loudness-based mixing. In level-based mixing, wideband scaling is applied to the audio channels and the scaled audio channels are summed. A wideband scale factor for each channel is selected to control the absolute level of the resulting mixed signal and the relative levels of the mixed channels in the mixed signal. In loudness-based mixing, one or more input signals are modified using frequency-dependent amplitude scaling. Frequency-dependent amplitudes are selected to provide the desired perceived absolute and relative loudness while preserving the perceived timbre of the input sounds.
[0048] The authoring tool allows for the creation of speaker channels and speaker channel groups, whereby metadata is associated with each speaker channel group. Each speaker channel group can be tagged according to content type, which is extensible via textual description. Content types include, but are not limited to, dialogue, music, and sound effects. Each speaker channel group is assigned unique instructions on how to upmix from one channel configuration to another, with upmixing defined as the generation of M audio channels out of N channels, where M>N. Upmix instructions include, but are not limited to: an enable / disable flag indicating if upmixing is allowed; an upmix matrix that controls the mapping between each input and output channel; and Default enabling and matrix settings are assigned based on the content type, for example, enabling upmixing only for music. Also, each speaker channel group is assigned unique instructions regarding how to downmix from one channel configuration to another, and downmixing is defined as the generation of Y audio channels out of X channels, where Y < X. The downmix instructions include, but are not limited to: a matrix that controls the mapping between each input and output channel; and The default matrix settings can be assigned based on the content type, for example, conversation, and downmix to the screen; Effects are downmixed from the screen. Each speaker channel is associated with a metadata flag that disables bus management during rendering.
[0049] Embodiments include a function that enables the generation of object channels and object channel groups. According to this invention, metadata is associated with each object channel group. Each object channel group can be tagged according to the content type. The content type can be extended via a text description. The content type includes, but is not limited to, conversation, music, and effects. Each object channel group is assigned metadata that describes how the object should be rendered.
[0050] Position information is provided indicating the desired apparent source location. The location can be described using an egocentric or allocentric frame of reference. Egocentric referencing is appropriate when the source location is referenced to the listener. For egocentric locations, spherical coordinates are useful for describing the location. Allocentric referencing is the typical frame of reference for cinema and other audio / visual presentations, where the source location is referenced relative to objects in the presentation environment, such as a visual display screen or room boundaries. Three-dimensional (3D) trajectory information is provided to enable position interpolation or for use in other rendering decisions, such as enabling "snap to mode." Size information is provided indicating the desired apparent perceived audio source size.
[0051] Spatial quantization is achieved by a "snap to closest speaker" control, which indicates the sound engineer's or mixer's intention to have an object rendered by only one speaker (at the expense of some spatial accuracy). The limits of allowed spatial distortion are indicated by the elevation and azimuth tolerance thresholds, beyond which the "snap" function does not occur. In addition to the distance thresholds, a crossfade rate parameter is indicated, which controls how quickly a moving object transitions from one speaker to another as the desired position moves between speakers. In one embodiment, some positional metadata uses dependent spatial metadata. For example, metadata can be automatically generated for "slave" objects, associating them with a "master" object to which they should follow. Time lags and relative speeds can be assigned to slave objects. A mechanism can be provided to define an acoustic centroid for a set or group of objects, allowing the objects to be rendered so that they are perceived as moving around other objects. In such cases, one or more objects rotate around a defined area, such as an object or dominant point, or a dry area of the room. Although ultimate location information is expressed as location relative to the room, as opposed to location relative to other objects, the acoustic centroid is used during the rendering stage to determine location information for each appropriate object-based sound.
[0052] When an object is rendered, it is assigned to one or more speakers by positional metadata and the location of the playback speakers. Additional metadata is associated with the object to restrict the speakers used. Constraints can be used to prohibit the use of indicated speakers or simply block indicated speakers (directing less energy to that speaker than would be the case if the constraint were not used). Constrained speaker sets include, but are not limited to, designated speakers or speaker zones (e.g., L, C, R, etc.), or speaker areas such as front wall, back wall, left wall, right wall, ceiling, floor, and room speakers. Similarly, in the process of defining a desired mix of multiple sound elements, it is possible to make one or more elements inaudible or "masked" due to the presence of other "masking" sound elements. For example, masked elements may be identified to the user via a graphical display when detected.
[0053] As described elsewhere, audio program descriptions can be adapted for rendering in a wide variety of speaker installations and channel configurations. When authoring an audio program, it is important to monitor the effect of rendering the program in expected playback configurations to ensure that the desired results are achieved. The present invention includes the ability to select a target playback configuration and monitor the results. The system can also automatically monitor the worst-case (i.e., maximum) signal level produced in each expected playback configuration and provide an indication if clipping or limiting occurs.
[0054] Figure 3 is a block diagram illustrating a workflow for generating, packaging, and rendering adaptive audio content according to one embodiment. The workflow 300 in Figure 3 is divided into three distinct task groups, labeled Generation / Authoring, Packaging, and Exhibition. In general, the hybrid bed and object model shown in Figure 2 allows most sound design, editing, premixing, and final mixing to occur similarly to how it is done today, without adding excessive overhead to the process. In one embodiment, adaptive audio functionality is provided in the form of software, firmware, or circuitry used in conjunction with sound production and processing equipment, which may be new hardware systems or updates to existing systems. For example, a plug-in application may be provided for a digital audio workstation, so existing panning methods in sound design and editing remain unchanged. In this way, both bed and object audio can reside within workstations in 5.1 or similar surround-enabled editing suites. Object audio and metadata are recorded in dubbing theaters during sessions in preparation for the premix and final mix stages.
[0055] As shown in FIG. 3 , the production or authoring task involves a user, e.g., a sound engineer in the following example, inputting mixing controls 302 into a mixing console or audio workstation 304. In one embodiment, metadata is collected on the mixing console surface, allowing channel strip faders, panning, and audio processing to coordinate with the bed or stem and audio objects. The metadata can be edited using either the console surface or the workstation user interface, and the sound is monitored using a rendering and mastering unit (RMU) 306. The bed and object audio data and associated metadata are recorded during the mastering session to generate a “print master.” This print master includes an adaptive audio mix 310 and any other rendered derivatives (such as a surround 7.1 or 5.1 theater mix) 308. Using existing authoring tools (e.g., a digital audio workstation such as ProTools), a sound engineer can label individual audio tracks in a mix session. An embodiment extends this concept by allowing users to label individual subsegments within a track to aid in locating or quickly identifying audio elements. The user interface to the mixing console that allows for the definition or creation of metadata can be implemented by graphical user interface elements, physical controls (eg, sliders or knobs), or any combination thereof.
[0056] During the packaging stage, the print master file is wrapped, hashed, and optionally encrypted using industry-standard MXF wrapping procedures to ensure the integrity of the audio content for delivery to a digital cinema packaging facility. This step can be performed by a digital cinema processor (DCP) 312 or any suitable audio processor depending on the final playback environment, such as standard surround sound in a theater 318, an adaptive audio-enabled theater 320, or other playback environment. As shown in Figure 3, processor 312 outputs the appropriate audio signals 314 and 316 depending on the exhibition environment.
[0057] In one embodiment, the adaptive audio print master includes an adaptive audio mix along with a standard DCI-compliant pulse code modulation (PCM) mix. The PCM mix can be rendered by the dubbing theater's rendering and mastering unit, or generated by a separate mix pass if desired. The PCM audio forms a standard main audio track file in the digital cinema processor 312, and the adaptive audio forms a sub-track file. Such track files conform to existing industry standards and are ignored by DCI-compliant servers that cannot use them.
[0058] In the example cinema playback environment, a DCP containing adaptive audio track files would be recognized as a valid package by the server, ingested by the server, and streamed to the adaptive audio cinema processor. Systems capable of using both linear PCM and adaptive audio files can switch between them as needed. For delivery to the exhibition stage, the adaptive audio packaging scheme allows a single type of package to be delivered to the cinema. The DCP package contains both PCM and adaptive audio files. It incorporates the use of security keys, such as Key Delivery Messages (KDMs), to enable secure delivery of movie content and other similar content.
[0059] As shown in Figure 3, an adaptive audio methodology is realized by allowing a sound engineer to express their intent regarding the rendering and playback of audio content through an audio workstation 304. By controlling input controls, the engineer can define where and how audio objects and sound elements are played depending on the listening environment. Metadata is generated at the audio workstation 304 in response to the engineer's mixing input 302, controlling spatial parameters (e.g., position, velocity, intensity, timbre, etc.) and providing rendering cues that define which speaker or speaker group in the listening environment will play each sound during the exhibition. The metadata is associated with each audio data at the workstation 304 or RMU 306 for packaging and transmission via the DCP 312.
[0060] The graphical user interface and software tools that provide engineer control of workstation 304 comprise at least a portion of authoring tool 106 of FIG.
[0061] Hybrid Audio Codec As shown in FIG. 1, processor 100 includes hybrid audio codec 108. This component comprises an audio encoding, distribution, and decoding system configured to generate a single bitstream containing both traditional channel-based audio elements and audio object-coded elements. The hybrid audio coding system is built around a channel-based coding system configured to generate a single (unified) bitstream that is simultaneously compatible with (i.e., decodable by) a first decoder configured to decode audio data encoded according to a first encoding protocol (channel-based) and one or more second decoders configured to decode audio data encoded according to one or more second encoding protocols (object-based). The bitstream may include both coded data (in the form of data bursts) that can be decoded by the first decoder (and ignored by any second decoders) and coded data (e.g., other data bursts) that can be decoded by one or more second decoders (and ignored by the first decoder). The decoded audio and related information (metadata) from the first and one or more second decoders can be combined such that both channel-based and object-based information are rendered simultaneously, creating a facsimile of the environment, channels, spatial information, and objects provided to the hybrid coding system (i.e., within the 3D space or listening environment).
[0062] The codec 108 generates a bitstream containing coded audio information and information about multiple sets of channel locations (speakers). In one embodiment, one set of channel locations is constant and used for a channel-based coding protocol, while the other set of channel locations is adaptive and used for an audio object-based coding protocol, such that the channel configuration of an audio object may change as a function of time (depending on where the object is located in the sound field). In this way, the hybrid audio coding system carries information about two sets of speaker locations for playback, one set being constant and the other being a subset of the other set. Devices that support legacy coded audio information decode and render audio information from the constant subset, while devices that can support a larger set decode and render additional coded audio information from the larger set that is assigned to different speakers in a time-dependent manner. Furthermore, the system does not rely on the simultaneous presence of a first and one or more second decoders within the system and / or device. Thus, legacy and / or existing devices / systems that include only decoders supporting the first protocol will produce a fully compatible sound field that is rendered via a conventional channel-based playback system, where unknown or unsupported portions of the hybrid bitstream protocol (i.e., audio information represented by the second encoding protocol) are ignored by system or device decoders supporting the first hybrid encoding protocol.
[0063] In another embodiment, the codec 108 is configured to operate in a mode in which the first encoding subsystem (supporting the first protocol) contains a combined representation of all sound field information (channels and objects) represented by both the first and one or more second encoders in the hybrid encoder, such that audio objects (typically carried by one or more second encoder protocols) are represented and rendered in decoders that support only the first protocol, thereby making the hybrid bitstream backward compatible with decoders that support only the protocols of the first encoder subsystem.
[0064] In yet another embodiment, codec 108 includes two or more encoding subsystems, each configured to encode audio data according to a different protocol, and combines the outputs of the multiple subsystems to generate a hybrid format (combined) bitstream.
[0065] One benefit of this embodiment is that hybrid-coded audio bitstreams can be carried across a wide range of content distribution systems, each of which traditionally supports only data encoded with a first encoding protocol, thereby eliminating the need for system and / or transport level protocol modifications / changes to support a hybrid coding system.
[0066] Audio coding systems typically utilize standardized bitstream elements to allow the transmission of additional (optional) data within the bitstream itself. This additional (optional) data is skipped (i.e., ignored) during decoding of the coded audio contained in the bitstream, but is used for purposes other than decoding. Different audio coding standards use unique nomenclature to refer to these additional data fields. Bitstream elements of this general type include, but are not limited to, auxiliary data, skip fields, data stream elements, fill elements, ancillary data, and substream elements. Unless otherwise specified, the use of the term "auxiliary data" in this document does not imply a particular type or format of additional data, but should be interpreted as a general term to include some or all of the embodiments related to the present invention.
[0067] A data channel enabled via an "ancillary" bitstream of a first encoding protocol in a combined hybrid coding system bitstream can carry one or more (independent or dependent) audio bitstreams (encoded with one or more second encoding protocols). The one or more second audio bitstreams can be divided into N sample blocks and multiplexed into the "ancillary data" field of the first bitstream. The first bitstream can be decoded by an appropriate (compliant) decoder. The ancillary data from the first bitstream can also be extracted and recombined with the one or more second audio bitstreams, decoded by a processor supporting the syntax of the one or more second bitstreams, and then combined and rendered together or independently. Furthermore, the roles of the first and second bitstreams can be reversed, such that blocks of data from the first bitstream are multiplexed onto the ancillary data of the second bitstream.
[0068] Bitstream elements associated with the second encoding protocol also carry information (metadata) characteristics of the underlying audio, including, but not limited to, desired sound source position, velocity, and size. This metadata is used in the decoding and rendering process to reproduce the proper (i.e., original) positions of associated audio objects carried in the adaptive bitstream. The above metadata may also be applied to audio objects contained in one or more secondary bitstreams in the hybrid stream, which may also be carried in bitstream elements associated with the first encoding protocol.
[0069] Bitstream elements associated with either or both of the first and second encoding protocols of the hybrid coding system carry / convey contextual metadata identifying spatial parameters (i.e., the essence of the signal properties themselves) and further information describing the underlying audio essence type in the form of audio classes carried in the hybrid-coded audio bitstream. Such metadata can indicate, for example, the presence of spoken dialogue, music, dialogue over music, applause, singing, etc., and can be used to adaptively modify interconnected pre- or post-processing modules upstream or downstream of the hybrid coding system.
[0070] In one embodiment, the codec 108 is configured to operate from a shared or common bit pool in which bits available for coding are "shared" among all or some of the encoding subsystems supporting one or more protocols. Such a codec dictates the bits available (from the common "shared" bit pool) among the encoding subsystems to optimize the overall audio quality of the combined bitstream. For example, during a first time interval, the codec may allocate more available bits to a first encoding subsystem and fewer available bits to the remaining subsystems, while during a second time interval, the codec may allocate fewer available bits to the first encoding subsystem and more available bits to the remaining subsystems. The decision on how to allocate bits among the encoding subsystems depends, for example, on a statistical analysis of the shared bit pool and / or an analysis of the audio content encoded by each subsystem. The codec allocates bits from the shared pool such that the combined bitstream constructed by multiplexing the outputs of the encoding subsystems maintains a constant frame length / bitrate over a specified time interval. It is also possible, in some cases, for the frame length / bitrate of the combined bitstream to vary over a specified time interval.
[0071] In another embodiment, the codec 108 generates an integrated bitstream that includes data encoded according to a first encoding protocol configured and transmitted as an independent substream of the encoded data stream (to be decoded by decoders supporting the first encoding protocol) and data encoded according to a second protocol sent as an independent or dependent substream of the encoded data stream (to be ignored by decoders supporting the first protocol). More generally, in one class of embodiments, the codec generates an integrated bitstream that includes two or more independent or dependent substreams, each substream containing data encoded according to a different or the same encoding protocol.
[0072] In yet another embodiment, the codec 108 generates a unified bitstream that includes data encoded according to a first encoding protocol configured and transmitted with a unique bitstream identifier (decoded by decoders supporting the first encoding protocol associated with the unique bitstream identifier) and data encoded according to a second protocol configured and transmitted with the unique bitstream identifier (ignored by decoders supporting the first protocol). More generally, in one class of embodiments, the codec generates a unified bitstream that includes two or more substreams (each substream containing data encoded according to a different or the same encoding protocol, each carrying a unique bitstream identifier). The above-described methods and systems for generating a unified bitstream provide the ability to unambiguously signal (to a decoder) which interleaving and / or protocol was used in the hybrid bitstream (e.g., signaling whether AUX data, SKIP, DSE, or a substream approach was used).
[0073] The hybrid coding system is configured to support deinterleaving / demultiplexing of bitstreams supporting one or more second protocols and reinterleaving / remultiplexing into a first bitstream (supporting a first protocol) at processing points found throughout the media distribution system. The hybrid codec is also configured to encode audio input streams of different sample rates into bitstreams, thereby providing a means for efficiently coding and delivering audio sources containing signals that inherently have different bandwidths. For example, dialogue tracks typically have a lower bandwidth than music or effects tracks. rendering In one embodiment, the adaptive audio system allows multiple tracks (e.g., up to 128) to be packaged, typically as a combination of beds and objects. The basic format of audio data for the adaptive audio system includes multiple independent monophonic audio streams. Each stream has associated metadata that specifies whether the stream is a channel-based stream or an object-based stream. Channel-based streams have rendering information encoded in the channel name or label. Object-based streams have location information encoded in a mathematical formula encoded in other associated metadata. The original independent audio streams are packaged as a single, ordered, serial bitstream containing all of the audio data. This adaptive data organization allows sounds to be rendered in an allocentric frame of reference. The final rendering location of the sound is based on the playback environment to correspond to the mixer's intent. In this way, sounds can be specified to emanate from the playback room's frame of reference (e.g., the center of the left wall) rather than a labeled speaker or speaker group (e.g., left surround). The object position metadata contains the appropriate allocentric frame of reference information needed to correctly reproduce the sound using the available speaker positions in the room configured to play the adaptive audio content.
[0074] The renderer takes the bitstream encoding the audio tracks and processes the content according to signal type. The bed is sent to an array. The array potentially requires different delay and equalization processing than individual objects. The process supports rendering these beds and objects to multiple (up to 64) speaker outputs. Figure 4 is a block diagram illustrating the rendering stages of an adaptive audio system according to one embodiment. As shown in system 400 of Figure 4, multiple input signals, such as up to 128 audio tracks with adaptive audio signals 402, are provided by components in the production, authoring, and packaging stages of system 300, such as RMU 306 and processor 312. These signals include channel-based beds and objects used by renderer 404. The channel-based audio (bed) and objects are input to level manager 406, which controls the output level or amplitude of different audio components. Certain audio components are processed by array correction component 408. The adaptive audio signal is passed through B-chain processing component 410. The B-chain processing component 410 generates multiple (e.g., up to 64) speaker feed output signals. Generally, the B-chain feed refers to the power amplifier, crossover, and speaker processed signals for the A-chain content that makes up the film stock soundtrack.
[0075] In one embodiment, the renderer 404 executes rendering algorithms that intelligently utilize the theater's surround speakers to their full potential. By improving the surround speaker's power handling and frequency response and maintaining the same monitoring reference level for each theater output channel or speaker, objects panned between the screen and surround speakers can maintain their own sound pressure level and better match timbre without increasing the theater's overall sound pressure level. While a properly defined surround speaker array typically has enough headroom to reproduce the maximum dynamic range available in a surround 7.1 or 5.1 soundtrack (i.e., 20 dB above reference level), a single surround speaker is unlikely to have the same headroom as a large, multi-way screen speaker. As a result, objects placed in the surround field may require more sound pressure than can be achieved using a single surround speaker. In these cases, the renderer distributes the sound across an appropriate number of speakers to achieve the required sound pressure level. The adaptive audio system improves the quality and power handling of the surround speakers, improving the fidelity of the rendering. The adaptive audio system allows each surround speaker to achieve improved power handling, while also optionally allowing for the use of smaller speaker cabinets, supports surround speaker bass management through the use of optional rear subwoofers, and adds side surround speakers closer to the screen than current practice, ensuring a smooth transition of objects from the screen to the surround.
[0076] By using metadata that specifies audio object location information in the rendering process, the system 400 provides a comprehensive and flexible way for content creators to move beyond existing systems. As previously mentioned, current systems generate and deliver audio that is fixed to a certain speaker location with only limited knowledge of the type of content carried in the audio essence (the portion of audio being played). The adaptive audio system 100 provides a new hybrid approach that includes the options of both speaker-location specific audio (left channel, right channel, etc.) and object-oriented audio elements with generalized spatial information, including but not limited to size and velocity. This hybrid approach provides a balanced approach between fidelity (provided by fixed speaker locations) and flexibility in rendering (of generalized audio objects). The system also provides additional useful information about the audio content that can be paired with the audio essence by the content creator at the time of content creation. This information provides powerful and detailed information about the audio's attributes that can be used in very powerful ways during rendering. Such attributes include, but are not limited to, content type (dialogue, music, effects, Foley, background / ambience, etc.), spatial attributes (3D position, 3D size, velocity), rendering information (snap to speaker location, channel weights, gain, bass management information, etc.).
[0077] The adaptive audio system described herein provides powerful information that can be used for rendering by a widely variable number of endpoints. Often, the optimal rendering method applied depends heavily on the endpoint device. For example, home theater systems and sound bars may have two, three, five, seven, or nine separate speakers. Many other types of systems, such as televisions, computers, and music docks, have only two speakers, and almost all commonly used devices (PCs, laptops, tablets, mobile phones, music players, etc.) have binaural headphone outputs. However, with traditional audio sold today (mono, stereo, 5.1, 7.1 channels), endpoint devices often must make simplistic decisions and compromises to render and play the audio being delivered in a specific channel / speaker format. Furthermore, little or no information is conveyed about the actual content being delivered (dialogue, music, ambience, etc.), and little or no information about the content creator's intent for the audio playback. However, the adaptive audio system 100 provides access to this information and potentially audio objects that can be used to create compelling, next-generation user experiences.
[0078] System 100, with its unique and powerful metadata and adaptive audio transmission format, allows content creators to incorporate the spatial intent of their mix into the bitstream using metadata such as position, size, and velocity. This allows for great flexibility in spatial reproduction of audio. From a spatial rendering perspective, adaptive audio allows the mix to adapt to the exact position of speakers in the room to avoid spatial distortions that occur when the geometry of the playback system is not the same as the geometry of the authoring system. Current audio playback systems, which only send audio from one speaker channel, do not know the intent of the content creator. System 100 uses metadata carried in the production and distribution pipeline. An adaptive audio-aware playback system uses this metadata information to play content that matches the content creator's original intent. Similarly, the mix can adapt to the exact hardware configuration of the playback system. Currently, rendering devices, such as televisions, home theaters, sound bars, and portable music player docks, have many different speaker configurations and types. Today, when these systems send specific channel audio information (i.e., left and right channel audio or multi-channel audio), they must process the audio to appropriately match the capabilities of the rendering device. One example is when standard stereo audio is sent to a soundbar with three or more speakers. Current audio playback, where only one speaker channel of audio is sent, does not reveal the intent of the content creator. By using metadata carried by the production and delivery pipeline, adaptive audio-enabled playback systems can use this information to play content in a way that matches the original intent of the content creator. For example, some soundbars have side-firing speakers to create an enveloping feeling. With adaptive audio, spatial information and content type (such as ambient effects) can be used by the soundbar to send only appropriate audio to these side-firing speakers.
[0079] Adaptive audio systems allow for unlimited interpolation in all front / back, left / right, up / down, and near / far dimensions. Current audio playback systems have no knowledge of how to handle audio where it is desired to position the audio so that the listener feels as if it is located between two speakers. Currently, audio that is assigned to only specific speakers introduces spatial quantization factors. With adaptive audio, the spatial positioning of the audio is known exactly and the audio playback system can play it accordingly.
[0080] For headphone rendering, the creator's intent is realized by matching head-related transfer functions (HRTFs) to spatial locations. When audio is played through headphones, spatial virtualization can be achieved by applying head-related transfer functions. Head-related transfer functions process the audio and add perceptual cues that create the feeling that the audio is emanating from 3D space rather than through headphones. The accuracy of spatial reproduction depends on the selection of appropriate HRTFs, which can vary based on multiple factors, including spatial location. Using spatial information provided by an adaptive audio system, one or a continuously variable number of HRTFs can be selected, significantly improving the playback experience.
[0081] The spatial information conveyed by adaptive audio systems can be used by content creators to create compelling entertainment experiences (film, television, music, etc.), as well as indicate where the listener is located relative to physical objects such as buildings or geographic points of interest, allowing users to interact with virtualized audio experiences relative to the real world, i.e., augmented reality.
[0082] Embodiments enable spatial upmixing by reading metadata only when object audio data is unavailable, providing enhanced upmixing. Knowing the location and type of all objects allows the upmixer to differentiate elements in channel-based tracks. Existing upmixing algorithms must infer information such as audio content type and the location of different elements in the audio stream to produce a high-quality upmix with minimal or no audible artifacts. Often, the inferred information is inaccurate or inappropriate. Adaptive audio allows the upmixing algorithm to use additional information obtained from metadata about audio content type, spatial location, velocity, audio object size, and so on to produce high-quality playback results. The system also spatially matches audio to video by accurately positioning audio objects on the screen to visual elements. In this case, a compelling audio / video playback experience is possible when the spatial location of the reproduced audio elements matches the image elements on the screen, especially for large screen sizes. One example is spatially matching dialogue in a film or television program with the people or characters speaking on the screen. With conventional speaker channel-based audio, there is no easy way to determine where dialogue should be spatially positioned to match the location of people or characters on screen. Using the audio information available with adaptive audio, such audio / visual alignment can be achieved. Visual position and audio spatial alignment can also be used for non-character / speaking objects such as cars, trucks, animations, etc.
[0083] The spatial masking process is facilitated by the system 100 because knowledge of the spatial intent of a mix through adaptive audio metadata means that the mix can be adapted to any speaker configuration. However, one runs the risk of downmixing objects in the same or nearly the same location due to the constraints of the playback system. For example, without surround channels, an object intended to be panned to the left rear would be downmixed to the left front. If a louder element also appears in the left front, the downmixed object would be masked and disappear from the mix. Using adaptive audio metadata, spatial masking is planned by the renderer, adjusting the spatial and / or loudness downmix parameters of each object so that all audio elements of the mix remain as perceptible as in the original mix. Because the renderer understands the spatial relationship between the mix and the playback system, it has the ability to "snap" objects to the nearest speaker instead of creating a phantom image between two or more speakers. This slightly distorts the spatial representation of the mix, but avoids unintended phantom images. For example, if the angular position of the left speaker in the mixing stage does not correspond to the angular position of the left speaker in the playback system, a snap to nearest speaker function can be used to avoid having the playback system reproduce a constant phantom image of the left channel of the mixing stage.
[0084] With regard to content processing, the adaptive audio system 100 allows content creators to create individual audio objects, add information about the content, and convey it to the playback system. This allows for greater flexibility in processing the audio before playback. From a content processing and rendering perspective, the adaptive audio system allows processing to be tailored to the object's context. For example, dialogue enhancement can be applied only to dialogue objects. Dialogue enhancement refers to a method of processing audio containing dialogue so that the audibility and / or intelligibility of the dialogue is enhanced and / or improved. In many cases, audio processing applied to dialogue is inappropriate for non-speech audio content (i.e., music, ambient effects, etc.) and can result in objectionable audible artifacts. In adaptive audio, audio objects can be appropriately labeled to contain only dialogue in a piece of content, allowing the rendering solution to selectively apply dialogue enhancement only to the dialogue content. Also, when an audio object is dialogue only (and not a mixture of dialogue and other content, as is often the case), the dialogue enhancement processing can process the dialogue exclusively (thereby limiting the processing to other content). Similarly, bass management (filtering, attenuation, gain) can be targeted to objects based on their type. Bass management refers to selectively isolating and processing only the bass (or below) frequencies in a piece of content. In current audio systems and delivery mechanisms, this is a "blind" process applied to all audio. With adaptive audio, audio objects for which bass management is appropriate can be identified via metadata, and rendering processes can be applied appropriately.
[0085] The adaptive audio system 100 also provides object-based dynamic range compensation and selective upmixing. While a traditional audio track has the same duration as the content itself, an audio object may occur for only a limited time within the content. Metadata associated with an object includes information about its average and peak signal amplitudes, and its onset or attack time (especially in the case of transitional material). This information allows a compressor to adapt its compression and time constants (attack, release, etc.) to better suit the content. For selective upmixing, content creators may choose to indicate in the adaptive audio bitstream whether an object should be upmixed or not. This information allows the adaptive audio renderer and upmixer to determine which audio elements can be safely upmixed while respecting the creator's intent.
[0086] Additionally, embodiments allow the adaptive audio system to select a preferred rendering algorithm from multiple available rendering algorithms and / or surround sound formats. Examples of available rendering algorithms include binaural, stereo dipole, ambisonics, wave field synthesis (WFS), multi-channel panning, and raw stems with positional metadata. Others include dual balance and vector-based amplitude panning.
[0087] Binaural delivery formats use a two-channel representation of the sound field in terms of signals at the left and right ears. Binaural information can be generated by in-ear recordings or synthesized using HRTF models. Reproduction of the binaural representation is typically done through headphones or with crosstalk cancellation. Reproduction through any speaker setup requires signal analysis to determine the relevant sound field and / or signal sources.
[0088] Stereo dipole rendering is a transaural crosstalk cancellation process that allows binaural signals to be reproduced on stereo speakers (e.g., at ±10° off center).
[0089] Ambisonics is a four-channel encoding (delivery format and rendering method) called B-format. The first channel, W, is an omnidirectional pressure signal; the second channel, X, is a directional pressure gradient containing front and back information; the third channel, Y, contains left and right, and Z contains up and down. These channels define a primary sample of the complete sound field at a point. Ambisonics uses all available speakers to recreate the sampled (or synthesized) sound field in the speaker array, so that when some speakers are pushing, others are pulling.
[0090] Wave Field Synthesis is a method of rendering sound reproduction based on the precise construction of a desired wave field by secondary sources. Based on Huygens' principle, WFS is implemented as an array of loudspeakers (tens or hundreds) surrounding the listening space and operating in a disciplined, phase-controlled manner to recreate each individual sound wave.
[0091] Multi-channel panning is a delivery format and / or rendering method known as channel-based audio, in which sound is presented as an equal number of individual sources via multiple speakers positioned at defined angles from the listener. Content creators / mixers can generate virtual images by panning signals between adjacent channels to provide directional cues; early reflections, reverb, etc. can be mixed across many channels to provide directional and environmental cues.
[0092] Raw stems with position metadata is a distribution format, also known as object-based audio. In this format, distinct "close-mic" sound sources are represented by position and environmental metadata. Based on the metadata, playback device, and listening environment, virtual sources are rendered.
[0093] The adaptive audio format is a hybrid of the multi-channel panning format and the Roshteem format. The rendering method in this embodiment is multi-channel panning. For audio channels, rendering (panning) is performed at authoring time, while for objects, rendering (panning) is performed at playback time.
[0094] Metadata and adaptive audio transmission formats As noted above, metadata is generated during the production phase to encode audio object location information, accompany audio programs, and assist in their rendering, specifically describing the audio program to enable rendering on a wide range of playback devices and environments. Metadata is generated for a given program and for the audio recorders and mixers that create, acquire, edit, and manipulate that audio during post-production. An important feature of adaptive audio formats is the ability to control how audio is translated into playback systems and environments that differ from the mix environment. In particular, some cinemas may have less capability than other mix environments.
[0095] The adaptive audio renderer is designed to make the most of available devices to recreate the mixer's intent. Additionally, the adaptive audio authoring tool allows the mixer to preview and adjust how the mix will be rendered in various playback configurations. All metadata values can be conditional on the playback environment and speaker configuration. For example, different mix levels of certain audio elements can be defined based on the playback configuration or mode. In one embodiment, the list of conditional playback modes is extensible and includes: (1) channel-based only playback: 5.1, 7.1, 7.1 (height), 9.1; (2) individual speaker playback: 3D, 2D (no height).
[0096] In one embodiment, metadata controls or governs different aspects of the adaptive audio content and is organized according to different types, including program metadata, audio metadata, and rendering metadata (for channels and objects). Each type of metadata includes one or more metadata items that provide values for properties referenced by an identifier (ID). Figure 5 is a table listing metadata types and associated metadata elements for an adaptive audio system, according to one embodiment.
[0097] The first type of metadata is program metadata, as shown in Table 500 of Figure 5. It includes multiple metadata elements that specify frame rate, track number, extensible channel description, and mix stage description. The frame rate metadata element specifies the frame rate of the audio content in frames per second (fps). The raw audio format does not include framing of the audio or metadata because the audio is supplied as full tracks (the length of a reel or entire feature) rather than audio segments (the length of an object). The raw format must carry all the information necessary for an adaptive audio encoder to be able to frame the audio and metadata, including the actual frame rate. Table 1 shows the IDs, example values, and descriptions of the frame rate metadata elements.
[0098] [Table 1] The Track Count metadata element indicates the number of audio tracks in a frame. An example adaptive audio decoder / processor can support up to 128 simultaneous audio tracks, while the adaptive audio format supports any number of audio tracks. Table 2 shows the ID, example values, and description of the Track Count metadata element.
[0099] [Table 2] Channel-based audio can be assigned to non-standard channels, and the extensible channel description metadata element allows the mix to use new channel positions. For each extended channel, the following metadata is provided, as shown in Table 3: [Table 3] The mix stage description metadata element specifies the frequency at which a speaker reproduces half the power of the passband. Table 4 shows the mix stage description metadata element IDs, example values, and descriptions, where LF = Low Frequency; HF = High Frequency; 3dB point = edge of speaker passband.
[0100] [Table 4] The second type of metadata is audio metadata, as shown in Figure 5. Each channel-based or object-based audio element consists of an audio essence and metadata. An audio essence is a monophonic audio stream carried by one of many audio tracks. Associated metadata describes how the audio essence is stored (audio metadata, e.g., sample rate) or how it should be rendered (rendering metadata, e.g., desired audio source location). Typically, audio tracks are continuous across the length of the audio program. The program editor or mixer is responsible for assigning audio elements to tracks. Track usage is expected to be coarse; that is, the median simultaneous track usage will be only 16 to 32. In a typical implementation, audio is transmitted efficiently using a lossless encoder. However, other implementations are possible, such as transmitting uncoded or lossy-coded audio data. In a typical implementation, the format has up to 128 audio tracks, where each track has a single sample rate and a single coding system. Each track lasts for the length of the feature (there is no explicit reel support). The mapping of objects to tracks (time multiplexing) is the responsibility of the content creator (mixer).
[0101] As shown in Figure 3, audio metadata includes elements of sample rate, bit depth, and coding system. Table 5 shows the ID, example values, and description of the sample rate metadata element.
[0102] [Table 5] Table 6 shows the bit depth metadata element IDs, example values, and descriptions (for PCM and lossless compression).
[0103] [Table 6] Table 7 shows the coding system metadata element IDs, example values, and descriptions.
[0104] [Table 7] The third type of metadata is rendering metadata, as shown in Figure 5. Rendering metadata specifies values that help the renderer match the original mixer's intent as closely as possible regardless of the playback environment. The set of metadata elements is different for channel-based and object-based audio. The first rendering metadata field selects between two types of audio: channel-based or object-based, as shown in Table 8.
[0105] [Table 8] Rendering metadata for channel-based audio includes position metadata elements that specify the audio source location as one or more speaker positions. Table 9 shows the IDs and values of the position metadata elements for the channel-based case.
[0106] [Table 9] The channel-based audio rendering metadata also includes rendering control elements that specify characteristics related to the playback of the channel-based audio, as shown in Table 10.
[0107] [Table 10] For object-based audio, the metadata contains similar elements as for channel-based audio. Table 11 gives the IDs and values of the object position metadata elements. Object position can be described in one of three ways: 3D coordinates, planar and 2D coordinates, or linear and 1D coordinates. The rendering method can be adapted based on the position type.
[0108] [Table 11] The IDs and values of the object rendering control metadata elements are shown in Table 12. These values provide additional means to control and optimize the rendering of object-based audio.
[0109] [Table 12-1] [Table 12-2] In one embodiment, the metadata described above and illustrated in Figure 5 is generated and stored as one or more files associated with or indexed to the corresponding audio content so that the audio stream can be processed by the adaptive audio system by interpreting the metadata generated by the mixer. Note that the above metadata is an example of IDs, values, and definitions, and other or additional metadata elements may be included for use by the adaptive audio system.
[0110] In one embodiment, two (or more) sets of metadata elements are associated with each channel and object-based audio stream. A first set of metadata is applied to the multiple audio streams in a first condition of the playback environment, and a second set of metadata is applied to the multiple audio streams in a second condition of the playback environment. Based on the condition of the playback environment, the second or subsequent set of metadata elements replaces the first set of metadata elements for a given audio stream. Conditions include room size, shape, material composition in the room, presence and density of occupants in the room, ambient noise characteristics, ambient light characteristics, and other factors that affect the sound or even mood of the playback environment.
[0111] Post-production and mastering The rendering stage 110 of the adaptive audio processing system 100 includes the post-production steps that lead to the generation of the final mix. In cinema applications, the three main categories of sounds used in a movie mix are dialogue, music, and effects. Effects consist of sounds that are not dialogue or music (e.g., ambient noises, background / scene noises). Sound effects may be recorded or synthesized by a sound designer, or may be sources obtained from an effects library. The subgroup of effects that includes specific noise sources (e.g., footsteps, doors, etc.) is known as Foley and is performed by Foley actors. Different types of sounds are marked and panned appropriately by the recording engineer.
[0112] Figure 6 illustrates an example post-production workflow for an adaptive audio system, according to one embodiment. As shown in diagram 600, all individual sound components—music, dialogue, Foley, and effects—are assembled in a dubbing theater during a final mix 606. A re-recording mixer 604 uses a premix (also known as a "mix minus") along with the individual sound objects and positional data to generate stems, for example, for dialogue, music, effects, Foley, and background sounds. In addition to forming the final mix 606, the music and all effects stems can be used as the basis for generating dubbed language versions of the movie. Each stem consists of a channel-based bed and multiple audio objects with metadata. The stems are then composited to form the final mix. Using object panning information from both the audio workstation and the mixing console, a rendering and mastering unit 608 renders the audio to the speaker locations in the dubbing theater. This rendering allows the mixer to hear how the channel-based bed and audio objects are composited and also provides the ability to render to different configurations. The mixer can control how the content is rendered into the surround channels using conditional metadata that defaults to the associated profile. In this way, the mixer retains complete control over how the movie will play in all scalable environments. A monitoring step is included after either or both of the re-recording step 604 and the final mix step 606, allowing the mixer to listen to and evaluate the intermediate content produced at each of these steps.
[0113] During the mastering session, the stems, objects, and metadata are assembled into an adaptive audio package 614. The adaptive audio package 614 is generated by the print master 610. This package also includes a backward-compatible (legacy 5.1 or 7.1) surround sound theater mix 612. The rendering / mastering unit (RMU) 608 can render this output as needed, eliminating the need for additional workflow steps in the creation of existing channel-based deliverables. In one embodiment, audio files are packaged using standard material exchange format (MXF) wrapping. The adaptive audio mix master file can also be used to generate other deliverables, such as consumer multichannel mixes and stereo mixes. Intelligent profiles and conditional metadata enable controlled rendering, which can significantly reduce the time required to create such mixes.
[0114] In one embodiment, a packaging system is used to generate a digital cinema package for a deliverable that includes an adaptive audio mix. The audio track files are locked together to help prevent synchronization errors with the adaptive audio track files. In some applications, additional track files are required during the packaging phase, such as adding a hearing impaired (HI) or visually impaired narration (VI-N) track to the main audio track file.
[0115] In one embodiment, the speaker array of the playback environment may include any number of surround sound speakers arranged and designed according to established surround sound standards. Any number of additional speakers for accurate rendering of object-based audio content may be arranged based on the requirements of the playback environment. These additional speakers are set up by a sound engineer, and this setup is provided to the system in the form of a setup file that the system uses to render the object-based components of the adaptive audio to one or more speakers across the speaker array. The setup file includes at least speaker designations, a mapping of channels to individual speakers, information about speaker groups, and a list of runtime mappings based on the relative position of the speakers in the playback environment. The runtime mapping is utilized by the system's snap-to feature, which renders point-source object-based audio content to the speaker closest to the perceived location of the sound intended by the sound engineer.
[0116] Figure 7 illustrates an example workflow for a digital cinema packaging process using adaptive audio files, according to one embodiment. As shown in diagram 700, audio files, including both adaptive audio files and 5.1 or 7.1 surround sound audio files, are input to wrapping / encryption block 704. In one embodiment, during creation of the digital cinema package in block 706, the PCM MXF files (with appropriate additional tracks added) are encrypted using SMPTE specifications, per existing practices. The adaptive audio MXFs are packaged as ancillary track files and optionally encrypted using a symmetric content key per SMPTE specifications. This single DCP 708 is sent to a Digital Cinema Initiative (DCI)-compliant server. Generally, unsuitable installations simply ignore the additional track files containing the adaptive audio soundtrack and use the existing main audio track file for standard playback. Installations with suitable adaptive audio processors can receive and play the adaptive audio soundtrack, reverting to standard audio tracks as needed, if applicable. The wrapping / encryption component 704 also provides direct input to the distribution KDM block 710 to generate the appropriate security keys for use by the digital cinema server. Other movie elements or files, such as subtitles 714 and images 716, are also wrapped and encrypted along with the audio file 702. In this case, in the case of image files 716, processing steps such as compression 712 are included.
[0117] With regard to content management, the adaptive audio system 100 allows content creators to create individual audio objects, add information about the content, and deliver it to the playback system. This allows for great flexibility in audio content management. From a content management perspective, adaptive audio methods enable several different functions. These include changing the language of content by simply replacing dialogue objects for space savings, download efficiency, geographic playback adaptation, etc. Films, television, and other entertainment programs are commonly distributed internationally. This often requires changing the language in the content depending on where it is being played (e.g., French for a film shown in France, German for a television program shown in Germany, etc.). For this reason, completely separate audio soundtracks are now being created, packaged, and distributed. With the inherent concept of adaptive audio and audio objects, the dialogue of content can be a separate audio object. This allows the language of content to be easily changed without having to update or change other elements of the audio soundtrack, such as music, effects, etc. This applies not only to foreign languages, but also to words that are inappropriate for certain audiences (e.g., children's television shows, movies for airlines, etc.), targeted advertising, etc.
[0118] Installation and equipment review The adaptive audio file format and associated processors allow for changes to how theater equipment is installed, calibrated, and maintained. With the introduction of more potential speaker outputs, in one embodiment, the adaptive audio system uses an optimized 1 / 12 octave band equalization engine. It can process up to 64 outputs to more accurately balance theater sound. The system also allows for scheduled monitoring of individual speaker outputs from the cinema processor output through to the sound played back to the auditorium. Local or network alerts can be generated so appropriate action can be taken. A flexible rendering system can automatically remove a failed speaker or amplifier from the playback chain and render around it, allowing the show to continue.
[0119] The cinema processor can connect to the digital cinema server over the existing 8x AES main audio connection and an Ethernet connection to stream adaptive audio data. Surround 7.1 or 5.1 playback uses the existing PCM connection. The adaptive audio data is streamed over Ethernet to the cinema processor for decoding and rendering, and the connection between the server and cinema processor allows for audio identification and synchronization. If any problems occur with the adaptive audio track playback, the sound reverts to Dolby Surround 7.1 or 5.1 PCM audio.
[0120] Although embodiments have been described with respect to 5.1 and 7.1 surround sound systems, it should be noted that many other current and future surround configurations can be used with embodiments, including 9.1, 11.1, and 13.1 and beyond.
[0121] The adaptive audio system is designed to allow both content creators and exhibitors to determine how to render sound content across different playback speaker configurations. The ideal number of speaker output channels to be used varies depending on the size of the room. Therefore, the recommended speaker placement depends on many factors, including the size, composition, seating configuration, environment, and average audience size. For illustrative purposes only, representative speaker configurations and layouts and examples are described herein and are not intended to limit the scope of the claimed embodiments.
[0122] The recommended speaker layout for adaptive audio systems is compatible with existing cinema systems, which is essential to avoid degrading the playback of existing 5.1 and 7.1 channel-based formats. To preserve the intent of the adaptive audio sound engineer and the mixer of 7.1 and 5.1 content, the positions of existing screen channels should not be significantly altered in an effort to accommodate new speaker locations. In contrast to using all 64 available output channels, adaptive audio formats can be accurately rendered to speaker configurations such as 7.1 in cinemas, allowing the format (and associated benefits) to be used in existing theaters without the need for amplifier or speaker changes.
[0123] The effectiveness of different speaker locations varies depending on the theater design, and there is currently no industry-specified ideal number or placement of channels. Adaptive audio is intended to be truly adaptable and capable of accurate playback in a variety of auditoriums, whether with a limited number of playback channels or a flexible configuration of many channels.
[0124] Figure 8 shows a top view 800 of a typical auditorium with a suggested speaker location layout for use with an adaptive audio system. Figure 9 shows a front view 900 of the same auditorium with a suggested speaker location layout on the screen. The reference position referred to below corresponds to a position on the centerline of the screen, two-thirds of the way from the screen to the rear wall. Standard screen speakers 801 are shown in their normal position relative to the screen. Studies of elevation perception at the screen plane have shown that additional speakers 804 behind the screen, such as left-center (Lc) and right-center (Rc) screen speakers (located in the left-extra and right-extra channel locations in 70mm film formats), are beneficial for smooth panning across the screen. Such optional speakers are recommended, especially for auditoriums with screens larger than 12 m (40 ft). All screen speakers should be angled toward the reference position. The recommended placement of the subwoofer 810 behind the screen remains the same, including maintaining symmetrical cabinet placement relative to the center of the room to prevent the excitation of standing waves. An additional subwoofer 816 may be placed behind the theater.
[0125] Surround speakers 802 should be individually wired to the amplifier rack and, if possible, individually amplified, using dedicated channels of power amplification matched to the speaker's power handling according to the manufacturer's specifications. Ideally, surround speakers should be specified to handle the SPL of each individual speaker and, if possible, have a wider frequency response. As a rule of thumb for an average-sized theater, surround speaker spacing should be 2 to 3 meters (6'6" to 9'9"), with left and right surround speakers positioned symmetrically. However, surround speaker spacing is best thought of as the angle from the listener between adjacent speakers, rather than using absolute distance between speakers. For optimal reproduction throughout the auditorium, the angular distance between adjacent speakers should be no more than 30°, as viewed from each of the four corners of the main listening area. Good results can be achieved with spacings up to 50°. For each surround zone, speakers should maintain equal linear spacing adjacent to the seating area, if possible. Linear spacing beyond the listening area, for example between the front row and the screen, can be slightly larger. Figure 11 is a diagram showing an example placement of top surround speakers 808 and side surround speakers 806 relative to a reference point, according to one embodiment.
[0126] The additional side surround speakers 806 should be placed closer to the screen than the currently recommended practice, which begins about one-third of the distance to the rear of the auditorium. These speakers are not used as side surrounds during playback of Dolby Surround 7.1 or 5.1 soundtracks, but they allow for smoother transitions and improved timbre matching when panning objects from the screen speakers to the surround zone. To maximize the impression of space, the surround array should be placed as low as possible, subject to the following constraints: the virtual placement of the surround speakers in front of the array should be close to the height of the acoustic center of the screen speakers, and just high enough to maintain sufficient coverage across the seating area depending on the speaker's directivity. The vertical placement of the surround speakers should form a straight line from front to back, as shown in Figure 10, and (generally) be tilted so that the relative elevation of the surround speakers is maintained above the listener and, as the seating elevation increases, toward the rear of the cinema. 10 is a side view of an example layout of suggested speaker locations for use with an adaptive audio system in a typical auditorium. In practice, this is most easily achieved by selecting the elevations of the front-most and rear-most surround speakers and placing the remaining speakers in a line between these points.
[0127] To provide optimal coverage of each speaker in the seating area, the side surrounds 806, rear speakers 816, and top surrounds 808 must be oriented to a reference position in the theater with defined guidelines for spacing, position, angle, etc.
[0128] Embodiments of the Adaptive Audio Cinema system and format enable a higher level of audience immersion than current systems by providing mixers with powerful new authoring tools and a new cinema processor with a flexible rendering engine that optimizes the sound quality and surround effects of the soundtrack for each room's speaker layout and characteristics, while maintaining backward compatibility with and minimizing the impact on current production and distribution workflows.
[0129] While the embodiments have been described with respect to an example and implementation of a cinema environment in which adaptive audio content is associated with film content used in a digital cinema processing system, it should be noted that the embodiments can also be implemented in non-cinema environments. The adaptive audio content, including object-based audio and channel-based audio, can be used with any associated content (such as associated audio, video, graphics, etc.) or may constitute standalone audio content. The playback environment can be any suitable listening environment, from headphones and near-field monitors to small or large rooms, cars, outdoor arenas, and concert halls.
[0130] Aspects of system 100 can be implemented in any suitable computer-based sound processing network environment that processes digital or digitized audio files. Portions of the adaptable audio system include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between the computers. Such networks may be configured over a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. In one embodiment where the network is the Internet, one or more machines are configured to access the Internet through a web browser program.
[0131] One or more components, blocks, processors, or other functional components are implemented by computer programs that control the execution of a processor-based computing device of the system. It should be noted that various functions disclosed herein can be described in terms of behavior, register transfers, logic components, and / or other characteristics using any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media on which such formatted data and / or instructions are embodied may be various forms of physical (non-transitory), non-volatile storage media, including, but not limited to, optical, magnetic, or semiconductor storage media.
[0132] Unless otherwise clearly required by context, throughout this specification and claims, words like "comprise," "comprising," and the like are inclusive and not exclusive or exhaustive; that is, meaning "including, but not limited to." Words in the singular or plural include the plural or singular, respectively. Also, words like "herein," "hereunder," "above," "below," and the like, refer to this application as a whole and not to particular portions of this application. The word "or," when used in reference to a list of two or more items, covers all interpretations of this word: any item in the list, any item in the list, and any combination of items in the list. While one or more implementations have been described by way of example and in terms of specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements that will be apparent to those skilled in the art. The scope of the appended claims, therefore, should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements. Additional notes regarding the above embodiment are provided. (Supplementary Note 1) A system for processing an audio signal, comprising: an authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and one or more sets of metadata associated with each audio stream and defining a playback location for each audio stream; the audio stream is identified as channel-based audio or object-based audio; the playback locations of the channel-based audio include speaker designations for a plurality of speakers in a speaker array, and the playback locations of the object-based audio include locations in three-dimensional space; a first set of metadata is applied by default to one or more of the plurality of audio streams, and when a condition of the playback environment matches the one condition of the playback environment, a second set of metadata is associated with the one condition of the playback environment and is applied to one or more of the plurality of audio streams instead of the first set; a rendering system coupled to the authoring component and configured to receive a bitstream encapsulating the plurality of monophonic audio streams and one or more data sets, and to render the audio streams to a plurality of speaker feeds corresponding to speakers of the playback environment according to the one or more metadata sets based on conditions of the playback environment. (Supplementary Note 2) Each metadata set includes metadata elements associated with each object-based stream, the metadata for each object-based stream defining spatial parameters controlling the playback of the corresponding object-based sound, including one or more of sound position, sound width, and sound velocity; each metadata set includes metadata elements associated with each channel-based stream, and the speaker array is configured in a defined surround sound configuration; a metadata element associated with each channel-based stream including a surround sound channel designation of a speaker in a speaker array according to a defined surround sound standard; 10. The system of claim 1. (Supplementary Note 3) The speaker array includes additional speakers for playback of object-based streams that are arranged in a playback environment related to setup instructions from a user based on conditions of the playback environment; the playback conditions depend on variables including the size and shape of the room of the playback environment, occupancy, and ambient noise; The system receives from a user a setup file that includes at least a list of speaker designations, a mapping of channels to individual speakers in the speaker array, information about speaker groups, and a run-time mapping based on the relative positions of speakers to the playback environment; 10. The system of claim 1. (Supplementary Note 4) The authoring component includes a mixing console having controls operable by a user to define a playback level of an audio stream including the original audio content; wherein the metadata elements associated with each object-based stream are generated automatically upon input by the user to the mixing console controls. 10. The system of claim 1. (Supplementary Note 5) The metadata set includes metadata that enables upmixing or downmixing of at least one of the channel-based audio stream and the object-based audio stream in response to a change from a first configuration of the speaker array to a second configuration of the speaker array. 10. The system of claim 1. (Supplementary Note 6) The content types are selected from the group consisting of dialogue, music, and effects, and each content type is embodied in a respective set of channel-based streams or object-based streams; sound components of each content type are transmitted to a defined group of speakers among one or more speaker groups designated in the speaker array; 10. The system described in Appendix 3. (Supplementary Note 7) The speakers of the speaker array are arranged at a plurality of positions within the playback environment; a metadata element associated with each object-based stream specifies that one or more sound components are to be rendered to speaker feeds for playback by the speakers closest to the intended playback location of the sound component as indicated by the position metadata; 10. The system described in Appendix 6. (Supplementary Note 8) The playback location is a spatial position relative to the playback environment or a screen within a plane that envelops the playback environment; The surfaces include a front surface, a rear surface, a left surface, a right surface, a top surface, and a bottom surface. 10. The system of claim 1. (Supplementary Note 9) The audio processing system further includes a codec coupled to the authoring component and the rendering component, configured to receive the plurality of audio streams and metadata, and generate a single digital bitstream including the plurality of audio streams in an ordered sequence. 10. The system of claim 1. (Supplementary Note 10) The rendering component further comprises means for selecting a rendering algorithm to be utilized by the rendering component, the rendering algorithm being selected from the group consisting of binaural, stereo dipole, Ambisonics, Wave Field Synthesis (WFS), multi-channel panning, low-level stereo with positional metadata, dual-balanced, and vector-based amplitude panning. 10. The system of claim 9. (Appendix 11) The playback location of each audio stream is independently defined relative to either the egocentric or allocentric frame of reference; an egocentric frame of reference is taken relative to a listener in said playback environment; the allocentric frame of reference is taken relative to the characteristics of the playback environment; 10. The system of claim 1. (Supplementary Note 12) A system for processing an audio signal, comprising: an authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and metadata associated with each audio stream and defining a playback location for each audio stream; the audio stream is identified as channel-based audio or object-based audio; the playback locations of the channel-based audio include speaker designations for a plurality of speakers in a speaker array, and the playback locations of the object-based audio include locations in three-dimensional space; each object-based audio stream is rendered on at least one speaker of the speaker array; a bitstream encapsulating the plurality of monophonic audio streams and metadata coupled to the authoring component, the bitstream encapsulating the plurality of monophonic audio streams and metadata, and rendering the audio streams into a plurality of speaker feeds corresponding to speakers in a playback environment; the speakers of the speaker array are arranged at a plurality of positions within the playback environment; A system in which metadata elements associated with each object-based stream specify that one or more sound components are rendered to speaker feeds for playback by speakers closest to the intended playback location of the sound components, such that the object-based stream is effectively snapped to the speakers closest to the intended playback location. (Supplementary Note 13) The metadata includes two or more metadata sets, and the rendering system renders the audio stream using one of the two or more metadata sets based on conditions of the playback environment; a first set of metadata is applied to one or more of the plurality of audio streams in a first condition of the playback environment, and a second set of metadata is applied to one or more of the plurality of audio streams in a second condition of the playback environment; each metadata set includes metadata elements associated with each object-based stream, the metadata for each object-based stream defining spatial parameters controlling playback of the corresponding object-based sound, including one or more of sound position, sound width, and sound velocity; each metadata set includes metadata elements associated with each channel-based stream, and the speaker array is configured in a defined surround sound configuration; a metadata element associated with each channel-based stream including a surround sound channel designation of a speaker in a speaker array according to a defined surround sound standard; 13. The system of claim 12. (Supplementary Note 14) The speaker array includes additional speakers for playback of object-based streams that are arranged in a playback environment related to setup instructions from a user based on conditions of the playback environment; the playback conditions depend on variables including the size and shape of the room of the playback environment, occupancy, and ambient noise; The system receives from a user a setup file that includes at least a list of speaker designations, a mapping of channels to individual speakers in the speaker array, information about speaker groups, and a run-time mapping based on the relative positions of speakers to the playback environment; an object stream rendered to a speaker feed for playback by a speaker closest to the intended playback location of the sound component is snapped to a single speaker of the additional speakers; 13. The system of claim 12. (Supplementary Note 15) The intended playback location is a spatial position relative to the playback environment or a screen within a plane enveloping the playback environment; The surfaces include a front surface, a rear surface, a left surface, a top surface, and a bottom surface. 15. The system of claim 14. (Supplementary Note 16) A system for processing an audio signal, comprising: an authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and metadata associated with each audio stream and defining a playback location for each audio stream; the audio stream is identified as channel-based audio or object-based audio; the playback location of the channel-based audio includes speaker designations for a plurality of speakers in a speaker array, and the playback location of the object-based audio includes a location in three-dimensional space relative to a playback environment including the speaker array; each object-based audio stream is rendered on at least one speaker of the speaker array; a rendering system coupled to the authoring component and configured to receive a first map to speaker audio channels including a list of a plurality of speakers and their respective locations in the playback environment, and a bitstream encapsulating a plurality of monophonic audio streams and metadata, and to render the audio streams to a plurality of speaker feeds corresponding to the speakers of the playback environment according to runtime mapping to the playback environment based on the relative positions of the speakers and conditions of the playback environment. (Appendix 17) The conditions of the playback environment depend on variables including the size and shape of the room, occupancy, material composition, and ambient noise of the playback environment. 17. The system of claim 16. (Supplementary Note 18) The first map is defined in a setup file including at least a list of speaker designations and mappings of channels to individual speakers of the speaker array, and information on speaker groupings. 18. The system of claim 17. (Supplementary Note 19) The intended playback location is a spatial position relative to a screen within the plane of the playback environment or an enclosure containing the playback environment; The surfaces include the front, rear, side, top, and bottom surfaces of the enclosure. 19. The system of claim 18. (Supplementary Note 20) The speaker array has speakers arranged in a defined surround sound configuration; a metadata element associated with each channel-based stream including a surround sound channel designation of a speaker in a speaker array according to a defined surround sound standard; the object-based stream is played by an additional speaker of the speaker array; The runtime mapping dynamically determines which individual speakers of the speaker array will play corresponding object-based streams during a playback process. 19. The system of claim 19. (Supplementary Note 21) A method for authoring an audio signal to be rendered, comprising: receiving a plurality of audio signals; generating a plurality of monophonic audio streams and one or more sets of metadata associated with each audio stream and defining a playback location for each audio stream; the audio stream is identified as channel-based audio or object-based audio; the playback location of the channel-based audio includes speaker designations for a plurality of speakers in a speaker array, and the playback location of the object-based audio includes a location in three-dimensional space relative to a playback environment including the speaker array; a first set of metadata is applied to one or more of the plurality of audio streams in a first condition of the playback environment, and a second set of metadata is applied to one or more of the plurality of audio streams in a second condition of the playback environment; and encapsulating the plurality of monophonic audio streams and one or more sets of metadata into a bitstream for transmission to a rendering system configured to render the audio streams to a plurality of speaker feeds corresponding to speakers of the playback environment according to the one or more sets of metadata based on conditions of the playback environment. (Supplementary Note 22) Each metadata set includes metadata elements associated with each object-based stream, the metadata for each object-based stream defining spatial parameters controlling the playback of the corresponding object-based sound, including one or more of sound position, sound width, and sound velocity; each metadata set includes metadata elements associated with each channel-based stream, and the speaker array is configured in a defined surround sound configuration; a metadata element associated with each channel-based stream including a surround sound channel designation of a speaker in a speaker array according to a defined surround sound standard; 22. The method described in Appendix 21. (Supplementary Note 23) The speaker array includes additional speakers for playback of the object-based stream arranged in a playback environment, and the method further comprises receiving setup instructions from a user based on conditions of the playback environment; the playback conditions depend on variables including the size and shape of the room of the playback environment, occupancy, and ambient noise; the setup instructions further include at least a list of speaker designations, mappings of channels to individual speakers of the speaker array, information about speaker groups, and runtime mappings based on the relative positions of speakers to the playback environment; 22. The method described in Appendix 21. (Supplementary Note 24) Receiving from a mixing console having user-operated controls to define a playback level of an audio stream containing the original audio content; and automatically generating, upon receipt of said user input, metadata elements associated with each generated object-based stream. 24. The method described in Appendix 23. (Supplementary Note 25) A method for rendering an audio signal, comprising: receiving a bitstream encapsulating a plurality of monophonic audio streams and one or more sets of metadata from an authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and one or more sets of metadata associated with each audio stream and defining a playback location for each audio stream; the audio stream is identified as channel-based audio or object-based audio; the playback location of the channel-based audio includes speaker designations for a plurality of speakers in a speaker array, and the playback location of the object-based audio includes a location in three-dimensional space relative to a playback environment including the speaker array; a first set of metadata is applied to one or more of the plurality of audio streams in a first condition of the playback environment, and a second set of metadata is applied to one or more of the plurality of audio streams in a second condition of the playback environment; and rendering the plurality of audio streams to a plurality of speaker feeds corresponding to speakers in the playback environment according to the one or more sets of metadata based on conditions of the playback environment. (Supplementary Note 26) A method for generating audio content including multiple monophonic audio streams processed by an authoring component, comprising: The monophonic audio stream includes at least one channel-based audio stream and at least one object-based audio stream, and the method comprises: indicating whether each of the plurality of audio streams is a channel-based stream or an object-based stream; associating with each channel-based stream a metadata element that defines a channel position at which the channel-based stream is to be rendered onto one or more speakers in a playback environment; associating with each object-based stream one or more metadata elements that define an object-based frame of reference for rendering the respective object-based stream to one or more speakers within the playback environment, relative to an allocentric frame of reference defined with respect to the size and dimensions of the playback environment; assembling the plurality of monophonic audio streams and associated metadata into a signal. (Supplementary Note 27) The reproduction environment includes an array of speakers arranged at defined locations and orientations relative to a reference point of an enclosure embodying the reproduction environment. 26. The method described in Appendix 26. (Supplementary Note 28) The first set of speakers of the speaker array comprises speakers configured according to a defined surround sound system; a second set of speakers of the speaker array comprising speakers configured according to an adaptive audio scheme; The method described in Appendix 27. (Supplementary Note 29) A step of defining an audio type of the set of monophonic audio streams; the audio type is selected from the group consisting of dialogue, music, and effects; transmitting the set of audio streams to a set of speakers based on an audio type of the set of audio streams; 29. The method described in Appendix 28. (Supplementary Note 30) The method further comprises automatically generating the metadata element by an authoring component implemented in a mixing console having a user-operable control for defining a playback level of the monophonic audio stream. 29. The method described in Appendix 29. (Supplementary Note 31) The method further comprises packaging, in the encoder, the multiple monophonic audio streams and associated metadata elements into a single digital bitstream. 31. The method described in Appendix 30. (Supplementary Note 32) A method for generating audio content, comprising: determining values for one or more metadata elements in a first metadata group associated with programming audio content for processing in a hybrid audio system that handles both channel-based and object-based content; determining values of one or more metadata elements of a second metadata group relating to storage and rendering characteristics of audio content in the hybrid audio system; and determining values of one or more metadata elements of a third metadata group related to audio source position and control information for rendering the channel-based and object-based audio content. (Supplementary Note 33) The audio source locations for rendering the channel-based audio content include names associated with speakers of a surround sound speaker system; the name defines the location of the speaker relative to a reference point in the playback environment; 32. The method described in Appendix 32. (Supplementary Note 34) The control information for rendering the channel-based audio content includes upmix and downmix information for rendering the audio content in different surround sound configurations; the metadata includes metadata for enabling or disabling upmix and / or downmix functions; 34. The method described in Appendix 33. (Supplementary Note 35) The audio source locations for rendering the object-based audio content have values associated with one or more mathematical functions that define intended playback locations for playback of sound components of the object-based audio content. 32. The method described in Appendix 32. (Supplementary Note 36) The mathematical function is selected from the group consisting of three-dimensional coordinates defined as x, y, and z coordinate values, a surface definition and a set of two-dimensional coordinates, a curve definition and a set of one-dimensional linear position coordinates, and a scalar position on a screen in the playback environment. The method described in Appendix 35. (Supplementary Note 37) The control information for rendering the object-based audio content includes values that define individual speakers or speaker groups within a playback environment on which the sound components are to be played. The method described in Appendix 36. (Supplementary Note 38) The control information for rendering the object-based audio content further includes a binary value that specifies a sound source to be snapped to the nearest speaker or nearest speaker group in the playback environment. The method described in Appendix 37. (Supplementary Note 39) A method for defining an audio transport protocol, comprising: defining values for one or more metadata elements in a first metadata group associated with audio content programming for processing in a hybrid audio system that handles both channel-based and object-based content; defining values for one or more metadata elements of a second metadata group relating to audio content storage and rendering characteristics of the hybrid audio system; and defining values of one or more metadata elements of a third metadata group related to audio source position and control information for rendering the channel-based and object-based audio content.
Claims
1. 1. A system for processing an audio signal, comprising: a rendering system, the rendering system comprising: receiving a bitstream including encoded audio data representing a plurality of mono audio streams, the bitstream further including metadata associated with each of the plurality of mono audio streams and indicating a playback position for the respective mono audio stream, wherein at least some of the plurality of mono audio streams are identified as object-based audio, the playback positions for the object-based mono audio streams including positions in three-dimensional space, and at least some other of the plurality of mono audio streams are identified as channel-based audio, the playback positions for the channel-based mono audio streams including designations of speakers in a speaker array; decoding the encoded audio data to provide the plurality of mono audio streams; rendering the plurality of mono audio streams into a plurality of speaker feeds corresponding to speakers in a playback environment, the speakers being located at particular positions in the playback environment, and one or more additional metadata elements associated with each object-based mono audio stream indicating whether rendering of the respective object-based mono audio stream into one or more particular speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based mono audio stream is not rendered into any of the one or more particular speaker feeds of the plurality of speaker feeds; A system configured to:
2. 2. The system of claim 1, wherein the metadata element associated with each object-based mono audio stream further indicates spatial parameters controlling the playback of the corresponding sound component, including one or more of sound position, sound width, and sound velocity.
3. 2. The system of claim 1, wherein the playback position for each of a plurality of object-based mono audio streams is independently specified with respect to either an egocentric frame of reference or an allocentric frame of reference, the egocentric frame of reference being taken with respect to a listener in the playback environment and the allocentric frame of reference being taken with respect to characteristics of the playback environment.
4. A method of authoring audio content for rendering, Receiving multiple audio signals generating a plurality of mono audio streams and metadata associated with each of the plurality of mono audio streams and indicating a playback position for the respective mono audio stream, wherein at least some of the plurality of mono audio streams are identified as object-based audio, the playback position for the object-based audio including a position in three-dimensional space, and at least some other of the plurality of mono audio streams are identified as channel-based audio, the playback position for the channel-based mono audio stream including a designation of a speaker in a speaker array; encoding the plurality of mono audio streams to provide encoded audio data; encapsulating the encoded audio data and the metadata in a bitstream for transmission to a rendering system configured to render the plurality of mono audio streams to a plurality of speaker feeds corresponding to speakers in a playback environment, the speakers being located at particular locations within the playback environment, and one or more additional metadata elements associated with each object-based mono audio stream indicating whether rendering of the respective mono audio stream to one or more particular speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based mono audio stream is not rendered to any of the one or more particular speaker feeds of the plurality of speaker feeds; A method comprising:
5. A method for rendering an audio signal, receiving a bitstream including encoded audio data representing a plurality of mono audio streams, the bitstream further including metadata associated with each of the plurality of mono audio streams and indicating a playback position for the respective mono audio stream, wherein at least some of the plurality of mono audio streams are identified as object-based audio, the playback positions for the object-based mono audio streams including positions in three-dimensional space, and at least some other of the plurality of mono audio streams are identified as channel-based audio, the playback positions for the channel-based mono audio streams including designations of speakers in a speaker array; decoding the encoded audio data to provide the plurality of mono audio streams; rendering the plurality of mono audio streams into a plurality of speaker feeds corresponding to speakers in a playback environment, the speakers being located at particular positions in the playback environment, and one or more additional metadata elements associated with each object-based mono audio stream indicating whether rendering of the respective object-based mono audio stream into one or more particular speaker feeds of the plurality of speaker feeds is prohibited, such that the respective object-based mono audio stream is not rendered into any of the one or more particular speaker feeds of the plurality of speaker feeds; A method comprising:
6. 6. The method of claim 5, wherein the metadata element associated with each object-based mono audio stream further indicates spatial parameters controlling the playback of the corresponding sound component, including one or more of sound position, sound width, and sound velocity.
7. 6. The method of claim 5, wherein the playback position for each of the plurality of object-based mono audio streams comprises a spatial position relative to a screen within a playback environment or a surface surrounding the playback environment, the surface including front, back, left, right, top and bottom, and / or is independently specified relative to either an egocentric or allocentric frame of reference, the egocentric frame of reference being taken relative to a listener in the playback environment and the allocentric frame of reference being taken relative to a characteristic of the playback environment.
8. 8. A non-transitory computer readable storage medium comprising a set of instructions that, when executed by a system for processing an audio signal, cause the system to perform the method of any one of claims 4 to 7.
9. A computer program which, when executed on a computer, causes the computer to carry out the method according to any one of claims 4 to 7.
Citation Information
Patent Citations
Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information
JP2009278381A
System for adaptively streaming audio objects
US20110040396A1