Audio signal processing system and method

The adaptive audio system addresses the limitations of current cinema sound systems by integrating channel-based and model-based audio methods, using metadata to describe sound positions, and enhancing audio quality and distribution, resulting in a more immersive and flexible audio experience.

JP7684356B2Active Publication Date: 2025-05-27DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023145272
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-04-20
Filing Date
2023-09-07
Publication Date
2025-05-27
Estimated Expiration
2032-06-27

AI Technical Summary

Technical Problem

Current cinema sound systems struggle to provide truly immersive, physical-like audio due to limitations in channel-based audio systems, which fail to accurately reproduce sound sources in different playback environments, leading to a dependent listener experience on seating position.

Method used

An adaptive audio system that combines channel-based and model-based audio scene description methods, using a new speaker layout and spatial description format to support multiple rendering techniques. This system transmits audio streams with metadata describing the mixer's intent, including desired sound positions, which can be represented as specific channels or 3D positions.

Benefits of technology

The adaptive audio system enhances audio quality in various rooms through improved room equalization and surround bass management, allowing mixers to freely address speakers without tone matching. It introduces dynamic audio objects, simplifies distribution, and optimizes playback of artistic intent across different theater configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684356000014
    Figure 0007684356000014
  • Figure 0007684356000015
    Figure 0007684356000015
  • Figure 0007684356000016
    Figure 0007684356000016
Patent Text Reader

Abstract

To provide systems and methods for processing an audio signal including a cinema sound format and a spatial description format.SOLUTION: An authoring component having a plurality of monophonic audio streams and associated metadata generates an adaptive audio mix indicating a playback location. A portion thereof is recognized as a channel-based audio, and another portion is recognized as an object-based audio. The playback location includes designation of a speaker and a location in a three-dimensional space. The metadata indicates whether rendering each of the monophonic audio streams to one or more certain speakers among a plurality of speakers is prohibited such that each of object-based monophonic audio streams is rendered to none of one or more certain speakers among a plurality of speaker feeds.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments generally relate to audio signal processing, and more specifically, to hybrid object and channel-based audio processing for use in cinema, home, and other environments.

Background Art

[0002] The subject matter described in the background art section should not be considered prior art merely because it is mentioned in the background art section. Similarly, the problems mentioned in the background art section, or related to the subject matter of the background art section, must not be assumed to have been previously recognized in the prior art. The subject matter of the background art section merely represents a plurality of different approaches and may itself be an invention.

[0003] Since the advent of sound-on-film, the technology used to capture the creative intent of the creators of movie sound tracks and accurately reproduce it in a cinema environment has evolved steadily. The basic role of cinema sound is to support the story shown on the screen. A typical cinema sound track includes many different sound elements where the elements on the screen correspond to the image, conversations, noises, sound effects that are emitted from different on-screen elements and combine with background music and ambient effects to form an overall audience experience. The creative intent of the creators and producers is expressed as a desire to play these sounds so that they correspond as closely as possible to what is shown on the screen with respect to sound source position, intensity, movement, and other similar parameters.

[0004] Current cinema authoring, distribution, and playback have limitations that constrain the generation of truly immersive, physical-like audio. Conventional channel-based audio systems send audio content in the form of speaker feeds to individual speakers in the playback environment, such as stereo or 5.1 systems. With the advent of digital cinema, new standards have emerged for sound on film, such as the incorporation of up to 16 channels of audio, giving content creators greater creativity and providing audiences with a more immersive and realistic acoustic experience. The introduction of 7.1 surround sound systems has provided a new format for increasing the number of surround channels by splitting the existing left and right surround channels into four zones, increasing the scope for sound designers and mixers to control the positioning of audio elements in the theater.

[0005] To further enhance the listener's experience, the reproduction of sound in a virtual 3D environment has become an area of research and development. The spatial representation of sound utilizes audio objects. An audio object is a parametric source description consisting of an audio signal and associated apparent source position (e.g., 3D coordinates), apparent source width, and other parameters. Object-based audio is increasingly being used in many current multimedia applications, such as digital movies, video games, simulations, and 3D videos.

[0006] As a means of distributing spatial audio, an extension beyond conventional speaker feeds and channel-based audio is essential, and there is significant interest in model-based audio descriptions. Model-based audio descriptions promise to give listeners / egocentric users the freedom to select a playback configuration that suits their individual needs and budgets and to render the audio in the configuration they have selected. At a high level, there are currently four main spatial audio description formats: speaker feed, in which audio is described as a signal intended for speakers at nominal speaker positions; Microphone feeds where audio is described as signals captured by virtual or physical microphones in a given array; Model-based descriptions where audio is described for sequences of audio events at described locations; and Binaural where audio is described by signals reaching the listener's ears. These four description formats are often associated with one or more rendering techniques for converting the audio signal into a speaker feed. Current rendering techniques include panning, where an audio stream is converted into a speaker feed using a set of panning laws and known or assumed speaker positions (generally rendered before distribution); Ambisonics, which converts microphone signals and feeds them for a scalable speaker array (generally rendered after distribution); WFS (wave field synthesis), where sound events are converted into appropriate speaker signals for synthesizing a sound field (generally rendered after distribution); and binaural, which generally uses headphones but also uses speaker and cross-talk cancellation and sends L / R (left / right) binaural signals to the left and right ears (rendered before or after distribution). Of these formats, the speaker feed format is the most common because it is simple and effective. The best sonic results (the most accurate and reliable) can be achieved by mixing / monitoring and direct distribution to the speaker feed. This is because there is no processing between the content creator and the listener. When the playback system is known in advance, the speaker feed description generally provides the highest fidelity. However, in many practical applications, the playback system is unknown. The model-based description seems to be the most adaptable. This is because it makes no assumptions about the rendering technique and can therefore be easily applied to any rendering technique. The model-based description captures spatial information efficiently, but becomes inefficient as the number of sound sources increases.

[0007] For many years, cinema systems have featured discrete screen channels in the form of left, center, right, and in some cases "inner left" and "inner right" channels. These discrete sources generally have the frequency response and power handling sufficient to accurately place sound in different areas of the screen and provide acoustic matching as the sound moves or pans between locations. In recent developments aimed at enhancing the listener experience, attempts have been made to accurately reproduce the location of sound for the listener. In a 5.1 system, the surround "zones" consist of an array of multiple speakers, all of which have the same audio information in each of the left surround zone and the right surround zone. Such arrays may be effective for "ambient" or diffused surround effects, but in everyday use, sound effects are emitted from randomly placed point sources. For example, in a restaurant, background music is playing throughout, while individual, hard-to-locate sounds, such as a person's voice from one point and the sound of a knife hitting a plate from another point, are emitted from multiple points. If such sounds could be placed directly around the audience, the sense of realism would be heightened, even if it is less obvious. Overhead sound is also an important component of the surround definition. In the real world, sound is emitted from all directions and not necessarily from a single horizontal plane. When sound is heard from above, i.e., from the "upper hemisphere", the sense of realism is enhanced. However, current systems cannot provide truly accurate reproduction of different audio types of sound in different playback environments. Attempting to accurately represent the location of sound using existing systems requires a great deal of processing, knowledge of the actual playback environment, and setup, and in most applications, the current rendering systems are not practical.

[0008] What is needed is a system that supports multiple screen channels to increase the definition of on-screen sound and conversations and improve audio-visual coherence, and the ability to accurately place sound sources anywhere in a surround zone to improve the audio-visual transition from the screen to the room. For example, when a character on the screen is looking towards the sound source in the room, the sound engineer (the "mixer") should have the ability to accurately place the sound so that it coincides with the character's line of sight and the effect is consistent among the audience. However, in conventional 5.1 or 7.1 surround sound mixes, the effect is highly dependent on the listener's seating position, which is a disadvantage in large listening environments. Increasing the surround resolution creates a new opportunity to use sound in a room-centric manner, as opposed to the conventional approach of creating content assuming a single listener is in the "sweet spot".

[0009] Apart from spatial issues, current state-of-the-art multi-channel systems also have problems with sound quality. For example, the timbral quality, such as the sound of steam leaking from a broken pipe, has to be reproduced by an array of multiple speakers. The ability to direct sound at a single speaker gives the mixer the opportunity to eliminate artifacts of array reproduction and deliver a more realistic experience to the audience. Conventionally, multiple surround speakers do not support the same full range of audio frequencies and the levels supported by large screen channels. Historically, this has been a problem for mixers, reducing the ability to freely move full-range sound from the screen to the room. As a result, theater owners have not felt the need to upgrade their surround channel configurations, hindering the spread of higher-quality sound equipment. [Cross-reference to related applications]

[0010] This application claims the benefit of U.S. Provisional Application No. 61 / 504,005, filed Jul. 1, 2011, and U.S. Provisional Application No. 61 / 636,429, filed Apr. 20, 2012, the entire disclosures of which are hereby incorporated by reference for all purposes. SUMMARY OF THE INVENTION

[0011] A system and method for a cinema sound format and a processing system including a new speaker layout (channel configuration) and an associated spatial description format are described. An adaptive audio system and format that support multiple rendering techniques are defined. Audio streams are transmitted with metadata that describes the "intent of the mixer" including the desired position of the audio stream. The position is represented as a specified channel (from within a predefined channel configuration) or as 3D position information. This channel and object format optimally combines channel-based and model-based audio scene description methods. Audio data for the adaptive audio system includes a plurality of independent monophonic audio streams. Each stream has associated metadata indicating whether the stream is a channel-based stream or an object-based stream. The channel-based streams have rendering information encoded by channel name.

[0012] The object-based stream has location information encoded by a mathematical formula encoded in another related metadata. The original multiple independent audio streams are packaged as a single serial bitstream containing all of the audio data. With this configuration, sound is rendered by an other-centered reference frame. The rendering location of the sound is based on the characteristics of the playback environment (e.g., room size, shape, etc.) to correspond to the intention of the mixer. The object position metadata includes appropriate other-centered reference frame information necessary to correctly reproduce the sound using the available speaker positions in a room configured to play adaptive audio content. Thereby, the sound is optimally mixed to fit a playback environment that may be different from the mix environment experienced by the sound engineer.

[0013] The adaptive audio system improves the quality of audio in different rooms, due to benefits such as improved room equalization and surround bass management, allowing the mixer to freely address speakers (whether on-screen or off-screen) without having to consider tone matching. The adaptive audio system adds the flexibility and power of dynamic audio objects to the conventional channel-based workflow. With these audio objects, the creator can control individual sound elements regardless of the playback speaker configuration including overhead speakers. Also, this system introduces new efficiency to the post-production process, whereby the sound engineer can efficiently capture all of their intentions and monitor or automatically generate surround sound 7.1 and 5.1 versions in real time.

[0014] This adaptive audio system simplifies distribution by encapsulating the essence and artistic intent of audio within a single track file in a digital cinema processor. This single track file can be faithfully reproduced in a wide range of theater configurations. The system provides optimal playback of artistic intent when mixing and rendering use a single inventory that is downwardly adapted to the same channel configuration and rendering configuration, i.e., when downmixing.

[0015] These and other advantages are provided through embodiments related to a cinema sound platform, which solve current system limitations and deliver an audio experience that goes beyond currently available systems.

Brief Description of the Drawings

[0016] In the following drawings, the same reference numbers are used to refer to the same elements. The following drawings illustrate various examples, but one or more implementations are not limited to the examples shown in the drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0017] An adaptive audio system and method, associated audio signals, and a data format supporting multiple rendering techniques are described. Aspects of one or more embodiments described herein can be implemented in an audio or audiovisual system that processes source audio information in a mixing, rendering, and playback system including one or more computers or processing devices that execute software instructions. Any of the embodiments described can be used alone or in combination with each other. Various embodiments are motivated by various deficiencies in the prior art, which are described in one or more places herein, but the embodiments need not necessarily address these deficiencies. In other words, different embodiments address different deficiencies described herein. One embodiment may only partially address a particular deficiency described herein, may address only one deficiency, or may not address any of these deficiencies.

[0018] For the purposes of this description, the following terms have the associated meanings: Channel or audio channel: A monophonic audio signal or audio stream and metadata whose position, such as left front or right top surround, is encoded as a channel ID. A channel object drives a plurality of speakers. For example, the left surround channel (Ls) feeds all the speakers in the Ls array.

[0019] Channel configuration: A predetermined set of speaker zones having associated nominal locations, such as 5.1, 7.1, etc.; 5.1 refers to a 6-channel surround sound audio system having left and right channels, a center channel, two surround channels, and a subwoofer channel; 7.1 refers to an 8-channel surround system that adds two additional surround channels to a 5.1 system. Examples of 5.1 and 7.1 configurations include Dolby® surround systems.

[0020] Speaker: An audio transducer or a set of multiple transducers that renders an audio signal.

[0021] Speaker zone: An array of one or more speakers that can be uniquely referenced, receives a single audio signal such as left surround as commonly seen in cinema, and is specifically excluded or included for object rendering.

[0022] Speaker channel or speaker feed channel: An audio channel associated with a specified speaker or speaker zone within a defined speaker configuration. A speaker channel is nominally rendered using the associated speaker zone.

[0023] Speaker channel group: A set of one or more speaker channels corresponding to a channel configuration (e.g., stereo track, mono track, etc.).

[0024] Object or object channel: One or more audio channels having a numerical source description, with an apparent source position (e.g., 3D coordinates), an apparent source width, etc. An audio stream having metadata in which the position is encoded as a 3D position in space. Audio program: A complete set of speaker channels and / or object channels and associated metadata describing a desired spatial audio presentation.

[0025] Allocentric reference: A spatial reference in which an audio object is defined with respect to features in a rendering environment, such as room walls or corners, standard speaker locations, screen locations (e.g., the front left corner of a room).

[0026] Egocentric reference: A spatial reference in which an audio object is often specified at an angle with respect to the (listener) viewer's perspective (e.g., 30° to the right of the listener).

[0027] Frame: A frame is an independently decodable segment into which an entire audio program is divided. The audio frame rate and boundaries are generally aligned with video frames.

[0028] Adaptive audio: Metadata for rendering an audio signal based on a channel-based and / or object-based audio signal and a playback environment.

[0029] The cinema sound format and processing system described herein, also referred to as an "adaptive audio system", utilizes new spatial audio description and rendering techniques to enhance audience immersion, increase artistic control, improve system flexibility and scalability, and ease installation and maintenance. Embodiments of the cinema audio platform include a plurality of individual components including mixing tools, packer / encoders, unpacker / decoders, in-theater final mix / rendering components, new speaker designs, and networked amplifiers. The system includes suggestions for new channel configurations for use by content creators and exhibitors. The system utilizes a model-based description that supports the following plurality of features: A single inventory that enables optimal use of available speakers, adapting both downward and upward to the rendering configuration; Improved sound envelopment including optimized downmixing that avoids inter-channel correlation; Increased spatial resolution by steer-thru arrays (e.g., audio objects dynamically assigned to one or more speakers in a surround array); Support for alternative rendering methods.

[0030] Figure 1 is a diagram showing a top-level overview of an audio generation and playback environment utilizing an adaptive audio system according to an embodiment. As shown in Figure 1, the comprehensive end-to-end environment 100 includes content creation, packaging, distribution, and playback / rendering components across a wide range of endpoint devices and use cases. The overall system 100 starts with content supplemented from a number of different use cases including different user experiences 112. The content capture element 102 includes, for example, cinema, TV, live broadcasts, user-generated content, recorded content, games, music, etc., including audio / visual content or pure audio content. As the content progresses through the system 100 from the capture stage 102 to the final user experience 112, it passes through multiple key processing steps through individual system components. These process steps include audio preprocessing 104, authoring tools and processing 106, such as encoding 108 by an audio codec that captures, for example, audio data, additional metadata, and playback information, and object channels. Various processing effects such as compression (lossy or lossless), encryption, etc. are applied to the object channels for efficient and secure distribution over various media. An appropriate endpoint-specific decoding and rendering process 110 is applied to reproduce and deliver the adaptive audio user experience 112. The audio experience 112 represents the playback of audio or audio / visual content by appropriate speakers and playback devices, and represents any environment such as a cinema, concert hall, outdoor theater, home or room, listening booth, automobile, game console, headphone or headset system, public address (PA) system, or other playback environment where the listener experiences the playback of the captured content.

[0031] Embodiments of system 100 include an audio codec 108, sometimes called a "hybrid" codec, that enables efficient distribution and storage of multi-channel audio programs. The codec 108 combines conventional channel-based audio data with associated metadata to create audio objects that facilitate the generation and distribution of audio adapted and optimized for rendering and playback in an environment that may differ from the mixing environment. This allows a sound engineer to encode their intentions regarding how the final audio will sound to the listener based on the listener's actual listening environment.

[0032] Conventional channel-based audio codecs operate under the assumption that an audio program is reproduced by a speaker array at a given location for the listener. To produce a complete multi-channel audio program, a sound engineer generally mixes a number of audio streams (e.g., dialogue, music, sound effects) to create an overall desired impression. Audio mixing decisions are generally made by listening to an audio program reproduced by a speaker array at a given location, such as a specific theater's 5.1 or 7.1 system. The final mixed signal serves as the input to the audio codec. In playback, a spatially accurate sound field is achieved only when the speakers are arranged in their given configuration.

[0033] With a new audio coding format called audio object coding, distinguishable sound sources (audio objects) are supplied as input to an encoder in the form of another audio stream. Examples of audio objects include conversation tracks, single instruments, individual sound effect sounds, and their point sources. Each audio object is associated with spatial parameters, including but not limited to sound position, sound width, and velocity information. The audio objects and associated parameters are encoded for distribution and storage. The final audio object mixing and rendering are performed at the receiving end of the audio distribution chain as part of the audio program playback. This step is an audio distribution system that can be customized according to the user's specific listening conditions based on knowledge of the actual speaker positions. The two coding formats, channel-based and object-based, operate optimally under different input signal conditions. Channel-based audio coders are generally more efficient in coding input signals that contain a high-density mixture of different audio sources. Conversely, audio object coders are more efficient in coding a small number of highly directional sound sources.

[0034] In one embodiment, these methods and the components of system 100 have an audio encoding, distribution, and decoding system configured to generate one or more bitstreams that include both conventional channel-based audio elements and audio object coding elements. Such a combined approach provides higher coding efficiency and rendering flexibility compared to either the channel-based approach or the object-based approach.

[0035] Other aspects of the described embodiments include backward-compatibly extending a given channel-based audio codec to include audio object coding elements. A new "extended layer" containing the audio object coding elements is defined and added to the "base" or "backward-compatible" layer of the channel-based audio codec bitstream. This approach enables one or more bitstreams containing the extended layer to be processed by legacy codecs, while providing an enhanced listener experience to users with new decoders. An example of the enhanced listener experience includes control of audio object rendering. An additional advantage of this approach is that audio objects may be added or modified anywhere in the distribution chain without having to decode / mix / recode the multi-channel audio encoded with the channel-based audio codec.

[0036] Regarding the reference frame, the spatial effect of the audio signal is important for providing an immersive experience to the listener. Sounds intended to emanate from a viewing screen or an area of the room should be emitted by speakers placed at the same relative location. Thus, the primary audio metadata of the sound event in the model-based description is the position, but other parameters such as size, direction, velocity, and acoustic dispersion can also be described. To convey the position, a model-based 3D audio space description requires a 3D coordinate system. The coordinate system used for transmission (Euclidean, spherical, etc.) is generally selected for convenience and compactness, but other coordinate systems may be used for the rendering process. In addition to the coordinate system, a reference frame is needed to represent the location of an object in space. The selection of an appropriate reference frame is an important factor for the system to accurately reproduce position-based sound in various different environments. In an other-centered reference frame, the sound source position is defined relative to features in the rendering environment such as the walls and corners of the room, standard speaker locations, screen locations, etc. In an egocentric reference frame, the location is represented from the listener's perspective, such as "in front of me, a little to the left". Chemical studies of spatial perception (audio and others) have shown that the egocentric perspective is most commonly used worldwide. However, in the case of cinema, for several reasons, the other-centered approach is generally more appropriate. For example, when there are relevant objects on the screen, the accurate location of the audio object is most important. Using an other-centered reference, for all listening positions and for any screen size, the sound is localized at the same relative position on the screen, e.g., one-third to the left of the center of the screen. Another reason is that mixers mix in an other-centered way, panning tools are configured in an other-centered frame (the walls of the room), and the mixer expects that this sound is rendered as on-screen, this sound is rendered as off-screen, or from the left wall, etc.

[0037] In a cinema environment, the ego - centric reference frame may be useful and more appropriate regardless of the use of the other - centric reference frame. These include non - diegetic sounds, i.e., sounds that are not part of the "diegetic world", such as mood music, for which an ego - centric uniform presentation is desirable. In other cases, it is the near - field effect that requires ego - centric expression (e.g., a mosquito buzzing in the listener's left ear). Currently, there is no means to render such a sound field without using headphones or very near - field speakers. Also, an infinitely distant sound source (and the resulting plane wave) appears to come from a fixed ego - centric position (e.g., 30° to the left), and such sounds are easier to describe in ego - centric terms than in other - centric terms.

[0038] In some cases, the use of the other - centric reference frame is possible as long as a nominal listening position is defined, while in some examples, ego - centric expressions that cannot yet be rendered are required. The other - centric reference is more useful and appropriate, but the audio representation should be flexible. This is because in certain applications and listening environments, many new features including ego - centric expressions are desirable. Embodiments of an adaptive audio system include a hybrid spatial description approach that uses ego - centric references to render diffused or complex multi - point sources (e.g., stadium crowds, ambiance, etc.) and other - centric model - based sound descriptions for optimal fidelity, including a recommended channel configuration to enhance spatial resolution and scalability.

[0039] System component Referring to FIG. 1, the original sound content data 102 is first processed by the preprocessing block 104. The preprocessing block 104 of the system 100 includes an object channel filtering component. In many cases, an audio object includes an individual sound source and enables independent sound panning. In some cases, such as when producing an audio program using natural or "production" sounds, it is necessary to extract individual sound objects from recordings that include multiple sound sources. Embodiments include methods for separating (isolating) independent source signals from more complex signals. Unwanted elements to be separated from the independent source signals include, but are not limited to, other independent sound sources and background noise. Also, reverberation is removed and a "dry" sound source is restored.

[0040] The preprocessor 104 also includes source separation and content type detection functions. This system provides automatic generation of metadata by analyzing the input audio. The position metadata is obtained from a multi-channel recording by analyzing the relative levels of the corresponding inputs between channel pairs. Detection of content types such as "speech" or "music" can be achieved, for example, by feature extraction and classification.

[0041] Authoring tool The authoring tool block 106 includes features that improve the authoring of audio programs by optimizing the input and encoding of the creative intent of sound engineers and enabling those engineers to realistically generate a final audio mix optimized for playback in any playback environment. This can be achieved by using audio objects and positional data that was encoded in relation to the original audio content. To correctly place sound in the audience area, sound engineers need to control how the sound will ultimately be rendered based on the actual constraints and characteristics of the playback environment. An adaptive audio system provides this control by allowing sound engineers to change the way audio content is designed and mixed using audio objects and positional data.

[0042] An audio object can be thought of as a group of sound elements that are perceived to emanate from a physical location in the audience area. Such objects can be static or moving. In the adaptive audio system 100, audio objects are controlled by metadata. This metadata is, among other things, the details of the position of the sound at a given point in time. When an object is monitored or played back in a theater, it is not necessarily output to a physical channel, but rather is rendered by positional metadata using the speakers that are there. The tracks in a session are audio objects, and standard panning data is similar to positional metadata. In this way, content placed on the screen is panned in much the same way as channel-based content, but content placed in the surround can be rendered to individual speakers if needed. The use of audio objects allows for the desired control of individual effects, but other aspects of a movie sound track function effectively in a channel-based environment. For example, many ambient effects and reverberations are actually benefited by being input into the speaker array. These can be treated as objects that have a width sufficient to fill the array, but it may be good to retain some of the channel-based functionality.

[0043] In one embodiment, an adaptive audio system supports "beds" in addition to audio objects. Here, the beds are effectively channel-based submixes or stems. These can be distributed individually or combined into a single bed according to the intent of the content creator for final playback (rendering). These beds can be generated in different channel-based configurations such as 5.1, 7.1, etc., and can be extended to larger formats such as 9.1 and arrays including overhead speakers.

[0044] FIG. 2 is a diagram showing the combination of channel and object-based data for creating an adaptive audio mix according to one embodiment. As shown in process 200, channel-based data 202 is 5.1 or 7.1 surround sound data provided, for example, in the form of pulse code modulation (PCM) data, but is combined with audio object data 204 to form an adaptive audio mix 208. The audio object data 204 is generated by combining elements of the original channel-based data with associated metadata that defines parameters regarding the location of the audio objects.

[0045] As conceptually shown in FIG. 2, the authoring tool provides the ability to generate an audio program that simultaneously includes a combination of speaker channel groups and object channels. For example, the audio program may optionally include one or more speaker channels organized into a plurality of groups (or tracks such as stereo or 5.1 tracks), descriptive metadata for one or more speaker channels, one or more object channels, and descriptive metadata for one or more object channels. Within an audio program, each speaker channel group and each object channel are represented using one or more different sample rates. For example, a digital cinema (D-cinema) application supports sample rates of 48 kHz and 96 kHz, but may support other sample rates. Furthermore, the capture, storage, and editing of channels having different sample rates can also be supported.

[0046] The generation of an audio program requires steps of sound design. This includes steps of combining sound elements as a sum of level-adjusted constituent sound elements to generate a new desired sound effect. The authoring tool of an adaptive audio system enables the generation of sound effects as a collection of sound objects in relative positions using a spatio-visual sound design graphical user interface. For example, a visual representation of a sounding object (e.g., a car) can be used as a template for assembling an object channel that includes audio elements (exhaust sound, tire sound, engine noise) as sounds and appropriate spatial positions (at the tailpipe, tires, bonnet). Individual object channels can be linked and manipulated as a group. The authoring tool 106 includes a plurality of user interface elements through which a sound engineer can input control information and view mix parameters to improve system functions. Also, the sound design and authoring process is improved by enabling the linking and manipulation of object channels and speaker channels as groups. One example is to combine object channels with individual dry sound sources having a set of speaker channels that include associated reverberation signals.

[0047] The audio mixing tool 106 supports the function of combining a plurality of audio channels, which is generally called mixing. A plurality of mixing methods are supported, including conventional level-based mixing and loudness-based mixing. In level-based mixing, wideband scaling is applied to the audio channels, and the scaled audio channels are added together. The wideband scale factor for each channel is selected to control the absolute level of the resulting mixed signal and the relative levels of the mixed channels in the mixed signal. In loudness-based mixing, one or more input signals are modified using frequency-dependent amplitude scaling. While preserving the perceived timbre of the input sound, frequency-dependent amplitudes are selected to provide the desired perceived absolute loudness and relative loudness.

[0048] The mixing tool can generate speaker channels and speaker channel groups. Thereby, the metadata is associated with each speaker channel group. Each speaker channel group can be tagged according to the content type. The content type can be extended via a text description. The content type includes, but is not limited to, conversation, music, and sound effects. Each speaker channel group is assigned unique instructions regarding how to upmix from one channel configuration to another, where upmixing is defined as the generation of M audio channels out of N channels, where M > N. The upmix instructions include, but are not limited to: an enable / disable flag indicating whether upmixing is permitted; an upmix matrix controlling the mapping between each input and output channel; and Default enabling and matrix settings are assigned based on the content type. For example, upmixing is enabled only for music. Also, each speaker channel group is assigned unique instructions regarding how to downmix from one channel configuration to another, and downmixing is defined as the generation of Y audio channels out of X channels, where Y < X. The downmix instructions include, but are not limited to: a matrix that controls the mapping between each input and output channel; and The default matrix settings can be assigned based on the content type, for example, conversation, and downmixed to the screen; Effects are downmixed from the screen. Each speaker channel is associated with a metadata flag that disables bus management during rendering.

[0049] Embodiments include a function that enables the generation of object channels and object channel groups. According to this invention, metadata is associated with each object channel group. Each object channel group can be tagged according to the content type. The content type can be extended via a text description. Content types include, but are not limited to, conversation, music, and effects. Each object channel group is assigned metadata that describes how the object should be rendered.

[0050] Position information indicating a desired apparent source position is provided. The position can be indicated using an egocentric or an other-centric reference frame. The egocentric reference is appropriate when the source position is referenced to the listener. In the case of an egocentric position, spherical coordinates are useful for the position description. The other-centric reference is the typical reference frame for cinema and other audio / visual presentations, where the source position is referenced to objects in the presentation environment such as the visual display screen or the boundaries of the room. 3D trajectory information is provided for use in other rendering decisions such as enabling interpolation of the position or enabling a "snap to mode". Size information indicating a desired apparent audio source size is provided.

[0051] The limits of the allowed spatial distortion by which spatial quantization is achieved by a "snap to closest speaker" control, indicating the intention by a sound engineer or mixer to render an object by just one speaker (at the expense of some spatial accuracy), can be indicated by the elevation and azimuth tolerance thresholds, above which the "snap" function does not occur. In addition to the distance threshold, a crossfade rate parameter is indicated to control how fast the moving object moves from one speaker to another when the desired position moves between speakers. In one embodiment, for certain location metadata, slave space metadata is used. For example, a "slave" object can be associated with a "master" object that the slave object should follow, and metadata can be automatically generated for the slave object. A time lag or relative speed can be assigned to the slave object. A mechanism can be provided for defining an acoustic center of gravity for a set or group of multiple objects, so that the objects can be rendered such that they are perceived to move around other objects. In such a case, one or more objects rotate around the dry area of a defined area or room, such as an object or a dominant point. The ultimate location information is represented as the location with respect to the room, as opposed to the location with respect to other objects, but the appropriate object-based sound location information is determined using the acoustic center of gravity at the rendering stage.

[0052] When an object is rendered, it is assigned to one or more speakers based on the location metadata and the location of the playback speakers. Additional metadata is associated with the object and limits the speakers used. By using constraints, the use of the indicated speakers can be prohibited, or the indicated speakers can simply be blocked (using less energy to the speakers than when not using the constraints). The set of speakers to be constrained includes, but is not limited to, a designated speaker or speaker zone (e.g., L, C, R, etc.), or a speaker area such as the front wall, rear wall, left wall, right wall, ceiling, floor, indoor speakers, etc. Similarly, in the process of defining the desired mix of multiple sound elements, one or more elements can be made inaudible or "masked" due to the presence of other "masking" sound elements. For example, a masked element can be identified to the user via a graphical display when detected.

[0053] As will be described separately, the audio program description can be adapted to a wide variety of speaker installations and rendering in channel configurations. When authoring an audio program, it is important to monitor the effect of rendering the program in the expected playback configuration to ensure that the desired result is achieved. This invention includes the function of selecting a target playback configuration and monitoring the result. Also, this system can automatically monitor the worst-case (i.e., highest) signal levels generated in each expected playback configuration and provide a display if clipping or limiting occurs.

[0054] Figure 3 is a block diagram showing a workflow for generating, packaging, and rendering adaptive audio content according to one embodiment. The workflow 300 of Figure 3 is divided into three distinguishable task groups labeled generation / authoring, packaging, and egibision and labeling. Generally, it is performed in the same way as most sound design, editing, premixing, and final mixing are done today by the hybrid model of bed and object shown in Figure 2, without adding excessive overhead to this process. In one embodiment, the adaptive audio functionality is provided in the form of software, firmware, or circuitry used with sound production and processing equipment, and such equipment may be a new hardware system or an update to an existing system. For example, a plugin application is provided for a digital audio workstation, and the existing panning methods in sound design and editing may remain unchanged. In this way, it is possible to place both the bed and the object within a workstation in a 5.1 or similar surround-capable editing room. Object audio and metadata are recorded in a dubbing theater in a session preparing for the premix and final mix stages.

[0055] As shown in FIG. 3, the production or authoring task includes an input by a user, e.g., a mixing control 302 by a sound engineer in the following example, to a mixing console or an audio workstation 304. In one embodiment, the metadata is integrated into the mixing console surface, whereby the faders, panning, and audio processing of the channel strip can cooperate with the bed or stem and the audio object. The metadata can be edited using either the console surface or the workstation user interface, and the sound is monitored using a rendering and mastering unit (RMU) 306. The audio data and associated metadata of the bed and the object are recorded during the mastering session to generate a "print master". This print master includes an adaptive audio mix 310 and any other rendered derivatives (such as a surround 7.1 or 5.1 theater mix) 308. Using existing authoring tools (e.g., digital audio workstations such as ProTools), a sound engineer can label individual audio tracks during a mixing session. Embodiments extend this concept by allowing a user to label individual sub - segments within a track to assist in the discovery or quick identification of audio elements. The user interface to the mixing console that enables the definition or generation of metadata can be implemented by graphical user interface elements, physical controls (e.g., sliders and knobs), or any combination thereof.

[0056] During the packaging stage, the print master file is wrapped, hashed, and optionally encrypted using industry-standard MXF wrapping procedures to ensure the integrity of the audio content before being sent to a digital cinema packaging facility. This step can be performed by a digital cinema processor (DCP) 312 or any suitable audio processor according to the final playback environment, such as the standard surround sound provided in theater 318, an adaptive audio-capable theater 320, or other playback environments. As shown in FIG. 3, processor 312 outputs appropriate audio signals 314 and 316 according to the egizibition environment.

[0057] In one embodiment, an adaptive audio print master includes an adaptive audio mix along with a standard DCI-compliant pulse code modulation (PCM) mix. The PCM mix can be rendered by a dubbing theater's rendering and mastering unit or generated by a separate mix path as needed. The PCM audio forms a standard main audio track file in digital cinema processor 312, and the adaptive audio forms an additional track file. Such track files comply with existing industry standards and are ignored by DCI-compliant servers that cannot use them.

[0058] In an example of a cinema playback environment, a DCP containing an adaptive audio track file is recognized as a valid package by a server, imported into the server, and streamed to an adaptive audio cinema processor. A system that can utilize both linear PCM and adaptive audio files can switch between them as needed. For distribution to the egizibition stage, a single type of package can be distributed to the cinema using an adaptive audio packaging scheme. The DCP package includes both PCM and adaptive audio files. Incorporating the use of security keys such as a key delivery message (KDM) enables secure delivery of movie content and other similar content.

[0059] As shown in FIG. 3, an adaptive audio methodology is realized by enabling a sound engineer to express his or her intentions regarding the rendering and playback of audio content using the audio workstation 304. By controlling the input controls, the engineer can specify where and how audio objects and sound elements are played according to the listening environment. Metadata is generated in the audio workstation 304 according to the engineer's mixing input 302, controls spatial parameters (e.g., position, velocity, intensity, timbre, etc.), and provides a rendering queue that specifies which speakers or speaker groups in the listening environment play each sound during exhibition. The metadata is associated with the respective audio data in the workstation 304 or the RMU 306 for packaging and transmission by the DCP 312.

[0060] The graphical user interface and software tools that provide control of the workstation 304 by the engineer have at least a part of the authoring tool 106 of FIG. 1.

[0061] Hybrid audio codec As shown in FIG. 1, the processor 100 includes a hybrid audio codec 108. This component has an audio encoding, distribution, and decoding system configured to generate a single bitstream that includes both conventional channel-based audio elements and audio object coding elements. The hybrid audio coding system is centered around a channel-based coding system configured to generate a single (integrated) bitstream that is simultaneously compatible (i.e., decodable by) with a first decoder configured to decode audio data encoded by a first encoding protocol (channel-based) and one or more second decoders configured to decode audio data encoded by one or more second encoding protocols (object-based). The bitstream can include both encoded data (in the form of data bursts) decodable by the first decoder (ignored by any of the second decoders) and encoded data (e.g., other data bursts) decodable by one or more of the second decoders (ignored by the first decoder). The decoded audio and associated information (metadata) from the first and one or more second decoders can be combined to generate an environment, channels, spatial information, and object facsimiles in which both channel-based and object-based information are simultaneously rendered and provided to the hybrid coding system (i.e., in a 3D space or listening environment).

[0062] Codec 108 generates a bitstream that includes the encoded audio information and information regarding a plurality of sets of channel positions (speakers). In one embodiment, one set of channel positions is fixed and is used for a channel-based encoding protocol, while another set of channel positions is adaptive and is used for an audio object-based encoding protocol such that the channel configuration of the audio object may vary as a function of time (depending on where in the sound field the object is placed). Thus, the hybrid audio coding system carries information regarding two sets of speaker locations for playback, one set being fixed and possibly a subset of the other set. Devices that support legacy-encoded audio information decode and render the audio information from the fixed subset, while devices that can support the larger set decode and render additional encoded audio information that is time-dependently assigned to different speakers from that larger set. Further, this system is independent of a first and one or more second decoders that may be present simultaneously within the system and / or device. Thus, legacy and / or existing devices / systems that include only decoders that support the first protocol create a fully compatible sound field that is rendered via a conventional channel-based playback system. In this case, the unknown or unsupported portions of the hybrid bitstream protocol (i.e., the audio information represented by the second encoding protocol) are ignored by a system or device decoder that supports the first hybrid encoding protocol.

[0063] In another embodiment, codec 108 is configured to operate in a mode in which a first encoding subsystem (supporting a first protocol) includes a combined representation of all sound field information (channels and objects) represented by both a first and one or more second encoders in a hybrid encoder. Thereby, an audio object (typically borne by one or more second encoder protocols) is represented and rendered in a decoder that supports only the first protocol, so that the hybrid bitstream is backward compatible with a decoder that supports only the protocol of the first encoder subsystem.

[0064] In yet another embodiment, codec 108 includes two or more encoding subsystems. Each of these subsystems is configured to encode audio data according to a different protocol, and the outputs of the multiple subsystems are combined to generate an (integrated) bitstream in a hybrid format.

[0065] One advantage of this embodiment is that a hybrid-encoded audio bitstream can be carried in a wide range of content delivery systems that each conventionally support only data encoded according to a first encoding protocol. Thereby, no modification / change of the system and / or transport-level protocol is required to support the hybrid coding system.

[0066] An audio encoding system generally utilizes standardized bitstream elements to enable the transmission of additional (optional) data within the bitstream itself. This additional (optional) data is skipped (i.e., ignored) during the decoding of the encoded audio contained in the bitstream but is used for purposes other than decoding. Different audio coding standards represent these additional data fields using unique terminologies. This general type of bitstream element includes, but is not limited to, auxiliary data, skip fields, data stream elements, fill elements, ancillary data, and substream elements. Unless otherwise specified, the use of the expression "auxiliary data" in this document should be interpreted as a general expression that includes some or all of the embodiments related to the present invention, rather than suggesting a particular type or format of additional data.

[0067] A data channel enabled via the "auxiliary" bitstream of a first encoding protocol in a combined hybrid coding system bitstream may carry one or more (independent or dependent) audio bitstreams (encoded with one or more second encoding protocols). The one or more second audio bitstreams can be divided into N sample blocks and multiplexed into the "auxiliary data" field of the first bitstream. The first bitstream is decoded by an appropriate (compliant) decoder. Also, the auxiliary data of the first bitstream is extracted, recombined into the one or more second audio bitstreams, decoded by a processor that supports the syntax of the one or more second bitstreams, and combined and rendered together or independently. Furthermore, it is possible to reverse the roles of the first and second bitstreams such that blocks of data of the first bitstream are multiplexed into the auxiliary data of the second bitstream.

[0068] The bitstream elements associated with the second encoding protocol also carry and convey the information (metadata) characteristics of the underlying audio. The information includes, but is not limited to, the desired sound source position, velocity, and size. This metadata is utilized in the decoding and rendering processes to reproduce the appropriate (i.e., original) positions of the associated audio objects carried in the adaptable bitstream. Also, the above metadata is applicable to the audio objects included in one or more second bitstreams in the hybrid stream and can also carry this within the bitstream elements associated with the first encoding protocol.

[0069] The bitstream elements associated with either or both of the first and second encoding protocols of the hybrid coding system carry / convey context metadata that identifies the spatial parameters (i.e., the essence of the signal properties themselves) and yet another type of information that describes the underlying audio essence type in the form of the audio classes carried in the hybrid-coded audio bitstream. Such metadata indicates the presence of, for example, spoken conversations, music, conversations over music, applause, singing voices, etc., and can be used to adaptively modify the interconnected pre- or post-processing modules upstream or downstream of the hybrid coding system.

[0070] In one embodiment, codec 108 is configured to operate on a shared or common bit pool in which the bits available for coding are "shared" among all or part of the encoding subsystems that support one or more protocols. Such a codec describes the bits available (from a common "shared" bit pool) among the encoding subsystems in order to optimize the overall audio quality of the integrated bitstream. For example, during a first time interval, the codec may allocate more available bits to a first encoding subsystem and fewer available bits to the remaining subsystems, and during a second time interval, the codec may allocate fewer available bits to the first encoding subsystem and more available bits to the remaining subsystems. The decision of how to allocate bits among the encoding subsystems depends, for example, on the statistical analysis of the shared bit pool and / or the analysis of the audio content encoded by each subsystem. The codec allocates bits from the shared pool such that the integrated bitstream, formed by multiplexing the outputs of the encoding subsystems, maintains a constant frame length / bit rate over a specified time interval. Also, in some cases, it is possible for the frame length / bit rate of the integrated bitstream to vary over a specified time interval.

[0071] In another embodiment, codec 108 generates an integrated bitstream that includes data encoded by a first encoding protocol that is transmitted as an independent substream of an encoded data stream (decoded by a decoder that supports the first encoding protocol), and data encoded by a second protocol that is sent as an independent or dependent substream of the encoded data stream (ignored by a decoder that supports the first protocol). More generally, in one class of embodiments, the codec generates an integrated bitstream that includes two or more independent or dependent substreams (each substream containing data encoded by a different or the same encoding protocol).

[0072] In yet another embodiment, codec 108 generates an integrated bitstream that includes data encoded by a first encoding protocol (decoded by a decoder that supports a first encoding protocol associated with a unique bitstream identifier) that is transmitted and composed of a unique bitstream identifier, and data that is transmitted and composed of a unique bitstream identifier and is encoded by a second protocol that is ignored by a decoder that supports the first protocol. More generally, in one class of embodiments, the codec generates an integrated bitstream that includes two or more sub-streams (each sub-stream includes data encoded by a different or the same encoding protocol, and each bears a unique bitstream identifier). The method and system for generating the integrated bitstream described above provides a function for unambiguously signaling (to the decoder) which interleaving and / or protocol was used in the hybrid bitstream (e.g., signaling whether to use AUX data, SKIP, DSE, or a sub-stream approach).

[0073] The hybrid coding system is configured to support de-interleaving / de-multiplexing of a bitstream that supports one or more second protocols and re-interleaving / re-multiplexing into a first bitstream (that supports the first protocol) at processing points found across the media delivery system. Also, the hybrid codec is configured to encode audio input streams of different sample rates into a bitstream. Thereby, a means is provided for efficiently coding and distributing sound sources that include signals having inherently different bandwidths. For example, dialogue tracks generally have an inherently lower bandwidth than music or effect tracks. Rendering In one embodiment, the adaptive audio system can package a plurality of (e.g., up to 128) tracks, typically as a combination of a bed and objects. The basic format of the audio data for the adaptive audio system includes a plurality of independent monophonic audio streams. Each stream has associated metadata that defines whether the stream is a channel-based stream or an object-based stream. The channel-based stream has rendering information encoded by a channel name or label. The object-based stream has location information encoded by a mathematical formula encoded in another associated metadata. The original plurality of independent audio streams are packaged as a single serial bitstream that includes all of the ordered audio data. This adaptive data configuration causes the sound to be rendered in an other-centered reference frame. The final rendering location of the sound is based on the playback environment so as to correspond to the intention of the mixer. Thus, the sound can be defined to emanate from a reference frame of the room being played (e.g., the center of the left wall) rather than from a labeled speaker or speaker group (e.g., left surround). The object position metadata includes the appropriate other-centered reference frame information necessary to correctly reproduce the sound, using the available speaker positions in the room configured to play the adaptive audio content.

[0074] The renderer takes a bitstream encoding an audio track and processes its content according to the signal type. The bed is sent to the array. The array may require latency and equalization processing potentially different from that of the objects here. This process supports the rendering of these beds and objects to multiple (up to 64) speaker outputs. FIG. 4 is a block diagram showing the rendering stage of an adaptive audio system according to one embodiment. As shown in system 400 of FIG. 4, multiple input signals such as up to 128 audio tracks having an adaptive audio signal 402 are provided by components in the production, authoring, and packaging stages of system 300 such as RMU 306 and processor 312. These signals include channel-based beds and objects utilized by renderer 404. The channel-based audio (beds) and objects are input to a level manager 406 that controls the output levels or amplitudes of different audio components. Certain audio components are processed by an array correction component 408. The adaptive audio signal is passed through a B-chain processing component 410. The B-chain processing component 410 generates multiple (e.g., up to 64) speaker feed output signals. Generally, the B-chain feed refers to signals processed by a power amplifier, crossover, and speakers for the A-chain content that makes up the sound track of a film stock.

[0075] In one embodiment, the renderer 404 executes a rendering algorithm that intelligently uses the theater's surround speakers to their fullest capabilities. By improving the power handling and frequency response of the surround speakers and maintaining the same monitoring reference level for each output channel or speaker of the theater, objects panned between the screen and the surround speakers can maintain their sound pressure levels and better match timbres without raising the overall sound pressure level of the theater. An appropriately defined array of surround speakers generally has sufficient headroom to reproduce the maximum dynamic range available in a surround 7.1 or 5.1 sound track (i.e., 20 dB above the reference level), but it is less likely that a single surround speaker will have the same headroom as a large multi-way screen speaker. As a result, objects placed in the surround field may require a greater sound pressure than can be obtained using a single surround speaker. In these cases, the renderer disperses the sound across an appropriate number of speakers to achieve the required sound pressure level. The adaptive audio system improves the quality and power handling of the surround speakers to improve the fidelity of the rendering. The adaptive audio system supports surround speaker bus management through the use of optional rear subwoofers such that each surround speaker can achieve improved power handling and optionally use a smaller speaker cabinet. Also, side surround speakers are added closer to the screen than current practice to ensure that objects smoothly transition from the screen to the surround.

[0076] By using metadata that defines the location information of audio objects in the rendering process, system 400 provides a comprehensive and flexible way for content creators to move beyond existing systems. As described above, current systems generate and distribute audio fixed to a certain speaker location with only limited knowledge about the type of content carried by the audio essence (the part of the audio being played). Adaptive audio system 100 provides a new hybrid approach that includes options for both specific speaker location audio (left channel, right channel, etc.) and object-oriented audio elements with generalized spatial information including, but not limited to, size and velocity. This hybrid approach provides a balanced approach between the fidelity provided by fixed speaker locations and the flexibility in rendering of generalized audio objects. Also, this system provides additional useful information about the audio content paired with the audio essence by the content creator during content production. This information provides powerful and detailed information about the attributes of the audio that can be used in a very powerful way during rendering. Such attributes include, but are not limited to, content type (conversation, music, effects, Foley, background / ambient, etc.), spatial attributes (3D position, 3D size, velocity), rendering information (snap to speaker location, channel weight, gain, bus management information, etc.).

[0077] The adaptive audio system described herein provides powerful information that can be used to render a widely variable number of endpoints. Often, the optimal rendering method applied depends heavily on the endpoint device. For example, a home theater system and a sound bar may have 2, 3, 5, 7, or 9 separate speakers. Many other types of systems, such as televisions, computers, and music docks, have only two speakers, and almost all commonly used devices (PCs, laptops, tablets, mobile phones, music players, etc.) have binaural headphone outputs. However, in the case of conventional audio (mono, stereo, 5.1, 7.1 channels) sold today, the endpoint device often has to make a simplified decision and compromise on rendering and playing the audio delivered in a specific channel / speaker format. Also, there is little or no information carried regarding the actual content being delivered (conversation, music, ambient, etc.), and little or no information regarding the intent of the content creator for audio playback. However, the adaptive audio system 100 provides this information and potentially access to audio objects. Using these, a next-generation user experience that attracts people can be created.

[0078] With system 100, content creators can incorporate the spatial intent of a mix into a bitstream using unique and powerful metadata and an adaptive audio transmission format, such as metadata for position, size, velocity, etc. This enables great flexibility in the spatial playback of audio. From the perspective of spatial rendering, with adaptive audio, the mix can be adapted to the exact position of the speakers in the room to avoid spatial distortion that occurs when the geometry of the playback system is not the same as that of the authoring system. In current audio playback systems where only one speaker channel of audio is sent, the intent of the content creator is unknown. System 100 uses the metadata conveyed in the production and distribution pipeline. An adaptive audio-aware playback system uses this metadata information to play the content in accordance with the original intent of the content creator. Similarly, the mix can be adapted to the exact hardware configuration of the playback system. Currently, there are many different speaker configurations and types in rendering devices such as televisions, home theaters, soundbars, portable music player docks, etc. These systems today have to process the audio and appropriately match it to the capabilities of the rendering device when sending specific channel audio information (i.e., left and right channel audio or multi-channel audio). An example is when standard stereo audio is sent to a soundbar with more than three speakers. In current audio playback where only one speaker channel of audio is sent, the intent of the content creator is unknown. By using the metadata conveyed by the production and distribution pipeline, an adaptive audio-aware playback system uses this information to play the content to match the original intent of the content creator. For example, a certain soundbar has side-firing speakers that create an immersive feeling. With adaptive audio, spatial information and content type (such as ambient effects) can be used by the soundbar to send only the appropriate audio to these side-firing speakers.

[0079] With an adaptive audio system, unlimited interpolation can be performed in all front / back, left / right, up / down, near / far dimensions. In current audio playback systems, there is no information on how to process audio that is desirably placed so that the listener feels as if they are between two speakers. Currently, for audio assigned to only a specific speaker, a spatial quantization factor is introduced. With adaptive audio, the spatial positioning of the audio is accurately known and can be played appropriately in the audio playback system.

[0080] Regarding headphone rendering, the creator's intention is realized by matching the head-related transfer function (HRTF) to the spatial position. When audio is played back on headphones, spatial virtualization can be realized by applying the head-related transfer function. The head-related transfer function processes the audio and adds perceptual cues that make it feel as if the audio is being emitted in 3D space rather than through headphones. The accuracy of spatial playback depends on the selection of an appropriate HRTF that can vary based on multiple factors including the spatial position. Using the spatial information provided by an adaptive audio system, one or a continuously varying number of HRTFs are selected, significantly improving the playback experience.

[0081] The spatial information conveyed by an adaptive audio system is used not only by content creators to generate engaging entertainment experiences (such as films, television, music, etc.), but can also indicate to the listener where they are located relative to physical objects such as buildings and points of geographical interest. This allows the user to interact with a virtualized audio experience related to the real world, i.e., augmented reality.

[0082] Embodiments enable spatial upmixing by performing enhanced upmixing by reading metadata only when object audio data is unavailable. By knowing the positions and types of all objects, an upmixer can differentiate elements in a channel-based track. Existing upmixing algorithms must infer information such as audio content type, and the positions of different elements in an audio stream, and generate a high-quality upmix with minimal or no audible artifacts. Often, the inferred information is inaccurate or inappropriate. In adaptive audio, additional information obtained from metadata regarding audio content type, spatial position, velocity, audio object size, etc., can be used by an upmixing algorithm to generate a high-quality playback result. Also, the system spatially matches audio to video by accurately positioning audio objects on the screen relative to visual elements. In this case, when the spatial location where an audio element is played matches an image element on the screen, an engaging audio / video playback experience is possible, especially when the screen size is large. An example is spatially aligning conversations in a film or television program with the person or character speaking on the screen. In normal speaker channel-based audio, there is no easy way to determine where a conversation should be spatially placed to match the location of a person or character on the screen. Such audio / visual alignment can be achieved using the audio information available in adaptive audio. Also, visual position and audio spatial alignment can be used for non-character / conversation objects such as automobiles, trucks, animations, etc.

[0083] Spatial masking processing is facilitated by system 100. This is because knowledge of the spatial intent of the mix with adaptive audio metadata means that the mix can be adapted to any speaker configuration. However, due to the constraints of the playback system, a person risks downmixing objects at the same location or almost the same location. For example, if there is no surround channel, an object intended to be panned to the left rear is downmixed to the left front, and if there is an element with a louder volume at the left front at the same time, the downmixed object is masked and disappears from the mix. Using adaptive audio metadata, spatial masking is scheduled by the renderer and the spatial and / or loudness downmix parameters of each object are adjusted, so that all audio elements of the mix remain as perceivable as the original mix. Since the renderer understands the spatial relationship between the mix and the playback system, instead of generating a phantom image between two or more speakers, it has a function of "snapping" the object to the nearest speaker. This slightly distorts the spatial representation of the mix, but can avoid phantom images not intended by the renderer. For example, if the angular position of the left speaker at the mixing stage does not correspond to the angular position of the left speaker of the playback system, the snap function to the nearest speaker can be used to avoid playing a certain phantom image of the left channel at the mixing stage on the playback system.

[0084] Regarding content processing, with the adaptive audio system 100, content creators can generate individual audio objects, add information regarding the content, and convey it to the playback system. This increases the flexibility of audio processing before playback. From the perspective of content processing and rendering, it is possible to adapt the processing to the object correspondence by the adaptive audio system. For example, conversation enhancement can be applied only to conversation objects. Conversation enhancement refers to a method of processing audio including conversation so that the audibility and / or clarity of the conversation is high and / or improved. In many cases, the audio processing applied to conversation is inappropriate for non-conversation audio content (i.e., music, ambient effects, etc.), and undesirable audible artifacts may occur. In adaptive audio, the audio object can be appropriately labeled so that a rendering solution selectively applies conversation enhancement only to conversation content when a content contains only conversation. Also, when the audio object is only conversation (and not a mixture of conversation and other content as is often the case), the conversation enhancement processing can process the conversation exclusively (thereby restricting the processing of other content). Similarly, bus management (filtering, attenuation, gain) can be targeted to objects based on their types. Bus management refers to selectively isolating and processing only the bus (the following) frequencies in a content. In current audio systems and distribution mechanisms, this is a "blind" process applied to all audio. In adaptive audio, audio objects for which bus management is appropriate can be identified by metadata and the rendering process can be appropriately applied.

[0085] In addition, the adaptive audio system 100 provides object-based dynamic range correction and selective upmixing. Conventional audio tracks have the same length as the content itself, but audio objects may occur only for a limited time within the content. The metadata associated with an object includes information about its average, peak signal amplitude, and its start or attack time (especially in the case of transition material). This information allows the compressor to adapt its compression and time constants (attack, release, etc.) to match the content. In the case of selective upmixing, the content creator may choose to indicate in the adaptive audio bitstream whether an object should be upmixed. This information allows the adaptive audio renderer and upmixer to determine which audio elements can be safely upmixed while respecting the creator's intent.

[0086] Also, according to an embodiment, the adaptive audio system can select a preferred rendering algorithm from a plurality of available rendering algorithms and / or surround sound formats. Examples of available rendering algorithms include binaural, stereo dipole, Ambisonics, Wave Field Synthesis (WFS), multi-channel panning, raw stems with position metadata, etc. Others include dual balance and vector-based amplitude panning.

[0087] The binaural delivery format uses a two-channel representation of the sound field with respect to the signals at the left and right ears. Binaural information can be generated by in-ear recording or synthesized using an HRTF model. Reproduction of the binaural representation is generally done using headphones or by using cross-talk cancellation. Reproduction with any speaker setup requires signal analysis to determine the associated sound field and / or signal source.

[0088] The stereo dipole rendering method is a transaural crosstalk cancellation process that enables a binaural signal to be reproduced by stereo speakers (e.g., at ±10° away from the center).

[0089] Ambisonics is a 4-channel encoded (distribution format and rendering method) called the B-format. The first channel W is an omnidirectional pressure signal; the second channel X is a directional pressure gradient containing front and back information; the third channel Y contains left and right, and Z contains up and down. These channels define the primary samples of the complete sound field at a point. Ambisonics uses all available speakers to reproduce the sampled (or synthesized) sound field within the speaker array such that when some speakers are pushing, other speakers are pulling.

[0090] Wave Field Synthesis is a rendering method for sound reproduction based on the accurate construction of a desired wave field by secondary sources. WFS is based on the principle of Huygens and is implemented as an array of (dozens or hundreds of) speakers that operate in a controlled phase-controlled manner to surround the listening space and reproduce each individual sound wave.

[0091] Multi-channel panning is a distribution format and / or rendering method, also known as channel-based audio. In this case, the sound is represented as the same number of individual sources by a plurality of speakers at a defined angle from the listener. The content creator / mixer can generate a virtual image by panning the signal between adjacent channels to give a direction cue; mix early reflections, reverberation, etc. into many channels to provide a direction and environmental cue.

[0092] Raw stems with position metadata is a delivery format, also known as object-based audio. In this format, distinguishable sound sources from "nearby microphones" are represented by position and environmental metadata. Based on the metadata, the playback device, and the listening environment, virtual sources are rendered.

[0093] The adaptive audio format is a hybrid of the multichannel panning format and the raw stems format. The rendering method of this embodiment is multichannel panning. For audio channels, rendering (panning) is performed during authoring, while for objects, rendering (panning) is performed during playback.

[0094] Metadata and Adaptive Audio Transmission Format As described above, metadata is generated during the production stage, encodes the position information of audio objects, and assists in the rendering of audio programs by accompanying the audio programs. Specifically, it describes the audio programs to enable the rendering of audio programs in a wide range of playback devices and playback environments. Metadata is generated for a given program and for the picture and mixer that generate, collect, edit, and manipulate that audio during post-production. An important feature of the adaptive audio format is the ability to control how audio is translated to different playback systems and environments from the mixing environment. Specifically, a certain cinema may have less functionality than the mixing environment.

[0095] An adaptive audio renderer is designed to make maximum use of the devices available for reproducing the intentions of the mixer. Further, with the adaptive audio authoring tools, the mixer can preview and adjust how the mix is rendered in various playback configurations. All metadata values can be conditioned on the playback environment and speaker configuration. For example, different mix levels for an audio element can be defined based on the playback configuration or mode. In one embodiment, the list of conditional playback modes is extensible and includes (1) channel-based only playback: 5.1, 7.1, 7.1 (height), 9.1; (2) individual speaker playback: 3D, 2D (no height).

[0096] In one embodiment, the metadata controls or governs different aspects of the adaptive audio content and is organized based on different types including program metadata, audio metadata, and rendering metadata (for channels and objects). Each type of metadata includes one or more metadata items that give values for the characteristics referenced by an identifier (ID). FIG. 5 is a table listing the metadata types and related metadata elements of an adaptive audio system according to one embodiment.

[0097] As shown in Table 500 of FIG. 5, the first type of metadata is program metadata. This includes a plurality of metadata elements that define the frame rate, track sub-clusters, and an extensible channel description, and a mix stage description. The frame rate metadata element defines the frame rate of the audio content in units of frames per second (fps). Raw audio formats do not include the framing of the audio or metadata. This is because the audio is supplied as a full track (the length of the reel or the entire feature) rather than an audio segment (the length of the object). The raw format needs to carry all the information necessary for the adaptive audio encoder to be able to frame the audio and metadata including the actual frame rate. Table 1 shows the ID, value examples, and descriptions of the frame rate metadata element.

[0098]

Table 1

[0099]

Table 2

Table 3

[0100]

Table 4

[0101] As shown in FIG. 3, the audio metadata includes sample elements, bit depth, and coding system. Table 5 shows the ID, value example, and description of the sample rate metadata element.

[0102] [Table 5] Table 6 shows the ID, value example, and description of the bit depth metadata element (for PCM and lossless compression cases).

[0103] [Table 6] Table 7 shows the ID, example values, and descriptions of the coding system metadata elements.

[0104] [Table 7] As shown in FIG. 5, the third type of metadata is rendering metadata. The rendering metadata defines values that help the renderer match as closely as possible the original mixer's intent regardless of the playback environment. A set of metadata elements differs between channel - based audio and object - based audio. The first rendering metadata field selects between two types of audio, namely channel - based or object - based, as shown in Table 8.

[0105] [Table 8] The rendering metadata for channel - based audio includes position metadata elements that define the audio source position as one or more speaker positions. Table 9 shows the IDs and values of the position metadata elements for the channel - based case.

[0106] [Table 9] Also, the rendering metadata for channel - based audio includes rendering control elements that define characteristics related to the playback of channel - based audio, as shown in Table 10.

[0107] [Table 10] In the case of object-based audio, the metadata contains elements similar to those of channel-based audio. Table 11 gives the IDs and values of the object position metadata elements. The object position is described in one of three ways: three-dimensional coordinates, a plane and two-dimensional coordinates, or a line and one-dimensional coordinates. The rendering method can be adapted based on the position information type.

[0108] [Table 11] Table 12 shows the IDs and values of the object rendering control metadata elements. These values provide additional means to control and optimize the rendering of object-based audio.

[0109] [Table 12-1] [Table 12-2] In one embodiment, the above metadata shown in FIG. 5 is generated and stored as one or more files related to or indexed to the corresponding audio content such that the audio stream is processed by interpreting the metadata generated by the mixer by an adaptive audio system. Note that the above metadata is an example of IDs, values, and definitions, and other or additional metadata elements may be included for use in an adaptive audio system.

[0110] In one embodiment, two (or more) sets of metadata elements are associated with each of a channel and an object-based audio stream. The first set of metadata is applied to a plurality of audio streams under a first condition of the playback environment, and the second set of metadata is applied to a plurality of audio streams under a second condition of the playback environment. Based on the conditions of the playback environment, for a given audio stream, the second or subsequent set of metadata elements replaces the first set of metadata elements. The conditions include room size, shape, indoor material composition, whether there are people in the room and their density, ambient noise characteristics, ambient light characteristics, and other factors that affect the sound or even the mood of the playback environment.

[0111] Post-production and mastering The rendering stage 110 of the adaptive audio processing system 100 includes post-production steps leading to the generation of a final mix. In a cinema application, the three main categories of sound used in a movie mix are dialogue, music, and effects. Effects consist of sounds that are not dialogue or music (e.g., ambient noise, background / scene noise). Sound effects may be recorded or synthesized by a sound designer, or may be sources obtained from an effects library. A subgroup of effects that includes specific noise sources (e.g., footsteps, doors, etc.) is known as Foley and is performed by Foley artists. Different types of sounds are appropriately marked and panned by a recording engineer.

[0112] FIG. 6 is a diagram showing an example of a post-production workflow of an adaptive audio system according to an embodiment. As shown in FIG. 600, all individual sound components of music, conversation, Foley, and effects are gathered in a final mix 606 at a dubbing theater. A re-recording mixer 604 uses a premix (also known as "mix minus") with individual sound objects and position data to generate stems, for example, as conversation, music, effects, Foley, and background sounds. In addition to forming the final mix 606, the music and all effect stems can be used as a basis for generating a dubbed-language version of the movie. Each stem consists of a channel-based bed and a plurality of audio objects with metadata. The stems are combined to form the final mix. Using object panning information from both an audio workstation and a mixing console, a rendering and mastering unit 608 renders the audio to the speaker locations of the dubbing theater. This rendering allows the mixer to hear how the channel-based bed and audio objects are combined and also provides the ability to render to different configurations. The mixer can control how the content is rendered to the surround channels using conditional metadata that defaults to the relevant profile. In this way, the mixer retains full control over how the movie is played back in all scalable environments. A monitoring step is included after either or both of the re-recording step 604 and the final mix step 606, allowing the mixer to listen to and evaluate the intermediate content generated at each of these steps.

[0113] During the mastering session, stems, objects, and metadata are gathered into an adaptive audio package 614. The adaptive audio package 614 is generated by the print master 610. This package also includes a backward compatible (legacy 5.1 or 7.1) surround sound theatrical mix 612. The rendering / mastering unit (RMU) 608 can render this output as needed. Thereby, additional workflow steps are not required in the generation of existing channel-based deliverables. In one embodiment, the audio files are packaged using standard material exchange format (MXF) wrapping. Also, using an adaptive audio mix master file, other deliverables such as consumer multi-channel mixes and stereo mixes can be generated. Controlled rendering can be done, which can significantly reduce the time required to generate such mixes by intelligent profiles and conditional metadata.

[0114] In one embodiment, a digital cinema package is generated for deliverables including adaptive audio mix using a packaging system. The audio track files are locked together to help prevent synchronization errors with the adaptive audio track files. In some fields, during the packaging phase, additional track files are required, such as the addition of a hearing impaired (HI) or visually impaired narrator (VI-N) track to the main audio track file.

[0115] In one embodiment, the speaker array of the playback environment may include any number of surround sound speakers arranged and designed according to established surround sound standards. Any number of additional speakers for the accurate rendering of object-based audio content may be arranged based on the conditions of the playback environment. These additional speakers are set up by a sound engineer, and this setup is provided to the system in the form of a setup file used by the system to render the object-based components of the adaptive audio to one or more speakers throughout the speaker array. The setup file includes at least a speaker designation, a mapping of channels to individual speakers, information regarding speaker groups, and a list of runtime mappings based on the relative positions of the speakers in the playback environment. The runtime mappings are utilized by the system's snap-to function to render point-source object-based audio content to the speaker closest to the perceived location of the sound intended by the sound engineer.

[0116] FIG. 7 is a diagram showing an example of a workflow of a digital cinema packaging process using an adaptive audio file according to an embodiment. As shown in FIG. 700, an audio file including both an adaptive audio file and a 5.1 or 7.1 surround sound audio file is input to a wrapping / encryption block 704. In one embodiment, when generating a digital cinema package in block 706, a PCM MXF file (with appropriate additional tracks added) is encrypted using the SMPTE specification according to existing practices. The adaptive audio MXF is packaged as an auxiliary track file and optionally encrypted using a symmetric content key according to the SMPTE specification. This single DCP 708 is sent to a Digital Cinema Initiative (DCI)-compliant server. Generally, inappropriate instructions simply ignore additional track files including the adaptive audio sound track and use the existing main audio track file for standard playback. Instructions with a suitable adaptive audio processor can, if applicable, revert to the standard audio track as needed and receive and play the adaptive audio sound track. Also, the wrapping / encryption component 704 inputs directly to a distribution KDM block 710 to generate an appropriate security key for use in a digital cinema server. Other movie elements or files such as subtitles 714 and images 716 are wrapped and encrypted together with the audio file 702. In this case, for the image file 716, processing steps such as compression 712 are included.

[0117] Regarding content management, with the adaptive audio system 100, content creators can generate individual audio objects, add information about the content, and convey it to the playback system. This provides great flexibility in audio content management. From the perspective of content management, the adaptive audio method enables multiple different functions. These include changing the language of the content simply by replacing the dialogue object for purposes such as space saving, download efficiency, and geographical playback adaptation. Films, television, and other entertainment programs are generally distributed internationally. This often requires changing the language in the content depending on where it is played (e.g., French for a film shown in France, German for a TV program shown in Germany, etc.). For this reason, today, completely independent audio sound tracks are generated, packaged, and distributed. With the original concept of adaptive audio and audio objects, the dialogue of the content can be an independent audio object. This allows the language of the content to be easily changed without updating or changing other elements of the audio sound track such as music and effects. This applies not only to foreign languages but also to inappropriate words for a certain audience (e.g., children's TV shows, movies for airlines, etc.), targeted advertisements, etc.

[0118] Installation and equipment consideration An adaptive audio file format and associated processors can change how theater equipment is installed, calibrated, and maintained. With the introduction of more potential speaker outputs, in one embodiment, the adaptive audio system uses an optimized 1 / 12 octave band equalization engine. To more accurately balance the sound of the theater, it can process up to 64 outputs. Also, the system can perform scheduled monitoring of individual speaker outputs from the cinema processor output through the sound reproduced in the auditorium. Local or network alerts can be generated so that appropriate actions are taken. A flexible rendering system can automatically exclude failed speakers or amplifiers from the playback chain and render around them to continue the show.

[0119] The cinema processor can connect to a digital cinema server with existing 8×AES main audio connections and Ethernet® connections to stream adaptive audio data. For playback of surround 7.1 or 5.1 connections, existing PCM connections are used. Adaptive audio data is streamed to the cinema processor via Ethernet® for decoding and rendering, and the connection between the server and the cinema processor allows the audio to be identified and synchronized. If any problems occur during the playback of adaptive audio tracks, the sound is reverted to Dolby Surround 7.1 or 5.1 PCM audio.

[0120] Although embodiments have been described with respect to 5.1 and 7.1 surround sound systems, it should be noted that many other current and future surround configurations, including 9.1, 11.1, and 13.1 and beyond, can be used with the embodiments.

[0121] An adaptive audio system is designed to allow both content creators and exhibitors to determine how to render sound content in different speaker configurations. The ideal number of speaker output channels used varies depending on the size of the room. Thus, the recommended speaker placement depends on many factors such as size, composition, seating configuration, environment, average audience size, etc. For illustrative purposes only, representative speaker configurations, layouts, and examples are described here, but are not intended to limit the scope of the claimed embodiments.

[0122] The recommended layout of speakers in an adaptive audio system is compatible with existing cinema systems, which is essential to not degrade the playback of existing 5.1 and 7.1 channel-based formats. In order to preserve the intent of the adaptive audio sound engineer and the intent of the mixers of 7.1 and 5.1 content, the positions of the existing screen channels should not be changed too much when introducing new speaker locations. In contrast to using all 64 available output channels, the adaptive audio format can be accurately rendered in a speaker configuration such as 7.1 in a cinema, so this format (and related benefits) can be used in existing theaters without changing the amplifiers or speakers.

[0123] The effectiveness of different speaker locations varies depending on the theater design, and currently there is no ideal number or placement of channels specified in the industry. Adaptive audio is intended to be truly adaptable and capable of accurate playback in various auditoriums, regardless of whether the number of playback channels is limited or many channels are in a flexible configuration.

[0124] FIG. 8 is a top view 800 showing an example layout of the suggested speaker locations for use with an adaptive audio system in a typical auditorium. Also, FIG. 9 is a front view 900 showing an example layout of the suggested speaker locations on the screen of that auditorium. The reference position referred to below corresponds to a position 2 / 3 of the distance from the screen to the rear wall on the center line of the screen. The standard screen speaker 801 is shown in its normal position relative to the screen. From studies of the perception of elevation in the screen plane, additional speakers 804 behind the screen, such as the left center (Lc) and right center (Rc) screen speakers (which are at the locations of the left and right extra channels of the 70mm film format), have been found to be beneficial for making a smooth pan across the screen. Such optional speakers are recommended, especially in auditoriums with screens larger than 12m (40ft.). All screen speakers should be angled towards the reference position. The recommended placement of the subwoofer 810 behind the screen remains unchanged, including maintaining a cabinet placement symmetric with respect to the center of the room and preventing standing wave excitation. An additional subwoofer 816 may be placed behind the theater.

[0125] The surround speakers 802 should be individually amplified, being individually wired to the amplifier rack and, if possible, using dedicated channels of power amplification that match the speaker's power handling according to the manufacturer's specifications. Ideally, the surround speakers should be defined to handle the SPL of each individual speaker and, if possible, have a wider frequency response. As a rough approach for an average-sized theater, the spacing of the surround speakers is 2 to 3 m (6'6" to 9'9"), and the left and right surround speakers should be arranged symmetrically. However, the surround speaker spacing is effectively considered as the angle from the listener between adjacent speakers, rather than using the absolute distance between the speakers. For optimal reproduction throughout the audience area, the angular distance between adjacent speakers should be 30° or less as viewed from each of the four corners of the main listening area. Good results can be obtained with spacing up to 50°. For each surround zone, the speakers should, if possible, maintain equal linear spacing adjacent to the seating area. The linear spacing beyond the listening area, for example between the front row and the screen, can be made slightly larger. FIG. 11 is a diagram showing an example of the arrangement of the top surround speaker 808 and the side surround speaker 806 with respect to a reference point according to one embodiment.

[0126] The additional side surround speaker 806 should be placed closer to the screen than the currently recommended practice of starting about 1 / 3 of the way to the back of the auditorium. These speakers are not used as side surrounds during the playback of Dolby Surround 7.1 or 5.1 sound tracks, but allow for smooth transitions and improved tonal matching when panning an object from the screen speakers into the surround zone. To maximize the sense of space, the surround array should be placed as low as possible under the following constraints: the virtual placement of the surround speakers in front of the array is close to the height of the screen speaker acoustic center and high enough to maintain sufficient coverage across the seating area depending on the speaker directivity. The vertical placement of the surround speakers, as shown in Figure 10, should form a straight line from front to back and (generally) be tilted such that the relative elevation of the surround speakers is above the listener and maintained towards the back of the cinema as the seating elevation increases. Figure 10 is a side view of an example layout of the suggested speaker locations for use with an adaptive audio system in a typical auditorium. In practice, this can be achieved most simply by selecting the elevations of the frontmost and rearmost surround speakers and placing the remaining speakers in a line between these points.

[0127] To provide optimal coverage of each speaker across the seating area, the side surrounds 806, rear speakers 816, and top surrounds 808 must be oriented towards the reference position of the theater under defined guidelines regarding spacing, position, angle, etc.

[0128] Embodiments of an adaptive audio cinema system and format provide a new and powerful authoring tool for a mixer and a new cinema processor with a flexible rendering engine that optimizes the sound quality and surround effects of a sound track according to the speaker layout and characteristics of each room, thereby enabling a higher level of audience immersion than current systems. Further, the system maintains backward compatibility with current production and distribution workflows and minimizes the impact on them.

[0129] Although embodiments have been described in the context of a cinema environment where adaptive audio content is associated with film content specified in a digital cinema processing system, it should be noted that the embodiments can also be implemented in non-cinema environments. Adaptive audio content, including object-based audio and channel-based audio, can be used with any associated content (such as associated audio, video, graphics, etc.) or can constitute stand-alone audio content. The playback environment can be any suitable listening environment, from headphones or near-field monitors to small or large rooms, cars, outdoor arenas, concert halls.

[0130] Aspects of system 100 can be implemented in a suitable computer-based sound processing network environment that processes digital or digitized audio files. Parts of the adaptable audio system include one or more networks including any desired number of individual machines, including one or more routers (not shown) that buffer and route data transmitted between computers. Such networks can be configured over different and diverse network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. In one embodiment where the network is the Internet, one or more machines are configured to access the Internet through a web browser program.

[0131] One or more components, blocks, processors, or other functional components are implemented by a computer program that controls the execution of a processor-based computing device of the system. Note that the various functions disclosed herein can be described using any number of combinations of data and / or instructions embodied in hardware, firmware, and / or various machine-readable or computer-readable media with respect to behavior, register transfers, logic components, and / or other characteristics. The computer-readable media in which such formatted data and / or instructions are embodied are various forms of physical (non-transitory), non-volatile storage media, including but not limited to, for example, optical, magnetic, or semiconductor storage media.

[0132] Unless the context clearly requires otherwise, throughout the specification and the claims, the words "comprise," "comprising," and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of "including, but not limited to." The singular or plural words are to include the plural or singular case, respectively. Also, the words "herein," "hereunder," "above," "below," and the like refer to the application as a whole and not to any particular portions of the application. The word "or" when used in reference to a list of two or more items covers all of the interpretations of that word as referring to any item in the list, any item in the list and any combination of items in the list. One or more implementations have been described by way of examples and with respect to specific embodiments, but of course, one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar configurations that will be apparent to those skilled in the art. Therefore, the appended claims should be construed as broadly as possible to include all such modifications and similar configurations. The above embodiments are appended. (Appendix 1) A system for processing an audio signal, An authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and one or more metadata sets associated with each audio stream and defining the playback location of each audio stream, wherein the audio streams are specified as channel-based audio or object-based audio, wherein the playback location of the channel-based audio includes speaker designations of a plurality of speakers in a speaker array, and the object-based audio includes a location in three-dimensional space, wherein a first set of metadata is applied by default to one or more of the plurality of audio streams, and when the conditions of the playback environment match the one condition of the playback environment, a second set of metadata is associated with one condition of the playback environment and is applied to one or more of the plurality of audio streams in place of the first set, A rendering system coupled to the authoring component and configured to receive a bitstream encapsulating the plurality of monophonic audio streams and one or more data sets and, based on the conditions of the playback environment, render the audio streams to a plurality of speaker feeds corresponding to the speakers of the playback environment by the one or more metadata sets. A system. (Appendix 2) Each metadata set includes metadata elements associated with each object-based stream, and the metadata for each object-based stream defines spatial parameters that control the playback of the corresponding object-based sound and has one or more of sound position, sound width, and sound velocity, Each metadata set includes metadata elements associated with each channel-based stream, and the speaker array is configured in a defined surround sound configuration, The metadata elements associated with each channel-based stream include the designation of the surround sound channels of the speakers in the speaker array according to a defined surround sound standard. The system described in Appendix 1. (Appendix 3) The speaker array includes additional speakers for playing an object-based stream arranged in a playback environment related to a setup command from the user based on the conditions of the playback environment. The playback conditions depend on variables including the size and shape of the room of the playback environment, the occupancy state, and ambient noise. The system receives from the user a setup file including at least a list of speaker designations, the mapping of channels to individual speakers of the speaker array, information regarding speaker groups, and runtime mapping based on the relative positions of the speakers in the playback environment. The system described in Appendix 1. (Appendix 4) The authoring component includes a mixing console having controls operable by the user to define the playback level of an audio stream including the original audio content. The metadata elements associated with each object-based stream are automatically generated upon input by the user to the mixing console controls. The system described in Appendix 1. (Appendix 5) The metadata set includes metadata that enables at least one of upmixing or downmixing of the channel-based audio stream and the object-based audio stream due to a change from a first configuration of the speaker array to a second configuration of the speaker array. The system described in Appendix 1. (Appendix 6) The content type is selected from the group consisting of conversation, music, and effects, and each content type is embodied in a respective set of channel-based streams or object-based streams. The sound components of each content type are sent to a defined speaker group among one or more speaker groups specified in the speaker array. The system according to Appendix 3. (Appendix 7) The speakers of the speaker array are arranged at a plurality of positions within the playback environment. The metadata elements associated with each object-based stream define that one or more sound components are rendered to the speaker feed for playback by the speaker closest to the intended playback location of the sound component indicated by the position metadata. The system according to Appendix 6. (Appendix 8) The playback location is a spatial position with respect to a screen in the playback environment or in a plane enclosing the playback environment. The plane includes a front surface, a rear surface, a left surface, a right surface, an upper surface, and a lower surface. The system according to Appendix 1. (Appendix 9) Further having a codec configured to be coupled to the authoring component and the rendering component, receive the plurality of audio streams and metadata, and generate a single digital bitstream including the plurality of audio streams in an ordered manner. The system according to Appendix 1. (Appendix 10) The rendering component further has means for selecting a rendering algorithm used by the rendering component, and the rendering algorithm is selected from the group consisting of binaural, stereo dipole, Ambisonics, Wave Field Synthesis (WFS), multi-channel panning, a Roth system with position metadata, dual balance, and vector-based amplitude panning. The system according to Appendix 9. (Appendix 11) The playback location of each audio stream is defined independently with respect to either a self-centered reference frame or an other-centered reference frame. The ego - centered reference frame is taken with respect to the listener in the playback environment, The other - centered reference frame is taken with respect to the characteristics of the playback environment, The system according to Appendix 1. (Appendix 12) A system for processing an audio signal, An authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and metadata associated with each audio stream and defining the playback location of each audio stream, The audio stream is specified as channel - based audio or object - based audio, The playback location of the channel - based audio includes speaker designations of a plurality of speakers in a speaker array, and the object - based audio includes a location in three - dimensional space, Each object - based audio stream is rendered by at least one speaker of the speaker array, Coupled to the authoring component, receiving a bitstream encapsulating the plurality of monophonic audio streams and metadata, and rendering the audio streams into a plurality of speaker feeds corresponding to the speakers in the playback environment, The speakers of the speaker array are arranged at a plurality of positions within the playback environment, A system, wherein a metadata element associated with each object - based stream defines that one or more sound components are to be rendered into a speaker feed by the speaker closest to the intended playback location of the sound components so that the object - based stream is effectively snapped to the speaker closest to the intended playback location. (Appendix 13) The metadata includes two or more metadata sets, and the rendering system renders the audio stream according to one of the two or more metadata sets based on the conditions of the playback environment, The first set of metadata is applied to one or more of the plurality of audio streams for the first condition of the playback environment, and the second set of metadata is applied to one or more of the plurality of audio streams for the second condition of the playback environment. Each metadata set includes metadata elements related to each object-based stream, and the metadata of each object-based stream defines spatial parameters that control the playback of the corresponding object-based sound, having one or more of sound position, sound width, and sound velocity. Each metadata set includes metadata elements related to each channel-based stream, and the speaker array is configured in a defined surround sound configuration. The metadata elements related to each channel-based stream include the designation of the surround sound channels of the speakers in the speaker array according to a defined surround sound standard. The system according to Appendix 12. (Appendix 14) The speaker array includes additional speakers for playing object-based streams arranged in the playback environment related to setup commands from the user based on the conditions of the playback environment. The playback conditions depend on variables including the size and shape of the room of the playback environment, occupancy status, and ambient noise. The system receives from the user a setup file that includes at least speaker designation, mapping of the channels to individual speakers of the speaker array, information regarding speaker groups, and a list of runtime mappings based on the relative positions of the speakers in the playback environment. The object stream rendered to the speaker feed for playback by the speaker closest to the intended playback location of the sound component snaps to a single speaker of the additional speakers. The system according to Appendix 12. (Appendix 15) The intended playback location is a spatial position with respect to the playback environment or a screen in the plane enclosing the playback environment. The surface includes a front surface, a rear surface, a left surface, an upper surface, and a lower surface. The system according to appendix 14. (Appendix 16) A system for processing an audio signal, An authoring component configured to receive a plurality of audio signals and generate a plurality of monophonic audio streams and metadata associated with each audio stream and defining the playback location of each audio stream, The audio stream is specified as channel-based audio or object-based audio, The playback location of the channel-based audio includes speaker designations of a plurality of speakers in a speaker array, and the object-based audio includes a location in a three-dimensional space with respect to a playback environment including the speaker array, Each object-based audio stream is rendered by at least one speaker of the speaker array, A rendering system coupled to the authoring component, receiving a first map of audio channels of speakers including a list of a plurality of speakers in the playback environment and their respective locations, and a bitstream encapsulating the plurality of monophonic audio streams and metadata, and configured to render the audio streams to a plurality of speaker feeds corresponding to the speakers of the playback environment based on a runtime mapping to the playback environment based on the relative positions of the speakers and the conditions of the playback environment. (Appendix 17) The conditions of the playback environment depend on variables including the size and shape of the room of the playback environment, occupancy status, material composition, and ambient noise. The system according to appendix 16. (Appendix 18) The first map is defined in a setup file including at least a list of speaker designations and a mapping of channels to individual speakers of the speaker array, and information regarding speaker grouping. The system according to appendix 17. (Appendix 19) The intended playback location is a spatial position with respect to a screen in the playback environment or an enclosure containing the playback environment, The surface includes the front, rear, side, top, and bottom surfaces of the enclosure, The system according to Appendix 18. (Appendix 20) The speaker array has speakers arranged in a defined surround sound configuration, The metadata element associated with each channel-based stream includes the specification of the surround sound channels of the speakers in the speaker array according to a defined surround sound standard, The object-based stream is played back by additional speakers of the speaker array, The runtime mapping dynamically determines which individual speakers of the speaker array play the corresponding object-based stream during the playback process, The system according to Appendix 19. (Appendix 21) A method for authoring an audio signal to be rendered, Receiving a plurality of audio signals, Generating a plurality of monophonic audio streams and one or more metadata sets associated with each audio stream, and defining the playback location of each audio stream, The audio stream is specified as channel-based audio or object-based audio, The playback location of the channel-based audio includes the speaker designations of a plurality of speakers in the speaker array, and the object-based audio includes a location in a three-dimensional space with respect to the playback environment including the speaker array, The first set of metadata is applied to one or more of the plurality of audio streams under the first condition of the playback environment, and the second set of metadata is applied to one or more of the plurality of audio streams under the second condition of the playback environment, For transmission to a rendering system configured to render the audio stream to a plurality of speaker feeds corresponding to speakers in the playback environment based on the conditions of the playback environment by the one or more metadata sets, encapsulating the plurality of monophonic audio streams and a set of one or more metadata in a bitstream. (Appendix 22) Each metadata set includes metadata elements related to each object-based stream, and the metadata of each object-based stream defines spatial parameters that control the playback of the corresponding object-based sound, and has one or more of sound position, sound width, and sound velocity. Each metadata set includes metadata elements related to each channel-based stream, and the speaker array is configured in a defined surround sound configuration. The metadata elements related to each channel-based stream include the designation of the surround sound channels of the speakers in the speaker array according to a defined surround sound standard. The method according to Appendix 21. (Appendix 23) The speaker array includes additional speakers for the playback of object-based streams arranged in the playback environment, and the method further includes receiving a setup command from a user based on the conditions of the playback environment. The playback conditions depend on variables including the size and shape of the room in the playback environment, occupancy status, and ambient noise. The setup command further includes at least a list of speaker designations, the mapping of channels to individual speakers of the speaker array, information regarding speaker groups, and runtime mapping based on the relative positions of the speakers in the playback environment. The method according to Appendix 21. (Appendix 24) Receiving from a mixing console having controls operated by a user to define the playback level of an audio stream including the original audio content. automatically generating metadata elements associated with each generated object-based stream when receiving the user input; The method according to Supplementary Note 23. (Supplementary Note 25) A method for rendering an audio signal, comprising: receiving a plurality of audio signals, and receiving, from an authoring component configured to generate a plurality of monophonic audio streams and one or more metadata sets associated with each audio stream and defining a playback location for each audio stream, a bitstream encapsulating the plurality of monophonic audio streams and the one or more metadata sets into a bitstream; the audio stream is specified as channel-based audio or object-based audio; the playback location of the channel-based audio includes speaker designations of a plurality of speakers in a speaker array, and the object-based audio includes a location in a three-dimensional space relative to a playback environment including the speaker array; a first set of metadata is applied to one or more of the plurality of audio streams under a first condition of the playback environment, and a second set of metadata is applied to one or more of the plurality of audio streams under a second condition of the playback environment; rendering the plurality of audio streams to a plurality of speaker feeds corresponding to speakers in the playback environment by the one or more metadata sets based on conditions of the playback environment. A method comprising: (Supplementary Note 26) A method for generating audio content including a plurality of monophonic audio streams processed by an authoring component, comprising: the monophonic audio streams include at least one channel-based audio stream and at least one object-based audio stream, and the method comprises: a step of indicating whether each of the plurality of audio streams is a channel-based stream or an object-based stream; a step of associating each channel-based stream with a metadata element defining a channel position for rendering the channel-based stream to one or more speakers in a playback environment; a step of associating each object-based stream with one or more metadata elements defining an object-based one for rendering the respective object-based stream to one or more speakers in the playback environment with respect to an other-centered reference frame defined with respect to the size and dimensions of the playback environment; assembling the plurality of monophonic audio streams and associated metadata into a signal. (Appendix 27) The playback environment includes an array of speakers arranged in a defined location and direction with respect to a reference point of an enclosure embodying the playback environment. The method according to Appendix 26. (Appendix 28) The first set of speakers of the speaker array has speakers configured according to a defined surround sound system. The second set of speakers of the speaker array has speakers configured according to an adaptive audio scheme. The method according to Appendix 27. (Appendix 29) defining an audio type of a set of the plurality of monophonic audio streams; the audio type is selected from the group consisting of conversation, music, and effects; transmitting the set of audio streams to a set of speakers based on the audio type of the set of audio streams. The method according to Appendix 28. (Appendix 30) Further comprising automatically generating the metadata element by an authoring component implemented in a mixing console having a user-operable control for defining a playback level of the monophonic audio stream. The method according to Appendix 29. (Appendix 31) Further comprising, in an encoder, packaging a plurality of monophonic audio streams and associated metadata elements into a single digital bitstream. The method according to Appendix 30. (Appendix 32) A method for generating audio content, comprising: determining values of one or more metadata elements in a first metadata group associated with programming of audio content for processing in a hybrid audio system that handles both channel-based and object-based content; determining values of one or more metadata elements in a second metadata group associated with storage and rendering characteristics of the audio content of the hybrid audio system; determining values of one or more metadata elements in a third metadata group associated with audio source location and control information for rendering the channel-based and object-based audio content. (Appendix 33) The audio source location for rendering the channel-based audio content includes a name associated with a speaker of a surround sound speaker system, wherein the name defines a location of the speaker relative to a reference point of the playback environment. The method according to Appendix 32. (Appendix 34) The control information for rendering the channel-based audio content includes upmixing and downmixing information for rendering the audio content in different surround sound configurations. The metadata includes metadata that enables or disables an upmixing and / or downmixing function. The method described in Supplementary Note 33. (Supplementary Note 35) The audio source position for rendering the object-based audio content has a value related to one or more mathematical functions that define the intended playback location of the sound components of the object-based audio content. The method described in Supplementary Note 32. (Supplementary Note 36) The mathematical function is selected from the group consisting of three-dimensional coordinates defined as x, y, and z coordinate values, a set of two-dimensional coordinates and a plane definition, a set of one-dimensional linear position coordinates and a curve definition, and a scalar position on the screen in the playback environment. The method described in Supplementary Note 35. (Supplementary Note 37) The control information for rendering the object-based audio content has a value that defines an individual speaker or speaker group in the playback environment where the sound component is played. The method described in Supplementary Note 36. (Supplementary Note 38) The control information for rendering the object-based audio content further includes a binary value that defines a sound source snapped to the nearest speaker or nearest speaker group in the playback environment. The method described in Supplementary Note 37. (Supplementary Note 39) A method for defining an audio transport protocol, defining values of one or more metadata elements in a first metadata group associated with programming of audio content for processing in a hybrid audio system that handles both channel-based and object-based content; defining values of one or more metadata elements in a second metadata group related to storage and rendering characteristics of the audio content of the hybrid audio system; A method having a step of defining values of one or more metadata elements of a third metadata group related to an audio source position and control information for rendering the channel-based and object-based audio content.

Claims

1. A system for processing an audio signal, comprising: a rendering system, wherein the rendering system receives a bitstream including encoded audio data representing a plurality of monaural audio streams, and further including metadata associated with each of the monaural audio streams and indicating a playback position of each respective monaural audio stream, wherein at least some of the plurality of monaural audio streams are identified as object-based audio, and the playback position of the object-based monaural audio stream includes a position in a three-dimensional space, and receiving decoding the encoded audio data to provide the plurality of monaural audio streams, and decoding rendering the plurality of monaural audio streams into a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are arranged at specific positions in the playback environment, and one or more additional metadata elements associated with each object-based monaural audio stream indicate whether rendering of each respective monaural audio stream to one or more specific speaker feeds of the plurality of speaker feeds is prohibited such that the respective object-based monaural audio stream is not rendered to any of the one or more specific speaker feeds of the plurality of speaker feeds, and rendering A system configured to perform the above.

2. The system according to claim 1, wherein the metadata element associated with each object-based monaural audio stream further indicates a spatial parameter for controlling playback of a corresponding sound component, including one or more of a sound position, a sound width, and a velocity.

3. The system according to claim 1, wherein the playback position for each of the plurality of object-based monaural audio streams is specified independently with respect to either a self-centered reference frame or an other-centered reference frame, the self-centered reference frame being taken with respect to a listener in the playback environment, and the other-centered reference frame being taken with respect to characteristics of the playback environment. A method of authoring audio content for rendering, implemented by a system for processing audio signals, comprising: Receiving a plurality of audio signals; Generating a plurality of monaural audio streams and metadata associated with each of the plurality of monaural audio streams, the metadata indicating the playback position of each monaural audio stream, wherein at least some of the plurality of monaural audio streams are identified as object-based audio and the playback position of the object-based audio includes a position in three-dimensional space; Encoding the plurality of monaural audio streams to provide encoded audio data; Encapsulating the encoded audio data and the metadata into a bitstream for transmission to a rendering system configured to render the plurality of monaural audio streams to a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are located at specific positions in the playback environment and one or more additional metadata elements associated with each object-based monaural audio stream indicate whether rendering of each monaural audio stream to one or more specific speaker feeds of the plurality of speaker feeds is prohibited; A method comprising the steps above. A method of rendering an audio signal, implemented by a system for processing audio signals, comprising: Receiving a bitstream comprising encoded audio data representing a plurality of monaural audio streams and further comprising metadata associated with each of the monaural audio streams, the metadata indicating the playback position of each monaural audio stream, wherein at least some of the plurality of monaural audio streams are identified as object-based audio and the playback position of the object-based monaural audio streams includes positions in three-dimensional space; Decoding the encoded audio data to provide the plurality of monaural audio streams, and decoding Rendering the plurality of monaural audio streams into a plurality of speaker feeds corresponding to speakers in a playback environment, wherein the speakers are arranged at specific positions in the playback environment, and one or more additional metadata elements associated with each object-based monaural audio stream indicate whether rendering each monaural audio stream to one or more specific speaker feeds among the plurality of speaker feeds is prohibited so that each object-based monaural audio stream is not rendered to any of the one or more specific speaker feeds among the plurality of speaker feeds, and rendering A method comprising.

6. The method according to claim 5, wherein the metadata element associated with each object-based monaural audio stream further indicates spatial parameters for controlling playback of a corresponding sound component, including one or more of sound position, sound width, and velocity.

7. The playback position for each of the plurality of object-based monaural audio streams includes a spatial position relative to a screen in the playback environment or a surface surrounding the playback environment, the surface including a front surface, a rear surface, a left surface, a right surface, an upper surface, and a lower surface, and / or is specified independently of either a self-centered reference frame or an other-centered reference frame, the self-centered reference frame being taken with respect to a listener in the playback environment, and the other-centered reference frame being taken with respect to characteristics of the playback environment, the method according to claim 5.

8. A non-transitory computer-readable storage medium including a series of instructions that, when executed by a system for processing audio signals, cause the system to execute the method according to any one of claims 4 to 7.

9. A computer program that, when executed by a computer, causes the computer to execute the method according to any one of claims 4 to 7.

Citation Information

Patent Citations

  • Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information

    JP2009278381A

  • JPP7348320B

  • System for adaptively streaming audio objects

    US20110040396A1