Screen-to-screen rendering of audio and encoding and decoding of audio for such rendering

By warping audio channels based on screen-related metadata, the method addresses the misalignment of audio and visual elements in non-cinema environments, ensuring accurate spatial reproduction across various playback systems.

JP2026012756APending Publication Date: 2026-01-27DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025173390
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-11-14
Filing Date
2025-10-15
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing audio content delivery systems struggle to accurately reproduce the spatial relationship between audio and visual elements in non-cinema environments due to mismatched speaker and screen configurations, leading to discrepancies in perceived audio and visual alignment.

Method used

The method involves warping audio channels based on screen-related metadata to distort audio positions relative to the playback system's display screen, allowing for smooth transitions between on-screen and off-screen positions, and generating speaker feeds that align with the intended audio-visual alignment.

Benefits of technology

This approach ensures consistent and accurate reproduction of audio-visual alignment across different playback environments, maintaining the intended spatial relationships between audio and visual elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012756000001_ABST
    Figure 2026012756000001_ABST
Patent Text Reader

Abstract

To provide rendering of audio to a screen and encoding and decoding of audio for such rendering.SOLUTION: In the system, a production unit 3 outputs a plurality of speaker channels of audio data, a plurality of object channels of audio data, and metadata. The object processing subsystem 9 receives the decoded speaker channels, object channels, and object related metadata delivered from the decoder 7, performs warping on the object channels using the screen related metadata, and outputs the resulting object channels and / or mixes to the rendering subsystem 11. Rendering subsystem 11 uses the rendering parameters to map the audio objects to the available speaker channels.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 61 / 904,233, filed November 14, 2013, the contents of which are incorporated herein by reference in their entirety.

[0002] Technical field of the invention The present invention relates to encoding, decoding, and rendering audio programs (e.g., soundtracks for movies or other audiovisual programs) with corresponding video content. In some embodiments, the programs are object-based audio channels that include at least one audio object channel, screen-related metadata, and typically also speaker channels. The screen-related metadata supports screen-relative rendering, in which sound sources indicated by the program (e.g., objects indicated by object channels) are rendered at positions relative to a playback system's display screen that are determined, at least in part, by the screen-related metadata. [Background technology]

[0003] Embodiments of the present invention relate to one or more aspects of an audio content generation and delivery pipeline (eg, a pipeline for generating and delivering audio content for an audiovisual program).

[0004] Such a pipeline implements the production of an audio program (typically an encoded audio program that represents audio content and metadata corresponding to the audio content). The production of an audio program may include audio production activities (audio capture and recording) and optionally also "post-production" activities (manipulation of the recorded audio). Live broadcast necessarily requires that all authoring decisions be made during audio production. In the production of films and other non-real-time programs, many authoring decisions may be made during post-production.

[0005] Audio content production and delivery pipelines optionally implement program remixing and / or remastering. In some cases, programs may require additional processing after content generation to repurpose the content for alternative use cases. For example, a program originally produced for playback in a cinema may be modified (e.g., remixed) to make it more suitable for playback in a home environment.

[0006] Audio content generation and distribution pipelines typically include an encoding stage. Audio programs may require encoding to enable distribution. For example, programs intended for home playback are typically data compressed to allow more efficient distribution. The encoding process may include stages of reducing the complexity of the spatial audio scene and / or reducing the data rate of the program's individual audio streams and / or packaging multiple channels of audio content (e.g., compressed audio content) and corresponding metadata into a bitstream having a desired format.

[0007] The audio content generation and delivery pipeline includes decoding and rendering stages (typically implemented by a playback system including a decoder). Ultimately, the program is presented to the end consumer by rendering the audio description into a loudspeaker signal based on the playback equipment and environment.

[0008] Exemplary embodiments of the present invention allow an audio program (e.g., the soundtrack of a movie or other program with audio and image content) to be played so that the location of an auditory image is reliably presented in a manner that is consistent with the location of a corresponding visual image.

[0009] Traditionally, in a cinema mixing room (or other audiovisual program authoring environment), the location and size of the display screen (referred to herein as the "reference" screen to distinguish it from the audiovisual program playback screen) is aligned with the front wall of the mixing environment, with the left and right edges of the reference screen coinciding with the locations of the left and right main screen loudspeakers. An additional center screen channel is generally positioned in the center of the reference screen / wall. In this way, the extent of the front wall, the front loudspeaker locations, and the screen location are consistently co-located. Typically, the reference screen is approximately as wide as the room, with the left, center, and right loudspeakers near the left, center, and right edges of the reference screen. This arrangement is similar to the typical arrangement of the screen and front speakers in an expected cinema theater playback location. For example, Figure 1 shows a diagram of the front wall of such a cinema theater, with the display screen S, left and right front speakers (L and R), and a front center speaker (C) mounted on or near the front wall. During playback of a movie, visual image B is displayed on screen S while an accompanying sound "A" is emitted from the speakers of the playback system (including speakers L, R, and C). For example, image B may be an image of a sound source (e.g., a bird or a helicopter), and sound "A" may be a sound intended to be perceived as emanating from that source. The movie is assumed to be authored and rendered such that sound A is perceived as emanating from a source location that coincides (or nearly coincides) with the location on screen S where image B is displayed, when the front speakers are located coplanar with screen S, the left and right front speakers (L and R) are at the left and right edges of screen S, and the center front speaker is near the center of screen S. Figure 1 assumes that screen S is at least substantially acoustically transparent, and that speakers L, C, and R are located behind screen S (but at least substantially within the plane of screen S).

[0010] However, during playback in a consumer's home (or on a mobile user's portable playback device), the size and position of the playback system's front speakers (or headset speakers) relative to each other and relative to the playback system's display screen need not match those of the front speakers and display screen in the program authoring environment (e.g., a movie theater mixing room). In such playback cases, the width of the playback screen is typically significantly smaller than the distance separating the left and right main speakers (the left and right front speakers or headset, e.g., the speakers in a pair of headphones). It is possible that the screen is not centered relative to the main speakers, or even in a fixed position (e.g., in the case of a mobile user wearing headphones and holding a display device). This can create a noticeable discrepancy between the perceived audio and visuals.

[0011] For example, Figure 2 is a diagram of a room's front wall (W') with a home theater system's display screen (S') mounted on (or near) the front wall, left and right front speakers (L' and R'), and front center speaker (C'). During playback (by the system of Figure 2) of the same movie described in the example of Figure 1, visual image B is displayed on screen S' while an accompanying sound A is emitted from the playback system's speakers (including speakers L', R', and C'). We assumed that the movie was authored for rendering and playback (by the movie theater playback system) so that sound A is perceived to emanate from a source location that coincides (or nearly coincides) with the location on the movie theater screen where image B is displayed. However, when the movie is played back by the home theater system of Figure 2, sound A will be perceived to emanate from a source location that is closer to the left front speaker L', which does not coincide or nearly coincide with the location on the home theater screen S' where image B is displayed. This is because the front speakers L', C', and R' of the home theater system have a different size and position relative to screen S' than the front speakers of the program authoring system have relative to the program authoring system's reference screen.

[0012] In the examples of Figures 1 and 2, the expected cinema playback system is assumed to have a well-defined relationship between its loudspeakers and screen. Thus, the content creator's desired relative positions of the displayed image and corresponding audio sources can be reliably reproduced (during playback in the cinema). For playback in other environments (e.g., in a home audio-video room), the expected relationship between loudspeakers and screen is typically not preserved, and thus the relative positions of the displayed image and corresponding audio sources (desired by the content creator) are typically not well reproduced. The relative positions of the displayed image and corresponding audio sources that are actually achieved during playback (other than in a cinema with an expected relationship between loudspeakers and screen) are based on the actual relative positions and sizes of the playback system's loudspeakers and display screen.

[0013] During the playback of an audiovisual program, for sounds that are rendered to be perceived at an on-screen location, the optimum auditory image location is independent of listener position. For sounds that are rendered to be perceived at an off-screen location (at a non-zero distance in a direction perpendicular to the plane of the screen), there is a possibility of parallax errors in the auditorily perceived location of the sound source, depending on the listener position. Methods have been proposed that attempt to minimize or eliminate such parallax errors based on known or assumed listener positions.

[0014] It is known to use high-end playback systems (e.g., in movie theaters) to render object-based audio programs (e.g., object-based programs presenting movie soundtracks). An object-based audio program that is, for example, a movie soundtrack, may present many different sound elements (audio objects) corresponding to on-screen images, dialogue, noises, and sound effects emanating from different positions on (or relative to) the screen, as well as background music and ambient effects (which may be represented by the program's speaker channels) to create the intended overall auditory experience. Accurate playback of such programs requires that the sounds be reproduced in a manner that corresponds as closely as possible to that intended by the content creator in terms of audio object size, position, intensity, movement, and depth.

[0015] Object-based audio programs represent a significant improvement over traditional speaker-channel-based audio programs because speaker-channel-based audio is more limited than object-channel-based audio with respect to the spatial reproduction of specific audio objects. The audio channels in a speaker-channel-based audio program consist of speaker channels only (no object channels), and each speaker channel typically determines the speaker feed for a specific, individual speaker in the listening environment.

[0016] Various methods and systems have been proposed for generating and rendering object-based audio programs. During the generation of an object-based audio program, it is typically assumed that any number of loudspeakers will be used for playback of the program, and that the loudspeakers used for playback (typically in a movie theater) will be located at any position in the playback environment, not necessarily in a (nominal) horizontal plane or any other predetermined arrangement known at the time of program generation. Object-related metadata included in the program typically indicates rendering parameters for rendering at least one object of the program at an apparent spatial position (in a three-dimensional volume) or along a trajectory, e.g., using a three-dimensional array of speakers. For example, an object channel of the program may have corresponding metadata that indicates a three-dimensional trajectory of apparent spatial positions at which the object (represented by that object channel) will be rendered. The trajectory may include a sequence of "floor" positions (in the plane of a subset of speakers of the playback environment assumed to be located on the floor or in other horizontal planes) and a sequence of "above-floor" positions (each determined by driving a subset of speakers assumed to be located in at least one other horizontal plane of the playback environment). Examples of object-based audio program rendering are described, for example, in commonly assigned US Patent No. 6,229,999, filed on Oct. 1, 2004, entitled "Object-Based Audio Programming."

[0017] The advent of object-based audio program rendering has significantly increased the amount of audio data to be processed and the complexity of the rendering that must be performed by rendering systems. This is in part because an object-based audio program may represent many objects (each with corresponding metadata) and may be rendered for playback by a system that includes many loudspeakers. It has been proposed to limit the number of object channels contained in an object-based audio program so that the intended rendering system is capable of rendering the program. For example, U.S. Provisional Patent Application No. 61 / 745,401, filed December 21, 2012, entitled "Scene Simplification and Object Clustering for Rendering Object-Based Audio Content," which lists Brett Crockett, Alan Seefeldt, Nocolas Tsingos, Rhonda Wilson, and Jeroen Breebaart as inventors and is assigned to the assignee of the present invention, describes methods and apparatus for so limiting the number of object channels in an object-based audio program by clustering input object channels to generate clustered object channels that are included in the program and / or by mixing audio content of input object channels with speaker channels to generate mixed speaker channels that are included in the program. It is contemplated that some embodiments of the present invention may be performed in conjunction with such clustering (e.g., in a mixing or remix facility) to generate object-based programs for delivery (together with screen-related metadata) to a playback system or for use in generating speaker-channel-based programs for delivery to a playback system. [Prior art documents] [Patent documents]

[0018] [Patent Document 1] PCT International Application No. PCT / US2011 / 028783, International Publication No. 2011 / 119041A2, published September 29, 2011 Summary of the Invention [Means for solving the problem]

[0019] Throughout this disclosure, "warping" at least one channel (e.g., an object channel or a speaker channel) of an audio program refers to processing the audio content (audio data) of each such channel to generate distorted audio content (or replacing each such channel with at least one other audio channel representing the distorted audio content), assuming the program has corresponding video content (e.g., the program may be the soundtrack of a movie or other audiovisual program). When the distorted audio content is rendered to generate a speaker feed that is used to drive playback speakers, the sound emitted from the speakers represents at least one audio element (intended by the content creator to be perceived at at least one predetermined position relative to a reference screen, e.g., a movie theater screen) with a perceived distorted position (which may be fixed or may change over time). The distorted position is "warped" in the sense that it is a predetermined position relative to the display screen of the playback system (rather than relative to the reference screen intended by the content creator). Typically, each distorted position is determined (at least in part) relative to the display screen of the playback system (sometimes referred to as the "playback screen") by metadata provided with (e.g., included in) the audio program (referred to herein as "screen-related" metadata). Each distorted position may be determined by the screen-related metadata and other data indicative of the playback system configuration (e.g., data indicative of the positions or positions and sizes and / or relationship(s) between the sizes and / or positions of the playback system's speakers and display screen). The distorted position(s) may, but need not, coincide with the actual playback screen.Some embodiments of the present invention allow smooth transitions between on-screen and / or off-screen (relative to the playback screen) distorted positions that change during playback.

[0020] In this document, the expression "off-screen distortion" of at least one channel of a program refers to a "distortion" of said at least one channel of the type in which the distorted position of at least one corresponding audio element (determined by the audio content of said at least one channel) is at a non-zero depth relative to the reproduction screen (i.e., has a non-zero distance from the reproduction screen in a direction at least substantially perpendicular to the plane of the reproduction screen).

[0021] In a first class of embodiments, the present invention is a method for rendering an audio program (e.g., an object-based audio program), comprising: (a) determining at least one warping degree parameter (e.g., by parsing the program to identify at least one said warping degree parameter indicated by screen-related metadata of the program, or by configuring a playback system to perform rendering that includes specifying at least one said warping degree parameter to the playback system); and (b) performing distortion on audio content of at least one channel of the program to a degree determined at least in part by the distortion degree parameter corresponding to said channel, each said distortion degree parameter indicating a maximum degree of distortion to be performed by the playback system on corresponding audio content of the program (e.g., a non-binary value indicating the maximum degree). In some embodiments of the first class, step (a) includes determining at least one off-screen distortion parameter (e.g., by parsing the program to identify at least one said off-screen distortion parameter indicated by screen-related metadata of the program), the off-screen distortion parameter indicating at least one characteristic of off-screen distortion of the corresponding audio content of the program by the playback system, and the distortion performed in step (b) includes off-screen distortion determined at least in part by the at least one said off-screen distortion parameter. For example, the off-screen distortion parameter may control the manner or degree of distortion (in a direction at least substantially parallel to the plane of the playback screen) or maximum distortion of the distorted position of the audio element as a function of depth (distance from the playback screen in a direction at least substantially perpendicular to the plane of the playback screen).In some embodiments, the distortion intensity parameter determined in step (a) indicates a maximum degree of distortion to be performed on corresponding audio content of the program in a plane at least substantially parallel to the plane of the playback screen (at a depth at least substantially perpendicular to the playback screen) and is thus an off-screen distortion parameter. In other embodiments, step (a) includes determining at least one distortion intensity parameter and at least one off-screen distortion parameter that is not a distortion intensity parameter. In some embodiments, the program exhibits at least two objects, and step (a) includes independently determining at least one distortion intensity parameter for each of at least two of the objects, and step (b) includes independently performing distortion on audio content exhibiting each of the objects to a degree determined at least in part by the at least one distortion intensity parameter corresponding to each of the objects.

[0022] In a second class of embodiments, the present invention is a method for generating (or decoding) an object-based audio program. The method includes determining at least one distortion degree parameter for at least one audio object and including in the program screen-related metadata indicating an object channel (indicating the object) and each of the distortion degree parameters for the object. Each of the distortion degree parameters indicates (e.g., is a non-binary value (e.g., a scalar value having any of many values ​​within a predetermined range) the maximum degree of distortion to be performed by a playback system on the corresponding object (e.g., in a plane parallel to the plane of a playback screen). For example, the distortion degree parameter may be a floating-point value ranging from a minimum value (indicating that no distortion should be performed) to a maximum value (indicating that full distortion should be performed) (e.g., an audio element position defined by the program to be at the right edge of the reference screen should be distorted to a distorted position at the right edge of the playback screen). Here, the range includes at least one intermediate value (greater than the minimum value but less than the maximum value) indicating that an intermediate degree of distortion (e.g., 50% of full distortion) should be performed (e.g., distorting an audio element position defined by the program to be at the right edge of the reference screen to a distorted position halfway between the right edge of the playback room and the right edge of the playback screen). In this context, full distortion may represent a distortion of the perceived position of the audio element in the plane of the playback screen such that the distorted position coincides with the playback screen, and an intermediate degree of distortion (or less than full distortion) may represent a distortion of the perceived position of the audio element in the plane of the playback screen such that the distorted position coincides with an area larger than (and including) the playback screen.

[0023] In some embodiments of the second class, the screen-related metadata indicates at least one distortion parameter for each of at least two objects of the program, each distortion parameter indicating a maximum degree of distortion to be performed on the corresponding object. For example, the distortion parameters may indicate different maximum degrees of distortion in or parallel to the plane of the reproduction screen for each object represented by different object channels. In another example, the distortion parameters may indicate different maximum degrees of distortion in a vertical direction in or parallel to the plane of the reproduction screen and different maximum degrees of distortion in a horizontal direction in or parallel to the plane of the reproduction screen for each object represented by different object channels.

[0024] In some embodiments of the second class, the screen-related metadata also indicates at least one off-screen distortion parameter that indicates at least one characteristic of off-screen distortion to be performed by the playback system on the corresponding audio content of the program (e.g., indicating the manner and / or degree of distortion to be performed in a plane at least substantially parallel to the plane of the playback screen as a function of distance in each plane at least substantially perpendicular to the plane of the playback screen). In some such embodiments, the screen-related metadata indicates one such off-screen distortion parameter for each of at least two objects shown by the program, each such off-screen distortion parameter indicating at least one characteristic of the off-screen distortion to be performed on each corresponding object. For example, the program may include off-screen distortion parameters for each object shown by a different object channel, the off-screen distortion parameters indicating the type of off-screen distortion to be performed on each corresponding object (i.e., the metadata may specify different types of off-screen distortion for the object(s) corresponding to each object channel). In some embodiments, at least one off-screen distortion parameter indicates a maximum degree of distortion to be performed in a plane at least substantially parallel to the plane of the playback screen (at a depth at least substantially perpendicular to the playback screen) for the corresponding audio content of the program, and thus the off-screen distortion parameter is a distortion degree parameter.

[0025] In a third class of embodiments, the present invention provides: (a) generating an object-based audio program; (b) generating, in response to the object-based audio program, a speaker channel-based program including at least one set of speaker channels intended for playback by loudspeakers positioned at predetermined positions relative to a playback screen, wherein generating the set of speaker channels includes distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter, each distortion degree parameter indicating a maximum degree of distortion (e.g., a non-binary value indicating the maximum degree, e.g., a scalar value having any value among many values ​​within a predetermined range) to be performed by a playback system on corresponding audio content of the object-based audio program (e.g., in a plane parallel to the plane of the playback screen).

[0026] In some embodiments of a third class, step (b) includes generating the speaker channel-based audio program to include two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program, and generating at least one other of which sets includes distorting (using the distortion parameter) audio content of the object-based audio program, the other of which sets being intended for playback by loudspeakers located at predetermined positions relative to a playback screen. In some embodiments of the third class, step (b) includes determining at least one off-screen distortion parameter (e.g., by parsing the object-based program to identify at least one of the off-screen distortion parameters indicated by screen-related metadata of the object-based audio program), the off-screen distortion parameter indicating at least one characteristic of off-screen distortion of the corresponding audio content of the object-based audio program by a playback system, and step (b) includes off-screen distortion determined at least in part by the at least one off-screen distortion parameter.

[0027] In some embodiments of a third class, the object-based audio program includes screen-related metadata indicating at least one of the distortion parameters (or at least one of the distortion parameters and at least one off-screen distortion parameter), and step (b) includes parsing the object-based audio program to identify the at least one of the distortion parameters (or the at least one of the distortion parameters and the off-screen distortion parameter).

[0028] The generation of the speaker channel-based program (according to a third class of embodiments) supports screen-to-screen rendering by a playback system that is not configured to perform decoding and rendering of object-based audio programs (but that is capable of decoding and rendering speaker channel-based programs). Typically, speaker channel-based programs are generated by a remix system that has knowledge of (or assumes) a particular playback system speaker and screen configuration. Typically, the object-based program (in response to which the speaker channel-based program is generated) includes screen-related metadata that supports screen-to-screen rendering of the object-based program by a suitably configured playback system (that is capable of decoding and rendering object-based programs).

[0029] In a fourth class of embodiments, the present invention is a method for rendering a speaker channel-based program including at least one set of speaker channels representing distorted content, the speaker channel-based program having been generated by processing an object-based audio program, the processing comprising distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion parameter to generate the set of speaker channels representing the distorted content, wherein each distortion parameter indicates a maximum degree of distortion (e.g., a non-binary value (e.g., a scalar value having any one of many values ​​within a predetermined range)) performed by a playback system on corresponding audio content of the object-based audio program (e.g., in a plane parallel to the plane of a playback screen). The rendering method comprises: (a) parsing the speaker channel-based program to identify speaker channels of the speaker channel-based program, including each of the sets of speaker channels exhibiting distorted content; (b) generating speaker feeds for driving loudspeakers positioned at predetermined positions relative to a playback screen in response to at least some of the speaker channels of the speaker channel-based program, including at least one said set of speaker channels showing distorted content.

[0030] In some embodiments of a fourth class, the speaker channel-based program was generated by processing the object-based audio program, the processing including by performing off-screen distortion of audio content of the object-based audio program to an extent determined at least in part by the at least one distortion degree parameter using at least one off-screen distortion parameter indicating at least one characteristic of off-screen distortion for corresponding audio content of the object-based program.

[0031] In some embodiments of a fourth class, the speaker channel-based audio program includes two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program and others of which sets are one said set of speaker channels representing distorted content, and step (b) includes selecting one of the sets that is one said set of speaker channels representing distorted content.

[0032] In some embodiments, a method of the present invention includes generating (e.g., in an encoder), decoding (e.g., in a decoder), and / or rendering an object-based audio program that includes screen-related metadata. The object-based program has corresponding video content (which may be, for example, the soundtrack of a movie or other audiovisual program) and includes at least one audio object channel, screen-related metadata, and typically also a speaker channel. The screen-related metadata includes metadata corresponding to each of at least one of the object channels (and optionally, metadata corresponding to each of at least one of the speaker channels). During rendering and playback of the object-based program, processing of the screen-related metadata (which typically includes data indicating the relationship(s) between a playback system's speaker(s) and the screen) allows for dynamic distortion of the perceived positions of on-screen audio elements (e.g., audio elements intended by a content creator to be perceived at predetermined positions on a movie screen during playback in a movie theater), such that the distorted positions have a predetermined size and position relative to the actual size and position of the playback system's display screen. The distorted position need not coincide with the actual display screen of the playback system, and exemplary embodiments of the present invention allow a smooth transition between the on-screen and off-screen perceived positions of audio elements whose positions change during playback of a program.

[0033] In some embodiments, an object-based audio program is generated, decoded, and / or rendered. The program includes at least one audio object channel and, optionally, at least one speaker channel (e.g., a collection or "bed" of speaker channels), each object channel representing an audio object or collection (e.g., a mix or cluster) of audio objects, with at least one object channel having (e.g., including) corresponding screen-related metadata. The bed of speaker channels may be a conventional mix (e.g., a 5.1-channel mix) of speaker channels of a type that may be included in a conventional speaker-channel-based broadcast program that does not include object channels. The method may include encoding audio data representing each of the object channels (and optionally the collection of speaker channels) to generate the object-based audio program. In response to an object-based audio program generated by exemplary embodiments of this class, the rendering step may generate speaker feeds representing the mix of audio content for each speaker channel and each object channel.

[0034] Aspects of the present invention include systems or devices configured (e.g., programmed) to implement any of the method embodiments of the present invention, and computer-readable media (e.g., disks) storing (e.g., non-transitory) code for implementing any of the method embodiments, or steps thereof. For example, systems of the present invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed and / or otherwise configured by software or firmware to perform any of a variety of operations on data, including the method embodiments, or steps thereof, of the present invention. Such a general-purpose processor may be or include a computer system including an input device, memory, and processing circuitry programmed (and / or otherwise configured) to perform the method embodiments (or steps thereof) of the present invention in response to data presented to it.

[0035] In one class of embodiments, the present invention is a system configured to generate an object-based audio program that presents at least one audio object channel (typically a collection of object channels) and at least one speaker channel (typically a collection of speaker channels). Each audio object channel represents an object or collection of objects (e.g., a mix or cluster) and typically includes corresponding object-relational metadata. The collection of speaker channels may be a conventional mix of speaker channels (e.g., a 5.1 channel mix) of the type that may be included in a conventional speaker channel-based broadcast program that does not include object channels. In response to the object-based audio program generated by an exemplary embodiment of the system, a spatial rendering subsystem may generate speaker feeds that represent the mix of speaker channels and audio content for each object channel.

[0036] In one class of embodiments, the present invention is an audio processing unit (APU) including a buffer memory (buffer) that stores (e.g., non-temporarily) at least one frame or other segment (including audio content) of an audio program generated by any embodiment of the method of the present invention. If the program is an object-based audio program, the stored segment typically includes audio content for both speaker channel bed and object channels and corresponding screen-related metadata. In another class of embodiments, the present invention is an APU including a buffer memory (buffer) that stores (e.g., non-temporarily) at least one frame or other segment of a speaker channel-based audio program, where the segment includes audio content for at least one set of speaker channels generated as a result of performing distortion of the audio content of the object-based audio program in accordance with embodiments of the present invention. The segment may include audio content for at least two selectable sets of speaker channels of the speaker channel-based program, where at least one of the sets is generated as a result of distortion in accordance with embodiments of the present invention.

[0037] Exemplary embodiments of the present invention are configured to implement real-time generation of encoded, object-based audio bitstreams for transmission (or otherwise delivery) to external rendering systems (e.g., devices). [Brief explanation of the drawings]

[0038] [Figure 1] A diagram of the front wall (W) of a movie theater with a display screen (S) and left and right front speakers (L and R) and a front center speaker (C) mounted on (or near) the front wall. [Figure 2]Diagram of a home theater system's display screen (S'), left and right front speakers (L' and R'), and front center speaker (C') mounted on (or near) the front wall (W') of a room. [Figure 3] FIG. 1 is a block diagram of an embodiment of a system configured to perform an embodiment of the method of the present invention. [Figure 4] 1 is a diagram of a playback environment including a display screen (playback screen S') and speakers (L', C', R', Ls and Rs) of the playback system. [Figure 4A] 5 is a diagram of the playback environment of FIG. 4, illustrating an embodiment in which the parameter EXP has a different value than the embodiment described with reference to FIG. [Figure 4B] 5 is a diagram of the playback environment of FIG. 4, illustrating an embodiment in which the parameter EXP has a different value than the embodiment described with reference to FIGS. 4 and 4A. [Figure 5] FIG. 2 is a block diagram of elements of a system configured to implement another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] Notation and Nomenclature Throughout this disclosure, including the claims, the expression performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation).

[0040] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0041] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0042] Throughout this disclosure, including the claims, the expressions "audio processor" and "audio processing unit" are used interchangeably in a broad sense to refer to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).

[0043] Throughout this disclosure, including the claims, the term "metadata" (e.g., in the phrase "screen-related metadata") refers to data that is separate and distinct from the corresponding audio data (the audio content of a bitstream that also includes the metadata). The metadata is associated with the audio data and indicates at least one feature or characteristic of the audio data (e.g., what type(s) of processing have been or should be performed on the audio data, or the trajectory of an object represented by the audio data). The association of the metadata with the audio data is time-synchronous. In this way, current (most recently received or updated) metadata may indicate that the corresponding audio data contemporaneously includes the results of audio data processing of the indicated characteristics and / or types.

[0044] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0045] Throughout this disclosure, including the claims, the following expressions have the following definitions.

[0046] Speaker and loudspeaker are used interchangeably to refer to any sound-producing transducer. This definition includes loudspeakers implemented as multiple transducers (e.g., a woofer and a tweeter).

[0047] Speaker Feed: An audio signal applied directly to a loudspeaker or to an amplifier and loudspeaker in series.

[0048] Channel (or "audio channel"): A monophonic audio signal. Such a signal can typically be rendered to be equivalent to applying the signal directly to a loudspeaker in a desired or nominal position. The desired position may be static, as is typically the case with a physical loudspeaker, or it may be dynamic.

[0049] Audio Program: A collection of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation).

[0050] Speaker Channel (or "Speaker Feed Channel"): An audio channel associated with a specified loudspeaker (in a desired or nominal position) or associated with a specified speaker zone within a defined speaker configuration. A speaker channel is rendered to be equivalent to applying the audio signal directly to the specified loudspeaker (in a desired or nominal position) or to a speaker within a specified speaker zone.

[0051] Object Channel: An audio channel that describes the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel determines a parametric audio source description (e.g., metadata describing the parametric audio source description is included within or provided along with the object channel). The source description may determine the sound emitted by the source (as a function of time), the apparent position of the source as a function of time (e.g., 3D spatial coordinates), and optionally at least one additional parameter characterizing the source (e.g., apparent source size or width).

[0052] Object-Based Audio Program: An audio program that includes a collection of one or more object channels (and optionally at least one speaker channel) and optionally associated metadata (e.g., metadata indicating the trajectory of the audio object(s) emitting the sound(s) represented by the object channels, or otherwise indicating the desired spatial audio presentation of the sound(s) represented by the object channels, or metadata indicating the identity of at least one audio object that is the source of the sound(s) represented by the object channels).

[0053] Rendering: The process of converting an audio program into one or more speaker feeds, or converting an audio program into one or more speaker feeds and then converting the speaker feeds into sound using one or more loudspeakers. (In the latter case, rendering is sometimes referred to herein as rendering "with" loudspeakers.) An audio channel can be trivially rendered ("at" the desired location) by applying the signal directly to a physical loudspeaker at the desired location. Alternatively, one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent (to a listener) to such a trivial rendering. In this latter case, each audio channel may be converted into one or more speaker feeds to be applied to a loudspeaker or loudspeakers at a known location, typically different from the desired location, so that the sound produced by the loudspeakers in response to the feed will be perceived to emanate from the desired location. Examples of such virtualization techniques include binaural rendering via headphones (e.g., using Dolby Headphone processing to simulate up to 7.1 channel surround sound for headphone wearers) and wave field synthesis.

[0054] Detailed Description of the Embodiments of the Invention An example embodiment of the system of the present invention (and the method performed by the system) will now be described with reference to FIGS.

[0055] 3 is a block diagram of an example audio processing pipeline (audio data processing system), one or more of the system's elements configured in accordance with an embodiment of the present invention. The system includes the following elements coupled together as shown: a capture unit, a production unit 3 (which includes an encoding subsystem), a delivery subsystem 5, a decoder 7, an object processing subsystem 9, a controller 10, and a rendering subsystem 11. In variations on the illustrated system, one or more of the elements are omitted or additional audio data processing units are included. Typically, elements 7, 9, 10, and 11 are included in a playback system (e.g., an end-user's home theater system).

[0056] Capture unit 1 is typically configured to generate PCM (time-domain) samples representing audio content and output the PCM samples, which represent multiple streams of audio captured by a microphone. Production unit 3 is configured to accept the PCM samples as input and output an object-based audio program representing the audio content. The program is typically or includes an encoded (e.g., compressed) audio bitstream. The encoded bitstream data representing the audio content is sometimes referred to herein as "audio data." When the encoding subsystem of production unit 3 is configured according to an exemplary embodiment of the present invention, the object-based audio program output by unit 3 represents (i.e., includes) multiple speaker channels of audio data ("beds" of speaker channels), multiple object channels of audio data, and metadata (including screen-related metadata corresponding to each object channel and, optionally, screen-related metadata corresponding to each speaker channel).

[0057] In a typical implementation, unit 3 is configured to output the object-based audio program generated therein.

[0058] In another implementation, unit 3 includes a remix subsystem coupled and configured to generate a speaker channel-based audio program (including speaker channels but no object channels) in response to an object-based audio program, and unit 3 is configured to output the speaker channel-based audio program. Remix subsystem 6 of the system of Figure 5 is another example of a remix subsystem coupled and configured to generate a speaker channel-based audio program (a program including speaker channels but no object channels "SP") in accordance with an embodiment of the present invention in response to an object-based audio program ("OP") generated by encoder 4 (of Figure 5) in accordance with an embodiment of the present invention.

[0059] The delivery subsystem of Figure 3 is configured to store and / or transmit (e.g., broadcast) programs (e.g., object-based audio programs or speaker channel-based audio programs generated in response to the object-based audio programs) generated by and output from unit 3. For simplicity, the system of Figure 3 will be described (and referenced) with the assumption that the programs generated by and output from unit 3 are object-based audio programs (unless it is clear from the context, description, or reference that the programs generated by and output from unit 3 are speaker channel-based audio programs).

[0060] In an exemplary embodiment of the system of Figure 3, subsystem 5 implements delivery of object-based audio programs to decoder 7. For example, subsystem 5 may be configured to store the programs (e.g., on a disk) and provide the stored programs to decoder 7. Alternatively, subsystem 5 may be configured to transmit the programs to decoder 7 (e.g., over a broadcast system or an Internet Protocol or other network).

[0061] Decoder 7 is coupled and configured to accept (receive or read) and decode programs delivered by delivery subsystem 5. If the programs are object-based programs and decoder 7 is configured according to an exemplary embodiment of the present invention, the output of decoder 7 in exemplary operation includes: a stream of audio samples indicating the bed of the speaker channels of said program (and optionally a corresponding stream of screen-related metadata); and A stream of audio samples representing the object channel of said program and a corresponding stream of screen-related metadata.

[0062] Object processing subsystem 9 is coupled to receive the decoded speaker channels, object channels, and object-related metadata of the delivered program (from decoder 7). Subsystem 9 is coupled and configured to use the screen-related metadata to perform warping on the object channels (or on a selected subset of the object channels, or on at least one mixture (e.g., cluster) of some or all of the object channels) and output the resulting object channels and / or mixtures to rendering subsystem 11. Subsystem 9 typically also outputs to rendering subsystem 11 the object-related metadata (parsed by decoder 7 from the program delivered by subsystem 5 and presented by decoder 7 to subsystem 9) corresponding to the object channels and / or mixtures it outputs to subsystem 11. Subsystem 9 is typically configured to pass the decoded speaker channels from decoder 7 through unchanged (to subsystem 11).

[0063] If the program delivered to decoder 7 is a speaker channel-based audio program (generated from an object-based program in accordance with an embodiment of the present invention), subsystem 9 may be implemented as (or replaced by) a simple speaker channel selection system configured to implement distortion in accordance with the present invention by selecting some of the speaker channels of the program (in a manner described in more detail below) and presenting the selected channels to rendering subsystem 11.

[0064] The distortion performed by subsystem 9 may be controlled, at least in part, by data provided to subsystem 9 by controller 10 (e.g., in response to user manipulation of controller 10 during system setup). Such data may indicate characteristics of the playback system speakers and display screen (e.g., the relative sizes and positions of the playback system screen and playback system speakers) and / or may include at least one distortion intensity parameter and / or at least one off-screen distortion parameter. The distortion performed by subsystem 9 is typically determined by at least one distortion intensity parameter and / or at least one off-screen distortion parameter indicated by the program's screen-related metadata (delivered to decoder 7) and / or at least one distortion intensity parameter and / or at least one off-screen distortion parameter provided to subsystem 9 by controller 10.

[0065] Rendering subsystem 11 of FIG. 3 is configured to render audio content determined by the output of subsystem 9 for playback through speakers (not shown) of a playback system. Subsystem 11 is configured to map audio objects determined by the object channels (or mixes) output by subsystem 9 to available speaker channels using rendering parameters output from subsystem 9 (e.g., spatial position and level values ​​indicated by object-relational metadata output from subsystem 9). Rendering subsystem 11 also receives any bed of speaker channels passed through subsystem 9. Typically, subsystem 11 is an intelligent mixer and is configured to determine speaker feeds for the available speakers. This determination involves mapping one or more objects (or mixes) to each of several individual speaker channels and mixing those objects (or mixes) with the "bed" audio content indicated by each corresponding speaker channel of the program's speaker channel bed.

[0066] Typically, the output of subsystem 11 is a collection of speaker feeds that are presented to and drive playback system loudspeakers (such as those shown in FIG. 4).

[0067] One aspect of the present invention is an audio processing unit (APU) configured to perform any embodiment of the method of the present invention. Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems (preprocessors), post-processing systems (postprocessors), audio bitstream processing systems, and combinations of such elements. Examples of APUs are production unit 3, decoder 7, object processing subsystem 9, and rendering subsystem 11 of FIG. 3. All of these exemplary APU implementations configured to perform an embodiment of the method of the present invention are contemplated and described herein.

[0068] In one class of embodiments, the present invention is an APU including a buffer memory (buffer) that stores (e.g., non-temporarily) at least one frame or other segment of an audio program (including audio content) generated by any embodiment of the method of the present invention. If the program is an object-based audio program, the stored segment typically includes audio content for the bed and object channels of the speaker channels and corresponding screen-related metadata. An example of such an APU is an implementation of production unit 3 of FIG. 3 that includes an encoding subsystem 3B (configured to generate an object-based audio program in accordance with an embodiment of the present invention) and a buffer 3A coupled to subsystem 3B, where buffer 3A stores (e.g., non-temporarily) at least one frame or other segment of the object-based audio program (including audio content for the bed and object channels of the speaker channels and corresponding screen-related metadata). Another example of such an APU is an implementation of decoder 7 of FIG. 3 that includes a buffer 7A and a decoding subsystem 7B (coupled to buffer 7A). Here, buffer 7A stores (e.g., in a non-transient manner) at least one frame or other segment of an object-based audio program (including the audio content of the bed and object channels of the speaker channels and the corresponding screen-related metadata) delivered from subsystem 5 to decoder 7. Decode subsystem 7B is configured to parse the program and perform any necessary decoding on the program.

[0069] In another class of embodiments, the present invention is an APU including a buffer memory (buffer) that stores (e.g., non-temporarily) at least one frame or other segment of a speaker channel-based audio program, where the segment includes audio content for at least one set of speaker channels generated as a result of performing distortion on the audio content of the object-based audio program in accordance with embodiments of the present invention. The segment may also include audio content for at least two selectable sets of speaker channels of the speaker channel-based program, where at least one of the sets is generated as a result of distortion in accordance with embodiments of the present invention. An example of such an APU is an implementation of production unit 3 of FIG. 3 that includes encoding subsystem 3B (configured to generate speaker channel-based audio programs in accordance with embodiments of the present invention, including by performing distortion on the audio content of the object-based audio program also generated by unit 3) and buffer 3A coupled to subsystem 3B. Here, buffer 3A stores (e.g., non-temporarily) at least one frame or other segment of a speaker channel-based audio program (comprising audio content of at least two selectable sets of speaker channels, where at least one of the sets is generated as a result of performing distortion according to an embodiment of the present invention on the audio content of the object-based audio program). Another example of such an APU is the implementation of decoder 7 of FIG. 3, which includes buffer 7A and decode subsystem 7B (coupled to buffer 7A). Here, buffer 7A stores (e.g., non-temporarily) at least one frame or other segment of a speaker channel-based audio program generated by an exemplary embodiment of unit 3, delivered to decoder 7 from unit 3 via subsystem 5.The decode subsystem 7B is configured to parse the program and perform any necessary decoding on the program. Another example of such an APU is an implementation of the remix subsystem 6 of FIG. 5 that includes a subsystem 6B (configured to generate a speaker channel-based audio program in accordance with an embodiment of the present invention, including by performing distortion on the audio content of the object-based audio program, typically including screen-related metadata, generated by the encoder 4 of FIG. 5) and a buffer 6A coupled to the audio processing subsystem 6B, where the buffer 6A stores (e.g., non-transitory) at least one frame or other segment of the speaker channel-based audio program generated by the subsystem 6B (including audio content for at least two selectable sets of speaker channels, where at least one of the sets is generated as a result of distortion in accordance with an embodiment of the present invention).

[0070] A typical embodiment of the present invention assumes that the playback environment is a unit cube with width along the x-axis, depth along the y-axis (perpendicular to the x-axis), and height along the z-axis (perpendicular to both the x-axis and the y-axis). The positions at which audio elements (sound sources) represented by an audio program (i.e., audio objects represented by object channels or sound sources represented by speaker channels) are rendered are identified in this unit cube using Cartesian coordinates (x, y, z). Each of the x and y coordinates ranges in the interval [0, 1]. For example, Figure 4 shows a diagram of a playback environment (room) including the display screen (playback screen S') and speakers (L', C', R', Ls, and Rs) of a playback system. The playback screen S' in Figure 4 has width W1 along the x-axis and its center is located along the vertical axis at the center of the room's front wall (the plane where y = 0). The back wall of the room (which has width W2) is the plane where y = 1. The front speakers L', C', and R' are positioned near the front wall of the room, the left surround speaker Ls is positioned near the left wall of the room (the plane where x=0), and the right surround speaker Rs is positioned near the right wall of the room (the plane where x=1).

[0071] Typically, the z coordinate of the playback environment is assumed to have a fixed value (nominally corresponding to ear level of the user of the playback system). Alternatively, the z coordinate of the rendering position can be allowed to vary (e.g., over the interval [-1,1], where the room is assumed to have width equal to 1, depth equal to 1, and height equal to 2) in order to render objects (or other sound sources) at positions that are perceived to be below or above ear level.

[0072] In some embodiments, screen parameterization and / or distortion is achieved using all or some of the following parameters (which may be determined during authoring and / or encoding, or may be indicated by screen-related metadata in the delivered program): Audio element (e.g. object) position relative to the reference screen; the degree of on-screen distortion (e.g., a parameter indicating the maximum degree of distortion performed in the plane of the reproduction screen or parallel to said plane). It is contemplated that the authoring may typically specify the distortion as a binary decision, and the encoding process may modify the binary decision into a continuous (or nearly continuous) variable ranging from no distortion to full (maximum) distortion; Desired off-screen distortion (e.g., one or more parameters indicating how or to what extent distortion in a plane at least substantially parallel to the plane of the playback screen should be performed as a function of distance at least substantially perpendicular to the plane of the playback screen). Authoring can define a parameter(s) indicating how or to what extent distortion should be performed as the perceived distorted position of an audio element moves away from the playback screen in a direction perpendicular to the plane of the playback screen. In some cases, such parameters are not delivered with the program (and can instead be determined by the playback system); the reference screen width for the reference room (or for the reference L / R speakers used during authoring). Typically, this parameter is equal to 1.0 for cinemas (i.e. for audiovisual programs authored for playback in cinemas); Reference screen center position relative to the reference room (or relative to the reference L / R speakers used during authoring). Typically, this parameter is equal to (0.5,0,0.5) for cinemas.

[0073] In some embodiments, screen parameterization and / or distortion is achieved using all or some of the following parameters (which are typically determined by the playback system, e.g., in a home theater setup): - playback screen width relative to the playback room (or relative to the playback system L / R speakers). For example, this parameter may have a default value of 1.0 (e.g., if the end user does not specify a playback screen size, the playback system will assume that the playback screen matches the playback room width, which effectively disables distortion); Desired off-screen distortion (e.g., one or more parameters indicating how or to what extent distortion in a plane at least substantially parallel to the plane of the playback screen should be performed as a function of distance at least substantially perpendicular to the plane of the playback screen). In some embodiments, the playback system (e.g., controller 10 in the embodiment of FIG. 3) is configured to allow custom settings that indicate how or to what extent distortion should be performed as a function of distance of the perceived distorted position of the audio element from the plane of the playback screen (in a direction at least substantially perpendicular to the plane of the playback screen). In typical embodiments, it is expected that the screen-related metadata of the program will indicate (e.g., include at least one off-screen distortion parameter that indicates) a fixed or default function (which can be replaced by an alternative function specified by the user, e.g., during playback system setup). This will at least partially determine how distortion should be performed as a function of distance of the perceived distorted position of the audio element from the plane of the playback screen. · playback screen aspect ratio (for example, with default value 1.0); Playback screen center position (for example, with default value (0.5,0,0.5)).

[0074] In some embodiments, distortion is achieved using other parameters (instead of or in addition to some or all of the parameters mentioned above) that may be indicated by the screen-related metadata of the delivered program. For example, for each channel (object channel or speaker channel) of the program (or for each of some of the channels of the program), one or more of the following parameters may be provided:

[0075] 1. Distortion Enable. This parameter indicates whether processing should be performed to distort the perceived position of at least one audio element determined by a channel. This parameter is typically a binary value indicating whether distortion should be performed. An example is the "apply_screen_warping" value described below.

[0076] 2. The degree of distortion (e.g., one or more floating-point values ​​or one or more other non-binary parameters, each having any of many different values ​​in the range [0, 1] or other predetermined range). Such distortion degree parameter(s) typically determine the maximum degree of distortion to be performed in the plane of the reproduction screen (or parallel to the plane) by modifying a function that controls distortion from positions in the plane of the reference screen (or parallel to the plane) to positions in the plane of the reproduction screen (or parallel to the plane). The distortion degree parameter (or parameter set) can be different for distortion along (or parallel to) the axis along which the reproduction screen is wide (e.g., the x-axis) and for distortion along (or parallel to) the axis along which the reproduction screen is high (e.g., the z-axis).

[0077] 3. Depth distortion (e.g., one or more parameters each having an arbitrary floating-point value in a predetermined range [1,N], e.g., N=2). Such parameter(s) (sometimes referred to herein as "off-screen distortion parameters") typically modify a function that controls the distortion of off-screen audio elements to control the degree of distortion or maximum distortion of audio element rendering positions as a function of distance (depth) from the plane of the playback screen. For example, such a parameter may control the degree of distortion (at least substantially parallel to the plane of the playback screen) of a sequence of rendering positions of audio elements that are intended to be perceived as "flying" from the playback screen (at the front of the playback room) to the rear of the playback room, or vice versa.

[0078] For example, in one class of embodiments, distortion is achieved using screen-related metadata included in the audio program (e.g., an object-based audio program), where the screen-related metadata indicates at least one non-binary value (e.g., a continuously variable or scalar value having any one of many values ​​within a predetermined range) that indicates the maximum degree of distortion to be performed by the playback system (e.g., the maximum degree of distortion to be performed in the plane of or parallel to the plane of the playback screen). For example, the non-binary value may be a floating-point value ranging from a maximum value (indicating that full distortion should be performed, e.g., an audio element position defined by the program to be at the right edge of the reference screen should be distorted to a distorted position at the right edge of the playback screen) to a minimum value (indicating that no distortion should be performed). In one example, a non-binary value at the midpoint of the range may indicate that half distortion (50% distortion) should be performed (e.g., an audio element position defined by the program to be at the right edge of the reference screen should be distorted to a distorted position halfway between the right edge of the playback room and the right edge of the playback screen).

[0079] In some embodiments of this class, the program is an object-based audio program that includes such metadata for each object channel of the program, the metadata indicating a maximum degree of distortion to be performed for each corresponding object. For example, the metadata may indicate, for each object represented by a different object channel, a different maximum degree of distortion in or parallel to the plane of the playback screen. As another example, the metadata may indicate, for each object represented by a different object channel, a different maximum degree of distortion in a vertical direction in or parallel to the plane of the playback screen (e.g., parallel to the z-axis in FIG. 4 ) and a different maximum degree of distortion in or parallel to the plane of the playback screen (e.g., parallel to the x-axis in FIG. 4 ).

[0080] In some embodiments of this class, the audio program also includes screen-related metadata indicating at least one characteristic of the off-screen distortion (and the distortion is achieved using the screen-related metadata) (e.g., indicating the manner or degree to which the distortion is performed in a plane at least substantially parallel to the plane of the reproduction screen as a function of distance at least substantially perpendicular to the plane of the reproduction screen). In some such embodiments, the program is an object-based audio program that includes such metadata for each object channel of the program, the metadata indicating at least one characteristic of the off-screen distortion to be performed for each corresponding object. For example, the program may include such metadata for each object channel indicating the type of off-screen distortion to be performed for each corresponding object (i.e., the metadata may specify different types of off-screen distortion for objects corresponding to each object channel).

[0081] Next, an example of how to process an audio program to implement distortion according to an embodiment of the present invention will be described.

[0082] In an exemplary method, the screen-related metadata for an audio program includes at least one distortion degree parameter (for each channel whose audio content is to be distorted) having a non-binary value indicating the maximum degree of distortion to be performed by the playback system on at least one audio element represented by that channel in the plane of or parallel to the plane of the playback screen, whereby audio elements that the program indicates should be rendered at positions (in the plane of the reference screen and relative to the reference screen) are rendered at distorted positions (in the plane of the playback screen and relative to the playback screen). Preferably, one or two such distortion degree parameters are included for each channel: one indicating a distortion factor (e.g., the value XFACTOR, described below) that controls how much distortion to be applied to at least one audio element represented by that channel in the horizontal direction (e.g., along the x-axis in FIG. 4 ) (i.e., the maximum degree of distortion to be applied) and / or one indicating a distortion factor that controls how much distortion to be applied to at least one audio element represented by that channel in the vertical direction (e.g., along the z-axis in FIG. 4 ) (i.e., the maximum degree of distortion to be applied). The program's screen-related metadata also indicates off-screen distortion parameters (e.g., the EXP values ​​described below) for each channel that control at least one characteristic of the off-screen distortion to be performed as a function of distance (of the distorted position of the corresponding audio element) perpendicular to the plane of the playback screen. For example, the off-screen distortion parameters may control the manner, degree, or maximum distortion of the distorted position of an audio element as a function of depth (distance along the y-axis in FIG. 4) perpendicular to the plane of the playback screen.

[0083] In an exemplary embodiment, the screen-related metadata for a program also includes a binary value (referred to herein as apply_screen_warping) for the program (or for each segment of a sequence of segments of the program). If the value of apply_screen_warping (for a program or its segment) indicates "off," no warping is applied to the corresponding audio content by the playback system. For example, warping can be disabled in this manner for audio content that should be rendered with perceived positions in the plane of (or coincident with) the playback screen, but that does not need to be closely associated with the visuals (e.g., audio content that is music or ambient sounds). If the value of apply_screen_warping (for a program or its segment) indicates "on," the playback system applies warping to the corresponding audio content as follows: The parameter apply_screen_warping is not an example of a "distortion degree" parameter of the type used and / or generated in accordance with the present invention.

[0084] The following description assumes that the program is an object-based program, and that each channel that undergoes distortion is an object channel that represents an audio object with an undistorted position determined by the program (which may be a time-varying position). It will be clear to one skilled in the art how to modify this description to implement distortion of speaker channels of a program where the speaker channel represents at least one audio element with an undistorted position determined by the program (which may be a time-varying position). The following description also assumes that the playback environment is as shown in Figure 4 and that the playback system is configured to generate five speaker feeds (for speakers L', C', R', Ls, and Rs shown in Figure 4) in response to the program.

[0085] In this exemplary embodiment, the playback system (e.g., subsystem 9 of the system of FIG. 3) receives from the program (e.g., from the program's screen-related metadata) the following values ​​indicating the undistorted positions of the objects (to be rendered at distorted positions to be determined by the playback system): Xs=(x-RefSXcenterpos) / RefSWidth where x is the undistorted object position along the horizontal (x or "width") axis relative to the left edge of the reference screen, RefSXcenterpos is the position of the center point of the reference screen along the horizontal axis, and RefSWidth is the width (along the horizontal axis) of the reference screen.

[0086] The playback system (e.g., subsystem 9 in the system of Figure 3) uses the program's screen-related metadata (and other data that indicates the playback system configuration) to determine the following values: Xwarp=Xs*SWidth+SXcenterpos YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] where Xwarp is the raw (unscaled) warped object position along the horizontal (x or "width") axis relative to the left edge of the playback system display screen ("playback screen"), Xs is the warped object position along the horizontal axis relative to the center point of the playback screen, SXcenterpos is the position of the center point of the playback screen along the horizontal axis, and SWidth is the width (along the horizontal axis) of the playback screen; YFACTOR is a depth distortion factor that indicates the degree of distortion along the horizontal (width) axis as a function of position along the depth axis (the y axis in Figure 4) perpendicular to the plane of the reproduction screen, y is the distorted object position along the depth axis, and EXP is a predetermined (e.g., user-selected) constant that is an example of what we call an "off-screen distortion" parameter; X' represents the warped object position along the horizontal axis relative to the left edge of the playback screen (a scaled version of the raw warped object position Xwarp) (so the warped object position in the horizontal plane of the playback environment is a point with coordinates X',y), and XFACTOR is the width axis distortion parameter indicated by the program's screen-related metadata (this may be determined during program authoring, mixing, remixing or encoding). XFACTOR is an example of what we refer to as a "warp factor" parameter.

[0087] The distortion of an undistorted object position (determined by the program) along the vertical (z or "height") axis to a distorted position along the vertical axis relative to the reproduction screen can be performed in a manner determined by a trivial modification of the above equation (replacing references to the horizontal or x-axis with references to the vertical or z-axis) to take into account the aspect ratio of the reference screen and the aspect ratio of the reproduction screen.

[0088] The parameter XFACTOR has a value ranging from 0 to 1 inclusive (i.e., one of at least three values ​​in this range, and typically one of many). The value of XFACTOR controls the degree to which distortion is applied along the horizontal axis. If XFACTOR=1, full distortion is performed along the horizontal axis (so that even if the object's undistorted position is off the playback screen, its distorted position will be on the playback screen). If XFACTOR=1 / 2 (or some other value less than 1), a reduced amount of distortion is performed along the x-axis (so that if the object's undistorted position is far off the playback screen, for example at the left front playback speaker, its distorted position may also be off the playback screen, for example halfway between the left front speaker and the left edge of the playback screen). For various reasons, it can be useful to set XFACTOR to a value less than 1 but greater than 0. For example, this may be the case when distortion is desired, but full distortion on a small playback screen is deemed undesirable, or when audio object position is only loosely coupled to display screen size (e.g., for diffuse sound sources).

[0089] The parameter YFACTOR is used to control the degree of distortion (along the horizontal and / or vertical axis) as a function of the distorted position of the audio object along the depth axis, and the value of the parameter YFACTOR is a function of the distorted position of the object along the depth axis. In the above example, this function is an exponential function. In alternative embodiments, other functions that are variations of or otherwise different from this exemplary exponential function are used to determine YFACTOR (e.g., YFACTOR may be the cosine or a power of the cosine of the distorted object position y along the depth axis). YFACTOR=y EXPIn the above example, when EXP is greater than 0 (as would be expected to be a typical choice), sounds with undistorted positions at the front of the playback room (i.e., on the playback screen) will be distorted more (in the x and / or z directions perpendicular to the depth axis) than sounds with undistorted positions further from the front of the room (i.e., on the back wall of the playback room). When EXP is greater than 0 and y=0 (i.e., when the object's distorted and undistorted positions are in the plane of the playback screen at the front of the playback room), then YFACTOR=0, and the distorted position along the horizontal "width" axis (X') is determined by the undistorted position along the width axis (x) and the parameters XFACTOR and Xwarp. If EXP is greater than 0 and y=1 (i.e., the object's distorted and undistorted positions are at the rear of the playback room), then YFACTOR=1 and the distorted position along the horizontal "width" axis (X') is equal to the undistorted position along the width axis (x), so effectively no distortion is performed on the object (along the width axis) in this case.

[0090] As a more specific example, the audio object A1 in FIG. 4 has its non-distorted position (and thus its distorted position) within the plane of the playback screen S' at the front of the playback room (i.e., y = y1 = 0). If EXP is greater than 0, then YFACTOR = 0 to perform a horizontal-axis distortion on object A1, and the distortion places the distorted position of object A1 at some position X' = x1, y = 0 that coincides with the playback screen S' (as shown, for example, in FIG. 4). The audio object A2 in FIG. 4 has its non-distorted position (and thus its distorted position) between the front and rear walls of the playback room (at 0 < y2 < 1). If EXP is greater than 0, then YFACTOR is greater than 0 to perform a horizontal-axis distortion on object A2, and the distortion places the distorted position of object A2 at some position X' = x2, y = y2 along the line segment between points T1 and T2 (as shown, for example, in FIG. 4). The separation between points T1 and T2 is W3 (as shown in FIG. 4), and since EXP is greater than 0, W3 satisfies W1 < W3 < W2. Here, W1 is the width of the screen S', and W2 is the width of the playback room. The specific value of EXP determines the value of W3. W3 is the range of widths within which an object at depth y = y2 can be mapped relative to the playback screen S' by the distortion. If EXP is greater than 1, then the distortion places the distorted position of object A2 at a position between curves C1 and C2 (shown in FIG. 4). Here, the separation (W3) between curves C1 and C2 is a function that increases exponentially with the depth parameter y (as shown in FIG. 4), and the separation W3 increases more rapidly as y has a larger value (with the increase in the value of y) and increases more slowly as y has a smaller value (with the increase in the value of y).

[0091] In another embodiment (described below with reference to FIG. 4A ) that is a variation on the exemplary embodiment described with reference to curves C1 and C2 in FIG. 4 , EXP is equal to 1. Thus, the distortion places the distorted position of object A2 at a position between two curves (e.g., curves C3 and C4 in FIG. 4A ), where the separation between curves C3 and C4 is a linearly increasing function of the depth parameter y. In another embodiment (described below with reference to FIG. 4B ) that is a variation on the exemplary embodiment described with reference to curves C1 and C2 in FIG. 4 , EXP is greater than 0 but less than 1. Thus, the distortion places the distorted position of object A2 at a position between two curves (e.g., curves C5 and C6 in FIG. 4B ), where the separation between curves C5 and C6 is a logarithmically increasing function of the depth parameter y. The separation between the curves increases more rapidly (as the value of y increases) for smaller values ​​of y and more slowly (as the value of y increases) for larger values ​​of y. Embodiments in which EXP is 1 or less are expected to be typical, since in such embodiments the distortion effect decreases more rapidly with increasing values ​​of y (i.e., with increasing distance of the distorted position from the screen) than when EXP is greater than 1. When EXP is less than 1, the distortion effect decreases rapidly as the distorted position begins to move away from the screen, and then decreases progressively more slowly as the distorted position progresses further from the screen, until the distorted position reaches the back wall where no distortion is performed.

[0092] Next, another class of embodiments is described in which a speaker channel-based audio program (including speaker channels but not object channels) is generated in response to an object-based program in a manner that includes a distortion step (e.g., using screen-related metadata). The speaker channel-based audio program includes at least one set of speaker channels that is generated as a result of distorting the audio content of the object-based program to a degree determined, at least in part, by a distortion parameter (and / or using an off-screen distortion parameter) and is intended for playback by loudspeakers located at predetermined positions relative to the playback system display screen. In some embodiments of this class, the speaker channel-based audio program is generated to include two or more selectable sets of speaker channels, at least one of which is generated as a result of the distortion and is intended for playback by loudspeakers located at predetermined positions relative to the playback system display screen. The generation of the speaker channel-based program supports screen-to-screen rendering by playback systems that are not configured to decode and render object-based audio programs (but are capable of decoding and rendering speaker channel-based programs). Typically, speaker channel-based programs are generated by a remix system that has knowledge of (or assumes) a particular playback system speaker and screen configuration. Typically, the object-based programs (in response to which the speaker channel-based programs are generated) include screen-related metadata that supports screen-to-screen rendering of the object-based programs by a suitably configured playback system (capable of decoding and rendering object-based programs).

[0093] This class of embodiments is particularly useful when it is desired to implement to-screen rendering but the available playback system(s) are not configured to render object-based programs. To implement to-screen rendering of an audio program that includes only speaker channels (no object channels), an object-based program that supports to-screen rendering is first generated in accordance with an embodiment of the present invention. Then, in response to the object-based program, a speaker-channel-based audio program (that supports to-screen rendering) is generated. The speaker-channel-based audio program may include at least two selectable sets of speaker channels, and the playback system may be configured to render a selected one of those sets of speaker channels to implement the to-screen rendering.

[0094] Common speaker channel configurations assumed by speaker-channel-based programs include stereo (for playback using two speakers) and 5.1 surround sound (for playback over five full-range speakers). In such channel configurations, speaker channels (audio signals) are, by definition, associated with loudspeaker positions, and the perceived location at which audio elements are rendered (as indicated by the audio content of the channels) is typically determined based on the assumed speaker positions in the playback environment or the assumed speaker positions relative to a reference listening position.

[0095] In some embodiments in which a speaker channel-based audio program is generated (in response to an object-based program), screen-related warping (scaling) functions enabled by the object-based program's screen-related metadata are utilized to generate speaker channels (for the speaker channel-based program) that are associated with loudspeakers having predetermined positions relative to the playback screen. Typically, a particular playback screen size and shape and position are assumed by the system generating the speaker channel-based program. For example, in response to an object-based program, a speaker channel-based program can be generated to include the following two sets of speaker channels (and optionally other speaker channels): a first set of conventional left (L) and right (R) front speaker channels for rendering audio elements at perceived positions determined relative to a reference screen (e.g., in a cinema mixing facility); and A second set of left and right front speaker channels, which may be referred to as "left screen" (Lsc) and "right screen" (Rsc), for rendering the same audio elements at perceived locations determined relative to the left and right edges of an assumed playback display screen (e.g., at a remix facility or remix stage of a mixing facility) (where the playback screen and playback system front speakers are assumed to have predetermined relative sizes, shapes, and positions).

[0096] Typically, the channels of a speaker-channel-based program that are produced as a result of distortion (e.g., the Lsc and Rsc channels) can be rendered to allow a closer proximity match between the image displayed on the playback screen and the corresponding rendered sound.

[0097] By selecting and rendering the usual left (L) and right (R) front speaker channels, the playback system can render the selected channels so that the determined audio elements are perceived to have undistorted positions. By selecting and rendering the "left screen" (Lsc) and "right screen" (Rsc) speaker channels, the playback system can render the selected channels so that the determined audio elements are perceived to have distorted positions (relative to the playback screen). However, the distortion is performed at the time of speaker-channel-based program generation (typically in response to an object-based program that includes screen-related metadata), not by the playback system.

[0098] Some embodiments of this class include: generating an object-based program with screen-related metadata (at the time and location of mixing); then generating a speaker-channel-based program from the object-based program using the screen-related metadata (at the time and location of a "remix," which can be in the same location as the original mixing, e.g., to create a recording for home use), including by performing screen-related distortion; and then delivering the speaker-channel-based program to a playback system. The speaker-channel-based program can include multiple selectable sets of channels. The multiple selectable sets include a first set of speaker channels (e.g., conventionally generated L and R channels) that are generated without performing distortion and that exhibit (when rendered) at least one audio element perceived to be in at least one undistorted position, and at least one additional set of speaker channels (e.g., Lsc and Rsc channels) that are generated as a result of distortion of the object-based program content and that exhibit (when rendered) the same audio element but perceived to be in at least one different (i.e., distorted) position. Alternatively, a speaker channel-based program may include only one set of channels (e.g., Lsc and Rsc channels) that exhibit (when rendered) at least one audio element that is produced as a result of the distortion and that is perceived at at least one distorted location, and no other set of channels (e.g., L and R channels) that exhibit (when rendered) the same audio element that is perceived at an undistorted location.

[0099] A speaker channel-based program generated from an object-based program according to an example embodiment includes five front channels: left (L), left screen (Lsc), center (C), right screen (Rsc), and right (R). The Lsc and Rsc channels are generated by performing a warp using the screen-related metadata of the object-based program. To render and play back a speaker channel-based program, a playback system may select and render the L and R channels to drive front speakers at the left and right edges of the playback screen, or may select and render the Lsc and Rsc channels to drive front speakers away from the left and right edges of the playback screen. For example, the Lsc and Rsc channels may be generated with the assumption that they will be used to render audio elements using front speakers located at +30 and -30 degrees azimuth relative to an expected user position, and the L and R channels may be generated with the assumption that they will be used to render audio elements using front speakers located at +15 and -15 degrees azimuth (at the left and right edges of the playback screen) relative to an expected user position.

[0100] For example, the system of Figure 5 includes an encoder 4 configured to generate an object-based audio program ("OP") including screen-related metadata according to an embodiment of the present invention. Encoder 4 may be implemented within or at a mixing facility. The system of Figure 5 also includes a remix subsystem 6 coupled to and configured to generate (according to an embodiment of the present invention) a speaker channel-based audio program ("SP") including speaker channels but not object channels in response to the object-based audio program generated by encoder 4. Subsystem 6 may be implemented within or at a remix facility or as a remix stage of a mixing facility (e.g., a mixing facility in which encoder 4 is also implemented). The audio content of the speaker channel-based program SP includes at least two selectable sets of speaker channels (e.g., one set including channels L and R as discussed above, and another set including channels Lsc and Rsc as discussed above), and subsystem 6 is configured, in accordance with an embodiment of the present invention, to generate at least one of those sets (e.g., channels Lsc and Rsc) as a result of distorting the audio content of the object-based program OP (generated by encoder 4) using screen-related metadata of the program OP (and also using other control data indicating the type and / or degree of distortion, typically not indicated by the screen-related metadata). The speaker channel-based program SP is output from subsystem 6 to delivery subsystem 5. Subsystem 5 can be identical to the above-discussed subsystem 5 of the system of FIG. 3.

[0101] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination thereof (e.g., as a programmable logic array). For example, the system of FIG. 3 (or subsystems 3 or 7, 9, 10, and 11 thereof) may be implemented in appropriately programmed (or otherwise configured) hardware or firmware, such as a programmed general-purpose processor, digital signal processor, or microprocessor. Unless otherwise specified, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs running on one or more programmable computer systems (e.g., a computer system implementing the system of FIG. 3 (or subsystems 3 or 7, 9, 10, and 11 thereof)). Each computer system has at least one processor, at least one data storage system (including volatile and nonvolatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and to generate output information that is applied to one or more output devices, in known fashion.

[0102] Each such program may be implemented in any desired computer language to communicate with a computer system, including machine, assembly, or high-level procedural, logical, or object-oriented programming languages, and in any case, the language may be a compiled or interpreted language.

[0103] For example, when implemented by a sequence of computer software instructions, various functions and steps of embodiments of the present invention may be implemented by a multi-threaded sequence of software instructions executed on suitable digital signal processing hardware, in which case various units, steps and functions of the embodiments may correspond to portions of the software instructions.

[0104] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., semiconductor memory or media, or magnetic or optical media) that, when read by a computer system, configures and operates the computer to perform the procedures described herein. The system of the present invention may also be implemented as a computer-readable storage medium configured with (i.e., having stored thereon) a computer program, which causes the computer system to operate in a specific, predefined manner to perform the functions described herein.

[0105] While implementations have been described by way of example and in terms of exemplary specific embodiments, it is to be understood that implementations of the invention are not limited to the disclosed embodiments. Rather, it is intended to cover various modifications and similar arrangements that will become apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.

[0106] Several aspects will be described. [Aspect 1] A method of rendering an audio program, comprising: (a) determining at least one skewness parameter; (b) performing distortion on the audio content of at least one channel of the program to a degree determined at least in part by the distortion degree parameter corresponding to the channel, each of the distortion degree parameters indicating a maximum degree of distortion to be performed by a playback system on the corresponding audio content of the program. method. [Aspect 2] The method of claim 1, wherein step (a) includes determining at least one off-screen distortion parameter, the off-screen distortion parameter indicating at least one characteristic of off-screen distortion to the corresponding audio content of the program by the playback system, and the distortion performed in step (b) includes off-screen distortion determined at least in part by at least one of the off-screen distortion parameters. Aspect 3 3. The method of claim 2, wherein the off-screen distortion parameters control the degree to which an undistorted position of an audio element along a width axis at least substantially parallel to the plane of a playback screen is distorted as a function of the distance of the distorted position at which the audio element is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 4 determining a value Xs indicating an undistorted position along the width axis of an audio element to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and Xwarp represents the raw warped position of the audio element along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the audio element along a depth axis at least substantially perpendicular to the plane of the playback screen; X′ represents the warped object position of the audio element along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 4. The method of any one of embodiments 1 to 3. Aspect 5 A method according to any one of aspects 1 to 4, wherein the program is an object-based audio program, and step (a) includes parsing the program to identify at least one of the distortion parameters indicated by screen-related metadata of the program. Aspect 6 The program indicates at least two objects, step (a) includes independently determining at least one distortion parameter for each of the objects, and step (b) includes: performing distortion independently on the audio content of each of the object channels to a degree determined at least in part by the at least one distortion degree parameter corresponding to each of the objects; The method of embodiment 5. Aspect 7 7. The method of any one of aspects 1 to 6, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 8 1. A method for generating an object-based audio program, comprising: (a) determining at least one distortion parameter for at least one object; (b) including in the program object channels indicating the objects and screen-related metadata indicating respective distortion parameters for the objects, each of the distortion parameters indicating a maximum degree of distortion to be performed on the object by a playback system; method. Aspect 9 The method of claim 8, wherein the program indicates at least two objects, and the screen-related metadata indicates at least one of the distortion degree parameters for each of at least two of the objects, each of the distortion degree parameters indicating a maximum degree of distortion to be performed on each corresponding object. Aspect 10 A method as described in aspect 8 or 9, wherein step (a) includes determining at least one off-screen distortion parameter for the at least one object, the off-screen distortion parameter indicating at least one characteristic of the off-screen distortion performed on the object by the playback system, and the screen-related metadata included in the program indicating each of the off-screen distortion parameters. Aspect 11 The method of aspect 10, wherein the off-screen distortion parameters control the degree to which the undistorted position of the object along a width axis at least substantially parallel to the plane of the playback screen is distorted as a function of the distance of the distorted position at which the object is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 12 determining a value Xs indicating an undistorted position along the width axis of the object to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and Xwarp represents the raw warped position of the audio element along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the audio element along a depth axis at least substantially perpendicular to the plane of the playback screen; X′ represents the warped object position of the audio element along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 12. The method of any one of embodiments 8 to 11. Aspect 13 13. The method of any one of aspects 8 to 12, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 14 (a) generating an object-based audio program; (b) generating, in response to the object-based audio program, a speaker channel-based program including at least one set of speaker channels intended for playback by loudspeakers positioned at predetermined positions relative to a playback screen, wherein generating the set of speaker channels includes distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter, each distortion degree parameter indicating a maximum degree of distortion to be performed by a playback system on corresponding audio content of the object-based audio program. Aspect 15 15. The method of claim 14, wherein step (b) includes generating the speaker channel-based audio program such that the speaker channel-based audio program includes two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program, and generating at least one other of which sets includes distorting the audio content of the object-based audio program to a degree determined, at least in part, by the distortion degree parameter, and wherein the other of the sets is intended for playback by a loudspeaker located at the predetermined position relative to the playback screen. Aspect 16 A method as described in aspect 14 or 15, wherein step (b) includes determining at least one off-screen distortion parameter, the off-screen distortion parameter indicating at least one characteristic of off-screen distortion by the playback system to the corresponding audio content of the object-based audio program, and step (b) includes off-screen distortion determined at least in part by at least one of the off-screen distortion parameters. Aspect 17 17. The method of claim 16, wherein the off-screen distortion includes distorting the undistorted position of the audio element along a width axis at least substantially parallel to the plane of the playback screen to a degree controlled by the off-screen distortion parameters as a function of the distance of the distorted position at which the audio element is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 18 the distortion step includes determining a value Xs indicating an undistorted position along a width axis at least substantially parallel to the plane of the playback screen of an audio object to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and Xwarp represents the raw warped position of the object along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the object along a depth axis at least substantially perpendicular to the plane of the reproduction screen; X' represents the warped object position of the object along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 18. The method of any one of embodiments 14 to 17. Aspect 19 A method described in any one of aspects 14 to 18, wherein the object-based audio program includes screen-related metadata indicating at least one distortion parameter, and step (b) includes parsing the object-based audio program to identify each of the distortion parameters indicated by the screen-related metadata. Aspect 20 20. The method of any one of aspects 14 to 19, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 21 1. A method of rendering a speaker channel-based program including at least one set of speaker channels representing distorted content, the speaker channel-based program having been generated by processing an object-based audio program, the processing comprising: distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter to generate the set of speaker channels representing distorted content, each distortion degree parameter indicating a maximum degree of distortion to be performed by a playback system on corresponding audio content of the object-based audio program, the rendering method comprising: (a) parsing the speaker channel-based program to identify speaker channels of the speaker channel-based program, including each of the sets of speaker channels exhibiting distorted content; (b) generating speaker feeds for driving loudspeakers positioned at predetermined positions relative to a playback screen in response to at least some of the speaker channels of the speaker channel-based program, including at least one of the sets of speaker channels showing distorted content; method. Aspect 22 The method of aspect 21, wherein the speaker channel-based program is generated by processing the object-based audio program, the processing including performing off-screen distortion of audio content of the object-based audio program to a degree determined at least in part by using the at least one distortion degree parameter and at least one off-screen distortion parameter indicating at least one characteristic of off-screen distortion for the corresponding audio content of the object-based program. Aspect 23 23. The method of claim 21 or 22, wherein the speaker channel-based audio program includes two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program and another of which sets is a set of speaker channels representing distorted content, and step (b) includes selecting one of the sets that is a set of speaker channels representing distorted content. Aspect 24 24. The method of any one of aspects 21 to 23, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 25 a first subsystem configured to parse a multi-channel audio program to identify channels of the program; a processing subsystem coupled to the first subsystem, the processing subsystem configured to perform distortion on audio content of at least one channel of the program to an extent determined at least in part by at least one distortion degree parameter corresponding to the channel, each of the distortion degree parameters indicating a maximum degree of distortion to be performed on corresponding audio content of the program by a playback system; system. Aspect 26 The system of aspect 25, wherein the distortion includes off-screen distortion determined at least in part by at least one off-screen distortion parameter, the off-screen distortion parameter indicating at least one characteristic of off-screen distortion by a playback system to the corresponding audio content of the program. Aspect 27 27. The system of claim 26, wherein the off-screen distortion includes distorting the undistorted position of the audio element along a width axis at least substantially parallel to the plane of the playback screen to a degree controlled by the off-screen distortion parameters as a function of the distance of the distorted position at which the audio element is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 28 determining a value Xs indicating an undistorted position along the width axis of an audio element to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and Xwarp represents the raw warped position of the audio element along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the audio element along a depth axis at least substantially perpendicular to the plane of the playback screen; X′ represents the warped object position of the audio element along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 28. The system of any one of embodiments 25 to 27. Aspect 29 The system of any one of aspects 25 to 28, wherein the program is an object-based audio program, and the first subsystem is configured to parse the program and identify at least one distortion parameter indicated by screen-related metadata of the program. Aspect 30 The system of aspect 29, wherein the program represents at least two objects, the first subsystem is configured to independently determine at least one distortion degree parameter for each of the objects, and the processing subsystem is configured to independently perform distortion on audio content representing each of the objects to a degree determined at least in part by the at least one distortion degree parameter corresponding to each of the objects. Aspect 31 A system described in any one of aspects 25 to 30, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 32 a first subsystem configured to generate an object-based audio program; a second subsystem coupled to the first subsystem configured to generate, in response to the object-based audio program, a speaker channel-based program including at least one set of speaker channels intended for playback by loudspeakers positioned at predetermined positions relative to a playback screen, the second subsystem configured to generate the set of speaker channels including by distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter, each distortion degree parameter indicating a maximum degree of distortion performed by a playback system on corresponding audio content of the object-based audio program. Aspect 33 The system of aspect 32, wherein the second subsystem is configured to generate the speaker channel-based audio program such that the speaker channel-based audio program includes two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program, and generating at least one other of which sets includes distorting the audio content of the object-based audio program to a degree determined, at least in part, by the distortion degree parameter, and the other of which sets is intended for playback by a loudspeaker located at the predetermined position relative to the playback screen. Aspect 34 The system of aspect 32 or 33, wherein the second subsystem is configured to generate the set of speaker channels, including by performing off-screen distortion of the audio content of the object-based audio program, the off-screen distortion being determined, at least in part, by at least one off-screen distortion parameter indicating at least one characteristic of the off-screen distortion. Aspect 35 35. The system of claim 34, wherein the second subsystem is configured to control, in response to the off-screen distortion parameters, the degree to which the undistorted position of the object along a width axis at least substantially parallel to the plane of the playback screen is distorted as a function of the distance of the distorted position at which the object is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 36 the second subsystem determines a value Xs indicative of an undistorted position along a width axis at least substantially parallel to a plane of the playback screen of an audio element to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and configured to perform the warping including by determining Xwarp represents the raw warped position of the object along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the object along a depth axis at least substantially perpendicular to the plane of the reproduction screen; X' represents the warped object position of the object along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 36. The system of any one of aspects 32 to 35. Aspect 37 37. The system of any one of aspects 32 to 36, wherein the object-based audio program includes screen-related metadata indicating the at least one distortion parameter, and the second subsystem is configured to parse the object-based audio program to identify each of the distortion parameters indicated by the screen-related metadata. Aspect 38 38. The system of any one of aspects 32 to 37, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 39 1. A system for rendering a speaker channel-based program including at least one set of speaker channels representing distorted content, the speaker channel-based program having been generated by processing an object-based audio program, the processing comprising: distorting audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter to generate the set of speaker channels representing distorted content, each distortion degree parameter indicating a maximum degree of distortion to be performed by a playback system on corresponding audio content of the object-based audio program, the system comprising: a first subsystem configured to parse the speaker channel-based program to identify speaker channels of the speaker channel-based program, including each of the set of speaker channels exhibiting distorted content; a rendering subsystem coupled to the first subsystem, configured to generate speaker feeds for driving loudspeakers positioned at predetermined positions relative to a playback screen in response to at least some of the speaker channels of the speaker channel-based program, including at least one of the sets of speaker channels showing distorted content; system. Aspect 40 The system of aspect 39, wherein the speaker channel-based audio program includes two or more selectable sets of speaker channels, at least one of which represents undistorted audio content of the object-based audio program and another of which is a set of speaker channels representing distorted content, and the first subsystem is configured to select one of the sets, which is the set of speaker channels representing distorted content, for rendering by the rendering subsystem. Aspect 41 41. The system of claim 39 or 40, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 42 Buffer memory and; at least one processing subsystem coupled to the buffer memory, the buffer memory stores at least one segment of an object-based audio program, the segment including audio content of at least one object channel representing at least one object, and screen-related metadata indicating at least one distortion degree parameter for at least one of the objects, each distortion degree parameter indicating a maximum degree of distortion to be performed on the object by a playback system; the processing subsystem is coupled and configured to perform at least one of rendering the object-based audio program using at least a portion of the screen-related metadata, generating the object-based audio program, or decoding the object-based audio program. Audio processing unit. Aspect 43 An audio processing unit as described in aspect 42, wherein the program indicates at least two objects, and the screen-related metadata indicates at least one distortion parameter for each of at least two of the objects, each distortion parameter indicating a maximum degree of distortion to be performed for each corresponding object. Aspect 44 An audio processing unit as described in aspect 42 or 43, wherein the segment of the object-based audio program stored in the buffer memory indicates at least one off-screen distortion parameter for the at least one object, the off-screen distortion parameter indicating at least one characteristic of off-screen distortion to be performed on the object by the playback system, and the screen-related metadata included in the program indicates each of the off-screen distortion parameters. Aspect 45 An audio processing unit as described in aspect 44, wherein the off-screen distortion parameters control the degree to which the undistorted position of the object along a width axis at least substantially parallel to the plane of the playback screen is distorted as a function of the distance of the distorted position at which the object is to be rendered at least substantially perpendicular to the plane of the playback screen. Aspect 46 determining a value Xs indicating an undistorted position along the width axis of the object to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and Xwarp represents the raw warped position of the audio element along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the audio element along a depth axis at least substantially perpendicular to the plane of the playback screen; X′ represents the warped object position of the audio element along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 46. ​​An audio processing unit according to any one of aspects 42 to 45. Aspect 47 47. The audio processing unit of any one of aspects 42 to 46, wherein the audio processing unit is an encoder and the processing subsystem is configured to generate the object-based audio program. Aspect 48 47. The audio processing unit of any one of aspects 42 to 46, wherein the audio processing unit is a decoder and the processing subsystem is configured to decode the object-based audio program. Aspect 49 49. An audio processing unit according to any one of aspects 42 to 48, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program. Aspect 50 Buffer memory and; at least one processing subsystem coupled to the buffer memory, the buffer memory stores at least one segment of a speaker channel-based audio program, the segment including audio content of at least one set of speaker channels of the speaker channel-based program intended for playback by loudspeakers positioned at predetermined positions relative to a playback screen, the set of speaker channels being generated in response to an object-based audio program, the generating including by distorting the audio content of the object-based audio program to a degree determined at least in part by at least one distortion degree parameter, each distortion degree parameter indicating a maximum degree of distortion to be performed by a playback system on corresponding audio content of the object-based audio program; the processing subsystem is configured to perform at least one of rendering the speaker channel-based audio program or decoding the speaker channel-based audio program; Audio processing unit. Aspect 51 An audio processing unit as described in aspect 50, wherein at least one of the segments of the speaker channel-based audio program stored in the buffer memory includes audio content of two or more selectable sets of speaker channels, at least one of which sets represents undistorted audio content of the object-based audio program, and at least one other of which sets represents audio content of the object-based audio program that was generated in response to the object-based audio program, including at least in part by distorting the audio content of the object-based audio program to a degree determined by the at least one distortion degree parameter. Aspect 52 52. The audio processing unit of embodiment 50 or 51, wherein the set of speaker channels is generated by a process including performing off-screen distortion determined at least in part by at least one off-screen distortion parameter. Aspect 53 An audio processing unit as described in aspect 52, wherein the off-screen distortion includes distorting the undistorted position of the audio element along a width axis at least substantially parallel to the plane of the playback screen to a degree controlled by the off-screen distortion parameters as a function of the distance of the distorted position at which the audio element is to be rendered, at least substantially perpendicular to the plane of the playback screen. Aspect 54 The set of speaker channels includes determining a value Xs that indicates an undistorted position along a width axis at least substantially parallel to the plane of a playback screen of an audio object to be rendered at a distorted position along the width axis; Xwarp YFACTOR=y EXP and X'=x*YFACTOR+(1-YFACTOR)*[XFACTOR*Xwarp+(l-XFACTOR)*x)] and is produced by a process involving the execution of distortions, including the determination of Xwarp represents the raw warped position of the object along the width axis relative to the edge of the playback screen; EXP is an off-screen distortion parameter, YFACTOR indicates the degree of distortion along the width axis as a function of the distorted position y of the object along a depth axis at least substantially perpendicular to the plane of the reproduction screen; X' represents the warped object position of the object along the width axis relative to the edge of the playback screen; XFACTOR is one of the distortion parameters. 54. An audio processing unit according to any one of aspects 50 to 53. Aspect 55 An audio processing unit described in any one of aspects 50 to 54, wherein the object-based audio program includes screen-related metadata indicating at least one distortion parameter, and the set of speaker channels was generated by a process including a step of parsing the object-based audio program to identify each of the distortion parameters indicated by the screen-related metadata. Aspect 56 56. The audio processing unit of any one of aspects 50 to 55, wherein the audio processing unit is a decoder. Aspect 57 57. An audio processing unit according to any one of aspects 50 to 56, wherein each of the distortion degree parameters is a non-binary value indicating the maximum degree of distortion performed by the playback system on the corresponding audio content of the program.

Claims

1. 1. A method for processing an object-based audio program, comprising: decoding the object-based audio program to generate a plurality of audio objects, each audio object including audio content and an audio object position; determining a distortion parameter indicated by metadata of the object-based audio program; performing distortion on at least one audio object of the object-based audio program to a degree determined at least in part by the distortion degree parameter and a position of the at least one audio object to determine at least one distorted audio object position, wherein the distortion degree parameter indicates a maximum degree of distortion to be performed on the audio object by a playback system, and possible values ​​for the degree determined at least in part by the distortion degree parameter are no distortion, the maximum degree of distortion, and one or more intermediate values ​​between no distortion and the maximum degree of distortion. method.

2. 2. The method of claim 1, wherein the distortion degree parameter is a non-binary value indicating the maximum degree of distortion performed on the audio object by the playback system.

3. 3. A system for processing object-based audio programs, comprising one or more processors configured to perform the method of claim 1 or 2.

4. 3. A storage medium having a software program adapted for execution on a processor for performing the method steps of claim 1 or 2 when executed on a computing device.

5. A computer program product having executable instructions for carrying out the method of claim 1 or 2 when run on a computer.

Citation Information

Patent Citations

  • PCT/US2011/028783

  • Proton, or mixed proton and electronic conducting thin films

    WO2011119041A1