Method and apparatus for reproduction of high-order ambisonics audio signal

The method adapts spatial audio content to different screen sizes by warping and distorting sound field coefficients, ensuring accurate sound localization and balance across varying screen sizes.

JP2025111780APending Publication Date: 2025-07-30DOLBY INTERNATIONAL AB
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025076515
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2012-03-06
Filing Date
2025-05-02
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Combining video playback on screens of various sizes with flexible and universal spatial audio representation is cumbersome due to the inflexible adaptation of spatial sound reproduction to speaker setups, leading to mismatches between sound and image positions, especially in environments with different screen sizes.

Method used

A method and apparatus that adapt spatial audio content, encoded as sound field decomposition coefficients, to match the positions of sound objects with their corresponding visible positions on different video screen sizes by applying a space warping process and distortion functions, ensuring accurate sound localization.

Benefits of technology

Ensures faithful playback of spatial audio by aligning sound object positions with video objects, maintaining sound balance and reducing playback errors across varying screen sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111780000001_ABST
    Figure 2025111780000001_ABST
Patent Text Reader

Abstract

To provide a method for decoding a Higher Order Ambisonics (HOA) signal that adapts a spatial audio content, represented as coefficients of a sound field decomposition, to video screens of various sizes, and matches the sound reproduction positions of objects on the screen to their corresponding visible positions.SOLUTION: A method includes HOA-decoding HOA-represented signals from a storage device 92 storing HOA-encoded signals in an HOA decoder 93, passing through a distortion stage 94 to a renderer 95, and outputting as speaker signals 91 for a set of speakers. The distortion stage 94 receives playback adaptation information 90 derived from the difference in width, possible height and possible curvature between the original screen and the current screen, and applies this to adapt the decoded HOA signals by warping them in the spatial domain.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for playing back an original higher-order ambisonics audio signal assigned to a video signal that is to be presented on a current screen but was generated for a different screen than the original one.

Background Art

[0002] One way to store and process the three-dimensional sound field of a spherical microphone array is the higher-order ambisonics (HOA) representation. Ambisonics uses orthonormal spherical functions to describe the region around a reference point in space, also known as the origin or sweet spot, and the sound field at that point. The accuracy of such a description is determined by the ambisonics order N. Here, a finite number of ambisonics coefficients describe the sound field. The maximum ambisonics order of a spherical array is limited by the number of microphone capsules. The number of microphone capsules must be O = (N + 1) 2 or more. The advantage of such an ambisonics representation is that the reproduction of the sound field can be individually adapted to almost any given speaker position arrangement.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Patent Document 5

Non-Patent Documents

[0004] [Non-Patent Document 1] Sandra Brix, Thomas Sporer, Jan Plogsties, "CARROUSO - An European Approach to 3D-Audio", Proc. of 110th AES Convention, Paper 5314, 12-15 May 2001, Amsterdam, The Netherlands [Non-Patent Document 2] Ulrich Horbach, Etienne Corteel, Renato S. Pellegrini and Edo Hulsebos, "Real-Time Rendering of Dynamic Scenes Using Wave Field Synthesis", Proc. of IEEE Intl. Conf. on Multimedia and Expo (ICME), pp.517-520, August 2002, Lausanne, Switzerland [Non-Patent Document 3] Richard Schultz-Amling, Fabian Kuech, Oliver Thiergart, Markus Kallinger, "Acoustical Zooming Based on a Parametric Sound Field Representation", 128th AES Convention, Paper 8120, 22-25 May2010, London, UK [Non-Patent Document 4] Franz Zotter, Hannes Pomberger, Markus Noisternig, "Ambisonic Decoding With and Without Mode-Matching: A Case Study Using the Hemisphere", Proc. of the 2nd International Symposium on Ambisonics and Spherical Acoustics, 6-7 May 2010, Paris, France

Summary of the Invention

Problems to be Solved by the Invention

[0005] Combining video playback on screens of various sizes while facilitating a flexible and universal representation of spatial audio that is almost independent of the speaker setup can be cumbersome because the spatial sound reproduction is not properly adapted.

[0006] Stereo and surround sound are based on discrete speaker channels, and there are very specific rules regarding where to place the speakers in relation to the video display. For example, in a theater environment, the center speaker is located at the center of the screen, and the left and right speakers are located on the left and right sides of the screen. Thereby, the speaker setup is inherently scaled with the screen. That is, for a small screen, the speakers are closer to each other, and for a large screen, the speakers are farther apart. This has the advantage that the sound mixing can be done in a very coherent way. That is, sound objects related to visible objects on the screen can be localized in a reliable manner between the left, center, and right channels. Thus, the listener's experience matches the creative intent of the sound artist from the mixing stage.

[0007] However, such advantages are also the disadvantages of channel-based systems. That is, the flexibility to change speaker placement is very limited. This disadvantage increases as the number of speaker channels increases. For example, 7.1 and 22.2 formats require precise placement of individual speakers and it is extremely difficult to adapt audio content to non-optimal speaker positions.

[0008] Another disadvantage of channel-based formats is that due to precedence effects, the ability to pan sound objects between the left, center, and right channels is limited, especially for large listening setups such as in a theater environment. For listening positions off-center, panned audio objects may "enter" the speaker closest to the listener. Thus, many movies have been mixed with important screen-related sounds, especially dialogue, mapped only to the center channel. Thereby, a very stable localization of such sounds on the screen is obtained, but at the cost that the spread of the overall sound scene is not optimal.

[0009] A similar compromise is typically made for the rear surround channels. The precise positions of the speakers reproducing those channels are mostly unknown in production, and since the density of those channels is rather low, typically only ambient sounds and uncorrelated items are mixed into the surround channels. Thereby, the possibility of significant playback errors in the surround channels can be reduced, but at the cost that discrete sound objects cannot be faithfully placed anywhere other than on the screen (and even further, as discussed above, in the center channel).

[0010] As described above, the combination of spatial audio with video playback on screens of various sizes can be cumbersome because the spatial sound reproduction is not properly adapted. Depending on whether the actual screen size matches the screen size used in production, the direction of the sound object may deviate from the direction of the visible object on the screen. For example, if mixing is performed in an environment with a small screen, the sound object (e.g., the actor's voice) associated with the screen object will be localized within a relatively narrow cone as seen from the position of the mixer. If this content is mastered for a soundfield-based presentation and played back in a theater environment with a much larger screen, a significant mismatch will occur between the wide field of view to the screen and the narrow cone of the sound object related to the screen. A large mismatch between the position of the visible image of the object and the position of the corresponding sound is cumbersome for the viewer and thus has a profound impact on the perception of the movie.

[0011] More recently, a parametric or object-oriented representation of an audio scene has been proposed that describes the audio scene by the composition of individual audio objects together with a set of parameters and characteristics. For example, in Non-Patent Documents 1, 2, etc., an object-oriented scene description has been proposed mainly for wave-field synthesis systems.

[0012] Patent Document 1 describes two different approaches to address the problem of adapting audio playback to a visible screen size. The first approach individually determines the playback position for each sound object depending on parameters such as its direction and distance to a reference point, aperture angles, and the positions of both the camera and the projection equipment. In practice, a tight coupling between such object visibility and sound mixing is not typical. On the contrary, some deviation of the sound mix from the related visible objects may actually be acceptable for artistic reasons. Furthermore, it is important to distinguish between direct sound and ambient sound. Finally, although not least important, the incorporation of physical camera and projection parameters is rather complex and such parameters are not always available. The second approach (see Claim 16) describes a pre - calculation of sound objects based on the above procedure, assuming a screen of a fixed reference size. This method requires a linear scaling of any position parameter (in Cartesian coordinates) to adapt the scene to a screen larger or smaller than the reference screen. However, this means that when adapting to a screen twice the size, the virtual distance to the sound object also doubles. This is merely a "breathing" of the acoustic scene with no change in the angular position of the sound object for the listener at the reference seat (i.e., the sweet spot). With this approach, it is not possible to generate faithful listening results for changes in the relative size (aperture angle) of the screen in angular coordinates.

[0013] Another example of an object - oriented sound scene description format is described in Patent Document 2. Here, an audio scene includes, in addition to various sound objects and their characteristics, information about the characteristics of the room to be reproduced and information about the opening angles in the horizontal and vertical directions of the reference screen. In a decoder, similar to the principle of Patent Document 1, the actual available screen position and size are determined, and the reproduction of sound objects is individually optimized to match the reference screen.

[0014] For example, in Patent Document 3, a sound - field - oriented audio format such as higher - order ambisonics HOA has been proposed for the universal spatial representation of a sound scene. Regarding recording and reproduction, the sound - field - oriented processing provides an excellent trade - off between universality and practicality. This is because, similar to the object - oriented format, it can be scaled to virtually any spatial resolution. On the other hand, in contrast to the completely synthetic representation required by the object - oriented format, there are some straightforward recording and generation techniques that allow for deriving a natural recording of an actual sound field. Clearly, since sound - field - oriented audio content does not contain information about individual sound objects, the mechanism introduced above for adapting the object - oriented format to different screen sizes is not applicable.

[0015] Currently, there are only a few publications available that describe means for manipulating the relative positions of individual sound objects included in a sound - field - oriented audio scene. The family of algorithms described in Non - Patent Document 3 requires decomposing the sound field into a limited number of discrete sound objects. The position parameters of these sound objects can be manipulated. This approach has the drawback that audio scene decomposition is prone to errors, and any error in the discrimination of audio objects is likely to lead to artifacts in sound rendering.

[0016] Many publications, such as Non-Patent Document 1 and Non-Patent Document 4 mentioned above, are related to the optimization of the "flexible playback layout" of HOA content playback. These techniques address the problem of using speakers at irregular intervals, but none of them aim to change the spatial composition of the audio scene.

Means for Solving the Problem

[0017] The problem solved by the present invention is to adapt spatial audio content expressed as coefficients of sound field decomposition to video screens of various sizes so that the sound playback positions of objects on the screen match the corresponding visible positions. This problem is solved by the method disclosed in claim 1. An apparatus using this method is disclosed in claim 2.

[0018] The present invention allows the playback of spatially sound field-directed audio to be systematically adapted to the playback of its linked visible objects. Thereby, significant essential requirements for faithful playback of spatial audio for movies are met.

[0019] According to the present invention, by applying a space warping process as disclosed in Patent Document 4 in combination with a sound field-directed audio format as disclosed in Patent Documents 3 and 5, the sound field-directed audio scene is adapted to various video screen sizes. One advantageous process is to encode the reference size of the screen (or the viewing angle such as the reference listening position) used in content production and transmit it as metadata together with the content.

[0020] [[ID=I8]] Alternatively, a fixed reference screen size is assumed for encoding and decoding, and the decoder knows the actual size of the target screen. The decoder distorts the sound field in such a way that all sound objects in the screen direction are compressed or stretched according to the ratio of the size of the target screen to the size of the reference screen. This can be achieved, for example, using a simple two-segment piecewise linear distortion function as described below. In contrast to the prior art above, this stretching is basically limited to the angular position of the sound item and does not necessarily lead to a change in the distance to the listening area of the sound object.

[0021] Some embodiments of the present invention are described below that allow for taking control over which parts of the audio scene are manipulated or not.

[0022] In principle, the method of the present invention is suitable for playing back a original high-order ambisonics audio signal assigned to a video signal that is to be presented on a current screen but was generated for a different original screen. The method comprises: · decoding the high-order ambisonics audio signal to provide a decoded audio signal; · receiving or establishing playback adaptation information derived from the difference in width and possibly height and possibly curvature between the original screen and the current screen; · adapting the decoded audio signal by distorting it in the spatial domain, wherein the playback adaptation information controls the distortion such that for a viewer of the current screen and a listener of the adapted decoded audio signal, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position on the screen of the associated video object; · rendering and outputting the adapted decoded audio signal for speakers.

[0023] In principle, the device of the present invention is suitable for reproducing the original high-order ambisonic audio signal assigned to a video signal that should be presented on the current screen but was generated for a different original screen. The device comprises: · means adapted to decode the high-order ambisonic audio signal to provide a decoded audio signal; · means adapted to receive or establish reproduction adaptation information derived from the difference in width and possibly height and possibly curvature between the original screen and the current screen; · means adapted to adapt the decoded audio signal by distorting it in the spatial domain, wherein the reproduction adaptation information controls the distortion such that, for a viewer of the current screen and a listener of the adapted decoded audio signal, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position of the associated video object on the screen; · means for rendering and outputting the adapted decoded audio signal for speakers.

[0024] Advantageous additional embodiments of the present invention are disclosed in the respective dependent claims.

Brief Description of the Drawings

[0025] Exemplary embodiments of the present invention will be described with reference to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0026] Figure 1 shows an exemplary studio environment with a reference point and a screen, and Figure 2 shows an exemplary cinema environment with a reference point and a screen. Different projection environments lead to different opening angles of the screen when viewed from the reference point. In the sound field-directed playback techniques of the current state of the art, audio content produced in a studio environment (opening angle 60°) does not match the screen content in a cinema environment (opening angle 90°). In order to allow adapting the content to the different characteristics of the playback environment, it is necessary to transmit the opening angle of 60° in the studio environment together with the audio content.

[0027] For clarity, these drawings simplify the situation to a 2D scenario.

[0028] In the higher-order ambisonics theory, a spatial audio scene is described via the coefficients A n m (k). For a volume without a source, the sound pressure is described as a function of spherical coordinates (radial distance r, inclination angle θ, azimuth angle φ, and spatial frequency k = ω / c (c is the speed of sound in air).

[0029]

Equation

[0030] The spatial composition of the audio scene can be distorted by the techniques disclosed in Patent Document 4.

[0031] The relative positions of the sound objects included in the two-dimensional or three-dimensional higher-order ambisonics (HOA) representation of the audio scene can be changed. Here, an input vector A in with magnitude O in determines the coefficients of the Fourier series of the input signal, and an output vector A out with magnitude O out correspondingly determines the coefficients of the Fourier series of the changed output signal. The input vector A of the input HOA coefficients in is decoded into the input signal s -1 in the spatial domain for regularly positioned speaker positions by using the inverse matrix Ψ1 of the mode matrix Ψ1 in such that s -1 = Ψ1 in A in is calculated. The input signal s in is distorted, encoded, and adapted to the output vector A out of the output HOA coefficients in the spatial domain by calculating A in = Ψ2s out Here, the mode vectors of the mode matrix Ψ2 are modified according to the distortion function f(φ) that maps the angles of the original speaker positions one-to-one to the target angles of the target speaker positions in the output vector A out

[0032] The modification of the speaker density can be canceled by applying a gain weighting function g(φ) to the virtual speaker output signal s in to give s out as a result. In principle, any weighting function g(φ) can be specified. One particularly advantageous variant is proportional to the derivative of the distortion function f(φ) ​

Mathematics

[0033]

Mathematics

[0034] Decoding, weighting, and warping / decode are usually performed by using a transformation matrix T = diag(w)Ψ2diag(g)Ψ1 warp ×O warp of size O. Here, diag(w) represents a diagonal matrix having the values of the window vector w as its main diagonal components, and diag(g) represents a diagonal matrix having the values of the gain function g as its main diagonal components. To reshape the transformation matrix T to obtain size O -1 ×O out ×O in the corresponding columns and / or rows of the transformation matrix T are removed so as to execute a spatial warping operation A out = TA in

[0035] ​Figures 3-7 illustrate spatial distortion in the two-dimensional (circular) case, showing an exemplary piecewise linear distortion function for the scenario in Figures 1 / 2 and its effect on the panning function of 13 exemplary regularly spaced loudspeakers. The system stretches the front sound field by a factor of 1.5 to accommodate the larger screen in a cinema. Sound items coming from other directions are therefore compressed. The distortion function f(φ) resembles the phase response of a discrete-time all-pass filter with a single real-valued parameter and is shown in Figure 3. The corresponding weighting function g(φ) is shown in Figure 4.

[0036] Figure 7 illustrates a 13x65 single-step transform distortion matrix T. The log-magnitude values of the individual coefficients of this matrix are shown by grayscale or shading based on the accompanying grayscale or shading bar. This example matrix is N orig = 6 input HOA order and N warp The eigenvalues are designed for an output order of 32. This higher output order is needed to capture most of the information that is spread by the transformation from lower order coefficients to higher order coefficients.

[0037] A useful property of this particular warping matrix is that a significant portion of it is zero, which allows for a significant saving of computational power when implementing this operation.

[0038] Figures 5 and 6 show the distortion characteristics of beam patterns generated by several plane waves. Both figures result from the same 13 input plane waves, all with the same amplitude "1", at φ positions 0, 2 / 13π, 4 / 13π, 6 / 13π, ..., 22 / 13π and 24 / 13π, resulting in 13 different angular amplitude distributions, i.e., the overdetermined regular decoding operation s=Ψ. -1It shows the result vector s of A. Here, the HOA vector A is either the original or distorted variant of the above set of plane waves. The numbers outside the circle represent the angle φ. The number of virtual speakers is considerably larger than the number of HOA parameters. The amplitude distribution or beam pattern for plane waves coming from the front is positioned at φ = 0.

[0039] Figure 5 shows the weights and amplitude distribution of the original HOA representation. All 13 distributions have a similar shape and are characterized by the same width of the main lobe. Figure 6 shows the weights and amplitude distribution for the same sound object, but after the distortion operation has been performed. The object is away from the front direction at φ = 0 degrees, and the main lobes around the front direction are wider. These modifications of the beam pattern are facilitated by the higher order N warp = 32 of the distorted HOA vector. A mixed - order signal with a locally varying order over space is generated.

[0040] To derive the distortion characteristics f(φ in )) suitable for adapting the playback of the audio scene to the actual screen configuration, additional information is sent or provided in addition to the HOA coefficients. For example, the following characteristics of the reference screen used in the mixing process can be included in the bitstream: · The direction of the screen center, · The width, · The height of the reference screen. These are all represented in polar coordinates measured from the reference listening position (also called the "sweet spot").

[0041] Furthermore, the following parameters may be required for special applications: · The shape of the screen, e.g., planar or spherical, · The distance of the screen, · Information about the maximum and minimum visible depths in the case of stereoscopic 3D video projection.

[0042] It is known to those skilled in the art how such metadata can be encoded.

[0043] In the following, it is assumed that the encoded audio bitstream includes at least the above three parameters, namely the direction, width, and height of the center of the reference screen. For clarity, it is further assumed that the center of the actual screen is the same as the center of the reference screen, for example, directly in front of the listener. Furthermore, it is assumed that the sound field is represented only in a 2D format (not in a 3D format), and changes in inclination in this regard are ignored (for example, when the selected HOA format does not represent the vertical component, or when the sound editor determines that the mismatch between the picture and the inclination of the sound source on the screen is small enough not to be noticed by casual observers). The transition to any screen position and to the 3D case is straightforward for those skilled in the art. Furthermore, for simplicity, it is assumed that the screen configuration is spherical.

[0044] Using these assumptions, only the width of the screen can vary between the content and the actual setup. In the following, a suitable two-segment piecewise-linear distortion characteristic is defined. The actual screen width is defined by the opening angle 2φ w,a (i.e., φ w,a represents the one-sided angle). The reference screen width is defined by the angle φ w,r , which is part of the meta-information delivered within the bitstream. For faithful reproduction of the sound object in the forward direction, i.e., on the video screen, all positions of the sound object (in polar coordinates) are multiplied by the factor φ w,a / φ w,r . Conversely, all sound objects in other directions are moved according to the remaining space. The distortion characteristic leads to the following results.

[0045]

Equation

[0046] Furthermore, the selected HOA representation also has provisions for tilt. If the sound editor is interested in the angle in the vertical direction where the screen is stretched, the screen corner height θ h (one-sided height) and related factors (e.g., the ratio of the actual height to the reference height θ h,a / θ h,r ) A similar formula based on can be applied as part of the distortion operator to the above tilt.

[0047] As another example, assuming a flat screen instead of a spherical screen in front of the listener, more sophisticated distortion characteristics may be required than in the above example. Again, this can relate to width-only or width+height distortion.

[0048] The above exemplary embodiments are fixed and have the advantage that the implementation is quite simple. On the other hand, no control of the adaptation process from the production side is allowed. The following embodiments introduce processing for more control in different ways.

Example

[0049] Example 1: Separation between sound related to the screen and other sounds There can be various reasons why such control techniques are required. For example, not all sound objects in an audio scene are directly coupled to visible objects on the screen, and it may be advantageous to manipulate direct sound in a different way from ambient sound. This distinction can be made by scene analysis on the rendering side. However, it can be significantly improved and controlled by adding additional information to the transmission bitstream. Ideally, the decision of which sound items should be adapted to the actual screen characteristics - and which sound items should be left untouched - should be entrusted to the artist performing the sound mixing.

[0050] There are different ways to transmit this information to the rendering process.

[0051] ● Two full sets of HOA coefficients (signals) are defined within the bitstream, one for describing objects related to visible items and the other for representing independent or ambient sounds. At the decoder, only the first HOA signal undergoes adaptation to the actual screen geometry, while the other remains untouched. Before playback, the manipulated first HOA signal and the unmodified second HOA signal are combined.

[0052] As an example, a sound engineer may decide to mix screen-related sounds such as dialogue or individual sound effect items into the first signal and ambient sounds into the second signal. In this way, the ambient sound remains the same regardless of which screen is used for the playback of the audio / video signal.

[0053] This kind of processing has the additional advantage that the HOA orders of the two constituent sub-signals can be individually optimized, such that the HOA order for the sound object related to the screen (i.e., the first sub-signal) is higher than that used for the peripheral signal components (i.e., the second sub-signal).

[0054] ● Through flags attached to the time-space-frequency tiles, it is defined whether the mapping of the sound is related to the screen or is independent. For this purpose, the spatial characteristics of the HOA signal are determined, for example, via plane-wave decomposition. Then each of the spatial region signals is input into time segmentation (windowing) and time-frequency conversion. Thereby, a three-dimensional set of tiles is defined, and those tiles can be individually marked by a binary flag stating, for example, whether the content of that tile should be adapted to the actual screen geometry. This sub-embodiment is more efficient than the previous sub-embodiment, but limits the flexibility to define which parts of the sound scene should or should not be manipulated.

Example

[0055] Example 2: Dynamic Adaptation In some applications, it may be necessary to change the reference screen characteristics being signaled in a dynamic manner. For example, audio content may be the result of concatenating content segments diverted from various mixes. In this case, the parameters describing the reference screen parameters change over time and the adaptation algorithm is changed dynamically. That is, for each change in the screen parameters, the applied distortion function is recalculated as appropriate.

[0056] Another application example results from mixing different HOA streams prepared for different sub - parts of the final visible video and audio scenes. In this case, it is advantageous to allow two or more (or three or more in the above Example 1) HOA signals, each having individual screen characteristics, within a common bitstream.

Example

[0057] Example 3: Alternative implementation Instead of distorting the HOA representation prior to decoding via a fixed HOA decoder, information about how to adapt the signal to the actual screen can also be integrated into the decoder design. This implementation is an alternative to the basic realization described in the above exemplary embodiments. However, this does not change the signaling of the screen characteristics within the bitstream.

[0058] In FIG. 8, the HOA - encoded signal is stored in the storage device 82. For presentation in a cinema, the HOA - represented signal from the device 82 is HOA - decoded in the HOA decoder 83, passes through the renderer 85, and is output as the speaker signal 81 for a set of speakers.

[0059] In FIG. 9, the HOA - encoded signal is stored in the storage device 92. For example, for presentation in a cinema, the HOA - represented signal from the device 92 is HOA - decoded in the HOA decoder 93, enters the renderer 95 through the distortion stage 94, and is output as the speaker signal 91 for a set of speakers. The distortion stage 94 receives the above - mentioned playback adaptation information 90 and uses it to appropriately adapt the decoded HOA signal.

[0060] Some supplementary notes are described below. 〔Supplementary Note 1〕 A method for playing an original high-order ambisonics audio signal assigned to a video signal that should be presented on the current screen but was generated for a different original screen, the method comprising: · Decoding the high-order ambisonics audio signal to provide a decoded audio signal; · Receiving or establishing playback adaptation information derived from the difference in width and possibly height and possibly curvature between the original screen and the current screen; · Adapting the decoded audio signal by distorting it in the spatial domain, wherein the playback adaptation information controls the distortion such that, for a viewer of the current screen and a listener of the adapted decoded audio signal, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position on the screen of the associated video object; · Rendering and outputting the adapted decoded audio signal for speakers. Method. [Appendix 2] The method according to Appendix 1, wherein the high-order ambisonics audio signal includes a plurality of audio objects assigned to a corresponding video object, and for a viewer and a listener of the current screen, the angle or distance of the audio object is different from the respective angle or distance of the video object on the original screen. [Appendix 3] The method according to Appendix 1 or 2, wherein the bitstream carrying the original high-order ambisonics audio signal also includes the playback adaptation information. [Appendix 4] The method according to any one of Appendices 1 to 3, wherein, in addition to the distorting, weighting by a gain function is performed to obtain a resulting uniform sound amplitude per opening angle. [Appendix 5] Two full coefficient sets of a high-order ambisonic audio signal are decoded, the first audio signal represents an object related to a visible object, the second audio signal represents an independent or ambient sound, and only the first decoded audio signal is subject to adaptation to the actual screen geometry by distortion, the second decoded audio signal is left untouched, and before playback, the adapted first decoded audio signal and the non-adapted second decoded audio signal are combined, the method according to any one of Appendices 1 to 4. [Appendix 6] The method according to Appendix 5, wherein the HOA orders of the first and second audio signals are different. [Appendix 7] The method according to any one of Appendices 1 to 6, wherein the playback adaptation information is dynamically changed. [Appendix 8] A playback device for a high-order ambisonic audio signal assigned to a video signal that should be presented on the current screen but was generated for a different original screen, the device comprising: · means adapted to decode the high-order ambisonic audio signal to provide a decoded audio signal; · means adapted to receive or establish playback adaptation information derived from the difference in width and possibly height and possibly curvature between the original screen and the current screen; · means adapted to adapt the decoded audio signal by distorting it in the spatial domain, wherein the playback adaptation information controls the distortion such that, for a viewer of the current screen and a listener of the adapted decoded audio signal, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position of the related video object on the screen; · means for rendering and outputting the adapted decoded audio signal for speakers. Device [Supplementary Note 9] The device according to Supplementary Note 8, wherein the high-order ambisonics audio signal includes a plurality of audio objects assigned to a corresponding video object, and for a viewer and a listener viewing the current screen, an angle or a distance of the audio object is different from an angle or a distance of the video object on the original screen [Supplementary Note 10] The device according to Supplementary Note 8 or 9, wherein a bitstream carrying the original high-order ambisonics audio signal further includes the playback adaptation information [Supplementary Note 11] The device according to any one of Supplementary Notes 8 to 10, wherein, in addition to distorting, weighting by a gain function is performed so as to obtain a resultant uniform sound amplitude per opening angle [Supplementary Note 12] The device according to any one of Supplementary Notes 8 to 11, wherein two full coefficient sets of high-order ambisonics audio signals are decoded, a first audio signal represents an object related to a visible object, a second audio signal represents an independent or ambient sound, only the first decoded audio signal undergoes adaptation to an actual screen geometry by distorting, the second decoded audio signal is left untouched, and before playback, the adapted first decoded audio signal and the unadapted second decoded audio signal are combined [Supplementary Note 13] The device according to Supplementary Note 12, wherein HOA orders of the first and second audio signals are different [Supplementary Note 14] The device according to any one of Supplementary Notes 8 to 13, wherein the playback adaptation information is dynamically changed [Supplementary Note 15] A method for generating digital audio signal data, the method comprising: - providing data of an original high-order ambisonics audio signal assigned to a video signal; · providing reproduction adaptation information data derived from the width of the original screen on which the video signal can be presented, and possibly the height and possibly the curvature; The reproduction adaptation information data is used to adapt, by distortion in a spatial region, the decoded version of the high-order ambisonics audio signal such that, for a viewer of the video signal on a current screen having a width different from the width of the original screen and a listener of the adapted decoded audio signal, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position of the related video object on the current screen. Method.

Claims

**Claim 1** A method for decoding an encoded higher order ambisonics (HOA) signal that describes a sound field, comprising: decoding, by one or more processors, the encoded HOA to obtain a first set of decoded HOA signals representing dominant components of the sound field and a second set of decoded HOA signals representing ambient components of the sound field; combining, by one or more processors, the first set of decoded HOA signals and the second set of decoded HOA signals to produce a combined set of decoded HOA signals; determining, by one or more processors, a transformation matrix for distorting the combined set of decoded HOA signals, the transformation matrix being based on a production screen size and a target screen size; a method. **Claim 2** A non-transitory computer-readable medium including instructions that, when executed by a processor, perform the method of claim 1. **Claim 3** An apparatus for decoding an encoded higher order ambisonics (HOA) signal that describes a sound field, comprising: an audio decoder that decodes the encoded HOA to obtain a first set of decoded HOA signals representing dominant components of the sound field and a second set of decoded HOA signals representing ambient components of the sound field; a combiner that integrates the first set of decoded HOA signals and the second set of decoded HOA signals to produce a combined set of decoded HOA signals; a processor that determines a transformation matrix for distorting the combined set of decoded HOA signals, the transformation matrix being based on a production screen size and a target screen size; an apparatus.

Citation Information

Patent Citations

  • Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information

    JP2009278381A

  • Audiovisual apparatus

    JP2011188287A

  • Method and apparatus for improved matching of auditory space to visual space in video viewing applications

    US20100328419A1

  • Method and apparatus for improved mactching of auditory space to visual space in video teleconferencing applications using window-based displays

    US20100328423A1

  • Method and device for decoding an audio soundfield representation for audio playback

    WO2011117399A1